Stop the Line

contact center AI operations agent management factory floor

By Damian Mathews & The Last Mile Team

If you’re running more than one AI agent in your contact center, you already have a factory floor.

The question is whether anyone on it can stop the line.

A case study circulating this week details how Toyota North America got to more than fifty AI agents in production. Building one used to take the company six engineers and six months. It now takes one engineer and four days, and every agent ships as a configuration file rather than a custom project.

What makes it worth your attention is the framework they built it on.

Toyota’s enterprise AI team mapped the entire platform onto the Toyota Production System, the manufacturing philosophy the company has been refining for a century.

The Andon board, the visibility system that shows at a glance what’s working and what’s broken, became their observability layer. Jidoka, automation with a human touch, became human-in-the-loop design. Kaizen became the continuous improvement cadence. Genchi Genbutsu, go and see for yourself, became trace-level debugging.

Read that as an operator rather than a technologist.

Every QA rubric, coaching cadence, adherence target, and improvement cycle in your contact center descends from the same source. Lean arrived in service operations from Toyota’s factory floor decades ago. You’ve been running a version of TPS for years without calling it that.

So the playbook for running AI agents at scale turns out to be the one your operation was already built on.

Start with the andon cord, because the mechanics matter more than the metaphor.

Toyota’s system uses two pulls. The first is a request for help. It alerts the team leader without stopping the line, which buys time to assess. If the issue gets resolved, a second pull keeps the line moving. If it doesn’t, the line stops.

Any worker can pull it. That’s the whole point.

Last week in Somebody Has to Watch It we covered OpenAI’s new protocol, which aims to raise an alert within thirty minutes and pauses the activity if the team can’t clear it as a false positive inside that window. That’s an andon cord with a stopwatch on it. Toyota worked the pattern out decades ago for cars.

Two pulls for every agent you run: one that raises a flag without stopping the work, and one that stops it.

Get that right and a bad agent day is a contained incident. Get it wrong and the only person who can stop the line is whoever wrote the deployment ticket, and they’re asleep.

One more line from the Toyota team is worth sitting with. Asked how they got so fast, they said they invested in security, authentication, identity, observability, a tool layer, and guardrails before they tried to scale the number of agents. “We did the boring stuff once.”

That’s the argument from The Pilot That Never Shipped, proven from the other direction. Toyota’s speed came from sequencing. They built the production environment first, which is what turned building inward into a four-day job.

Fifty agents as config files also answers the problem we wrote about earlier this summer, when three people in our own company turned out to be quietly building the same tool. Standardize the plumbing and duplication becomes throughput.

This is what a run model looks like. Plan, build, run, optimize, on a cadence, with a visibility layer good enough that the person closest to the work can see a defect and act on it. Kaizen with different labels.

Your vendor ships the line. The cord is yours to install.

Who in your operation can stop an agent right now, without asking permission?

— Damian

 
 

 

 

Here’s what went down this week.

Bleeding Edge

Early signals you should keep on your radar.

The right to reach a human is about to become a routing requirement. California’s AB 1609 was amended on the Senate floor on 20 August and ordered to third reading on 24 August; the Legislature adjourns sine die on 31 August. The scope is narrower than the “right to talk to a human” framing suggests: businesses above $500 million in national revenue, paid services only, across a 10-hour window of business hours. Free tiers are out. The duty is a good-faith effort to reach a live agent within 15 minutes, with cumulative hold under an hour. The same amendment dropped the pattern-or-practice predicate and split penalties into $5,000 for a first violation and $10,000 after, so single violations are now actionable. Read it as a design requirement, not a disclosure one. Containment targets acquire a legal ceiling, and a 15-minute clock has to be instrumented and evidenced.

The displacement argument stopped being a forecast. Goldman Sachs Research put US call-center employment 39% below its historical trend on 19 August, with Canada 33% below and Germany 27% below, and estimated that each 10% rise in AI exposure cuts annual staffing growth by 0.2 percentage points, entry-level roles absorbing most of it. This is sell-side in-house research, not a commissioned survey. Pair it with last week’s correction: Challenger has AI cited in 112,713 cuts through July, 24% of the total, against total cuts down 41% year over year and hiring plans up 25%. Both are true. The aggregate labor market is not collapsing; the one occupation our clients staff is already far below trend. That matters if you are defending a business case that assumes future attrition will pay for the AI program

Leading Edge

Proven moves you can copy today.

Amazon Connect will now answer the question instead of charting it. AWS shipped conversational analytics for Connect Customer AI on the 21st of August: supervisors can query more than 150 self-service, agent and queue metrics in plain language and get back prioritized recommendations with confidence scores and projected impact. It is GA in every Region where Connect Customer AI Agents run, so there is no rollout gate. This lands a week after the routing-step reporting we covered last issue, and the two work together: one shows where contacts stall, the other lets you interrogate why without a BI analyst in the loop. Pull up any in-flight custom dashboard work and check you are not about to build what AWS just gave you.

Somebody is running a fully autonomous voice operation at real volume. xAI disclosed on 25 August that Grok Voice handles more than 15,000 inbound Starlink calls a day and completes more than 3,000 orders a week, diagnosing faults and dispatching replacements with no human in the loop. Earlier baselines put containment at 70% and conversion at 20%. All of it is vendor-disclosed and unaudited, and Starlink is a favorable case: high volume, a narrow product surface, diagnostics that reduce to a decision tree. Treat it as an existence proof, not a benchmark you are behind on. It is still the most concrete answer available to the board question about what full autonomy looks like at scale.

Off the Ledge 

Hype and headaches we’re steering clear of.

Colorado’s chatbot law probably does not cover your chatbot. The ADMT Act does. Two corrections to the rulemaking we flagged last issue. The date: the formal period closes 26 October, but comments must land by 4 September to shape the revised draft due 23 September. September buys influence, October buys a record. The scope: proposed Rule 8 exempts narrow, task-based outputs and services not designed to simulate emotional companionship, putting most routine service bots outside the Chatbot Safety Act entirely. The exposure sits in the ADMT Act, which has no carve-out and attaches wherever a system materially influences a consequential decision on lending, insurance, housing, employment, education or health care. A bot that triages toward any of those is in scope. The AG is seeking comment on how duties split between deployers and vendors — the contract question in every CCaaS deal you sign next year.

Human review of automated decisions just got a price. The Dutch data protection authority fined Uber €825 million on 23 August over automated suspension and deactivation of driver accounts between 2018 and 2022 without adequate human review, the second-largest GDPR penalty on record. Deputy chair Monique Verdier put it plainly: a computer should not make decisions on its own that carry consequences of that size. Uber is appealing, so the figure may move. The transferable point has nothing to do with ride-hailing. Automated QA scoring, adherence monitoring, coaching flags and BPO deactivation workflows are the same category of system making the same category of decision about people. If anything in your WFM stack feeds discipline or termination, human review needs documenting rather than asserting.

See you next week!

Sorry, no content found.