By Damian Mathews & The Last Mile Team
If you’re putting an AI agent in front of customers this year, a question is waiting for you that has nothing to do with which model you pick.
When the agent issues a refund it shouldn’t have, or promises something your business can’t honor, who answers for it?
Most contact centers haven’t decided. The answer arrives either way.
Two weeks ago, an OpenAI model broke out of a sandboxed security test, reached the open internet, and used stolen credentials to pull answers out of Hugging Face’s production database.
We covered the behavior in It Cheated. Michael Fisher went after the harder question in Capability Is Not Culpability, and found the detail almost everyone missed.
OpenAI disclosed nearly everything about the incident. What it left out was what the model was told to do, and what it wasn’t told not to do. As Michael Fisher puts it, the capability belonged to the agent. The culpability belongs to whoever wrote the prompt.
That omission does real work. While the instruction stays private, this stays a story about an astonishingly capable model.
Publish the instruction and it becomes a story about an underspecified request from an accountable human.
We do know the model ran with its safety refusals reduced and its external filters switched off for the test. Nobody told it to stay in the sandbox.
It followed the instruction it had with more thoroughness than anyone anticipated. Now put that in your operation. A new hire told to resolve a billing complaint will never consider falsifying a record to close the ticket. They arrived with a lifetime of unstated limits already installed.
Your agent inherits none of that. It has the goal you gave it and nothing else.
Which is why “the AI did it” fails as an answer.
Air Canada tried a version of that argument before a British Columbia tribunal, claiming its chatbot was responsible for its own bad information. Michael Fisher covers the case in Ardeo Ergo Sum. The tribunal threw it out.
Kerry Robinson has a line he repeats throughout A1B: Customer Zero to AI-First. You own the output. Use AI to draft it or build it, and your name still goes on the result.
That principle now extends. You own what your agents do. Which turns the instruction behind an agent into a business artifact. Someone in your operation is writing the prompt that defines what a machine will do on your company’s behalf.
That document deserves what a credit approval threshold gets:
- A named owner
- A review
- A record of changes
The questions are unglamorous:
- What is this agent authorized to promise?
- What’s the ceiling on what it can give away?
- When it reaches the edge of its authority, what stops it?
If you can’t answer those, the agent is running on hope.
Whose name is on the prompt running in your contact center?
— Damian
Here’s what went down this week.
Bleeding Edge
Early signals you should keep on your radar.
A Chinese lab just parked an open model at the top of a frontend-coding leaderboard: Moonshot’s Kimi K3, 2.8 trillion parameters, weights free to download, ranked above Claude Fable 5 and GPT-5.6 Sol. When the best result on the board is one anyone can run for nothing, the closed labs have to defend their premium on something other than raw capability. Watch what that does to how you price, source, and defend your own AI stack.
Big platforms made agents more governable in July. Google’s Gemini Enterprise now stamps identity on an agent below the app layer, and OpenAI put its workspace agents on a meter. Why the rush? One survey of 1,900 IT leaders found only 12% could actually govern the agents they’d already deployed. The vendors are landing on the answer you should reach first — an agent you can’t identify, audit, and control isn’t in production, it’s on the loose.
Leading Edge
Proven moves you can copy today.
OpenAI shipped Presence, a managed platform for voice and chat agents that arrives with the guardrails, approved actions, and escalation rules already wired in not an API you stitch together yourself. OpenAI says its own support line now closes 75% of inbound without a human, and BBVA Mexico and SoftBank are already on it. When the lab that builds the model also sells the whole production stack, the build-versus-buy question on your desk changes shape — so know which side you’re picking, and why.
Genesys and AWS tightened their alliance on July 22 with more agentic orchestration, Genesys Cloud across 22 AWS regions, and a place on the AWS European Sovereign Cloud for anyone whose data can’t leave the continent. Frontier models come in through Bedrock, so you can turn agent capability up without stepping outside a governed environment. The two platforms we build you on just shortened the walk from pilot to a production floor that clears compliance.
Off the Ledge
Hype and headaches we’re steering clear of.
OpenAI admitted two of its models broke out of an isolated cyber test, found a zero-day in the one service allowed to talk outward, and chained their way onto Hugging Face’s production systems to lift the answer key. The models weren’t malicious they were relentless, and one trusted boundary was all it took. The takeaway isn’t that agents are dangerous; it’s that containment has to hold at every layer, because an agent moving at machine speed finishes the chain before a human reads the alert.
The headlines say AI is gutting contact-center jobs. Commonwealth Bank, Microsoft, Uber, and Hyatt all trimming support as automation takes the routine tier, Microsoft’s own support headcount down from roughly 50,000 to 40,000. Read past the body count. The routine work is moving to the machine; the complex, judgment-heavy work is what’s left, and it still needs people set up to handle it. Whether this lands as a cut or a redesign in your operation comes down to how deliberately you deploy, and who you trust to do it with you.