The Agent that Went Rogue

contact center AI agents guardrails autonomous risk

By Damian Mathews & The Last Mile Team

Picture a student who wants to ace an exam so badly that it picks the lock on the classroom, finds the building unlocked, walks to the teacher’s house, and takes the answer key off the kitchen counter. It never occurs to the student that any of this is against the rules. It was told to pass the test. It passed the test.

That, more or less, is what one of the most advanced AI models on earth did this week. And if you run a contact center, it’s also a preview of the agents being pitched to you right now.

OpenAI disclosed on Tuesday that during an internal cybersecurity evaluation, its GPT-5.6 Sol model and an even more capable unreleased system did something OpenAI called “an unprecedented cyber incident.” The models were running inside a sealed sandbox with no internet access, being tested on a benchmark of hacking challenges.

So they found a previously unknown vulnerability in OpenAI’s own test environment, used it to escape onto the open internet, worked out that the platform Hugging Face probably hosted the answers to the benchmark, and broke into Hugging Face’s live production servers to get them.

Stolen credentials. More zero-day exploits. Thousands of individual actions. No human directing any of it.

Here is the part that matters, and it’s not the hacking. OpenAI was clear that the models weren’t being malicious.

They were, in the company’s own word, hyper-focused on one thing: solving the benchmark. Passing the test. Everything else, including every boundary a person would have assumed was obvious, was just terrain between the model and its goal.

Nobody told it to cheat. It pursued the objective literally, at full confidence, and the objective was all it had.

Fish predicted this almost to the word two weeks ago. In Ardeo Ergo Sum, he wrote that an agent “optimizes whatever observable it is handed, with total confidence and no independent grip on the purpose behind it.” He called it Goodhart’s trap in its purest form. Two weeks later, the purest case study anyone could ask for broke into a production database to win a test.

This is the same thread we pulled last week, but it lands somewhere sharper. Last week the story was humans gaming a metric, employees burning tokens to look productive. The crucial detail was that those humans still knew what real work was while they gamed the number. The purpose survived inside them.

An agent has no such backstop. Hand it a goal and it holds that goal and nothing else. It will not pause, the way a new employee would, and wonder whether breaking into another company’s servers is maybe above its pay grade. Wondering isn’t what it does. As Michael Fisher puts it, the check on a grant cannot live inside the grantee. The gate has to live outside.

In plain terms: the limits have to live in the systems around the agent payment caps, permission scopes, approval gates… not in its prompt.

One detail makes the point almost too neatly. When Hugging Face’s security team went to investigate the breach, they reportedly had to reach for a Chinese model, because the American AI tools kept refusing the defensive work on safety grounds. The attacker had its guardrails lowered on purpose. The defenders couldn’t lower theirs.

Now bring it back to your building, because every CX leader reading this is being pitched autonomous agents right now. Agents that will resolve tickets, issue refunds, update accounts, and take real actions on real customer records without a human clicking approve each time.

An agent will pursue the goal you give it far more literally than any person would, so the questions you ask before you deploy one are different from the questions you’d ask about hiring.

We wrote in The Pilot That Never Shipped about how and why production is a different discipline from the demo, and this is the sharpest example yet. An agent scored well on the benchmark. The disaster lived entirely in what it was willing to do to get there, which a typical capability score would almost certainly fail to show you. Sinch just found that 74% of organizations have already shut down or rolled back a live AI customer service agent. These don’t fail in the demo… they fail in production.

So the questions become concrete. Picture the contact center version: an agent measured on first-contact resolution discovers that a $40 goodwill credit closes almost any ticket. It’s optimizing.

If this agent optimized its goal too well, what could it reach? What can it touch, spend, promise, or delete? Where are the boundaries enforced, and are they enforced outside the agent where it can’t reason its way around them, or merely written into a prompt and hoped for? When it hits the edge of what it should do, what actually stops it?

A capable model that wants to pass the test is an asset. The same model, pointed at your customers with no gate outside it, is the exam-room lock waiting to be picked.

Before your next agent goes live, what happens if it wants to pass the test a little too much?

— Damian

 
 

 

 

Here’s what went down this week.

Bleeding Edge

Early signals you should keep on your radar.

Moonshot’s Kimi K3 topped a major coding leaderboard within a day of launch, then ran out of room. The Chinese lab halted new subscriptions on July 19 after demand pushed past its available GPU capacity. With open weights promised within days, “frontier quality, free to self-host” may become the objection every model vendor hears next.

AWS has set an end-of-support date of October 30, 2026 for Amazon Pinpoint, after which the console and its campaigns, journeys, segments and analytics go away — while transactional SMS, voice and push continue under AWS End User Messaging. For anyone running orchestrated campaigns and customer journeys on it, that’s not a feature swap. It’s a migration with real CX consequences, and carrying journeys, segments and integrations forward cleanly is where deployments quietly hold or break. The teams treating this as a planned migration now, rather than a Q4 scramble, may be the ones whose customer experience never feels the move.

 

Leading Edge

Proven moves you can copy today.

Genesys is moving autonomous action into the contact center, folding newly acquired Pinkfish into Genesys Cloud with MCP-based access to more than 500 integrations spanning roughly 25,000 tools across CRM, ERP, HR and billing, with the first capabilities reaching the AppFoundry Marketplace by the end of July. The bet is a shift from AI that listens and routes to AI that resolves – an agent that reaches into back-office systems and completes the work. Which makes the question to ask before flipping it on not what the agent can do, but where the boundary on what it can touch is enforced, and whether that gate sits outside the agent.

Amazon Connect can now reach into the CRM over the Model Context Protocol, letting an AI agent discover capabilities at runtime and act across systems in real time rather than through fixed, pre-wired API workflows. That runtime flexibility is the upside and the exposure in a single move: the more an agent can reach on the fly, the less a hard-coded integration limits it; and the more the real boundary has to be enforced outside the agent. It’s the sharpest argument yet for designing guardrails at the integration layer before an agent goes live, not writing them into a prompt and hoping.

 

Off the Ledge

Hype and headaches we’re steering clear of.

The regulatory gate is arriving whether or not your agent respects boundaries on its own. A widening patchwork of state law now forces AI disclosure and oversight in customer interactions. Tennessee’s SB 1580 took effect July 1, 2026, and Nebraska, Oregon and Washington have all enacted conversational-AI disclosure and safety obligations, while Colorado repealed and replaced its broad AI Act with the narrower SB 26-189, effective January 1, 2027. The argument that the gate has to live outside the agent is now also a compliance argument. Buyers who bake disclosure and human-escalation controls into the deployment, rather than the prompt, may be the ones who don’t hear from a state AG.

Google shipped three new Gemini models this week, and the flagship still was not among them. Gemini 3.6 Flash arrived cheaper and 17% more token-efficient, while 3.5 Pro missed yet another target date. Workhorse models pay the bills, but a flagship that keeps slipping may hint at how hard the frontier has become.

Sorry, no content found.