Somebody has to watch it

contact center AI monitoring agent oversight escalation

By Damian Mathews & The Last Mile Team

Your agent is live. It’s talking to customers right now.

Who is watching it? What does that cost you every month?

Most contact centers can’t answer the second question, because nobody has ever put a number on it. This week OpenAI did.

In a post explaining why it is deliberately slowing its own model development down, the company disclosed the run rate on its supervision:

“Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored.”

The company with the largest safety organization in the industry spends a fifth of its compute watching its own models work.

Read that as an operator rather than a safety researcher. It isn’t a project cost. It’s a run rate. Supervision scales with volume, and it doesn’t stop when the implementation does.

We wrote in Half Can’t Prove It that organizations with full visibility into what AI costs to operate are five times more likely to prove a return. Monitoring is part of what it costs to operate. Leave it out of the model and the business case was never real.

Now look at your own numbers. Most CX budgets carry a line for build and a line for licensing. Oversight gets absorbed by a QA team that was already at capacity before the agent arrived, sampling a handful of interactions a week using a method designed for human agents.

There is now a number for what that costs you. Gartner research reported this week examined 432 AI use cases in customer service. A quarter produce a return. A quarter produce a negative one. Eleven percent break even. 42% have unclear ROI with leaders saying they simply don’t know the value produced. The largest group isn’t the wins or the losses. It’s the work nobody can account for.

That column has company: 56% of service leaders expect their incentives to be tied to AI outcomes this year. A majority of this profession is about to be paid on a number it cannot currently produce.

 
The protocol is the portable part

The rest of OpenAI’s post is worth reading for a reason that has nothing to do with frontier models.

Their monitoring escalates to a human. They aim to issue an alert within 30 minutes of concerning activity surfacing. And if the team cannot conclusively determine within 30 minutes that a flag is a false positive, they are expected to pause the activity.

Three things, none of them exotic:

  • Someone who gets paged
  • A time limit on deciding
  • Standing permission to pause it

Get that right and an incident is a bad half hour. Get it wrong and the first person to notice your agent misbehaving is a customer, and the second is a journalist.

 
The third one is no longer optional

Standing permission to pause used to be a maturity marker. This month it moved closer to a regulatory expectation.

The EU AI Act’s Article 50 transparency duties became enforceable on 2 August. And Colorado’s Automated Decision-Making Technology Act, effective 1 January 2027, requires deployers to permit “meaningful human review and reconsideration” of consequential decisions, to the extent commercially reasonable — with the Attorney General’s proposed rules, filed 11 August, defining what a qualified reviewer actually looks like.

Your customers were already asking for it. In Gartner’s survey of 3,566 customers, 87% said access to a human agent is essential when a company uses GenAI for service. Eric Keller, the analyst who ran it, put the operational conclusion plainly: “Service leaders should not use GenAI as a mandatory first step for every issue.”

Regulators and customers converging on the same design requirement inside a fortnight. The escalation path is becoming something you may have to evidence, not just something you’re glad you built.

 
Drift is the quiet part

Drift in a contact center doesn’t announce itself. The knowledge base moves and nobody re-tests the agent against it. A model version updates and the tone shifts half a degree. The contact mix changes after a product launch and the agent starts confidently answering a question it was never given an answer for.

None of this is unfamiliar work. You already run QA, coaching cycles, and monthly reviews for human agents, because you already know quality decays without a cadence.

What changes is the sample size. Sampling was always a compromise forced by the cost of human review — you listened to eight calls because you couldn’t listen to eight thousand. That constraint is loosening.

Waterfield Tech builds the observability and guardrails into everything we do by enabling companies to automatically analyze and coach up to 100% of their interactions for both AI and humans. When full coverage is available, a 2% sample stops being a methodology and starts being a choice.

The vendors selling you agents are not selling you the supervision.

Who gets paged when your agent does something strange at 2am?

— Damian

 
 

 

 

Here’s what went down this week.

Bleeding Edge

Early signals you should keep on your radar.

Disclosure and escalation now have dates on them. The EU AI Act’s Article 50 transparency obligations became enforceable on 2 August: interactive AI systems must tell users they are dealing with AI “from the start of the first interaction in a clear and distinguishable manner,” meeting accessibility requirements — which means voice channels need an audible equivalent, not a line in a footer. Penalties run to €15 million or 3% of worldwide annual turnover, and systems already on the market have until 2 December for machine-readable marking. Colorado followed on 11 August, when the Attorney General filed proposed rules for the ADMT Act and the Chatbot Safety Act, both effective 1 January 2027; comments close 26 October. Two jurisdictions, two different obligations, one shared implication: turn-one disclosure and a working escalation path are becoming things you evidence rather than things you intend.

Your customers are asking someone else about you. Gartner surveyed 3,566 B2B and B2C customers in February and March and found them roughly three times more likely to use a third-party GenAI assistant — ChatGPT, Gemini, Copilot — than a company’s own chatbot. Half say interactions are easier with GenAI, and 58% of users have used it to complete a task, rising to 74% in B2B. Containment and resolution rates only measure the traffic that reaches you. The metric on your dashboard is describing a shrinking share of how customers actually resolve questions about your business, and nothing in your reporting stack will tell you by how much.

AI triage is now the front door for more than a third of interactions. Metrigy’s second-quarter AI Consumer Experience Index, a 1,000-consumer US study run in partnership with NiCE, puts 36.3% of all contact center interactions as starting with AI triage. Consumer satisfaction with AI agents rose 11.8% quarter on quarter, the largest movement of any metric it tracks, and trust rose 7.6%. Two numbers underneath the headline are the useful ones: 59% would accept proactive outbound calls from an AI agent, but only where escalation is available, and 95% of consumers aged 65 and over would still choose a human. The 36.3% figure is a benchmark worth holding your own front door against. (Vendor-sponsored research — read the satisfaction gains accordingly; the escalation condition is the durable finding.)

Leading Edge

Proven moves you can copy today.

Amazon Connect can now show you where contacts stall. AWS added routing-step and agent-proficiency reporting to Connect dashboards on 17 August, letting supervisors filter agents by assigned proficiency, group metrics by routing step, and see how many contacts are queued at each step. The immediate use is diagnostic: find the step where contacts pile up waiting for a narrowly skilled agent, then widen the criteria. Available in commercial regions and GovCloud (US-West) wherever Connect is offered, so there’s no rollout gate for public-sector operations.

Zero data retention, on the record. OpenAI published its zero-data-retention posture for frontier models on 19 August: for eligible API customers, prompts and model responses are not retained after a request is processed, “customer content is not available to OpenAI personnel for review,” and enterprise data is not used for training unless customers explicitly opt in. In ZDR deployments content stays on infrastructure the customer controls, with customer-held encryption keys in development and a technical whitepaper due in September. Data retention is the most common stall point in contact center AI security review, particularly in regulated verticals and outsourced operations carrying client data-residency terms. This is a document you can hand to your CISO.

Genesys puts an auditable standard on the platform. In its FY26 report on 11 August, Genesys disclosed ISO/IEC 42001 certification — the AI management-system standard — for the Genesys Cloud platform, alongside the work of its cross-functional AI Ethics Board. CX Today’s analysis of the same report adds the operational figures that matter more day to day: platform availability of 99.995% in FY26, under 30 minutes of annual downtime, AI-powered conversations up more than 120% year over year, and roughly 20% of Genesys Cloud new-business ACV coming from AI. ISO 42001 is starting to appear in enterprise RFPs and client audit packs. A certified platform is one fewer thing you have to argue.

Off the Ledge 

Hype and headaches we’re steering clear of.

The AI ROI number is a measurement problem, not a technology problem. Gartner research covered this week analyzed 432 AI customer service use cases: 25% producing ROI, 25% producing negative returns, 11% breaking even — and 42% with unclear ROI, where leaders say they simply don’t know the value produced. Teams are running nearly five use cases each on about 13% of functional budget, and more than 75% of leaders plan to increase AI investment this year. The failure mode isn’t the models. It’s expanding scope before instrumenting the baseline, which is how you end up in the 42% with no way to argue your way out. Worth noting the primary research is client-only; the figures reach the rest of us through trade press, so cite it as such.

Check the layoff number before you quote it. A figure of roughly 205,000 AI-attributed US job cuts is circulating widely, sourced to a job-board tracker and relayed through trade press. The primary series says something different. Challenger, Gray & Christmas reported on 6 August that AI has been cited in 112,713 cuts year to date (about 24% of the total) with 10,970 in July, or 33% of that month. Total announced cuts are 477,033 YTD, down 41% from 806,383 in the same period of 2025. Services cuts are down 55% year over year; retail is down 84%. The real story is more useful than the one being repeated: AI is becoming the stated reason for a shrinking pool of cuts, 184,538 of them cumulatively since tracking began in 2023. That’s a workforce-planning input you can model. The 205,000 figure is not.

California’s AI disclosure law probably doesn’t cover the channel you’re worried about. The California AI Transparency Act became operative on 2 August for generative systems with over a million monthly visitors or users, requiring permanent embedded provenance metadata, an offered visible label, and a free public detection tool at $5,000 per violation, with each day a discrete violation, enforceable by the Attorney General, city attorneys or county counsel. Read the scope carefully, because most teams have it backwards: the disclosure duties attach to image, video and audio content. AI-generated text is outside scope. AI-generated audio is inside it. Your text chatbot is likely out. Your synthetic voice agent, AI-generated outbound voice messages, and any cloned voice in your IVR are likely in. There is also a 96-hour clock requiring a licensor to revoke a license after discovering a licensee disabled latent disclosure, which belongs in your vendor paper rather than in a compliance memo nobody reads.

See you next week!

Sorry, no content found.