AI · August 19, 2026
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows . In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior . Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month. Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers , essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once. The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much
What happened
New research from VentureBeat Intelligence's VB Pulse survey shows enterprises are moving to strip human oversight from AI deployment decisions even as confidence in automated evaluation grows — despite nearly half of respondents reporting that AI agents or LLM-powered features which passed internal testing went on to cause customer-visible problems in production.
Among the 108 enterprises surveyed, trust in automated evaluation more than doubled month over month, rising from 5% in June to 13% in July. Concern that test results poorly reflect real-world performance also eased, falling from 29% to 19% over the same period. Yet 49% of respondents said an AI agent or feature that had cleared company evals subsequently created a problem visible to customers — essentially flat versus June's 50% — and 24% said this had happened more than once.
The pattern suggests organisations are responding to AI failures not by adding more human checkpoints before release, but by leaning further into automated evaluation and reducing the people positioned to catch the next miss before customers do.
Why it matters
This is fundamentally a story about how enterprises are operationalising AI governance under pressure. As agentic AI and LLM-powered features move faster into production, the gap the research surfaces isn't just between testing and real-world performance — it's between rising institutional confidence in automated checks and the persistent, largely unchanged rate at which those checks fail to catch customer-facing errors. For leaders building AI operating models, this points to evaluation infrastructure maturing in perception faster than it is in practice.
For experience and service-design teams specifically, the implication is that customers are increasingly the last line of defence for AI errors that internal processes were supposed to catch — a shift with direct consequences for trust, complaint volumes and the workload placed on frontline and support teams.
By the numbers
- 13% of surveyed enterprises said they trust automated AI evaluation in July, up from 5% in June.
- 19% cited poor alignment between tests and real-world results as their top concern in July, down from 29% in June.
- 49% of respondents said an AI agent or feature that passed company testing later caused a customer-visible problem, versus 50% in June.
- 24% said such a failure had occurred more than once.
- 108 enterprises were surveyed as part of VentureBeat Intelligence's VB Pulse research.
The Renascence take
The data describes a classic behavioural pattern: confidence rising faster than competence, because automation reduces the visible friction of oversight even when the underlying error rate hasn't moved. Removing humans from the loop doesn't remove the risk — it just relocates where the failure gets caught, and increasingly that location is the customer.
Most organisations reading this will focus on the automation-versus-oversight debate and miss the real signal: a near-halved concern score sitting next to an unchanged failure rate is a trust calibration problem, not a testing problem. Confidence should track evidence, not convenience — and right now the evidence hasn't earned the confidence. A customer-obsessed operator doesn't ask "can we remove the human checkpoint," they ask "what is the cost of the failures our evals are still missing, and who bears it." Until customer-visible incident rates actually fall, cutting human review isn't efficiency — it's transferring risk onto the people least equipped to absorb it.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.