AI · 2026年8月25日
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows . In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior . Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month. Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers , essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once. The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much
What happened
New survey data from VentureBeat Intelligence shows enterprises that have already been burned by AI agents failing after passing internal testing are moving to remove human oversight from deployment decisions faster, not slower. Trust in automated evaluation among the 108 enterprises surveyed rose sharply in July, with 13% saying they trust automated evaluation, up from 5% the month before, even as the underlying problem persists.
Nearly half of respondents, 49%, said an AI agent or LLM-powered feature that had passed company testing later caused a customer-visible problem in production — essentially flat against 50% in June. Almost a quarter, 24%, said this had happened more than once. At the same time, concern over misalignment between test results and real-world performance fell ten points month over month, from 29% to 19%, suggesting confidence is growing even as the failure rate is not improving.
Why it matters
The data points to a widening gap between how enterprises measure AI readiness and how AI agents actually behave once exposed to real customers and real edge cases. Rather than responding to repeated failures by strengthening human review, many organisations appear to be doing the opposite — leaning further into automated evaluation and stepping back human checkpoints, on the assumption that better testing infrastructure will eventually close the gap.
For leaders running AI and digital transformation programmes, this is a governance signal as much as a technical one. Evaluation frameworks that look rigorous on paper can create false confidence, and removing human judgement from deployment gates at the same moment failure rates remain unchanged raises the stakes for any customer-facing agent, chatbot or automated decision system already in production.
By the numbers
- 13% of enterprises surveyed in July said they trust automated evaluation, up from 5% in June.
- 49% of respondents said an AI agent or LLM feature that passed testing went on to cause a customer-visible problem.
- 24% said this kind of failure had occurred more than once.
- 19% cited poor alignment between tests and real-world results as their top concern in July, down from 29% in June.
- 108 enterprises were surveyed as part of the VentureBeat Intelligence research.
The Renascence take
The uncomfortable pattern here is behavioural, not technical: rising confidence is being generated by the existence of a testing process, not by evidence that the process works. That is a classic automation bias — the score on a dashboard becomes a substitute for the outcome it was meant to predict.
Passing an eval is not the same as earning trust with a customer, and treating it as such is a service-design failure dressed up as a data point. The organisations getting this right will keep a human in the loop precisely where an agent's failure would be visible to a customer — not because automation can't be trusted in general, but because customer-facing moments are where the cost of being wrong is paid by someone else. Confidence should scale with evidence of real-world performance, not with the smoothness of the testing pipeline.
来源
本简报由我们的新闻编辑部撰写,综合了以下媒体的报道。点击链接可查看原始报道。