关于

行为经济学与人类体验的交汇点,由此诞生了 Renascence 咨询公司。

现正招聘

加入我们的团队,一起重塑世界体验品牌的方式。

查看开放岗位 →

公司

与我们共同成长

联系我们

服务

为企业品牌提供全面的客户体验和管理咨询。

所有服务

探索 Renascence 提供的全方位客户体验和管理咨询服务。

浏览所有服务 →

核心业务

专家

解决方案

转化为可衡量的成果。" is smooth, authoritative, and perfectly captures the source meaning and tone.通过结构化解决方案,将 CX 愿景转化为可衡量的成果。

所有解决方案

探索我们提供的所有 CX 解决方案。

浏览解决方案 →

战略与治理

设计与交付

文化与体验

行业

跨越十年的客户体验转型,赋能本地区最具代表性的行业。

所有行业

了解我们如何服务各大行业。

浏览行业 →

建筑环境

金融与科技

人员与流动性

产品

Renascence专有的工具、平台和AI,赋能客户体验转型。

所有产品

探索 Renascence 的完整产品生态。

浏览产品 →

人工智能与技术

学习与游戏

平台与工具

AI 产品

观点

Renascence:客户体验前沿的洞察、研究和对话。

阅读体验日志关于CX、行为与转型的文章和研究。观看与收听体验蓝图我们的CX与行为视频播客。精选CX新闻CX行业重要新闻,去芜存菁。

最新文章

最新剧集

最新新闻

中心

免费工具、模板和资源,助您提升客户体验实践。

新 · 宣言

破釜沉舟。十大美德。绝无借口。——阅读我们为勇敢的咨询顾问撰写的宣言。

开始阅读 →

AI 工具

免费工具

学习资料

文化

AI · 2026年8月25日

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows . In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior . Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month. Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers , essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once. The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much

新闻中心
精选简报 · 3 分钟阅读
分享分享至 X分享至领英

What happened

New survey data from VentureBeat Intelligence shows enterprises that have already been burned by AI agents failing after passing internal testing are moving to remove human oversight from deployment decisions faster, not slower. Trust in automated evaluation among the 108 enterprises surveyed rose sharply in July, with 13% saying they trust automated evaluation, up from 5% the month before, even as the underlying problem persists.

Nearly half of respondents, 49%, said an AI agent or LLM-powered feature that had passed company testing later caused a customer-visible problem in production — essentially flat against 50% in June. Almost a quarter, 24%, said this had happened more than once. At the same time, concern over misalignment between test results and real-world performance fell ten points month over month, from 29% to 19%, suggesting confidence is growing even as the failure rate is not improving.

Why it matters

The data points to a widening gap between how enterprises measure AI readiness and how AI agents actually behave once exposed to real customers and real edge cases. Rather than responding to repeated failures by strengthening human review, many organisations appear to be doing the opposite — leaning further into automated evaluation and stepping back human checkpoints, on the assumption that better testing infrastructure will eventually close the gap.

For leaders running AI and digital transformation programmes, this is a governance signal as much as a technical one. Evaluation frameworks that look rigorous on paper can create false confidence, and removing human judgement from deployment gates at the same moment failure rates remain unchanged raises the stakes for any customer-facing agent, chatbot or automated decision system already in production.

By the numbers

  • 13% of enterprises surveyed in July said they trust automated evaluation, up from 5% in June.
  • 49% of respondents said an AI agent or LLM feature that passed testing went on to cause a customer-visible problem.
  • 24% said this kind of failure had occurred more than once.
  • 19% cited poor alignment between tests and real-world results as their top concern in July, down from 29% in June.
  • 108 enterprises were surveyed as part of the VentureBeat Intelligence research.

The Renascence take

The uncomfortable pattern here is behavioural, not technical: rising confidence is being generated by the existence of a testing process, not by evidence that the process works. That is a classic automation bias — the score on a dashboard becomes a substitute for the outcome it was meant to predict.

Passing an eval is not the same as earning trust with a customer, and treating it as such is a service-design failure dressed up as a data point. The organisations getting this right will keep a human in the loop precisely where an agent's failure would be visible to a customer — not because automation can't be trusted in general, but because customer-facing moments are where the cost of being wrong is paid by someone else. Confidence should scale with evidence of real-world performance, not with the smoothness of the testing pipeline.

来源

本简报由我们的新闻编辑部撰写,综合了以下媒体的报道。点击链接可查看原始报道。

分享分享至 X分享至领英

保持CX领先

获取信号,而非噪音。

塑造客户体验的故事——以及《期刊》和“体验之梭”——尽在您的收件箱。