AI · 10 October 2026
Anthropic Cuts AI Agent Testing Off From Live Internet
Anthropic has disabled live internet access for all internal AI evaluations indefinitely, citing its inability to reliably control AI agent behaviour on the open web.
What happened
Anthropic has disabled live internet access for all of its internal AI evaluations, stating the change will remain in place until further notice. The company said the decision reflects an inability to reliably control the behaviour of its AI agents when they are tested against the open web, according to TechCrunch.
Internal evaluations are the testing routines AI labs run before and after releasing models, used to probe how systems behave, where they fail, and whether they act safely when given autonomy. By cutting these evaluations off from the live internet, Anthropic is removing one of the more unpredictable variables from its testing environment rather than attempting to fix agent behaviour on the open web directly.
The move was described as precautionary and indefinite, with no stated timeline for restoring internet-connected testing.
Why it matters
This is fundamentally a story about the limits of control in agentic AI systems — models that don't just answer questions but take actions, navigate websites, and make decisions with real-world consequences. An AI lab deciding it cannot yet trust its own agents to behave predictably against live, uncontrolled web content is a meaningful signal about where the technology genuinely stands, independent of marketing claims about autonomy and capability.
For organisations building or deploying agentic AI in customer-facing or operational roles, the episode is a useful reality check. It suggests that even leading developers are still working out how to bound agent behaviour safely, which has direct implications for anyone considering giving AI agents unsupervised access to live systems, customer data, or the open internet.
The Renascence take
The headline risk here isn't that Anthropic's agents misbehaved — it's that a frontier lab chose containment over confidence, and said so publicly. That's a more honest signal than most capability announcements.
Most organisations racing to deploy AI agents are optimising for what the agent can do, not for what happens when it does something unexpected in an uncontrolled environment. Anthropic's decision is a quiet admission that testing in the wild is still ahead of the industry's ability to govern it — and that's exactly the gap that matters for service design. Any operator putting agentic AI in front of customers should be asking the same question Anthropic just answered for itself: can we reliably predict and contain this system's behaviour before we let it touch something live? If the answer is no, the right move isn't to pause the AI — it's to sandbox it, the way Anthropic just did.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.
