AI · 24 August 2026
AI is becoming AI's biggest customer as agentic token usage jumps 14x on OpenRouter
AI agents have consumed more tokens than humans on OpenRouter since February 6, 2025. Agentic usage has grown 14x since then, while human usage is up just 2.8x. Nearly 70 percent of agent token consumption comes from cheap cached prompts, though, so actual costs are rising far more slowly than the raw numbers suggest. The article AI is becoming AI's biggest customer as agentic token usage jumps 14x on OpenRouter appeared first on The Decoder .
What happened
Token consumption by AI agents has overtaken human usage on OpenRouter, the model-routing platform used by developers to access multiple large language models. According to data reported by The Decoder, agentic token usage has grown fourteenfold since 6 February 2025, while human-driven token usage over the same period rose just 2.8 times.
The shift means autonomous AI agents — software that calls models repeatedly to plan, reason and execute tasks without a human prompting each step — are now the platform's dominant source of demand. Notably, close to 70 percent of that agentic token volume comes from cached prompts, which are far cheaper to serve than fresh queries, meaning the cost curve is rising much more gently than the raw consumption numbers imply.
Why it matters
This is a structural signal about how AI infrastructure is actually being used, not just how it's being marketed. When agents become the primary "customer" of a model platform, the design priorities shift: latency, caching efficiency, tool-calling reliability and cost-per-task start to matter more than single-response quality. Organisations building or buying agentic AI need to understand that usage patterns — and therefore economics — look nothing like traditional chatbot or assistant deployments.
For technology and operations leaders, the caching detail is the more important story. It suggests that much of today's "agentic surge" is repetitive, structured work — agents re-running similar prompts, checking state, or looping through multi-step tasks — rather than each token representing novel reasoning. That has direct implications for infrastructure planning, vendor pricing models, and how enterprises forecast the true cost of scaling agentic workflows.
By the numbers
- 14x growth in agentic token usage on OpenRouter since 6 February 2025
- 2.8x growth in human token usage over the same period
- ~70 percent of agentic token consumption comes from cached prompts
The Renascence take
The headline number invites a simple story — "agents are exploding" — but the caching figure tells a more useful one about how that growth is actually composed.
Most organisations watching agentic AI adoption are counting the wrong thing: raw token volume, when what actually determines cost, reliability and customer impact is what proportion of that volume is genuinely new reasoning versus repetitive, cached execution. A service or product team deploying agents should be tracking the ratio of cached-to-fresh tokens as closely as they track resolution rates, because it reveals whether their agents are getting smarter or just looping harder. The behavioral lesson is the same one that applies to human service operations: volume without visibility into what's driving it is a vanity metric, not a strategy.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.