AI · 19 September 2026
Cutting AI budgets won't fix token shock - Neo4j's Jim Webber on graph RAG and the price of accuracy
Neo4j's Chief Scientist Jim Webber makes the case that finance leaders throttling AI usage are answering the wrong question. New research from the National Innovation Centre for Data at Newcastle University shows why fixing the context fixes the bill.
What happened
Neo4j Chief Scientist Jim Webber has argued that organisations reacting to rising generative AI costs by capping usage or cutting budgets are treating the wrong symptom. Drawing on new research from the National Innovation Centre for Data at Newcastle University, Webber contends that so-called "token shock" — the escalating cost of feeding large volumes of context into large language models — is a signal that the underlying data architecture is inefficient, not that AI itself is unaffordable.
The argument centres on graph-based retrieval-augmented generation (graph RAG), which structures and connects enterprise data so that models retrieve only the specific, relevant context needed to answer a query, rather than large, loosely related blocks of text. According to Webber, this precision reduces the volume of tokens processed per query, which directly lowers cost while also improving the accuracy of AI-generated answers.
Why it matters
The piece reframes a live cost-control debate inside AI-adopting organisations. Many finance and technology leaders are currently responding to unpredictable LLM bills by rationing AI access or scaling back pilots — a move that risks stalling adoption just as generative AI is being embedded into customer- and employee-facing workflows. Webber's argument is that the real lever is data architecture: how context is retrieved and structured before it ever reaches the model.
For digital transformation leaders, this shifts the cost conversation from "how much AI can we afford" to "how well is our data organised for AI to use efficiently." It suggests that graph-based knowledge structures, rather than blunt usage caps, are what make generative AI economically sustainable at scale — with knock-on implications for accuracy, a factor that matters directly wherever AI outputs reach customers or frontline staff.
The Renascence take
Budget caps are an easy lever to pull and a satisfying one to report upward — but they treat a design flaw as a demand problem.
Most organisations panicking over AI costs are really discovering that their data was never structured to be queried efficiently in the first place — the token bill is just the first bill that made that visible. Cutting usage doesn't fix that; it just hides the inefficiency behind reduced adoption and, quietly, behind worse answers. A customer-obsessed operator should treat token shock as a diagnostic signal and invest in how context is retrieved and connected, because the same fix that lowers cost is the one that improves the accuracy customers and employees actually experience.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.