AI · July 31, 2026
Specialised AI Training Data: Ex-OpenAI Researcher's $100B Bet
Former OpenAI researcher Andrew Ho is founding a start-up on the thesis that AI labs will spend $100B+ on specialised training data as raw scaling hits its limits.
What happened
Andrew Ho, a former OpenAI researcher, is leaving the company to found a start-up focused on specialised training data for large language models (LLMs), betting that AI laboratories will collectively spend more than $100 billion acquiring targeted data sets. Ho, working alongside Cambridge researcher Adam Hunt, argues that the current scaling paradigm — simply feeding models more compute and raw text — is hitting a ceiling.
The core problem Ho and Hunt identify is one of uneven capability development: rather than becoming broadly more capable, leading LLMs are growing increasingly specialised, improving markedly at coding and mathematics while stagnating or even regressing in other domains. Their contention is that brute-force scaling cannot resolve this imbalance, and that purpose-built, high-quality data pipelines will become the decisive competitive input for the next generation of AI systems.
Why it matters
For customer experience and service-design practitioners, this development carries a pointed implication. Most enterprise deployments of LLMs today lean on general-purpose models to handle the full breadth of customer interactions — from complaint resolution and empathy-heavy conversations to technical troubleshooting. If those models are quietly regressing in the very domains that require nuanced, context-sensitive language (precisely the territory of great service), organisations relying on off-the-shelf AI may be building on an eroding foundation without realising it.
From a behavioural-economics perspective, there is also a capability illusion risk: operators and customers alike may attribute consistent, confident-sounding outputs to genuine competence, when the underlying model has in fact become narrower. Trust calibration — knowing when to rely on AI and when to escalate to a human — becomes far harder when degradation is invisible. Service designers will need to build explicit verification loops and domain-specific evaluation into their AI governance frameworks, not assume that a model's headline benchmark scores translate to frontline CX performance.
By the numbers
- $100 billion+ — Ho's projected total spend by AI labs on specialised training data collection, the investment thesis underpinning his new venture.
The Renascence take
The industry conversation around AI and CX has been dominated by model size and speed-to-deployment. Ho's departure from OpenAI to chase the data problem is a signal that insiders believe the real scarcity — and therefore the real value — is shifting upstream, to the quality and specificity of what models are trained on. Most CX leaders are not yet asking their AI vendors the right questions about this.
The organisations that will win on AI-powered experience are not those who adopted the largest model earliest, but those who invested in the right data to shape model behaviour for their specific customer context. Scaling gives you a capable generalist; curated data gives you a reliable specialist. In service design, reliability beats brilliance every time — a model that consistently handles a complaint well is worth far more than one that occasionally dazzles and frequently drifts. Customer-obsessed operators should be auditing their AI deployments now for domain-specific regression, not waiting for a customer to notice it first.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.