AI · August 1, 2026
Inkling Small: Thinking Machines Lab Launches Compact Open-Weights Model
Thinking Machines Lab releases Inkling Small, an open-weights reasoning model under one-third the size of its predecessor that outperforms it on coding and reasoning benchmarks.
What happened
Thinking Machines Lab, the AI research company founded by former OpenAI chief technology officer Mira Murati, has released its second model: Inkling Small. The open-weights reasoning model is less than one-third the size of its predecessor, Inkling, yet outperforms it on several coding and reasoning benchmarks — a direct demonstration of the lab's stated focus on efficiency over raw scale.
Inkling Small is released under an open-weights licence, meaning developers can download and deploy the model directly rather than accessing it solely through an API. The release positions Thinking Machines Lab as a competitor in the fast-growing segment of compact, high-performance models — a space increasingly crowded by offerings from Mistral, Meta and Google.
Why it matters
For CX practitioners and service designers, the significance of Inkling Small lies less in the benchmark numbers and more in what the efficiency-over-size philosophy enables operationally. Smaller, capable models are cheaper to run, faster to respond and easier to deploy on-premise or within tightly governed enterprise environments — all of which lower the barrier to embedding genuine reasoning capability into customer-facing workflows, from intelligent triage to real-time agent assistance.
From a behavioural economics standpoint, latency is a trust signal. Customers and frontline staff alike form rapid judgements about AI tools based on response speed; a compact model that answers quickly and accurately will be adopted and trusted far more readily than a larger model that is marginally more capable but noticeably slower. Inkling Small's design philosophy, if it delivers on its benchmarks in production, is therefore directly aligned with the conditions that drive human acceptance of AI assistance.
The Renascence take
The industry conversation around AI has been dominated by parameter counts and leaderboard rankings — metrics that matter to researchers but tell operators almost nothing about whether a model will actually improve a customer interaction. Thinking Machines Lab's move with Inkling Small reframes the question in a way that is far more useful for anyone building or managing service experiences.
Most organisations chasing AI for CX are solving the wrong problem: they are asking "which model is most powerful?" when they should be asking "which model will my team actually use, and will customers notice the difference?" Smaller, faster, open-weights models reduce deployment friction — and reduced friction is the single most reliable predictor of adoption. The behavioural principle here is implementation intention: people follow through on using tools when the path of least resistance leads to the desired behaviour. Customer-obsessed operators should be evaluating Inkling Small not against GPT-4 on a benchmark sheet, but against their current escalation rates, handle times and agent satisfaction scores in a live pilot.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.