AI · 22 September 2026
Grok 4.7 launches cheap but trails Claude and GPT-6 on benchmarks
xAI's Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index, versus 53 for both Claude Fable 5.1 and GPT-6, but undercuts rivals sharply on price.
What happened
xAI has released Grok 4.7, positioning it as the company's most capable model to date. Independent evaluation on the Artificial Analysis Intelligence Index puts it at 46 points, placing it mid-pack and meaningfully behind Claude Fable 5.1 and GPT-6, which both score 53. The gap widens further on agentic coding benchmarks, where Grok 4.7 trails its rivals by a larger margin than on general intelligence tasks.
The model's standout attribute is price. xAI has brought Grok 4.7 to market at notably lower cost than comparable frontier models from Anthropic and OpenAI, effectively trading top-tier benchmark performance for aggressive affordability.
Why it matters
This is a story about how AI vendors are starting to compete on more than raw capability. As the leading labs converge on similar architectures, price is emerging as a distinct axis of differentiation — and Grok 4.7 shows xAI leaning into that lever rather than chasing benchmark supremacy outright.
For enterprises building AI into products, workflows or customer-facing tools, this widens the practical decision space. Not every use case needs frontier-level reasoning or agentic coding sophistication; many need "good enough" performance at scale, reliably and cheaply. Grok 4.7's positioning suggests a maturing market where buyers can increasingly match model tier to task rather than defaulting to whichever model tops the leaderboard.
By the numbers
- 46 points — Grok 4.7's score on the Artificial Analysis Intelligence Index.
- 53 points — the score achieved by both Claude Fable 5.1 and GPT-6 on the same index.
- Two — the number of named rival models (Claude Fable 5.1, GPT-6) that Grok 4.7 trails in the benchmark comparison.
The Renascence take
The headline gap in benchmark scores tells only part of the story. What matters more for operators is whether the price-to-capability trade-off Grok 4.7 offers actually holds up in production, where cost, latency and "good enough" accuracy often outweigh marginal intelligence gains.
Most coverage will fixate on the seven-point gap to Claude and GPT-6, but that framing misses the more useful question: what job is the model actually being asked to do? A support-triage bot or a first-draft content assistant rarely needs frontier-level agentic coding — it needs consistent, low-cost output at volume. Leaders evaluating AI vendors should resist leaderboard anchoring and instead map their own use cases to the minimum viable intelligence required, then let price and reliability decide. The organisations that win here won't be the ones chasing the smartest model; they'll be the ones that correctly diagnose which tasks are commodity work and which genuinely need the premium tier.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.