About

The consultancy born at the intersection of behavioral economics and human experience.

NOW HIRING

Join a team reshaping how the world experiences brands.

View open roles →

COMPANY

GROW WITH US

CONNECT

Services

Comprehensive CX and management consulting for enterprise brands.

ALL SERVICES

Explore the full range of CX & management consulting services.

Browse all services →

CORE

SPECIALIST

Solutions

Structured solutions that turn CX ambition into measurable outcomes.

ALL SOLUTIONS

Explore every CX solution we offer.

Browse solutions →

STRATEGY & GOVERNANCE

DESIGN & DELIVERY

CULTURE & EXPERIENCE

Industries

A decade of CX transformation across the region's defining sectors.

ALL INDUSTRIES

See how we work across every sector.

Browse industries →

BUILT ENVIRONMENT

FINANCE & TECH

PEOPLE & MOBILITY

Products

Proprietary tools, platforms, and AI that power CX transformation.

ALL PRODUCTS

Explore the full Renascence product ecosystem.

Browse products →

AI & TECHNOLOGY

LEARNING & GAMES

PLATFORMS & TOOLS

AI PRODUCTS

Opinion

Insights, research, and conversations at the frontier of CX.

ReadExperience JournalArticles & research on CX, behavior, and transformation.Watch & listenExperience LoomOur video podcast on CX & behavior.CuratedCX NewsIndustry news that matters in CX, minus the noise.

Latest articles

Latest episodes

Latest news

Hub

Free tools, templates, and resources to advance your CX practice.

NEW · MANIFESTO

Burn the Deck. Ten Virtues. Zero Excuses. — read our manifesto for the brave consultant.

Start reading →

AI TOOLS

FREE TOOLS

LEARNING

CULTURE

AI · August 16, 2026

PerceptionBench Reveals AI Models Still Struggle to 'See' Images

Moonshot AI's new PerceptionBench benchmark shows no frontier multimodal model exceeds 60% accuracy at pure visual perception, with GPT-5.6 Sol narrowly leading rivals.

R
Renascence Newsdesk
Curated briefing · 2 min read

What happened

Moonshot AI has released PerceptionBench, a new benchmark designed to isolate how well multimodal AI models can actually "see" an image, independent of their logical reasoning ability. The results show that no frontier model currently clears 60 percent accuracy on the test, with GPT-5.6 Sol taking the lead by only a narrow margin over rival systems.

The benchmark's key finding is that many errors previously attributed to flawed reasoning actually originate earlier in the pipeline, at the point where a model reads and interprets the image itself. By separating perception from reasoning, PerceptionBench suggests that a meaningful share of multimodal AI mistakes are visual in origin rather than logical.

Why it matters

Multimodal AI systems are increasingly positioned as capable of interpreting images, screenshots, documents and video alongside text, underpinning use cases from automated quality checks to visual search and document processing. PerceptionBench's findings indicate that the weak link in many of these systems may not be their reasoning chains but their more basic ability to correctly perceive what is in front of them.

For organisations building AI-enabled products and services on top of these models, this distinction matters for where investment and testing should be focused. If visual misreads are driving downstream errors, then better prompting or reasoning scaffolds will not fix the underlying problem — the models need to get better at looking before they get better at thinking.

By the numbers

  • 60 percent accuracy is the ceiling no frontier model has yet reached on PerceptionBench.
  • GPT-5.6 Sol currently leads the benchmark, though only by a narrow margin over other tested models.

The Renascence take

It is tempting to treat AI errors as reasoning failures because that is the more interesting, more fixable-sounding story. PerceptionBench is a useful corrective: it shows that a chunk of what looks like poor judgement is actually a model failing to register basic visual facts correctly in the first place.

This is a familiar service-design lesson wearing new clothes: before you fix how a system decides, check whether it is even perceiving the situation accurately. Many customer-facing AI deployments — visual claims processing, document verification, in-store computer vision — will inherit this same weakness, and teams that only test reasoning outputs will miss it. Operators piloting multimodal AI in service workflows should build perception-specific test cases separate from end-to-end accuracy checks, and treat "the model got the answer wrong" as two distinct failure modes rather than one.

Sources

This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.

FAQ

Questions we get on this topic

PerceptionBench is a benchmark released by Moonshot AI that isolates a multimodal AI model's raw visual perception ability from its logical reasoning, testing how accurately it reads and interprets images independent of the analysis it performs afterward.

GPT-5.6 Sol currently leads the benchmark, but only by a narrow margin over other tested frontier multimodal models.

No frontier model has yet exceeded 60 percent accuracy on PerceptionBench, indicating that visual perception remains a significant weak point across current AI systems.

It suggests that many multimodal AI errors stem from misreading images rather than flawed reasoning, meaning organisations using AI for tasks like document processing, visual search or quality checks should test perception accuracy separately from overall output accuracy.

Stay ahead of CX

Get the signal, not the noise.

The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.