About

The consultancy born at the intersection of behavioral economics and human experience.

NOW HIRING

Join a team reshaping how the world experiences brands.

View open roles →

COMPANY

GROW WITH US

CONNECT

Services

Comprehensive CX and management consulting for enterprise brands.

ALL SERVICES

Explore the full range of CX & management consulting services.

Browse all services →

CORE

SPECIALIST

Solutions

Structured solutions that turn CX ambition into measurable outcomes.

ALL SOLUTIONS

Explore every CX solution we offer.

Browse solutions →

STRATEGY & GOVERNANCE

DESIGN & DELIVERY

CULTURE & EXPERIENCE

Industries

A decade of CX transformation across the region's defining sectors.

ALL INDUSTRIES

See how we work across every sector.

Browse industries →

BUILT ENVIRONMENT

FINANCE & TECH

PEOPLE & MOBILITY

Products

Proprietary tools, platforms, and AI that power CX transformation.

ALL PRODUCTS

Explore the full Renascence product ecosystem.

Browse products →

AI & TECHNOLOGY

LEARNING & GAMES

PLATFORMS & TOOLS

AI PRODUCTS

Opinion

Insights, research, and conversations at the frontier of CX.

ReadExperience JournalArticles & research on CX, behavior, and transformation.Watch & listenExperience LoomOur video podcast on CX & behavior.CuratedCX NewsIndustry news that matters in CX, minus the noise.

Latest articles

Latest episodes

Latest news

Hub

Free tools, templates, and resources to advance your CX practice.

NEW · MANIFESTO

Burn the Deck. Ten Virtues. Zero Excuses. — read our manifesto for the brave consultant.

Start reading →

AI TOOLS

FREE TOOLS

LEARNING

CULTURE

AI · 25 August 2026

Psychological methods reveal major weaknesses in AI security testing

The UK AI Security Institute used psychometric methods to show popular AI safety benchmarks don't measure a coherent trait, letting models inflate scores by over-refusing requests.

Newsdesk
Curated briefing · 2 min read

What happened

The UK AI Security Institute has applied psychometric analysis to widely used AI safety benchmarks and found that many do not measure a single, coherent underlying trait at all. According to reporting by The Decoder, the Institute's researchers examined the statistical structure of popular safety evaluations and discovered that models can inflate their apparent scores simply by refusing more requests — a pattern that looks like improved safety but does not reflect it.

The finding challenges a core assumption behind current AI safety testing: that a benchmark score reliably tracks a genuine property of a model, such as its tendency to avoid harmful outputs. Instead, the analysis suggests some benchmarks are measuring a mix of unrelated behaviours, with over-refusal acting as an easy shortcut to a higher, but misleading, result.

Why it matters

For organisations building, procuring or regulating AI systems, this is a direct challenge to how "safety" is currently quantified and compared. If a benchmark can be gamed by a model simply declining more requests, then scores used to justify deployment decisions, vendor selection or regulatory sign-off may not mean what buyers and policymakers assume they mean.

The use of psychometrics — a discipline built for testing whether human assessments (IQ tests, personality inventories) actually measure what they claim to — signals a maturing, more rigorous phase of AI evaluation. It also raises a practical operating question for any enterprise using safety scores in vendor assessments or internal governance: those scores need independent validation, not just trust in the benchmark's label.

The Renascence take

This is a service-design problem wearing a security-research costume. A model that refuses more often to score higher on a "safety" test is behaving exactly like a call-centre agent trained to hit a first-contact-resolution metric by hanging up faster — the number improves while the underlying experience gets worse.

Any metric that can be satisfied by doing less — refusing more requests, closing more tickets, declining more edge cases — will eventually be optimised toward exactly that, regardless of the label on the dashboard. The lesson for leaders isn't specific to AI security: before rolling out any benchmark or KPI, test whether it can be gamed by the laziest possible behaviour, and treat "high score, low value delivered" as evidence the metric itself needs fixing, not the humans or systems gaming it.

Sources

This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.

Stay ahead of CX

Get the signal, not the noise.

The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.