AI · July 28, 2026
AI Chatbot Hallucinations in Medical Diagnosis: Carnegie Mellon Findings
Carnegie Mellon researchers found LLMs including GPT-5, Gemini and Claude fabricated diagnoses in 18% of cases when medical images were withheld, raising urgent service-design and patient-safety concerns.
What happened
Researchers at Carnegie Mellon University's School of Computer Science have published findings warning that large language models (LLMs) — including GPT-5, Gemini and Claude — are unreliable tools for medical self-diagnosis, occasionally fabricating clinical information rather than flagging the limits of their knowledge.
The study, reported by The National, included a telling experiment in which Siddharth Vohra, a master's student at Carnegie Mellon's Robotics Institute, submitted medical queries to popular LLMs while deliberately omitting a diagnostic image. Instead of asking the user to supply the missing visual information — as a clinician would — the models generated a diagnosis anyway in roughly one in five cases, drawing inferences from demographic details such as the patient's age, gender and race.
The pattern points to a broader vulnerability: when pressed for answers, these models can produce confident-sounding but fabricated clinical conclusions, a behaviour researchers describe as hallucination. Patients have consequently been advised to avoid relying on AI chatbots for at-home diagnosis.
Why it matters
Healthcare is one of the highest-stakes service environments there is, and the Carnegie Mellon findings expose a critical gap between how AI tools present themselves and how they actually perform under pressure. From a service-design perspective, the problem is not simply technical error — it is a failure of the system's honesty architecture. A well-designed service acknowledges uncertainty; these models, by contrast, defaulted to a plausible-sounding answer rather than admitting ignorance, a behaviour that erodes the foundational trust patients must have in any diagnostic interaction.
For behavioural economists, the dynamic is equally concerning. Patients seeking reassurance are already primed to accept confident answers — a well-documented tendency known as automation bias. When an AI delivers a fluent, demographically tailored response, users are far less likely to question its validity than they would a hesitant human answer. The result is a service interaction that feels authoritative while potentially directing people away from timely, accurate care.
By the numbers
- 18% of cases saw LLMs fabricate a diagnosis when a required medical image was intentionally omitted from the query.
- 3 leading AI models tested: GPT-5, Gemini and Claude.
- Nearly 1 in 5 AI-generated diagnoses in the study were assessed as potentially false or misleading.
The Renascence take
Most commentary on this study will focus on AI accuracy as a technology problem to be patched in the next model update. That framing misses the deeper service-design failure: the models were never designed with an explicit "I don't know" pathway — and that omission is a choice, not an oversight.
The most dangerous moment in any service interaction is not when a system gets something wrong — it is when a system gets something wrong with complete confidence. Operators deploying AI in health-adjacent contexts must treat epistemic humility as a core design requirement, not a nice-to-have. That means building explicit uncertainty signals into every patient-facing interface, and measuring how often the system declines to answer as a positive quality metric. A chatbot that says "I cannot assess this without more information" is delivering better customer experience than one that invents a diagnosis — even if it feels less impressive in a demo.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.