AI · 6 September 2026
AI’s attribution problem gets worse as models scale
Diffusion models are becoming sophisticated enough that they can reproduce an image even when they don’t have access to the original. In a series of ‘what if’ scenarios, researchers associated with MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL) swapped out different training datasets to test the impact on image outputs when original image data was completely removed. It turns out that, at sufficient scale, nothing changed. The researchers call the phenomenon “ attribution decay ”: The more data a diffusion model is trained on, and the larger it gets, the less individual inputs matter. “If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output,” Zheng Dai, lead author on the work, explained in an MIT blog post . These findings could have significant ramifications when it comes to resolving growing concerns about intellectual property (IP) and copyright infringement. Models can recreate images even i
What happened
Researchers associated with MIT's Computer Science & Artificial Intelligence Laboratory (CSAIL) have found that diffusion models — the architecture behind many image-generation systems — can reproduce a given image even after the original training data for that image has been removed entirely. In a series of controlled experiments, the team swapped different training datasets in and out to test what happened to model outputs when specific source images were excluded. At sufficient scale, the outputs barely changed.
The researchers describe this as "attribution decay": as a diffusion model is trained on more data and grows larger, the influence of any single input diminishes to the point of becoming statistically undetectable. Lead author Zheng Dai summarised the logic in an MIT blog post, noting that if removing a piece of data leaves the model's output unchanged, that data effectively had no measurable effect on the result.
Why it matters
The finding cuts to the heart of an unresolved question in generative AI: whether a model's output can be reliably traced back to the specific data it was trained on. As diffusion models scale up, this research suggests that traditional notions of attribution — identifying which training images "caused" a given output — become progressively harder to establish, even when a model can still faithfully recreate a protected or original image.
That has direct implications for the intellectual property and copyright disputes now working through courts and regulatory bodies worldwide. If attribution genuinely decays with scale, arguments that hinge on proving a specific image was used in training — or wasn't — may need to be rethought, both by rights holders seeking redress and by AI developers seeking to demonstrate compliance.
The Renascence take
Most coverage of this research will focus on the legal headache it creates. The more interesting story is what it reveals about trust and transparency as design problems, not just legal ones.
Attribution decay is really a transparency decay problem: as systems scale, the ability to explain "why" an output happened shrinks — and that same dynamic quietly erodes trust in any AI-driven service, not just image generators. Organisations deploying generative or decision-making AI in customer-facing roles should treat explainability as a design requirement from day one, not a forensic exercise bolted on after a dispute arises. The lesson for service leaders is blunt: if you can't explain how an AI-powered experience produced a specific outcome, you don't actually control that experience — you're hoping it behaves.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.