AI · August 19, 2026
AI’s attribution problem gets worse as models scale
Diffusion models are becoming sophisticated enough that they can reproduce an image even when they don’t have access to the original. In a series of ‘what if’ scenarios, researchers associated with MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL) swapped out different training datasets to test the impact on image outputs when original image data was completely removed. It turns out that, at sufficient scale, nothing changed. The researchers call the phenomenon “ attribution decay ”: The more data a diffusion model is trained on, and the larger it gets, the less individual inputs matter. “If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output,” Zheng Dai, lead author on the work, explained in an MIT blog post . These findings could have significant ramifications when it comes to resolving growing concerns about intellectual property (IP) and copyright infringement. Models can recreate images even i
What happened
Researchers affiliated with MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) have found that diffusion models — the architecture behind most modern image-generation AI — become harder to audit as they scale up. In a series of experiments, lead author Zheng Dai and colleagues systematically removed specific images from training datasets and then tested whether the models could still reproduce outputs resembling those images. At sufficient model and dataset scale, removing an individual training image made no measurable difference to what the model could generate.
The team has labelled this effect "attribution decay": as diffusion models are trained on ever-larger datasets, the influence of any single input shrinks toward the point of being statistically undetectable. Dai's framing, as reported by Computerworld and InfoWorld, is straightforward — if withdrawing a piece of data leaves the output unchanged, that data point cannot be said to have caused the output.
Why it matters
The finding cuts to the heart of one of generative AI's thorniest open questions: whether a model's output can be traced back to specific training inputs. As copyright and intellectual-property disputes over AI-generated imagery multiply, plaintiffs and defendants alike have leaned on the idea that removing or excluding protected works from training data should demonstrably change what a model can produce. This research suggests that assumption breaks down as models grow — meaning the standard technical test for infringement becomes progressively weaker precisely at the scale where commercial models now operate.
For organisations building or deploying generative AI, the implications extend beyond litigation. Attribution decay complicates any governance framework that relies on tracing outputs to sources — whether for licensing, compliance, content provenance, or simply explaining to a client or regulator how a system arrived at a given result. It also raises a harder question for AI vendors: if influence can't be isolated to individual inputs, accountability has to be built into how systems are trained and monitored, not just argued after the fact.
By the numbers
- Zheng Dai is the lead author of the CSAIL research documenting the attribution decay effect.
- Two independent outlets — Computerworld and InfoWorld — have reported on the findings, indicating early but growing attention in the trade press.
The Renascence take
Most coverage of this research will focus on the legal headache it creates for rights-holders. The more interesting story is what it reveals about the limits of explainability in systems we increasingly trust to make decisions on our behalf — not just generate images, but recommend, personalise and automate.
Attribution decay is really a trust-design problem wearing a legal costume. If a system's own creators can't isolate why an output looks the way it does, no amount of after-the-fact disclosure or "AI transparency" messaging will satisfy a customer, regulator or court asking a simple question: why this result, for me, right now? Organisations deploying generative AI in customer-facing contexts should stop treating explainability as a compliance checkbox to be solved later and start treating it as a design constraint from day one — choosing architectures, audit trails and human-in-the-loop checkpoints that make provenance traceable, even if that means accepting smaller or more curated models where accountability genuinely matters more than raw capability.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.