HomeScienceMIT Study Finds AI Images Cannot Be Tr
SCIENCE

MIT Study Finds AI Images Cannot Be Traced to Training Data

MIT researchers discovered that scaling training datasets causes the influence of individual images on AI-generated outputs to decay to zero.

WHAT YOU NEED TO KNOW
  • MIT CSAIL researchers Zheng Dai and David Gifford published the study on attribution decay in Nature Communications on August 18, 2026.
  • The team tested diffusion ensembles across seven datasets containing from 256 to over 160,000 images, alongside 1,282 brute-force retrained models.
  • The research was funded by Schmidt Futures and focused on diffusion models rather than large language models.

MIT researchers found that generative diffusion models trained on large datasets produce outputs that cannot be traced to individual training images, according to a study published Tuesday in Nature Communications.

Former MIT CSAIL researcher Zheng Dai and MIT Professor David Gifford identified a phenomenon called attribution decay, where expanding training data reduces the influence of any single training example on a specific generated output. At large scales, removing a single image, all works by a specific artist, or every photo of a subject leaves the resulting image practically unchanged.

Diffusion ensembles

Dai and Gifford built an architecture called a diffusion ensemble to test data removal directly without costly retraining from scratch. The system combines smaller components trained on separate data slices, allowing researchers to switch off components that processed specific inputs to observe true counterfactual outputs.

The researchers trained 24 ensembles on datasets ranging from 256 to more than 160,000 images across seven public datasets, including CIFAR-10, CelebA, MetFaces, and ArtBench. The team confirmed that the counterfactual radius—the maximum difference caused by removing a single input—shrank along an inverse power law as dataset size increased. The team also verified the decay pattern by brute-force retraining 1,282 separate models.

Copyright questions

Gifford stated that the findings bear directly on whether generated outputs qualify as derivative works under copyright law. He argued that outputs unaffected by individual training examples challenge standard fair use analyses and suggested companies must update their architectures to prove outputs do not derive from individual works.

Cornell Law School professor James Grimmelmann noted that attribution methods will likely fail for complex generative models, requiring courts and technologists to develop alternative approaches for evaluating copying.

Schmidt Futures supported the research, which focused on diffusion models used in audiovisual generation and protein structure modeling rather than large language models.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →