mimile
Back to feed

MIT CSAIL researchers find attribution decays as diffusion models scale

AI digest

This digest was compiled by AI from multiple sources — links to the originals are below.

MIT CSAIL researchers find attribution decays as diffusion models scale

MIT CSAIL researchers find that as diffusion models scale, individual training examples exert less measurable influence on outputs, a phenomenon they call attribution decay. The finding raises attribution challenges for AI copyright, auditing, and governance.

Key Facts

  • MIT CSAIL researchers call the phenomenon “attribution decay”: the larger a diffusion model and the more data it trains on, the less individual inputs matter.
  • In experiments, removing original image data from training datasets left outputs unchanged at sufficient scale.
  • Zheng Dai, lead author, said in an MIT blog post that if removing a piece of data does not change the output, that piece of data did not affect the output.
  • Stability AI and Midjourney face an ongoing class action lawsuit filed by several artists in federal court in California over alleged scraping of billions of copyrighted images.
  • Getty Images partly won trademark claims against Stability AI in England and Wales after some AI-generated images closely resembled its work.

Attribution Decay

MIT CSAIL researchers ran a series of “what if” scenarios, swapping out training datasets to test image outputs when original image data was completely removed. At sufficient scale, removing any individual piece of data left model outputs unchanged. The more data a diffusion model is trained on and the larger it gets, the less individual inputs matter, a phenomenon the researchers call attribution decay. Zheng Dai, lead author, said in an MIT blog post that if taking away a piece of data does not change the model’s output, that data did not affect the output. Modern generative diffusion models replicate statistical patterns in large training datasets to produce image, video, and audio outputs.

IP and Legal Cases

Stability AI and Midjourney are defendants in an ongoing class action lawsuit filed by several artists in federal court in California. The artists argue the models scraped billions of their copyrighted images without consent. Getty Images brought claims against Stability AI in the High Court of Justice Business and Property Courts of England and Wales; those claims were struck down, but Getty partly won trademark claims because some AI-generated images closely resembled its work. The researchers wrote that developing a method to attribute generated outputs to influential training data would greatly advance understanding of and ability to regulate these models.

2 sources

MIT CSAIL researchers find attribution decays as diffusion models scale