Back to feed

Cara art platform hit by three scrapes, 12 million works posted online

2 min
Cara art platform hit by three scrapes, 12 million works posted online

This digest was compiled by AI from multiple sources — links to the originals are below.

Artist platform Cara was scraped three times starting August 13, with one scraper posting a 12-terabyte archive of 12 million works on Reddit. The attacks spiked server fees and alarmed the platform's 1.5 million artists. The first scraper later regretted the action and agreed to collaborate with Cara founder Jingna Zhang on an open-source protection tool.

Key Facts

  • Cara, an image-sharing platform for artists, has about 1.5 million users.
  • Beginning August 13, Cara was subjected to three major scrapes.
  • The first scraper posted a 12-terabyte archive of 12 million works on the subreddit r/DefendingAIArt.
  • A second scraper pulled about 8.5 million links and metadata from Cara and uploaded them to Hugging Face.
  • The first scraper later agreed to collaborate with Cara founder Jingna Zhang on an open-source tool to protect artists.

The Scrapes

Cara, an image-sharing platform for artists, was subjected to three major scrapes beginning August 13. The first incident came to light when the individual responsible posted a 12-terabyte archive of 12 million works on the subreddit r/DefendingAIArt. The redditor, MandarinDawnPoppy994, wrote in a since-deleted post that the process cost him less than $10. A second scraper pulled about 8.5 million links from Cara, along with metadata like usernames, titles, and tags, and uploaded them to Hugging Face. Zhang described the subsequent attacks as 'copycat' actions.

Platform Response

Cara founder Jingna Zhang said the platform learned of the first scrape through user tags, as the scraper was 'gloating and looking for other people to join him' on Reddit. The scrapes spiked Cara's server fees and alarmed creators who had migrated from platforms like Instagram. Zhang noted that 'laws are not caught up on' protections against such data harvests, meaning scrapers can often justify the activity as technically legal. Zhang is separately part of two ongoing class actions brought by visual artists against Stability AI, Midjourney, and Google over alleged training on copyrighted work.

Unexpected Collaboration

The first scraper later regretted the stunt and agreed to collaborate with Zhang on a new open-source tool to protect artists. Cara offers protective features including Glaze, a tool meant to mask the style of images picked up by scrapers. Despite the collaboration, other scrapers continued to take advantage of Cara's vulnerabilities and minimal resources.

1 source

Time · lag behind first