Thursday Jul 23

Suno's Leak Hands RIAA Its Case

23JUL
55M LEAKEDSUNO FILESPROOF FOUND

Suno's own code just became evidence against it. Leaked files show Suno scraped YouTube, Deezer, and Genius for training audio. RIAA suits now have itemized proof, with a German ruling due July 31.

The trigger traces to a November 2025 breach, but the scraping details only surfaced last week via 404 Media.

Suno's own pipeline logs show 113,000 hours pulled from YouTube Music alone, plus Genius lyrics and Pond5 audio.

That turns vague copyright claims into a receipt. Three separate lawsuits, including one in Germany, now have dates on the calendar.

full brief & sources

Why this matters

  • Turns years of 'AI music companies probably scrape everything' suspicion into an itemized paper trail.
  • Feeds directly into three live lawsuits, with a German court ruling due within days.
  • Raises the bar for what 'clean training data' actually has to mean for AI music tools.

🔍 What happened

  • A November 2025 breach of Suno's source code became public last week via 404 Media.
  • Leaked pipeline docs show scraping from YouTube Music, Deezer, Genius, Pond5, Jamendo, and podcasts.
  • Annotations tie 113,879 hours to YouTube Music alone, plus over 62,000 hours from Pond5.
  • The same leak reportedly exposed 55 million user emails and Stripe payment records.
  • RIAA's US label suits, Germany's GEMA case, and a Sony fair-use hearing all cite the evidence.

💬 Smart takes

  • 404 Media: reported the leak traces to a hacker known as "ellie.191" and a supply-chain breach from November 2025.
  • Skeptic: scraping publicly streamed audio isn't automatically illegal, the legal fight is still about fair use, not just the fact of scraping.

🧭 Where this goes

  1. LikelyGermany's Munich court ruling on July 31 sets an early precedent for AI training data in the EU.
  2. Likelymore AI music and video tools face similar leak-driven discovery in the next year.
  3. PossibleSuno settles with labels before the US case reaches a verdict.
  4. Wild Cardthe leak triggers a broader industry standard requiring disclosed training data sources.

🥄 The Spoon Take

For a year, 'AI music companies scraped everything' was an assumption. Now it's a spreadsheet. That changes the legal fight from 'prove it' to 'explain it,' a much harder position for any AI company sitting on undisclosed training data.

🤔 Pushback

A leak proves scraping happened, not that it was illegal, courts still have to rule on fair use.