Sunday, 2 August 2026 Archypedia index online
ArchypediaA
The living archive of world news
Technology

Suno scraped millions of songs from YouTube and other platforms via hack

A security breach at Suno has confirmed the systematic scraping of millions of audio files and lyrics for training its AI models. The leaked internal logs provide new evidence for record labels currently suing the company over alleged copyright infringement.

Suno scraped millions of songs from YouTube and other platforms via hack
Suno scraped millions of songs from YouTube and other platforms via hack

A significant security breach has revealed the internal operational methods of Suno, a generative AI music platform currently embroiled in major copyright litigation. Hacked data obtained by 404 Media confirms that the company systematically scraped millions of audio files and lyrics from various internet services to populate its training libraries. The breach, which took place in November 2025, involved a hacker operating under the pseudonym ellie.191, who gained unauthorized access to company source code and records from 2023 and 2024.

The leaked files provide a detailed look at the scale of Suno’s data collection efforts. According to the recovered source code and internal logs, the platform targeted a wide range of repositories to train its models. The volume of material identified in the datasets is extensive, with YouTube Music serving as a primary target. At the time the relevant file was last updated, Suno had consumed 2,013,545 music clips from that platform alone.

Related imagery

Image via techcrunch.com
Image via techcrunch.com
Image via theverge.com
Image via theverge.com
Image via 404media.co
Image via 404media.co

Data Collection Scope

The leaked material quantifies the duration and volume of content ingested by Suno's systems across several platforms:

Source Platform Quantity/Hours
YouTube Music 113,879 hours
Pond5 62,117 hours
International Music Score Library Project (IMSLP) 19,514 hours
Genius 17,615 hours
Deezer 12,287 hours
Jamendo 3,726 hours
Freesound 410 hours
MuseScore (lyrics) 103 hours

Further investigation of the code indicates that Suno attempted to harvest approximately one million hours of podcast content via PodcastIndex. The company also reportedly employed a third-party service, Bright Data, to bypass platform protections. This tactic was specifically used to hunt for a cappella vocal tracks on YouTube to refine the platform's voice generation capabilities.

Legal Context and Company Response

Suno has confirmed the incident occurred in November 2025 but has consistently characterized the event as a limited security incident that was quickly contained. The company maintains that the exposed source code was outdated and that no sensitive personal information was compromised. According to a spokesperson, Suno does not have access to full credit card numbers via Stripe, and the company determined that individual notification of customers was not required under applicable privacy laws.

This incident surfaces while Suno faces multiple copyright infringement lawsuits from major record labels, including Sony Music Entertainment and Universal Music Group, as well as the Recording Industry Association of America (RIAA). The plaintiffs have alleged that Suno engaged in stream ripping to circumvent technical measures on platforms like YouTube. Suno has acknowledged that its training data includes essentially all accessible music of reasonable quality on the open internet, but it argues that such activity is protected by the fair use doctrine of US copyright law.

In November 2025, Warner Music Group became the first major label to move from litigation to a licensing partnership with the platform. This agreement permits the use of specific artist voices and likenesses, provided the artists grant explicit permission. Conversely, other figures in the industry, such as Kenneth Blume, have publicly criticized the company’s data sourcing practices, arguing that the platform operates by exploiting the creative labor of artists.

Despite the public outcry and the legal pressure, Suno maintains that its objective is to facilitate original creation. The company stated that it does not use artist names in its training metadata and has implemented safeguards to block prompts that attempt to replicate specific copyrighted works or existing artist identities.

Transparency record

Evidence behind this report

This report synthesizes 9 distinct sources. Open the source ledger below to compare the underlying coverage.

Prepared under the Archypedia Editorial Policy by the Niko Vale editorial desk profile. AI-assisted tools may support drafting and verification; public accountability remains with Archypedia. Report an error.