Skip to main content
The Vergecast

How to train your data

26 min episode · 2 min read
·
Alex Reisner

Episode

26 min

Read time

2 min

Topics

Productivity, Leadership, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Training data defines model capability: A model's output is determined almost entirely by what it trains on — not just its architecture. A music model trained on 1950s jazz produces 1950s jazz; one trained on hip-hop produces hip-hop. Reisner argues the most accurate name for any AI model is a description of its training dataset, not a product name like GPT or Claude.
  • Competitive secrecy masks legal exposure: AI companies cite competitive advantage as the reason training data stays secret, but a second motive is equally significant — much of the data was scraped without creator consent. Authors, musicians, and video creators frequently discover their work was used only after the fact, and companies prefer to avoid that conversation entirely rather than address it proactively.
  • Data laundering via nonprofits and universities: AI companies routinely route data acquisition through academic institutions and nonprofits — such as Common Crawl and the European organization LION, which holds 12 million songs scraped from YouTube — allowing companies to claim distance from the scraping while funding those organizations directly and benefiting from the resulting datasets.
  • YouTube is the dominant scraping target due to weak technical barriers: YouTube appears repeatedly across AI training datasets because download tools work reliably and easily, and because digital rights management protections are absent compared to platforms like Spotify. YouTube states scraping violates its terms of service but has taken no effective technical or legal action to stop it over multiple years.
  • Synthetic data does not replace human-generated content: Research on "model collapse" demonstrates that training AI models on their own outputs causes rapid quality degradation. Reisner states AI companies deliberately use vague language when claiming synthetic data works. The next actual data frontier is paying human creators directly — one company has already paid creators over $10 million to produce content specifically for AI training purposes.

What It Covers

Atlantic staff writer Alex Reisner joins The Vergecast to examine AI training data — what it is, where it comes from, and why companies guard it so closely. The conversation covers Common Crawl, YouTube as a data source, music datasets, synthetic data limitations, and the emerging paid creator economy for AI training.

Key Questions Answered

  • Training data defines model capability: A model's output is determined almost entirely by what it trains on — not just its architecture. A music model trained on 1950s jazz produces 1950s jazz; one trained on hip-hop produces hip-hop. Reisner argues the most accurate name for any AI model is a description of its training dataset, not a product name like GPT or Claude.
  • Competitive secrecy masks legal exposure: AI companies cite competitive advantage as the reason training data stays secret, but a second motive is equally significant — much of the data was scraped without creator consent. Authors, musicians, and video creators frequently discover their work was used only after the fact, and companies prefer to avoid that conversation entirely rather than address it proactively.
  • Data laundering via nonprofits and universities: AI companies routinely route data acquisition through academic institutions and nonprofits — such as Common Crawl and the European organization LION, which holds 12 million songs scraped from YouTube — allowing companies to claim distance from the scraping while funding those organizations directly and benefiting from the resulting datasets.
  • YouTube is the dominant scraping target due to weak technical barriers: YouTube appears repeatedly across AI training datasets because download tools work reliably and easily, and because digital rights management protections are absent compared to platforms like Spotify. YouTube states scraping violates its terms of service but has taken no effective technical or legal action to stop it over multiple years.
  • Synthetic data does not replace human-generated content: Research on "model collapse" demonstrates that training AI models on their own outputs causes rapid quality degradation. Reisner states AI companies deliberately use vague language when claiming synthetic data works. The next actual data frontier is paying human creators directly — one company has already paid creators over $10 million to produce content specifically for AI training purposes.

Notable Moment

Reisner flatly contradicts a widely held industry assumption: when asked whether synthetic data represents the next training frontier, he says the evidence shows the opposite. Models trained on their own outputs degrade quickly because AI functions as an averaging machine, and human-generated content contains qualities that AI outputs simply do not replicate.

Know someone who'd find this useful?

Episode Transcript

Hello, and welcome to the Vergecast, the flagship podcast of music that sounds eerily but not exactly like other music. I'm your friend David Pearce. And today on the show, we're talking about training data. It's the raw materials of everything that is AI. And we probably don't understand it or talk about it enough. I'm talking to Alex Reisner, who's a staff writer at The Atlantic. And over the last couple of years, he has spent a lot of time investigating training data and how all of these books and all of these articles and all of these YouTube videos and all of these songs are compiled into these gigantic datasets that AI companies then use to form the basis of their models. I think the way that these models are created and the sources of these data has a lot to do with the way that we feel about AI, particularly generative AI as a creative expression. And understanding how this data works, where it comes from, and how it gets used, I think is really important. So we're gonna have Alex on. We're gonna dig into it. I'm very excited about it. But first, here's everything else happening on the Verge today. This is Ninety Seconds on the Verge for Thursday, 06/25/2026. Apple just raised the prices of a huge number of its most important products. IPads and MacBooks in particular are now up anywhere from a 100 to $500, and even, like, the Apple TV is way up. It's now $200, roughly the price of 900 Amazon Firesticks. We knew this was coming, but this is still a big moment. Apple is as well managed a supply chain company as anybody and exists at really high margins. If it can't keep prices down in this era of AI driven shortages of memory and storage, nobody can. And I suspect this is going to get even worse. My advice from yesterday still very very much holds, which is Prime Day is not over. Go get your deals while you still can. And here's a way to help you pay for all of those more expensive gadgets. Go get some money from Disney. If you have YouTube TV or DirecTV stream, you might be eligible for a cash payout from a new $50,000,000 settlement. The case started in 2022, and the argument was essentially that because Disney owns ESPN and Hulu, which are both powerful streaming services in their own right, Disney was able to drive up the prices of all of its rivals to unnecessary heights. Heights. Live TV is lucrative and competitive and is going to keep being messy in this way. This will not end things. But hey, go get some of that $50,000,000. And finally, a quick PSA to advertisers everywhere. Yes, I get it. The whole idea of using AI creative and AI targeting to put your AI ads in front of AI people is exciting, I guess. But maybe check your ads …

Get the full transcript (5,215 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The Vergecast transcripts →

You just read a 3-minute summary of a 23-minute episode.

Get The Vergecast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • The conversation covers Common Crawl, YouTube as a data source, music datasets, synthetic data limitations, and the emerging paid creator economy for AI training.
  • Data laundering via nonprofits and universities: AI companies routinely route data acquisition through academic institutions and nonprofits — such as Common Crawl and the European organization LION, which holds 12 million songs scraped from YouTube

More from The Vergecast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Tech Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The Vergecast.

Every Monday, we deliver AI summaries of the latest episodes from The Vergecast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime