In newly unsealed court filings, a senior Microsoft scientist privately described OpenAI’s mass scraping of online journalism to train ChatGPT as “the largest theft of labor in human history,” exposing a stark gap between Big Tech’s public defense of AI training and what some of its own experts feared behind closed doors.
The phrase appears in internal memos from Brent Hecht, Microsoft’s director of applied science, unearthed as part of The New York Times’ landmark copyright lawsuit against OpenAI and Microsoft over the use of its reporting to build generative AI models. Hecht warned colleagues that using millions of news articles as training fuel without compensation would be seen as “an astonishing theft of unprecedented proportions,” and he worried it could trigger a “doom loop” where AI systems cannibalize the very publishers they depend on.
The Times’ suit, filed in 2023, accuses OpenAI and Microsoft of copying vast amounts of Times content, including paywalled stories, to train ChatGPT and Microsoft’s Copilot products without permission. Newly unredacted documents claim OpenAI and Microsoft built training datasets by mass-scraping news sites, allegedly bypassing paywalls and even stripping copyright notices, turning that secret sauce into evidence against them. One filing says a mid-training dataset alone contained more than 91,000 copies of Times and related publishers’ works, underscoring the scale of what news organizations argue is wholesale appropriation of their labor.
Internally, Microsoft’s own risk assessments paint a grim picture of what happens if AI answer engines replace clicks to publishers with synthetic summaries. Hecht’s presentation reportedly warned that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained,” predicting that fewer readers and less ad revenue would mean fewer reporters and lower-quality news — which in turn degrades future AI models trained on that shrinking corpus. It’s a self-inflicted content crisis: kill the news ecosystem, and you eventually starve your models.
OpenAI executives were hardly sanguine about the fallout either, according to the filings. One OpenAI leader, Nick Turley, allegedly described the company’s approach as an “existential threat to publishers,” bluntly acknowledging that newsrooms could be gutted by AI systems that ingest their work, answer users’ questions, and never send those users back to the source. The documents also suggest OpenAI engineers explored ways to get around paywalls for training data, raising the stakes on whether the court views this as aggressive fair use or deliberate infringement.
Microsoft’s top leadership, meanwhile, is trying to draw a clearer line in the sand. In a deposition cited by the Times, CEO Satya Nadella said that “anything that is paywalled should be licensed by anyone who wants to use it,” and added that if he had known OpenAI trained on paywalled content he would have pushed to retrain the models. A Microsoft spokesperson has stressed that Hecht’s memos don’t represent official company policy, but the filings show the company was warned internally that relying on scraped news might make “a complete mockery of the idea of ‘fair use.’”
For geek-culture fans who live on specialist sites, wikis, and forums, the implications are immediate. The same scraping pipelines that pulled in national and international news almost certainly harvested movie reviews, game guides, tabletop homebrew blogs, and anime commentary that now help power AI search and chat experiences. One analysis cited in the filings claims Microsoft’s Copilot answer box cut click-throughs to Times stories by up to 93 percent compared to standard Bing search, a statistic that should send shivers down the spines of small entertainment outlets already fighting for every view. If AI front-ends become the default way people ask “What happened in the latest MCU show?” or “How do I beat that Elden Ring boss?”, the human writers who built that knowledge base may vanish from the business model.
Courts and regulators are only beginning to grapple with that tension. The Times case has already spawned accusations that OpenAI withheld or obscured evidence about how it trained its models, suggesting that full transparency on training data usage remains elusive. Privacy and data-protection authorities, including a joint investigation summarized by Canadian regulators, have started probing how systems like ChatGPT source and process public content, but the copyright questions at the heart of the news lawsuits go further: should scraping an entire generation’s creative output be treated as fair use, or as mass unpaid labor? For now, the filings show that even inside Microsoft and OpenAI, some of the people building AI are worried the industry is devouring the internet’s creative and journalistic commons faster than it can be rebuilt — and that the bill for that “largest theft of labor” may soon come due.








