Microsoft Staff Called AI Scraping Biggest Labor Theft in History

Unsealed court documents from the NYT vs. OpenAI lawsuit reveal Microsoft employees debated whether AI training on news content was the largest theft of labor in human history.

Eliza Crichton-Stuart

Eliza Crichton-Stuart

Updated

Microsoft Staff Debated AI Scraping as Massive Labor Theft
Advertisement

"Millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions."

That line did not come from a critic on the outside. It came from inside Microsoft.

Court documents unsealed this week in the New York Times copyright lawsuit against OpenAI and Microsoft have pulled back the curtain on some uncomfortable internal conversations. The filings reveal that Microsoft employees were asking hard questions about AI training practices years before the public debate caught up, including whether the whole operation amounted to the largest theft of labor in human history.

EXCLUSIVE OFFER

Use code TRYHARD33 at checkout to redeem this offer before code expires again.

Get 33% Off GAMES+ Subscription

What the unsealed documents actually say

The lawsuit, which the Times filed in late 2023, has since been joined by eleven other publishers. Judge Sidney Stein of the Southern District of New York is currently weighing summary judgment motions, and documents are being unsealed as part of that process.

The memo that produced the most striking language came from Brent Hecht, a Microsoft director of applied science who also held a position at Northwestern University. In a 2023 internal document, Hecht wrote that large AI models "are a product that destroys its supply chain," and warned that the public would eventually view the mass ingestion of their work as theft on an unprecedented scale.

Microsoft has been quick to distance itself from those words. The company says Hecht was employed specifically to "present divergent and asymmetric perspectives" and was not a decision maker. His memos, Microsoft argues, do not represent company views.

Here's the thing, though: the fact that someone was paid to raise those concerns, and that those concerns were documented, is exactly the kind of evidence that tends to matter in court.

Nadella's testimony and the paywall question

Satya Nadella, Microsoft's CEO, testified that "anything that is paywalled should be licensed by anyone who wants to use it." He went further, saying that had he known OpenAI was training on paywalled content, he would have exercised Microsoft's contractual right to require a retrain of the models. A company spokesman later said Nadella "spoke to broad principles" about how people find and consume information, which reads as a fairly significant walk-back.

The gap between what Nadella said under oath and what the company's spokesman clarified afterward is worth paying attention to as the case moves forward.

OpenAI's internal conversations were just as candid

The documents do not let OpenAI off the hook either. An OpenAI staffer told president Greg Brockman about building a workaround to bypass the Times paywall during training. Brockman's response: "ah nice."

Nick Turley, who ran the ChatGPT product team, wrote in June 2023 that AI posed an "existential threat" to publishers. By February 2024, he was writing that AI products "will get more and more substitutive as they get better." An OpenAI engineer, writing in February 2023, noted that "no matter how prominently we show the links, users won't click" on sources, which cuts directly against the argument that chatbots drive traffic back to publishers.

What most players miss here is that these are not leaked documents from a whistleblower. They are internal communications that OpenAI and Microsoft created, stored, and are now being required to produce in litigation. The companies knew what they were writing at the time.

The "doom loop" warning

Beyond the labor theft framing, Hecht's 2023 memo raised a separate concern that has since become a live debate in AI research circles. The argument was that training on AI-generated content creates a feedback loop that degrades model quality over time, as the outputs of one generation of models become the training data for the next.

This is not a fringe concern. Researchers have been studying what happens to model outputs when synthetic data saturates the training pool, and the results are not encouraging. The key here is that someone inside Microsoft was flagging this risk internally at the same time the company was publicly bullish on AI's potential.

The fair use argument and what publishers say

Both Microsoft and OpenAI are arguing that training on published content constitutes fair use, transforming articles into something new rather than substituting for the originals. That argument has not been tested at trial yet.

Steven Lieberman, who represents the New York Daily News and seven other plaintiff publishers, summarized the plaintiffs' read on the unsealed documents plainly: "The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."

A 2020 memo from Jack Clark, then OpenAI's policy director, warned that the company was "creating systems that substitute for the labor of the people that define the culture of society" and would "become the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet." Clark later left to co-found Anthropic.

The Times, as a plaintiff in its own case, declined to comment. OpenAI did not respond to requests for comment.

What this means for the broader AI and gaming ecosystem

This case matters well beyond the publishing industry. The same questions about training data, consent, and compensation are being asked across creative fields, including game development, concept art, voice acting, and narrative writing. If the courts ultimately rule that mass AI scraping does not constitute fair use, the implications for how AI tools are built and licensed would ripple across every creative industry.

For gamers and developers watching AI tools become standard parts of production pipelines, the outcome of this case could shape what those tools look like and who gets paid when they are used. You'll want to keep an eye on Judge Stein's summary judgment ruling, which will signal whether this goes to trial or gets resolved before the full evidentiary record is tested.

For more coverage of the industry stories that affect how games get made, check out our gaming guides and keep watching this space as the lawsuit progresses through the courts.

Reports

updated

September 18th 2026

posted

September 18th 2026

Advertisement