Microsoft privately called AI scraping “theft” while doing it: Report

Newly unsealed filings in the New York Times lawsuit show both companies knew their training practices threatened publishers — and pressed ahead anyway

Staff Writer
Microsoft
Image: Reuters

Article summary

AI Generated

Newly unredacted filings in the New York Times' copyright lawsuit against OpenAI and Microsoft show that a senior Microsoft executive privately called AI data scraping "the largest theft of labor in human history" — while both companies were actively building training datasets from paywalled news content. Internal documents also show Microsoft's own data found its Copilot tool cut New York Times click-through rates by as much as 93%.

Key points

  • Microsoft exec privately called AI scraping "the largest theft of labor in human history"
  • OpenAI datasets contained over 91,692 copies of works from major news publishers
  • Microsoft's Copilot cut New York Times click-through rates by up to 93%, internal data showed

Subscribe to our free newsletter to continue reading.

Newsletters

Newly unredacted court filings in the copyright lawsuit The New York Times filed against OpenAI and Microsoft in 2023 have exposed a striking gap between what both companies said publicly and what they recorded internally about AI training practices.

The documents show that Microsoft’s director of Applied Science, Brent Hecht, described the scraping of news content in a January 2023 internal memo as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” That memo was written while both companies were actively building training datasets from scraped news content, including material behind paywalls.

Microsoft CEO Satya Nadella also testified in deposition this year that “anything that is paywalled should be licensed by anyone who wants to use it for grounding or training,” and said that had he known OpenAI scraped paywalled material, he would have invoked Microsoft’s right to require OpenAI to retrain its models.

The filings detail how OpenAI and Microsoft obtained the content: scraping it from the Bing Index, pulling millions of articles from Common Crawl, and building datasets named Project Mango and Project Taxi. Project Mango alone reportedly contains copies of at least 160,903 unique works from the named news publishers. OpenAI’s mid-training datasets are said to include more than 91,692 copies of works from the New York Times, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset contained more than 2 million documents from nytimes.com.

The filings also describe what appears to be a deliberate effort to circumvent paywalls. When OpenAI researcher Nick Ryder told company president Greg Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.” Employees also allegedly stripped copyright notices from training data because researchers “wouldn’t want model outputting” them to users.

The documents also contain admissions that cut against OpenAI’s fair use defence in the litigation. OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communications that publishers face an “existential threat” from the chatbot, which is “largely substitutive” and will “get more and more substitutive” as the technology improves. Nadella agreed under oath that conversing with chatbots has substituted the need to visit underlying sources.

Advertisement

Microsoft’s own data showed that its Copilot answer engine caused click-through rates for the New York Times domain to drop as much as 93% compared to traditional Bing search. A presentation by Hecht described this as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the Microsoft document states.

A separate Microsoft document acknowledged a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

“The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Steven Lieberman, counsel for the New York Daily News, said in a statement.

OpenAI and Microsoft did not respond to requests for comment. The underlying exhibits in the case remain sealed. The new information comes from The New York Times’ legal brief rather than those exhibits.

The broader legal question — whether training AI on copyrighted material constitutes fair use — has not been definitively settled. Courts have generally sided with AI companies so far, and earlier this month the Trump administration filed a brief supporting OpenAI’s position on unlicensed training data.

Advertisement