Home / Publications / Blog / ForIntel Research / Unsealed NYT v. Microsoft/OpenAI filings: Microsoft exec called AI training scraping the largest theft of labor in human history

Unsealed NYT v. Microsoft/OpenAI filings: Microsoft exec called AI training scraping the largest theft of labor in human history

Newly unsealed filings in The New York Times's copyright suit against Microsoft and OpenAI include a Microsoft scientist's own words: AI training-data scraping as "the largest theft of labor in human history." The admission points at a fixable design flaw.

By Foragentis teamPublished 2026-09-183 min read

On September 17, 2026, newly unredacted filings became public in The New York Times Company v. Microsoft Corporation, et al., the copyright suit The Times filed against Microsoft and OpenAI on December 27, 2023 in the U.S. District Court for the Southern District of New York. Per TechCrunch's reporting on the material, Brent Hecht, Microsoft's Director of Applied Science, wrote in an internal memo that AI training-data scraping was "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." Engadget's independent read of the same unsealed material corroborates both lines and adds a third: Hecht called the resulting AI models "a product that destroys its supply chain."

He was not alone. Satya Nadella, Microsoft's CEO, testified that paywalled content "should be licensed" for training and said he would have "required OpenAI to retrain" had he known about the paywall scraping at the time. He also acknowledged that chatbots "substituted" for visiting the original websites, which is the plain-language version of the Times's core legal claim. Nick Turley, who leads ChatGPT at OpenAI, wrote that publishers face an "existential threat" from AI products that are "largely substitutive." Greg Brockman, OpenAI's president, called the models "excellent at news" and replied "ah nice" when a researcher described a plan to get around paywalls.

The filings also put numbers on the practice. Per TechCrunch, OpenAI's training datasets contained more than 91,692 copies of works from The Times, The Daily News, and the Center for Investigative Reporting. More than 2 million documents from nytimes.com turned up in the Common Crawl dataset used for training. An internal dataset called "Project Mango" held more than 160,903 unique news publisher works. Microsoft Copilot's launch tracked with a 93% drop in click-through rate to the Times's domain. The filings allege that employees deliberately stripped copyright notices from the training data before it went in.

What makes this filing different from the usual AI copyright story is that the harshest assessment did not come from a plaintiff's lawyer or an outside critic. It came from inside the company being sued, in writing, months before there was any courtroom reason to say it. That is worth sitting with: the people building these systems already had the vocabulary for what was happening. They called it theft privately and kept building.

Strip the legal argument away and look at what actually broke. A system trained on someone else's reporting produced answers that substituted for that reporting, without carrying any pointer back to where it came from, without compensation, and, per the filings, with the original attribution deliberately removed. That is not a side effect of how large language models work. It is a design choice about whether an output is required to carry a traceable link back to a checkable source. Nadella's own testimony makes the point without meaning to: he says he would have required a retrain if he had known. The fix he is describing after the fact is exactly the kind of requirement that belongs at the start of a system's design, not as a correction issued once a lawsuit forces the question.

That is the design bet behind ForIntel, Foragentis's knowledge-products line. Every report we produce is built so that a claim without a verifiable source row does not make it into the output. Provenance is a structural requirement of how the report gets assembled, not a footnote added after the fact or a fix bolted on once someone objects.

How the court rules on damages, fair use, and the DMCA claims in this case is a separate question from the one worth asking today: was the extraction-without-attribution the only way to build a useful AI system? The filings themselves suggest the people building it did not think so. If you are evaluating an AI vendor and want to know whether their outputs trace back to real, checkable sources, that is exactly the question ForIntel was built to answer: https://foragentis.com