The copyright lawsuit brought by The New York Times against OpenAI and Microsoft gained a newly unsealed document on September 17. It is a legal brief filed by the Times side that had previously been on the docket in redacted form; the unredacted version quotes remarks a Microsoft executive made in an internal presentation in January 2024.
The executive is Brent Hecht, Microsoft's director of applied sciences. According to the brief, he described training large language models on news content as:
"an astonishing theft of unprecedented proportions" ... "the largest theft of labor in human history"
The brief also states that Hecht warned OpenAI that scrubbing which content in its systems came from plaintiffs such as the Times could amount to an "accidental cover up."
The numbers cited in the brief
The Times' filing lays out the makeup of several datasets, compiled from material produced during discovery in the case:
- An interim OpenAI training dataset contained more than 91,692 copies of works from The New York Times, the Daily News, and the Center for Investigative Reporting;
- A dataset derived from Common Crawl contained more than 2 million files from nytimes.com;
- A dataset code-named Project Mango contained at least 160,903 distinct works from news publishers.
Agence France-Presse reported a separate top-line figure: OpenAI scraped more than 10 million articles, with close to a third coming from The New York Times. That count uses a different methodology from the figures above; the brief's own numbers take precedence.
The alleged conduct falls into three categories: bypassing paywalls to scrape content, large-scale crawling carried out through Bing's indexing, and stripping copyright management information before material entered the training sets. That last item falls under a separate provision of U.S. copyright law, and if proven, the damages calculation differs from ordinary copying infringement.
How the two companies have responded
A Microsoft spokesperson told AFP that Hecht's remarks were "the personal view of one employee" and did not represent the company's position. TechCrunch reported that neither OpenAI nor Microsoft responded to its requests for comment. As of publication, neither company has addressed the specific allegations about dataset scale, paywalls, or copyright notice removal, and Hecht himself has not commented publicly.
Where the three-year case stands
The Times filed the suit in December 2023 in the U.S. District Court for the Southern District of New York; the Daily News and the Center for Investigative Reporting later joined as co-plaintiffs. This site reported in June on Times publisher Sulzberger publicly accusing AI companies of "brazenly stealing content" — but that was a statement from the newspaper's own side. What gives this new filing its weight is that the wording comes from an internal document produced by an employee of one of the defendants.
Among comparable lawsuits, the one that has gone furthest is the authors' class action against Anthropic. In that case, the judge drew a line between training on legally purchased books and training on books downloaded from pirate libraries — the latter went straight to a damages phase and ended in a settlement of roughly $1.5 billion. The Times case's central allegations, about scraping paywalled content and stripping copyright notices, sit closer to the part of that earlier case that was decided unfavorably for the defendant.
How much weight a judge ultimately gives this presentation is not something the brief itself can settle. It shows that someone inside Microsoft saw things this way; whether it shows the two companies knowingly did wrong will depend on further discovery and trial. No trial date has been set.
Sources: TechCrunch, The Washington Post, CocoLoop, Agence France-Presse, Futurism; the unredacted New York Times legal brief was checked against the Hecht quotes and the three dataset figures, and Microsoft's response is per AFP's reporting.