VMTech
Discuss a project

Unsealed NYT filing details claims over OpenAI and Microsoft training data

Unsealed NYT filing details claims over OpenAI and Microsoft training data

Newly unredacted filings in The New York Times copyright lawsuit against OpenAI and Microsoft set out allegations about how news content was acquired for AI training and how generative AI products could affect publishers. The filing says OpenAI’s mid-training datasets contained more than 91,692 copies of works published by the NYT, Daily News and Center for Investigative Reporting.

It also states that a Common Crawl-derived dataset contained more than 2 million documents from nytimes.com. The underlying exhibits remain sealed, and much of the newly public material is drawn from the Times’ brief. OpenAI and Microsoft did not respond to requests for comment cited in the report.

Internal statements and traffic concerns

The filing attributes unusually blunt language to Microsoft and OpenAI personnel. A January 2023 memo by Microsoft Director of Applied Science Brent Hecht reportedly called the practice “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” The quotations are presented in the filing without their original context.

Microsoft’s internal analysis cited by the Times said its Copilot “answer engine” reduced click-through rates to the NYT domain by as much as 93% compared with traditional Bing search. A January 2024 Microsoft presentation described a possible “doom loop,” in which products that reduce publisher traffic could hurt both the web and model performance.

The filing also cites internal comments from OpenAI Head of ChatGPT Nick Turley describing an “existential threat” to publishers from products that are “largely substitutive.” Microsoft CEO Satya Nadella testified that paywalled material should be licensed for grounding or training, and said he would have required OpenAI to retrain models if he had known paywalled information had been scraped and used for training.

Allegations over collection and use

The Times alleges that OpenAI and Microsoft gathered material at scale through sources including the Bing Index, Common Crawl and datasets such as WebText and WebText2. It says OpenAI supplied its full GPT-3 training dataset to Microsoft for evaluation of commercial implementations, while Microsoft provided data to OpenAI through initiatives named Project Taxi and Project Mango.

Project Mango allegedly contained copies of at least 160,903 unique works from the news publishers. The filing further alleges that OpenAI staff discussed a way to bypass the NYT paywall without detection and that copyright notices were stripped from training data to avoid models returning those notices to users.

Those claims arrive as the legal status of training on copyrighted work remains unresolved. Courts have generally been receptive to AI companies’ fair-use arguments, but the Times’ filing focuses on whether the use substitutes for the originals or harms their market. The broader dispute also sits alongside litigation involving publisher claims against OpenAI and Microsoft over publishers’ claims against OpenAI and Microsoft.

What organisations should examine

For organisations building AI search, grounding or retrieval products, the filings underline the operational importance of tracing content sources and assessing product effects on the suppliers of that material. Licensing terms, paywall controls, copyright metadata and referral patterns are practical issues to review before a system presents answers that could displace visits to the original publisher.

#aiethics#copyright#openai#microsoft
Open analytics
On the site 5 views
min read 4 17.09.2026
Instagram

Unsealed NYT filing details claims over OpenAI and Microsoft training data

Open the post on Instagram ↗