Microsoft and OpenAI’s Web Scraping Debate Exposes AI’s News Problem

Microsoft and OpenAI executives described AI web scraping in unusually blunt terms, calling it a threat to journalism and a challenge to the idea of fair use. Their comments appear in unredacted court documents tied to a lawsuit filed by The New York Times.
The documents show that the debate was not limited to outside critics. Brent Hecht, Microsoft’s director of Applied Science, called OpenAI’s work “the largest theft of labor in human history.” He also described scraping news for AI training as “an astonishing theft of unprecedented proportions” and said the plan made “a complete mockery of the idea of fair use.”
Executives warned that AI could replace news websites
Hecht’s concerns focused on what happens after an AI model learns from news articles. If a chatbot gives users the information they want without sending them to the original publisher, the news site loses the visit, and the publisher loses a chance to earn money from that reader.
Microsoft recorded 83–93 percent drops in click-through rates for some news plaintiffs and 51–94 percent drops for others. Those figures point to a direct problem for publishers: even when an AI product includes links, users may not follow them to the original stories.
Satya Nadella, Microsoft’s CEO, described chatbots as stealing clicks from news sites by “giving you the information right there on the website on the AI platform versus needing to go to the underlying source.” A software engineer at OpenAI made the concern even clearer, writing, “no matter how prominently we show the links, users won’t click.”
Nick Turley, an OpenAI executive, said AI training on news content posed an “existential threat to publishers.” He wrote that products trained on news content could substitute for news providers, calling AI products “largely substitutive” to journalism and predicting they “will get more and more substitutive as they get better.”
That concern also appeared in a Microsoft document, which described a “doom loop” that would hurt “the performance of our models and the entire web at the same time.” The idea is direct: if AI tools reduce the traffic that supports publishers, fewer publishers may be able to create new work, leaving the web with less original material for future models to learn from.
Questions over training data and permission
Hecht acknowledged the central fairness issue in plain language: “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.” His January 2024 internal document called the practice “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” A January 2023 internal memo also addressed the same conflict over scraped news and AI training.
The documents describe large collections of news material in OpenAI’s datasets. They contained more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com, while OpenAI obtained 1.8 million articles from the NYT from a third party.
Those numbers do not settle the legal question by themselves, but they show the scale of the material involved. They also explain why the dispute reaches beyond one publisher or one model. The argument concerns whether companies can use large amounts of published work to build AI products that may answer questions without sending readers back to the people who created the work.
Nadella testified that “anything that is paywalled should be licensed by anyone who wants to use it for grounding or training.” That position sets out a clear rule for paid news, even as the documents show tension between the need for permission and the way AI datasets were assembled.
A revealing moment inside OpenAI
Another exchange involved Greg Brockman, OpenAI’s president. When told about a hack to get around the NYT paywall, Brockman responded, “ah nice.” The brief response appears alongside the wider discussion about access to news, paywalls, and the use of published material in AI systems.
Taken together, the statements show executives recognizing two connected risks. The first is that AI companies may use journalism without payment or permission. The second is that the resulting products may give readers less reason to visit news sites, weakening the system that produces the content AI models depend on.
The court documents therefore present a conflict at the heart of AI development. OpenAI’s systems rely on vast amounts of written material, while publishers depend on readers reaching their websites. The executives’ own words show how easily those goals can collide, especially when AI products become better at providing answers instead of directing users to the original source.
Based on
- Microsoft executive called OpenAI’s web scraping the ‘largest theft of labor in human history’ — engadget.com
- Microsoft exec called AI scraping the “largest theft of labor in human history” – Ars Technica — arstechnica.com
- Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal | TechCrunch — techcrunch.com




