Unsealed Court Filings Expose Microsoft and OpenAI Fears Over AI’s Impact on News

Image: News.bloomberglaw
Main Takeaway
Unsealed court filings show Microsoft and OpenAI executives acknowledged that AI tools copied news content, displaced publisher traffic, and threatened journalism’s business model.
Jump to Key PointsSummary
What the unsealed filings show
Newly unsealed filings in The New York Times’ copyright case portray Microsoft and OpenAI executives privately acknowledging that AI systems were built with large volumes of journalistic work and could displace the publishers that produced it. A Microsoft executive described the practice as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history,” according to the filing and multiple accounts of its contents.
The documents also describe internal recognition that products such as ChatGPT and Microsoft Copilot competed directly with news websites. OpenAI executive Nick Turley wrote that the company’s products were “largely substitutive,” while Microsoft CEO Satya Nadella testified that conversations with chatbots had replaced visits to publisher sites, Bloomberg Law reported. The statements sit alongside Microsoft and OpenAI’s public legal position that model training qualifies as lawful, transformative fair use.
How publishers say scraping worked
The filings allege that Microsoft and OpenAI used restricted Times material in datasets, bypassed paywalls, and stripped copyright notices. The material reportedly includes more than 91,000 copies of works from The Times and related publishers in OpenAI mid-training datasets, according to a detailed account cited by Msukhareva.substack.
Those allegations matter because the case concerns both model training and the products built on that training. Publishers argue that copying articles without permission supplied the raw material for systems that can summarize or reproduce reporting, while chatbot answers reduce the clicks that fund the original work. Ground cited Microsoft data showing Copilot cut New York Times referrals by 93%, although the broader filings and publisher arguments frame traffic substitution as part of a wider economic threat.
The internal fear of a news loop
The documents describe a feared cycle in which AI tools consume journalism, answer questions without sending readers to publishers, and weaken the revenue base needed to produce new reporting. OpenAI and Microsoft personnel reportedly recognized that the systems could threaten news organizations and journalists even as the companies defended their data practices in court.
That concern reaches beyond copyright ownership. News publishers depend on subscriptions, advertising, licensing, and referral traffic, while generative AI systems can deliver an answer inside a closed interface. If readers stop visiting the original article, publishers lose both audience data and commercial opportunities. Ars Technica described the internal concern as a “doom loop,” and Bloomberg Law reported that the admissions extended to books as well as newspaper articles.
Why the admissions matter legally
The statements give publishers evidence for an argument that the AI companies understood the commercial effect of their systems and the value of the material used to train them. They also help plaintiffs connect data collection, product design, and traffic loss into one claim about market harm.
The admissions don't decide whether the companies infringed copyright. Courts still must assess questions including fair use, the legal significance of intermediate copying, the nature of model outputs, and whether publishers can prove measurable economic injury. Microsoft and OpenAI maintain that their AI products are transformative, while the newly public communications provide the publishers with a sharper contrast between private assessments and public defenses.
The pressure on AI companies
The filings increase pressure on Microsoft and OpenAI to explain how training data was obtained, how copyrighted works were handled, and whether safeguards prevented verbatim reproduction. They also raise questions for other model developers, because the dispute targets a business practice shared across the generative AI industry rather than a single chatbot feature.
For publishers, the case strengthens the push for licensing agreements, attribution, crawler controls, and compensation tied to AI usage. For developers, it reinforces the need for documented data provenance and systems that distinguish licensed content from material gathered through access restrictions. The dispute has already shifted from an argument about technical training methods to a fight over who captures the economic value of professional knowledge.
What happens next in the case
The next stage will turn the newly unredacted evidence into arguments over summary judgment and, if necessary, trial. The New York Times and other publishers will use the executive statements, dataset figures, and traffic data to argue that the companies knew their systems substituted for original reporting and threatened the market for it.
Microsoft and OpenAI will contest the legal weight of those statements and the interpretation of the underlying data. The outcome will shape how courts evaluate AI training, publisher licensing, chatbot citations, and claims that generative systems replace visits to original websites. It will also influence negotiations between model developers and news organizations, which now have a clearer record of the economic risk at the center of the fight.
Key Points
Microsoft executive called AI scraping the largest theft of labor in human history, unsealed filings show.
OpenAI executive described ChatGPT products as largely substitutive for publisher websites and original reporting.
OpenAI datasets reportedly contained more than 91,000 copies of Times-related works.
Microsoft data cited in court filings showed Copilot cut New York Times referrals by 93%.
Publishers argue AI-generated answers threaten traffic, subscriptions, advertising, and journalism funding.
Questions Answered
The Microsoft executive called mass AI scraping “the largest theft of labor in human history.” The statement appears in newly unsealed filings connected to The New York Times’ copyright lawsuit against Microsoft and OpenAI.
OpenAI executive Nick Turley wrote that the company’s products were “largely substitutive.” Microsoft CEO Satya Nadella also testified that chatbot conversations had replaced visits to publisher sites.
OpenAI mid-training datasets reportedly contained more than 91,000 copies of works from The New York Times and related publishers. The figure is cited in accounts of the newly unsealed court documents.
Microsoft data cited in the litigation reportedly showed Copilot cut New York Times referrals by 93%. Publishers use the figure to argue that chatbot answers create measurable market harm.
The court will evaluate the unsealed evidence alongside arguments over fair use, copying, model outputs, and market harm. The ruling could influence AI training licenses, publisher compensation, attribution, and crawler policies.
Source Reliability
40% of sources are highly trusted · Avg reliability: 72
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems