Newly unsealed parts of the copyright lawsuit The New York Times filed against OpenAI and Microsoft three years ago contain statements from company executives that characterize AI training practices as theft and warn of an existential threat to news publishers. According to the filing, a top Microsoft executive privately described the firms’ AI training as “theft,” and OpenAI leadership said its models posed an “existential threat” to the journalists whose work trained them. The documents also allege that the companies bypassed paywalls undetected, built training datasets via mass scraping, and deliberately stripped copyright notices from the data before it reached models. One internal Microsoft presentation from January 2024, written by director of Applied Science Brent Hecht, described a observed 93% drop in click‑through rates for The New York Times’ domain when using Microsoft’s Copilot “answer engine” compared with traditional Bing search, calling the trend a “doom loop” that would hurt model performance and the web. In a deposition earlier this year, Microsoft CEO Satya Nadella testified that anything behind a paywall should be licensed for grounding or training, adding that if he had known OpenAI had scraped paywalled material he would have required the company to retrain its models. OpenAI’s head of ChatGPT, Nick Turley, wrote in an internal message that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and will become more so as they improve. OpenAI President Greg Brockman described the models as “excellent at news,” and Nadella agreed under oath that conversing with chatbots has substituted for visiting the original source. A Microsoft document cited in the filing warns of a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” The scale of the copying is highlighted by figures showing that OpenAI’s mid‑training datasets contain more than 91,692 copies of works from the New York Times, Daily News, and Center for Investigative Reporting, while a Common Crawl‑derived dataset includes over two million documents from nytimes.com alone. In a January 2023 memo, Hecht called the practice “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” The filings also detail how OpenAI and Microsoft exchanged data through projects named Taxi and Mango, with the latter assembling a dataset of at least 160,903 unique works from the news publishers. Employees allegedly devised ways to circumvent nytimes paywalls, and when researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied “ah nice.” The companies also allegedly pulled millions of articles from the open repository Common Crawl and removed copyright notices from training data to avoid model output of such notices. Steven Lieberman, counsel for the New York Daily News, said the evidence shows the firms knew their actions were wrong. OpenAI and Microsoft did not respond to requests for comment.

