On September 17, 2026, a federal court in Manhattan unsealed a filing in The New York Times' copyright case against OpenAI and Microsoft, putting internal company documents, emails and deposition testimony on the public record. The line drawing the most attention is not a legal argument. It is a prediction about public opinion, written inside Microsoft.
The filing is the news organizations' motion for summary judgment, the request that Judge Sidney Stein resolve central issues without a trial, and it is part of the case the publishers filed in 2023. Much of the material had been sealed or redacted at the companies' request until now. You can read the filing itself on CourtListener.
What the documents say
An internal Microsoft document attributed to Brent Hecht, the company's director of applied science, warned that "millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions." The filing describes that language as perhaps "the largest theft of labor in human history".
The scale is part of the argument. OpenAI's mid-training datasets contained more than 91,692 copies of works published by the Times, the New York Daily News and the Center for Investigative Reporting, and a dataset built from Common Crawl held more than 2 million documents from nytimes.com. The filing also describes a method for getting around the Times' paywall. In the exchange it cites, OpenAI researcher Nick Ryder described the workaround and Greg Brockman, the company's president and co-founder, replied, "ah, nice".
Two more items cut against the public posture. Microsoft chief executive Satya Nadella testified that anything paywalled should be licensed for training or grounding. And Nick Turley, OpenAI's head of ChatGPT, wrote in notes cited in the filing that publishers faced "an existential threat from their products, which are largely substitutive and will get more and more substitutive as they get better".
A separate Microsoft document went further still. It said the company's AI content strategy had started a "doom loop" that would hurt the performance of its models and the entire web at the same time. The plaintiffs' motion also cites Microsoft's own traffic data showing that click-through rates from Bing Chat were 87 to 93 percent lower for the Times' sites than for traditional Bing search. Other summaries of the same material give wider ranges, including 83 to 93 percent for some news plaintiffs and 51 to 94 percent for a broader set, so the precise number is worth handling with care.
The companies' answer
Microsoft says the Hecht passages reflect one employee's individual perspective, are not a legal analysis and do not represent the company's views. A spokesperson pointed to the company's filings, which argue that these transformative uses are consistent with copyright law and that Copilot is not a substitute for publishers' journalism. OpenAI disputes the plaintiffs' reading as well.
The news organizations say the documents eviscerate the fair-use defense. Steven Lieberman, counsel for the New York Daily News and seven sister papers, put it plainly: "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong".
Why it matters
Fair use is decided by four statutory factors: the purpose of the copying, the nature of the work, how much was taken and the effect on the market for the original. Internal documents about copying at scale, paywall workarounds and falling referral traffic speak most directly to the fourth factor, market harm. A company's private sense that its conduct looks like theft is not itself the legal test, and courts do not rule on embarrassment. But evidence that a defendant understood the harm it was causing can matter, because market harm and the question of whether a use substitutes for the original are central to the analysis.
That is the tension the filing sharpens. The public case rests on transformation, the idea that models learn from works rather than reproduce them. The internal record describes substitution, the idea that the products replace the journalism they were built on. Both can be argued at once, and the judge will have to decide which frame fits the evidence.
The stakes reach past this lawsuit. If courts find that training on scraped news infringes, the practical fix is either a licensing market or a statute that sets the terms. The documents do not settle which, and the companies dispute that any of it changes the outcome. But they do make that choice harder to avoid.
Should the answer be a court-ordered licensing market, or a law that sets the terms for training data before the next generation of models is built?