Microsoft disclosed in new court filings that fewer than 1 percent of 8.2 million analyzed Copilot chat logs reproduced 16 or more words in common with publishers’ content. The findings were submitted as part of Microsoft’s request for summary judgment in a consolidated copyright lawsuit brought by plaintiffs including The New York Times, the Center for Investigative Reporting, and book authors.
According to Microsoft, an expert for CIR found 51 instances of substantial overlap, while an authors’ expert identified just 24 responses with at least 30 matching words across 8.2 million conversations. Microsoft argues the data demonstrates that LLM outputs rarely substitute for original works, supporting its claim that training AI models on copyrighted text constitutes fair use.
The New York Times rejected the claims, asserting that discovery evidence proves Microsoft and OpenAI misappropriated its content to build competing commercial products. The legal dispute continues under a single judge as both sides await a ruling on summary judgment.
Why it matters
Provides quantitative evidence to defend fair use claims in ongoing AI copyright litigation.
Summary judgment outcome could set a critical legal precedent for commercial model training rights.
Demonstrates low rates of direct text verbatim regurgitation in deployed LLM applications.
Source: theverge.com



