Abstract:
In response to a joint copyright infringement lawsuit initiated by news publishers and book authors, Microsoft submitted the latest defense materials to the court, insisting that its artificial intelligence chatbot Copilot almost never copies verbatim or substantially reproduces reports from media such as The New York Times and copyrighted literary works. This stance aims to build a solid "fair use" legal defense for Microsoft and OpenAI, and push the court to end related cases through summary judgment as soon as possible.

As part of the discovery process in the lawsuit, Microsoft released a database of 8.2 million transcripts of actual Copilot conversations to experts hired by the plaintiff news organizations. Microsoft pointed out that the data set was not randomly selected, but was specifically screened based on keywords related to the plaintiff's news website. In theory, it is a sample set that is "most likely to contain allegedly infringing content."
However, data analysis results revealed extremely low repetition frequencies. Among these 8.2 million filtered records, only 59,545 records contained at least 16 consecutive words that overlapped with the news content used for model training, accounting for less than 1% of the total records. In addition, in the lawsuit involving the non-profit investigative reporting organization "Center for Investigative Reporting" (Center for Investigative Reporting), the original expert only identified 51 cases in a massive sample that constituted "substantial overlap" with the organization's material.
In another consolidated case brought by a group of writers over book copyrights, test data showed a similar trend. Independent experts hired by the plaintiff’s writers conducted a review of the 212 books involved in the case, and ultimately found generated records with at least 30 word overlaps in only 10 books, with a total of only 24 such matching samples.
Microsoft argued to the court that these detailed data just confirmed its core legal defense, that is, using copyrighted materials to train large language models is "fair use" under the framework of U.S. copyright law. Microsoft emphasized that the purpose of using copyrighted materials by Copilot and the generative AI system behind it is fundamentally different from that of the original publication. Its fundamental mechanism is to absorb and reorganize language patterns to generate new content, rather than circulate as a substitute for the original work. Even if sporadic text fragments reappear occasionally, it cannot erase the "transformative use" characteristics of the large language model training process.
Previously, a number of media organizations and creator groups, including the New York Times, the Center for Investigative Reporting, and the Authors Guild, jointly filed a lawsuit accusing Microsoft and OpenAI of illegally grabbing millions of original articles and books without authorization to train their commercial AI platforms, and directly competing with original authors by generating highly overlapping content, severely impacting the business ecosystem of independent journalism and literary creation. In order to improve the efficiency of judicial proceedings, related lawsuits involving news publishers and writers have been consolidated and handed over to the same federal judge.
At present, Microsoft has officially applied to the court for summary judgment, hoping to use the above data to prove that the plaintiff’s substantive infringement accusation lacks a universal factual basis, thereby resolving this key industry lawsuit at an early stage regarding the legality of the underlying training of generative artificial intelligence. As of now, the New York Times, the Center for Investigative Reporting and the Writers Guild of America have not publicly commented on the latest statistics and court statements submitted by Microsoft.
Comments