Internal documents from Microsoft executives are exposed, claiming that AI training may be "the largest labor theft in human history."

📅 2026-09-18

Abstract:

Controversy surrounding the legality of generative artificial intelligence training data has heated up again recently. A court document unsealed in the New York Times copyright lawsuit shows that a senior research director at Microsoft once described in an internal discussion that the current training method of large language models may be "the largest labor theft in human history." Related statements have attracted widespread attention in the industry.

The remarks exposed this time came from Brent Hecht, director of applied science at Microsoft. According to the latest public legal documents, when discussing Microsoft, OpenAI and other artificial intelligence companies' use of massive Internet content to train models, Hecht once called this process "an astonishing theft of unprecedented scale" and further stated that this may be "the largest labor theft in human history."

This content appears in a copyright lawsuit filed by The New York Times and other news publications against Microsoft and OpenAI. The plaintiffs believe that the two companies used news reports and other copyrighted content to train artificial intelligence models without authorization and use them to develop commercial products.

Microsoft and OpenAI have always adhered to their "fair use" stance. The two companies believe that the model training process is similar to students reading books to learn knowledge and falls within the scope of fair use. They also argue that AI-generated content has sufficiently transformed the original material to not constitute direct copying.

However, some internal communication records in the newly unsealed documents seem to be in clear contrast to this public position.

The documents show that Hecht once believed that if they were to win the lawsuit, the companies involved might need to "completely make a mockery of the concept of fair use." This statement is regarded as important evidence by the plaintiff and is used to question the copyright defense logic that artificial intelligence companies have long adhered to.

At the same time, the document also disclosed some internal discussions from OpenAI. One of the records shows OpenAI co-founder Greg Brockman admitting that when the model was exposed to New York Times content, ChatGPT was indeed able to predict and continue sentences in related articles. This type of content is used by plaintiffs to prove that the model has the ability to reproduce the original work, rather than just abstracting the learned information.

The lawsuit documents also revealed that Hecht’s concerns about the long-term development of the artificial intelligence industry are not limited to the copyright issue itself. He once proposed the so-called "doomsday loop" concept in another internal presentation document.

Following this logic, artificial intelligence answer engines are gradually reducing users’ need to visit news websites, resulting in a decline in media traffic and business revenue. As more and more content production organizations lose their viability, the sources of high-quality information for artificial intelligence systems to learn in the future will also decrease, which will ultimately harm the model and the entire Internet ecosystem.

He pointed out in internal documents that it is extremely rare for a product to threaten the economic foundation of its key suppliers, but the artificial intelligence industry is currently in this situation. Since news organizations and content creators are important sources of information for AI models, when these groups are impacted, the entire content supply chain is at risk.

The plaintiff also cited Microsoft’s internal data and claimed that AI answering systems such as Copilot have had a significant impact on the traffic of some news websites. In some cases, the proportion of users clicking to enter the New York Times page dropped by more than 90% compared with traditional search results.

In fact, the dispute over the source of training data has become one of the most important legal challenges in the current artificial intelligence industry.

OpenAI has previously stated publicly that it is an almost impossible task to train advanced artificial intelligence systems without any contact with copyrighted content. Former Meta executive Nick Clegg also made a similar statement, believing that if traditional copyright rules are strictly followed, many current AI systems will be difficult to maintain development.

At the same time, many companies, including Microsoft, Meta and other technology companies, continue to face criticism from the outside world due to issues with the source of training data. Some companies have been accused of using user content, code libraries, articles, images and other online resources to train models, often without explicit authorization or compensation from the relevant creators.

As the New York Times lawsuit continues, more and more internal documents are being made public. Industry insiders believe that these materials are not only related to the outcome of the copyright litigation faced by Microsoft and OpenAI, but may also affect the future global regulatory direction on artificial intelligence training, intellectual property protection, and the distribution of rights and interests of content producers.

Currently, the relevant cases are still under review. Regardless of the final verdict, this legal dispute surrounding the legality of artificial intelligence training is becoming an important milestone in determining the future development rules of the AI ​​industry.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet