Abstract:
American technology companies Microsoft and OpenAI are once again facing new copyright lawsuits. A number of local news organizations, led by Emmerich Newspapers, recently filed a lawsuit in court, accusing the two companies of long-term and systematic unauthorized use of their copyrighted news content to train artificial intelligence models and profit from it in the commercialization process.

The plaintiff stated that Microsoft and OpenAI obtained tens of thousands of news articles without permission or payment of any compensation, and used the content to train large language models. Relevant agencies believe that the two companies ultimately created huge commercial value through these AI products, while the original content providers failed to obtain any revenue.
In fact, similar controversies have continued to arise in recent years. Previously, large media organizations such as the New York Times had filed lawsuits against Microsoft and OpenAI for similar reasons, accusing them of using copyrighted content to train artificial intelligence systems.
According to the complaint, the plaintiff pointed out that Microsoft and OpenAI have created and continue to create billions of dollars in revenue, but content producers "have not received a penny." They also claimed that some content that originally had access restrictions such as paywalls was obtained and used for training without authorization.
The lawsuit documents specifically mentioned that the articles involved were originally embedded with copyright management information, including the author's signature, name of the publishing institution, copyright statement, terms of use, etc. This information is used to clarify ownership and copyright ownership of the work. The plaintiffs allege that Microsoft and OpenAI removed this copyright information before training the model, thereby weakening the copyright identification of the work.
In addition, the plaintiff also claimed that its news content has been output by the AI system in a nearly verbatim repetition form many times in the past few years, indicating that relevant materials have been deeply incorporated into the model training data.
These local publications believe that technology companies’ failure to pay news organizations properly could become the “death knell for local journalism.” In recent years, local media in the United States have experienced challenges such as declining print newspaper sales and shrinking advertising revenue. The ability of artificial intelligence to provide answers directly to users has further exacerbated the operating pressure faced by traditional news organizations.
The main parties involved in the lawsuit are independent newspaper and magazine publishers. Among them, Emmerich Newspaper Group is a large private newspaper group in Mississippi, USA, with its business scope covering parts of Louisiana and Arkansas. In many of the markets it serves, the group is the only local news provider. The plaintiff believes that if similar media are forced to scale down or even close, local residents’ access to news information will be significantly affected.
In addition to Emmerich Newspaper Group, organizations involved in the lawsuit also include Ojai Media, Coopwood Publishing, Coopwood Magazine, Coopwood Media Group and Coopwood Newspapers and other media organizations.
The plaintiff believes that the actions of Microsoft and OpenAI violated the U.S. Copyright Law and the Digital Millennium Copyright Act (DMCA), constituting copyright infringement and related illegal activities.
The complaint also points out that on the one hand, the two companies are accused of neglecting the copyright rights of publishers, and on the other hand, they actively use copyright and legal means to protect their products, including through licensing agreements, paid services, and legal measures to protect their codes, models, and systems from being used at will by others.
The plaintiff specifically mentioned that not long ago, OpenAI publicly complained about competitors using its generated content to train models, believing that this harmed its own interests. The litigants argue that this approach stands in stark contrast to the conduct with which they are accused.
In terms of demands, these local media organizations hope that the court will order Microsoft and OpenAI to pay damages and related compensation, and also request the court to issue an injunction to force the two companies to delete all content involving the plaintiff’s copyrighted works from the model training data.
As more and more news organizations, publishers and content creators launch legal challenges against generative artificial intelligence, disputes surrounding the legality of AI training data, content authorization mechanisms and copyright compensation models continue to become one of the most concerning legal issues in the development of the global artificial intelligence industry.
Comments