Abstract:
While large model manufacturers are competing to burn money for hundreds of billions of computing power infrastructure, a more secretive "copyright mine" is approaching. In mid-August, news broke that Apple planned to pay hundreds of millions of dollars to purchase news corpus, causing a stir in the content industry. In the past two years, copyright lawsuits surrounding AI training corpus have spread around the world. In the press, the New York Times sued OpenAI and Microsoft, accusing the two companies of using millions of articles without authorization to train large models such as ChatGPT.
The literary community, including two-time Pulitzer Prize winner John Carreiro and six other writers, simultaneously took Anthropic, OpenAI, Google, Meta, xAI and Perplexity to court, accusing them of "deliberate theft".
Previously, according to media reports, Anthropic was accused of abusing its works to train the AI chatbot Claude. After a settlement, Anthropic still had to pay US$1.5 billion in compensation. It is also the largest known settlement in a U.S. copyright case. Domestically, in November last year, the Shanghai Jinshan District People's Court pronounced Shanghai's first copyright infringement case on a large artificial intelligence model in the first instance, involving issues of the right to copy and the right to disseminate information network characters in the "Fights Against the Sphere" series of animations.
How can the "copyright dispute" behind the iteration of large models be resolved? In the face of controversy, how to define the boundary between "fair use" and "infringement" legally? The reporter talked with relevant lawyers, intellectual property professors, publishers and other content practitioners about this.
Is "pay-as-you-go" a fair deal or a pricing power game?
Apple’s move has attracted attention not only because it plans to spend hundreds of millions of dollars to buy news rights, but also because it proposes a “pay-as-you-go” floating remuneration mechanism. That is, the publisher is paid when the partner's content is actually used by Siri.

Some people believe that the “pay-per-use” model proposed by Apple this time is a new paradigm for content licensing in the AI era.
From the legal perspective of copyright protection,
"Pay-as-you-go is not a new paradigm. Apple Pay is still a U.S. royalty generated from the use of copyrighted works by AI. The types of works involved and authorized content are determined in accordance with the current U.S. Copyright Law. Pay-as-you-go itself is one of the common licensing fee pricing models in copyright law."
Li Yanbing, executive director of the Competition Law and Intellectual Property Research Center of Shanghai University of Finance and Economics, told reporters. At the same time, she is also a distinguished researcher at the Chinese Modernization Institute of Shanghai University of Finance and Economics. Her main research directions are intellectual property and competition law, data and artificial intelligence law.Can "pay-as-you-go" completely avoid the risk of infringement?
Fu Gang, senior partner and lawyer at Beijing Dacheng (Shanghai) Law Firm, particularly emphasized: Copyright law focuses on the actual use of works by companies and whether these behaviors are licensed, rather than whether the fee is paid in one lump sum or in installments.
“
Even if the contract stipulates pay-per-view when Siri displays news, if the source’s rights are flawed or the scope of authorization does not fully cover the use scenarios, there may still be risks of infringement or breach of contract.
"Fu Gang said.However, Huang Yangyang, a partner and lawyer at the Shanghai branch of Beijing Zhonglun Wende Law Firm, pointed out that under the framework of the agreement, "Apple has basically avoided the risk of copyright infringement, unless the news it cites itself infringes the copyright of a third party, but generally similar agreements will stipulate terms that exempt the user from liability."
Li Yanbing reminded that pay-as-you-go is only a pricing method for license fees and is a contract clause. It only agrees on the number of times and unit price, and only binds the parties to the contract. It does not involve the qualitative issue of AI's use of works.
Even if it becomes an industry standard, it cannot solve the definition problem of replication and deduction, and also brings new measurement and identification problems. Judging from the public information, the details of such measurement calibers are usually not disclosed.
"For example, a query calls up five news reports from three publishers and generates an abstract. Is this used five times, three times or once? If a technical failure occurs in the middle and the call is made again, but only an abstract is generated in the end, how is this usage calculated? "Usage", is the usage behavior for different links and purposes such as crawling, training, retrieval, citation, output, etc., classified and calculated separately?" Even if it is pay-as-you-go, the current definition is still unclear.
This also means that the choice of technical model will directly affect the collection of license fees.
At the actual operational level, for large AI model companies and content providers, "Which side does the bargaining power lie on?" is a core issue that everyone is concerned about.
Li Yanbing believes that whether it is a fixed annual fee system or a payment based on usage, there is no merit in the two pricing methods.
What really determines the outcome is bargaining power: "Whoever controls the measurement data has the right to interpret pricing."
The difference is that under the fixed licensing (such as fixed annual fee) model, the risk of "whether the content is used" is borne by the platform, while pay-as-you-go transfers this risk to the media or publishers.Generally speaking, "Measurement data is controlled by the platform. Content publishers cannot verify how many times their content has been called, nor do they know how the total amount is distributed. The collective bargaining power is weakened." Li Yanbing said.
Lawyer Huang Yangyang also believes that content parties are at a disadvantage in proving whether the content is infringing. At present, the copyrights of high-quality Chinese content are scattered in the hands of a large number of different rights holders. Publishers, news media and other institutions are mostly independent and scattered. The collective bargaining power of copyright holders is weak. For AI platforms, the authorization process is complicated and the transaction costs are high.
A practitioner from a publishing house revealed, "From the perspective of corpus training, many people in the publishing house are not aware of this problem. They are not very interested in this issue and are relatively passive. This is the actual reaction state of the publishing house."
He told reporters that if someone takes a fancy to your copyright and is willing to pay the fee, everyone will share the benefits and risks, which will do more good than harm to the publishing house. But "what should be" and "what actually is" are two different things, and there is a split between them.
Apple’s purchase of news corpus caused market fluctuations. How to analyze it?
The news that Apple plans to invest up to hundreds of millions of dollars to acquire the copyright of news content has caused quite a shock in the industry. On the day the news came out, the A-share "AI corpus" sector rose sharply. Duke Culture once reached its daily limit of 20cm, followed by CITIC Publishing, China Publishing, Haitian Ruisheng, etc.
西南证券传媒首席分析师苟宇睿认为,这类“海外映射”行情,具备情绪与主题资金轮动特征。
长期来看,AI正推动传媒内容产业,从“消费品"转向“数据+IP资产“双重定价。定价逻辑方面,国内正从“规模竞争“转向“价值和质量竞争“,政策在推动“以词元为基础的数据集价值体系”确权。

Image source: Oriental Fortune’s “AI corpus” concept
The biggest impact on the valuation of the A-share media sector is that "the market has shifted from 'looking at channels and issuance volume' to 'looking at data asset quality'. AI corpus narrative provides a new, growth-oriented valuation narrative." Gou Yurui told reporters. However, he also reminded that the implementation of the Token value system will be more thematically flexible in the short term. Whether it can be settled into a long-term valuation center still needs to be observed whether subsequent authorized transactions and policies can be implemented in real terms.

Image source: Apple's Fiscal Year 2025 10-K Annual Report U.S. Securities and Exchange Commission (SEC)
Going back to the news of Apple purchasing news copyrights, this "active compliance" posture can also be found in its 2025 annual report. In its 2025 annual report, Apple listed AI as a risk factor. Apple believes that complex technologies such as AI may expose users to harmful or inaccurate content, may also involve obtaining and using copyrighted materials for training, and reproducing copyrighted content in model output, resulting in product modifications, litigation costs, licensing fees and regulatory penalties.
“This shows that Apple has regarded AI copyright, content security and product liability as front-end mid- and long-term operating risks.” Lawyer Fu Gang believes that the most direct implication for large domestic model manufacturers is that AI corpus compliance should cover the complete process of data acquisition, cleaning, training, fine-tuning, indexing, retrieval, output and complaint handling.
Spending hundreds of millions of dollars to buy news rights will the content industry be revalued in the future?
"What is truly scarce in the future may not be 'more data', but high-quality data resources that can be used continuously, stably, and compliantly by large models. The data competition of large model manufacturers will gradually shift from early public capture to establishing long-term authorized cooperation with news media, publishing agencies, professional databases and vertical industry content providers, and reducing compliance risks through rights clearance, source marking, usage scope agreements and call audits." Lawyer Fu Gang told a reporter from the Science and Technology Innovation Board Daily.
Gou Yurui also has a similar view. He believes that there are many popular texts, public books, and open web pages and the supply is sufficient. What is really scarce is high-quality, exclusive, compliant, and traceable professional vertical data, such as medical textbooks and cases, legal cases, academic journals, in-depth industry reports, authoritative news, and multi-modal real physical interaction data. High-quality, real-time content is becoming one of the important resources for AI competitions.
What content assets are expected to benefit? Gou Yurui’s summary is: “scarcity + exclusivity + compliance + traceability + real-time”, which mainly includes five categories: professional vertical knowledge (academic/reference book/publishing company), authoritative ancient books/reference book (publishing company), real-time authoritative news (mainstream media), online literature/film and television IP (online literature platform company), and data service/processing company.
At the same time, the revenue structure and valuation system of corpus-related manufacturers may also face reconstruction in the future. For example, for content parties, the ideal path is to add the incremental revenue line of "data licensing/API call sharing/data processing services" in addition to traditional publishing/distribution revenue. Traditional publishing is expected to switch from "dividends alone/PE valuation" to the dual logic of "dividends + data/copyright revaluation". However, direct corpus licensing revenue has not yet been realized on a large scale.
“不过,兑现节奏或许并不会非常快,存在几个约束条件,首先权属复杂、二是工程壁垒:从版权库到"可用数据资产",中间要经历数字化、清洗、标注、脱敏,成本不低,所以“重估”更可能发生在少数专业、独家、稀缺、可溯源的垂类内容持有方,而非全行业。”苟宇睿表示。
IDC data shows that China’s artificial intelligence basic data service market will reach 6.262 billion yuan in 2025, a year-on-year increase of 27.8%, and the growth rate exceeds market expectations. The market size is expected to further grow to 7.834 billion yuan in 2026, with a compound annual growth rate of 19.6% from 2025 to 2030.
According to industry insiders, at present, domestic corpus data is mainly "data localization + private deployment", and large-scale paid authorization transactions have not yet been disclosed. The following three types of practices are in parallel.
1. Self-built/co-developed proprietary models: such as CITIC Publishing’s “Kuafu AI” digital intelligence publishing platform, China Publishing’s promotion of the “Chinese Ancient Books Big Model”, Chinese online self-developed “Xiaoyao” big model, Zhongyuan Media and Tencent’s cooperation on the “Yu Education Big Model”, etc.;
2. Jointly build a national-level genuine corpus: The first batch of 22 institutions jointly built a high-quality artificial intelligence corpus, adhering to "authorization first, use later", covering publishing, media, copyright, technology and other fields; People's Daily Online took the lead in establishing the "Mainstream Value Corpus Ecological Alliance";
3. Professional data service outsourcing.
Is it "fair use" or infringement? Undecided global controversies
Many lawyers told reporters that when using large artificial intelligence model training corpus, the use of the copyright holder's works is "fair use" or infringement. The main legal dispute is whether the training use of artificial intelligence large models constitutes fair use in the sense of copyright law.
Both parties also disagree on whether it is an infringement.
AI companies believe that after learning language rules, knowledge relationships and expression patterns, the model has formed a new technical tool and expression method. There is a strong conversion from corpus output to text output, and the two cannot be equated.
The copyright owner believes that model training is usually for commercial purposes and requires complete and large-scale reproduction of the work. Some data may also come from pirated websites or unauthorized data sets; the model may not only reproduce the original work, but may also generate a large amount of alternative content, weakening the original publishing, subscription and licensing markets. This is "infringement".
“For large model manufacturers, the ‘conversion’ of training cannot be used as a natural defense; for rights holders, infringement cannot be presumed simply because the work is used for training.” This is the judgment of lawyer Fu Gang.
Huang Yangyang also has a similar view. She believes that copying corpus without authorization during the training phase does not mean that the output content must be infringing. Just because the training is legal, it does not mean that the output summary and paraphrased content are accurate.
Regarding the definition of "fair use" and infringement, lawyer Fu Gang believes that at least the following issues should be considered: how the data was obtained, what copying behaviors were implemented during the training process, whether the model can stably reproduce the protected expression, whether the final product competes with the original work or the business of the right holder, and whether the company has taken measures to filter, remove duplication, prevent reproduction, and handle complaints.
When it comes to actually determining whether there is infringement, the situation is more complicated.
"Even in the same AI business, different links may reach different conclusions. For example, the court may believe that the training model itself has certain transformability, but at the same time, it may believe that obtaining and long-term preservation of works from pirated channels does not constitute fair use; it may also believe that there is a lack of evidence of infringement in the training stage, but the model release, content display or final output constitutes infringement." Fu Gang further explained.
On July 20 this year, the "Butts v. Anthropic case" that lasted for nearly two years officially ended with a settlement of US$1.5 billion. The judge believed that Anthropic's use of copyrighted books to train Claude's large language model was "typical transformative use" in nature and constituted fair use. However, Anthropic’s fair use defense of using pirated books downloaded from the Shadow Library for training was rejected and was deemed to constitute infringement.
Lawyer Fu Gang believes that the Bartz v. Anthropic case is of reference significance because it clarified that "the training purpose is transformative" and "the data source is legal" are two independent issues. Just because the back-end use is innovative, it does not automatically eliminate the risk of infringement of data obtained by the front-end.
However, law professor Li Yanbing also reminded that the Bartz v. Anthropic case ended in a settlement of US$1.5 billion, so there was no binding precedent in the determination of the case. The only case currently on appeal, Thomson Reuters v. ROSS, is still pending before the Third Circuit Court.
She pointed out that,
So far, no federal appeals court in the United States has made a substantive judgment on whether AI training constitutes fair use. Regarding the determination of the fair use defense, existing judgments of the U.S. District Court also conflict with each other. There are also disputes over whether and what kind of copyright is infringed. This legal uncertainty is difficult to eliminate in the short term.
With improper copyright protection, will AI be “no rice left in the pot”?
If copyright protection is improper, many interviewees also expressed concerns to reporters about "fishing out of the lake". The more unscrupulously AI devours copyrighted content, the more original supply shrinks; the more original supply shrinks, the more AI can only rely on "copying itself and recycling itself" to survive.
A study released by the Pew Research Center shows that more than one-third of web pages published since the advent of ChatGPT may contain traces of AI generation. The findings suggest that a large number of web pages may be creating a cycle of bots reading bot-generated content.
Lawyer Huang Yangyang reminds, "Knowledge is solidified in the weights of model parameters, and answers are generated based on statistical rules learned through training. Knowledge will not be updated automatically, and it is easy to produce hallucinations and output wrong content."
An anonymous content publishing practitioner revealed a more hidden crisis: "Many upstream authors use a large number of human-machine collaborations to create, and the works they obtain are very AI-like. The production efficiency is getting higher and faster, and their originality issues may cause greater trouble in terms of copyright."
Focusing on the field of news, lawyer Huang Yangyang also believes that the current large-scale AI model training does not pay much attention to the issue of copyright infringement of news reporting content corpus. Moreover, due to the large number of news reports, the high quality of textual expressions, and their widespread publication on the Internet, they are relatively easy to obtain and have become the hardest hit area for disputes over copyright infringement of content corpus.
Huang Yangyang believes that except for a very small amount of simple descriptions of news events, the vast majority of news reports are copyrighted works. The facts of news events themselves are not protected by copyright, but reporters’ written expressions, narrative arrangements, and interview results are protected works. There is still a risk of infringement when AI large model training uses news report content corpus.
She further pointed out: "Compliance with large models requires that copyright, content safety, and product responsibility be evaluated together. We cannot only focus on the effect of the model and treat copyright as a post-issue."
What is the current actual situation? "Most large model companies have not proactively proposed to purchase the copyright holder, and have concealed the source of the training corpus data before the large model products are commercialized, and have performed technical processing on the large model output results to prevent clues and traces from being exposed." Liu Chunquan, a partner and lawyer at Shanghai Duan & Duan Law Firm, told reporters. This also brings identification difficulties to current copyright protection and proof.
In actual cases, some judges have interpreted it from the perspective of "tolerance of innovation" as constituting fair use to encourage the advancement of new technologies.
Lawyer Liu Chunquan believes that artificial intelligence companies scanning and copying works and then applying them to databases to train large models is commercial behavior and should be found to infringe the right to copy the works.
"Corpus training of large models is a purely commercial act that is different from public welfare undertakings such as scientific research. Although corpus training of large models is not like publishing or online dissemination to directly collect commercial benefits from reader sales, the output of large models is based on the corpus training of calling work elements and combining them for output. This process uses works to realize commercial benefits through tokens and other methods, and does not distribute them to the copyright owner. If the court determines that fair use is equivalent to generosity to others (authors) to help large model companies save costs, so fair use lacks rationality."
From the perspective of the balance of interests, lawyer Liu Chunquan believes that the development of new technologies certainly requires the tolerance of the legal system, but the development of artificial intelligence must also be considered. The use of large models has greatly replaced the creative labor of writers and other writers, reducing the market demand for existing works. Although it is a new technology and new business, it is still necessary to consider the protection of the inheritance of human works by the traditional copyright system.
Is there a set of rules that can be implemented for corpus trading?
From the perspective of my country's current Copyright Law, fair use is mainly stipulated in specific situations such as personal study, appropriate quotation, news reporting, teaching and scientific research listed in Article 24. It also requires that the normal use of the work shall not be affected, and the legitimate rights and interests of the right holder shall not be reasonably damaged.
Many lawyers told reporters that my country's current laws do not clearly stipulate an exception for text and data mining that is generally applicable to commercial large model training, but they cannot simply copy the concept of "transformative use" in American law and believe that as long as the work is used for model training, it certainly constitutes fair use.
"There is a high probability that the relevant adjudication rules in the future will continue to be gradually refined in the direction of stages, scenarios, and responsible entities. The focus of judgment has shifted from abstract discussion of "whether AI can learn works" to a full-process review of data sources, training methods, output results, and market impact. The judiciary still needs to make specific judgments based on data sources, purpose of use, scale of replication, technical necessity, output results, and market impact."
At the same time, a series of domestic management measures and supporting policies have also been introduced.
Article 7 of the "Interim Measures for the Management of Generative Artificial Intelligence Services" requires the use of data and basic models from legal sources; if intellectual property rights are involved, the intellectual property rights enjoyed by others according to law must not be infringed.


Image source: "Interim Measures for Generative Artificial Intelligence Service Management"
In November 2025, the national standard "Network Security Technology Basic Security Requirements for Generative Artificial Intelligence Services" was officially implemented. This national standard is my country's first national standard for the security of generative artificial intelligence services.

Article 4.1.3 of the standard states that open source corpora must have a license agreement or authorization document; self-collected corpora must have collection records, and corpora that others have clearly stated cannot be collected should not be collected; commercial corpora must have a legally binding contract and require the provider to issue commitments and certification materials.

Image source: "Network Security Technology Basic Security Requirements for Generative Artificial Intelligence Services"
In May this year, the "National Genuine High-Quality Corpus" launched in Shenzhen by 22 institutions established the principle of "authorization first, use later" and introduced blockchain traces. The corpus is still in the initial stage of co-construction, and the scale of the corpus, licensing price and sharing plan have not been disclosed.
Law professor Li Yanbing said that compliance obligations have been implemented, but the market infrastructure and tradable markets necessary to fulfill the obligations are still under construction, and there are still actual institutional gaps. Specifically, standardized authorization templates for corpus transactions, open and transparent pricing mechanisms, normalized authorization channels between copyright owners and AI companies, and authorization trading platforms that can be reused on a large scale are all in the early stages of exploration.
However, Li Yanbing said that as technology matures and cases increase, judicial adjudication will gradually reach a consensus through typing and legal interpretation methods. "As far as AI's identification of the use of works is concerned, there are no legal structural problems caused by AI. Modern copyright law plus legal doctrine interpretation methods are sufficient to deal with it. This is a normal phenomenon when disruptive technologies appear, and requires normal update and adaptation processes and time."
Comments