Abstract:
With OpenAI's high-profile announcement that its internal artificial intelligence system has successfully overcome one of the "Millennium Prize problems" that has plagued the mathematics community for nearly a century - the existence and smoothness of the Navier-Stokes equations (Navier-Stokes equations), a huge controversy surrounding the source of underlying training data for cutting-edge models, the ownership of scientific research intellectual property rights, and academic ethics quickly detonated in the global academic community and technology industry.
As artificial intelligence gradually invades the forefront of basic scientific research, where do the high-order mathematical reasoning models of giants such as OpenAI draw key nutrients, and whether scientific researchers inadvertently become "data fertilizer" for the private models of technology giants when they use commercial AI tools on a daily basis.

The controversy originated from a 165-page mathematical proof and Lean formal code certificate released by OpenAI. According to OpenAI's disclosure, it deployed an undisclosed internal model that is more powerful than the newly released GPT-6 Astra, mobilized about 10,000 autonomous AI agents capable of collaborative communication, consumed hundreds of billions of output tokens and supercomputing resources worth approximately US$22.5 million in 88 hours, and finally derived a mathematical proof that the incompressible fluid equation produces a singularity in a limited time. OpenAI also made it clear that it will not receive the $1 million bonus established by the Clay Mathematics Institute. However, just before the release of the OpenAI results, Tristan Buckmaster, a professor at the Courant Institute of Mathematical Sciences at New York University, and Levent Alpöge, a researcher at competitor Anthropic, had just disclosed their breakthrough precursor results on related fluid dynamics problems such as the forced Euler equation.
Buckmaster immediately issued a public statement, raising strong doubts about OpenAI’s independent breakthrough. Buckmaster pointed out that during the year-long collaborative research, he and Alperger used OpenAI's code generation tool Codex and Anthropic's Claude model in depth to deduce formulas and draft intermediate drafts. All unpublished core derivation manuscripts, subdivided mathematical entry paths, and specific boundary condition settings were completely input into the Codex session log. Buckmaster revealed that after learning that OpenAI had learned about relevant external research rumors, he directly asked OpenAI whether the internal problem-solving model had been exposed to or trained from their private interaction data on Codex; although OpenAI technical staff insisted that the model did not directly access specific user data in real time, they avoided answering the key question of "whether it had been incorporated into the model pre-training or fine-tuning data set." Even more controversially, Buckmaster also accused OpenAI core researcher Sébastien Bubeck of having set out conditions for co-authorship in private communications, but requested that Alperger, an Anthropic employee, be completely excluded from co-authorship.
Faced with growing doubts about academic plagiarism and data misappropriation, OpenAI officially issued a statement denying that it had snooped on the manuscripts of the two people through any informal channels before public release. Bubeck also emphasized that the mathematical proof derived by OpenAI is essentially different from the Buckmaster team in terms of technical structure and specific conclusions. However, OpenAI rarely included an intriguing disclaimer in its official response: Officials "cannot rule out the possibility that de-identified anonymous data generated from users' use of its products has objectively helped improve model performance." This statement instantly caused an uproar in the academic circle, and was interpreted by many scholars as a disguised admission by OpenAI that the private interaction records of top scholars stored in commercial developer tools have most likely been involved in the continuous iteration pipeline of its underlying model.
This turmoil revealed the hidden reality of "data depletion" and "reverse absorption" faced by cutting-edge artificial intelligence laboratories when building high-order mathematical capabilities. Most of the traditional pre-training data comes from public web crawling, but the vast majority of public mathematics corpus on the Internet is only at the level of primary and secondary school competitions or undergraduate basic textbooks, and cannot support the model to understand profound geometric analysis of partial differential equations or algebraic topology. In order to equip the model with the deep logical reasoning and rigorous deduction capabilities of the world's top mathematicians, the AI laboratory, in addition to hiring professional mathematicians to write formal Lean code and synthetic data, relies heavily on cutting-edge research pain points, unpublished problem-solving hypotheses and advanced prompt word flows input by real top scholars in development tools (such as the Codex platform). Once the user agreement contains a clause that allows the platform to use de-identified interaction data for algorithm optimization, the original brain work of scholars will be quietly digested and amplified by corporate computing power clusters in a legal gray area.
Experts in the philosophy of science and intellectual property law pointed out that the Navier-Stokes dispute was not a squabble between individual scholars and business giants, but rather marked a structural ethical crisis that must be faced when the paradigm of human scientific discovery undergoes a fundamental fission. Over the past hundreds of years, the honor and incentive system of academia has been based on strict peer review, publication priority, and standardized literature citation ethics; however, in the face of industrialized "violent search" that involves tens of thousands of intelligent agents cooperating in parallel and costing tens of millions of dollars, the traditional defense lines of academic discovery are becoming vulnerable. If commercial technology giants not only control the leading computing power hegemony, but also use the productivity software they control to unilaterally absorb the unpublished ideas of cutting-edge scholars and deny substantial attribution, the foundation of the global academic community's trust in open scientific research tools will be completely shaken. This will also force major universities and independent scientific research institutions to begin to re-examine the huge potential risks of transferring core scientific research secrets in unverified cloud AI tools.
Comments