Was the millennium problem that Wei Shen focused on solved by Claude?

📅 2026-09-06

Abstract:

It turns out that Qiu Chengtong, the first Chinese Fields Medal winner and professor at Tsinghua University, is also a New Wisdom reader! Yesterday, we reported that the proof of Fermat’s Last Theorem was formalized for the first time by Claude, and the key to this was Tsinghua Yao Class alumnus Peng Tianyi’s Prove2Me system based on Claude. To our surprise, we noticed a

related screenshot circulating on the Internet:

Mr. Yau Shing-tung is offering a reward of NT$100,000 for formal verification of Perelman’s Poincare Conjecture proof!


The Poincaré conjecture is a famous mathematical theorem about the shape of three-dimensional space in topology. It states that if in a closed three-dimensional space, every closed curve can shrink to a point, the space must be a three-dimensional sphere. In 2006, the mathematical community generally recognized that Perelman had completed the proof of the Poincaré conjecture.

At present, we have not yet completely confirmed the authenticity of the screenshots (it is said that screenshots are not allowed to be shared outside according to regulations), but since AI can complete the formal proof of Fermat's last theorem, it is very likely that AI can also formalize the proof of Poincaré's conjecture.

Almost at the same time, another Millennium Puzzle, the Navier-Stokes equation, with a $1 million reward, is rumored to have been solved by Anthropic's Claude, and the proof is being sent to experts for review.


This difficult direction is also the main focus of Wei Shen (Wei Dongyi). Only related phased results won him the second prize of the National Natural Science Award.

If the rumors are true, the mathematical reasoning capabilities of Claude’s yet-to-be-released internal flagship model will far exceed all public models, including GPT-6 Astra.

Tao Zhexuan responded that although the authenticity of the rumors cannot be confirmed at present, he rigorously said that it is "not completely impossible" for AI to overcome such problems.


But he then spent thousands of words deducing a question: What would happen to mathematics if AI really solved such difficult problems in a black box way.

The conclusion is beyond the expectations of many people: it may do more harm than good.

Navier-Stokes equation supports the engineering world

But it has stumped the mathematical community for nearly two hundred years

The main communication node of the rumor is the "prediction" thrown out by technology blogger Andrew Curran on X: Claude solved the N-S equation,

Anthropic plans to announce before IPO

.


https://x.com/AndrewCurran_/status/2096062392442724805

Subsequently, multiple sources crossed over: OpenAI employees relayed the news to the outside world, and mathematician Elliot Glazer followed the rumor chain and found that the initial version claimed that Anthropic solved "two" Millennium Problems, which was later corrected to one;

The whistleblower "Brother Strawberry" @iruletheworldmo said that he has conquered "multiple";


https://x.com/iruletheworldmo/status/2096262005149548951

Some people estimate that if true, the model capability will exceed 170 ECI.


Epoch AI latest release: GPT-6 Astra’s ECI is 169

Another mathematics professor told Glazer that the version he heard from a reliable source was that "the Hodge conjecture was solved."

However, all sources can be traced back to reports from OpenAI employees, and there is no independent primary source.


https://x.com/ElliotGlazer/status/2096298696438906934

As of September 6, the Clay Mathematics Institute still lists the N-S equation as an unsolved problem, with the $1 million prize unclaimed.

Don’t kill the chicken that lays the golden eggs

Tao Zhexuan's core point of view can be summarized as the following chain of reasoning:

The value of the famous conjecture lies in the tools generated by the solution process, and the answer itself is secondary →

If AI gives answers in a black box manner, all insights gained in the process will be locked →

The problem is technically "solved" but the mathematics stops growing.

Take the N-S equation as an example: Over the past few decades, mathematicians have spawned a series of foundational tools such as Leray-Hopf weak solution theory, Gagliardo-Nirenberg-Ladyzhenskaya inequality, and Beale-Kato-Majda blasting criterion in the solution process, which constitute the infrastructure of modern partial differential equations.

Fermat's Last Theorem illustrates the same problem: more than three hundred years of verification have given rise to modular form theory and elliptic curve theory, which continue to play a role in cryptography and coding theory.

These conjectures are like the chicken that lays the golden egg. AI violently proves that it takes away the egg, but may kill the chicken.

Even new flagship models, including GPT-6 Astra, have begun to use Current Depth to replace Transformer as a reasoning method. The cost is that the thinking process enters a black box, and we cannot obtain the scaffolding to climb scientific peaks in a clear chain of thinking.

Tao Zhexuan used the prime number spacing problem to construct a counterfactual thought experiment to support: On the real timeline, Zhang Yitang's breakthrough in 2013 (the upper bound was 70 million) triggered the birth of textbook-level tools such as Polymath8 collaboration and the Maynard sieve method, and the upper bound dropped all the way to 246.

But if in 2005 there had been an AI company pursuing a brute force solution, the Maynard sieve method might have been buried in hundreds of pages of proofs generated by AI and never been discovered. Zhang Yitang continued to work in a sandwich shop, and the entire field was "emptied out."

A graduate student’s comment under the post condensed this sentiment: “When chatting with classmates, a very common feeling is that everything I do is no longer important.”

However, Terence Tao clearly gave a positive case.

He cited Erdos problem #1026: AI used for collaboration, both solving problems and advancing human understanding.

What he criticizes are some deliberate choices: not cooperating with experts in the field, not writing papers for peer review, just to achieve a benchmark score.

Anthropic and OpenAI use cutting-edge mathematics as benchmarks and large-scale PR to endorse their IPO valuations of RMB 10 trillion.

Mortals will eventually be left behind

But Terence Tao himself admitted one thing: even if the N-S equation is solved in an ideal way, the final structure will be "extremely complex, and it is impossible for humans to complete the verification on their own." The proof document formalized in Lean may become the largest of its kind.

This judgment has been fully confirmed elsewhere.

On September 4, Anthropic announced that Claude completed the Lean formalization of Fermat’s Last Theorem in 11 days: 13 million lines of code, 29,500 intermediate theorems, and the output of 6 billion Tokens.

Reviewer Kevin Buzzard of Imperial College London commented that this was a "remarkable achievement in automatic formalization", with the proof process "relying on no additional assumptions other than mathematical axioms".

But no human can read the 13 million lines of Lean code line by line, and human verification has degenerated into: running the compiler and seeing if it reports an error.

This is the beginning of ASI.

AI’s mathematical output is exceeding the limits of human understanding.

Currently humans can also design heuristic solutions, choose strategies, and determine which paths are worth exploring.

But when the reasoning ability of the model continues to rise, and even top mathematicians cannot understand its intermediate steps, the model of "human beings accumulating intuition in the process of exploration" cherished by Terence Tao will no longer be possible in a physical sense.

Mortals with limited intelligence will eventually be left far behind by ASI.

Well said, no one is going to stop

Anthropic and OpenAI are clearly not going to slow down due to Terence Tao’s warning.

The AI ​​mathematical arms race in 2026 is already in full swing: In August, an unreleased research model from Anthropic tried to prove the Riemann Hypothesis, but failed. However, in the process, it pushed the lower bound of the proportion of the zero point of the Riemann zeta function that satisfies the conjecture from 41.6% to 67.2%, breaking the best record maintained by humans.

The most outrageous detail is: Jarred Sumner, the Anthropic employee who drove this breakthrough, had no mathematical background. He just kept prompting the model to "keep trying."

Also in August, OpenAI announced that Astra had solved 10 outstanding mathematical and theoretical computer science problems for at least ten years, including the first construction of a non-sofic group (a problem that has remained unsolved since Gromov proposed it in 1999), all with Lean 4 formal proofs and a total computing power cost of about $2,000.


https://openai.com/zh-Hans-CN/index/ten-advances-in-mathematics/

Google DeepMind is also moving forward.

In May, AlphaProof Nexus solved 9 of 353 Erdos open problems and proved 44 of 492 OEIS conjectures, all with Lean machine verification.

In the same month, OpenAI overturned the unit distance conjecture proposed by Erdos in 1946. Fields Medal winner Tim Gowers said that he would "without hesitation recommend this paper to be published in Annals of Mathematics (the most authoritative journal in mathematics)."

On the prediction market Manifold, the contract for "2026 AI Solving the Millennium Prize Puzzle" is currently quoted at 38%.

Rumors, evidence and bets are all piling up in the same direction.

The puzzle with a $1 million reward can be solved, but if it is solved in a way that stops mathematics from growing, then the thing that $1 million buys may not be worth the price.

Goodhart's Law

It has been proven again

: Once an indicator becomes the only goal, it loses its meaning in measuring value.

But to support an IPO valuation of more than a trillion dollars, Anthropic and OpenAI cannot just prove that they are "safe", they must prove that they have "achieved AI that comprehensively surpasses human levels", that is, ASI.

Is there anything that can give a company a greater sense of "miracle" than solving pure mathematical problems that have troubled mankind for hundreds of years (Fermat's last theorem, N-S equation, Poincaré's conjecture, etc.)?

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet