AlphaGo core author: Ten years later, LLM still hasn’t learned the essence of the “god’s move” in that chess game

📅 2026-10-05

Abstract:

March 2016, Four Seasons Hotel Seoul. Midway through the second game of the human-machine battle, AlphaGo dropped a sunspot on the fifth line. The on-site commentary was at a loss for words, and some people suspected that there was a bug in the program, because according to the common sense of professional chess players, this move was almost a free throw.


Everyone knows the subsequent story. AlphaGo not only won that game, but also defeated Lee Sedol with a total score of 4:1. After the game, Lee Sedol said something that is often quoted: "I originally thought AlphaGo was just a machine based on probability calculation. But after seeing that hand, I changed my mind. AlphaGo must be creative."

That's the famous move 37.

Ten years later, the man who built AlphaGo stood up and said: today's AI does not have the capabilities that that game of chess relies on.


The person who said this is Thore Graepel, one of the core authors of AlphaGo and has been engaged in machine learning research for a long time. Recently, he made an intriguing decision - to leave Google DeepMind to start a business to do something he thinks is more important: to equip machines with real reasoning capabilities.


That poignant point comes from an article he recently published in MIT Technology Review. The title of the article is very straightforward, called "Don't be fooled - LLMs can't reason." After reading this article, Turing Award winner Yann LeCun expressed his appreciation, emphasizing that "real reasoning must involve search, and LLM does not have this ability."


Why did Thore Graepel say that? Let’s take a look at what is written in this article.

Move 37 was not a “flash of inspiration”,

It is the product of reasoning

Many people describe the 37th move as a "mythical moment of machine intuition" - in a flash of lightning, the AI ​​saw a masterful move beyond the thousand-year history of human chess. Graepel said this is the biggest misunderstanding about AlphaGo.


To understand this, we must first dismantle the internal structure of AlphaGo. It consists of two systems: one is a policy network, which learns from massive human chess records to predict "how the master will play in this situation." This is the intuitive part. The other is a search mechanism, which does not care whether a certain chess move "looks like it was played by a human", but explicitly builds a game tree containing thousands of branches to deduce the future that each candidate position leads to.

The key point is: in the eyes of the strategy network, move 37 is unremarkable - the probability of a professional human chess player playing this move is only about 1 in 10,000. If you only listen to your intuition, this hand will not be chosen at all. It was the search mechanism that discovered after careful calculation that this "unbelievable" move led to a better ending.

Graepel borrows Kahneman's famous framework to explain this. Human thinking is divided into two modes: System 1 is fast, intuitive, and effortless; System 2 is slow, step-by-step, and carefully weighed. AlphaGo happens to be a rare combination of these two modes on the machine: the neural network provides the intuition of "this move has a chance" and "this situation is to be won", and the search mechanism is responsible for testing these intuitions in future offensive and defensive changes. It won't work without any half: intuition alone will never make the 37th move; violent search alone won't be able to screen out the astronomical possibilities of Go.

There is another comparison here that is often overlooked. Deep Blue defeated Kasparov in 1997 by relying on rules hard-coded by humans to evaluate 200 million chess positions per second and look six to eight steps ahead. The complexity of Go is on another level entirely. The value of a piece depends on the growth and decline of the momentum and the field after dozens of rounds. Even if only a small part of all changes are calculated, a supercomputer would have to calculate it over billions of years. Therefore, AlphaGo must learn to "see the situation clearly at a glance" and even create moves that humans have never thought of. It is not won by crushing computing power, which is crucial.

Large language model:

A system trained to the extreme 1

Let’s look at today’s AI.

The working principle of large language models is to predict the next token over and over again. This is essentially System 1 in action: fast, associative, and capable of stunning pattern completion in almost every field humans have ever written about, but without the "think before you answer" part.

Not long after ChatGPT was born, the industry realized that language fluency alone could not support real usefulness. Everyone is familiar with the remedy plan: let the model not rush to answer, first generate intermediate steps, break the problem apart, and bring the intermediate results along - that is, the chain of thought.

The improvement of the thinking chain is real, especially in mathematics and coding. But Graepel pointed out a fact that is easily obscured by excitement: these intermediate reasoning steps are still the same process of "predicting the next token", it just iterates a little longer before giving the final answer. It does not introduce a truly independent reasoning mechanism like AlphaGo's search. Intuition is stretched, but it does not turn into prudence.

He further listed three shortcomings to explain why what the chatbot does falls short of "the kind of reasoning recognized by scientists."

First, the model does not have a clear, continuously updated, and inspectable cognitive state. What hypotheses it is considering at the moment, how much confidence it places in various explanations, what evidence it is weighing, and what questions remain unanswered. These things should be systematically revised as new information comes in, but there is no such "open ledger" within the model.

Second, the model does not achieve the separation between "what to know" and "how to use knowledge". Knowledge and reasoning are tightly entangled in the weights of the neural network, and there is no independent and explicitly expressed belief system.

Third, and most subtle: Thought chains look like thoughtful deliberation, but research has shown that models often make them up after the fact. The answer is obtained by walking one way, but what is reported to you is another way. A "story that sounds reasonable" and a "reasoning that really happened" are two different things.

Whether the reasoning is true or false,

Is it that important?

If you are just chatting and chatting, it may not matter whether the reasoning is true or false. But in the high-stakes fields we really care about — medicine, engineering, scientific research — what a system reaches is just as important as how it reaches it. When something goes wrong, such as an error in diagnosis or treatment, we need to be able to pinpoint where the error lies: is there a problem with the reasoning process, is it accepting invalid evidence, or making a wrong assumption? A system that can only make up stories after the fact cannot be trusted or held accountable in such scenarios.

This is exactly why Graepel left DeepMind. He believes that we need a new machine reasoning path, and the inspiration comes from AlphaGo ten years ago.

AlphaGo maintains a record of all its knowledge of the current situation at any time, which is the game tree. The tree contains all the changes it has considered, all possible futures, and every move and every situation is marked with the judgment of the neural network. As the deduction progresses, it continuously updates the tree, and finally determines the move based on the information in the tree.

Graepel believes that the same is true for general reasoning systems: maintaining a cognitive state that clearly expresses what is confirmed, what is in doubt, what has been ruled out, and what questions remain open. And "reasoning" is a series of "actions" that change this cognitive state - deriving conclusions, dismantling problems, and most importantly: deciding what questions to ask next, what calculations to do, and what experiments to run.

Of course, open-world reasoning is much harder than playing chess. The state on the Go board is completely visible and the rules are fixed; in the real world, the current situation is only partially known, the available actions are numerous and complex, and the consequences of actions are full of randomness or even completely unknown.

But the good news is that the progress of LLM and other neural models over the years has just provided parts for this: LLM can propose a path to solve the problem based on known information and available resources; it can interact with tools through APIs and codes; it can help evaluate whether a conclusion is supported by existing evidence. Most importantly, there must be an independent "referee" in the system who evaluates each reasoning move based on "how much uncertainty is resolved by this step" - and only updates beliefs when supported by evidence. Once the rules are established, the system can accumulate certified knowledge, learn from past reasoning experiences, and improve its own reasoning strategies.

Graepel gave this picture a vivid expression: the scientific method on steroids. There is only one goal, to produce knowledge that can withstand scrutiny.

Larger System 1,

System 2 will not automatically grow

At the end of the article, he made his attitude clear: Trustworthy machine intelligence cannot be achieved by making System 1 bigger. Scale sharpens intuition, but not prudence.

The reason why move 37 is important is because a machine defended a position, weighed the possible futures, and then selected a move that even its own "gut feeling" would have probably rejected.

Human society needs more such "hands of God" in the fields of drug discovery, new materials, climate, and medical diagnosis. The chessboard there looked nothing like Go, and no one even handed us the rules. Such insights can only come from systems that truly reason: conclusions drawn from a chain of auditable evidence, inferences, and belief revisions, rather than from a convincing story concocted after the fact.

Ten years ago, AlphaGo used its 37th hand to tell the world that machines can surpass human intuition.

Ten years later, one of its creators is reminding us: Don’t let the smooth language fool you. That real leap has not happened again to LLM so far.

You said LLM cannot reason,

Is it tenable?

After this opinion article was published, it caused quite a controversy on X.

The most trenchant question came from University of Toronto professor David Duvenaud (one of the authors of the NeurIPS best paper on neural ODE). He mentioned: Graepel’s three charges against LLM – no checkable cognitive state, no separation of knowledge and reasoning, and fabrication of explanations after the fact – are equally valid for humans. "According to this standard, can I be considered good at reasoning?"


Graepel responded: You can't reason well on your own; if you are given paper and pen, it will be much better; and if you are asked to strictly follow the scientific method, you will be considered an excellent reasoner. The trick is never to follow the stream of consciousness, but to structure the process – otherwise no one can escape cognitive biases.


Duvenaud immediately caught the implication of this sentence: "It sounds like we can achieve real reasoning according to your definition as long as we equip LLM with similar tools - isn't this what you intend to do?"


Graepel's "paper-and-pencil enhancement theory" also attracted doubts from other netizens. Netizen KingBroken pointed out that paper and pen are essentially a plug-in for working memory. It allows you to reason about more data at the same time, but it does not change the existence of reasoning ability "out of nothing"; but it has an additional benefit: the written words and numbers happen to be checkable evidence of human "cognitive status".


Immediately afterwards, netizen Daniel Grant asked a more pointed question: According to this, the human reasoning method is equivalent to "existing model + plug-in system", so are you building a model with stronger internal reasoning capabilities than both? Or am I missing some nuance?


Graepel has not yet responded to these queries.

Another voice felt that this discussion was unnecessary. Netizen visimo-dino said bluntly, "I can't believe there are still people writing this kind of article": LLM has already performed well in too many things, and the statement "cannot reason" is either ridiculous or irrelevant, and hundreds of similar articles have been published.


Graepel didn't dwell on it, just throwing back three questions: Are they reliable in areas where they can't be verified? Has it produced a powerful new scientific theory? Do you dare let it dictate your medical treatment? The other party was frank, admitting "not yet", but blaming the existing reinforcement learning methods rather than the LLM architecture itself - the implication is that it will be solved in time. This is probably the real difference between the two groups: one side thinks that what is lacking is time, and the other side thinks that what is lacking is structure.


Another comment asked what many people are thinking: Why doesn’t DeepMind support this research? Graepel's answer is straightforward: cutting-edge laboratories are betting on scaling, and auditable reasoning cannot emerge from scaling alone. It requires a different architecture. “Sometimes radical new ideas are more difficult to implement within large organizations.” The disappointment of the questioner Robert Sperry was palpable: of all the laboratories, he had expected DeepMind to be the one working hardest to find answers beyond Transformer.


Some people also pointed the finger at a more fundamental place. Commentator Katherine Moore left only one sentence: "Language itself is the wrong tool, and reliable reasoning cannot be made with it." Graepel took this more radical position seriously and raised an open question: Reasoning does require some kind of knowledge representation. If natural language is too vague, can formal language be used to reason about the real world? What would such a language look like? . He did not give an answer, but this may be the question that some people will answer when starting their own businesses.


Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet