Gemini 4 Pro is suspected of leaking, and "AI slowdown" is empty talk again

📅 2026-09-19

Abstract:

In the past few months, OpenAI and Anthropic took turns pushing their flagship models forward, but Google seemed unusually quiet. Flash is updated almost every three weeks, with 3.6, 3.7, and 3.8 moving forward continuously, but there has been no movement for Pro, which represents the highest capability limit. Until these two days, a model with the name gemini-3.8-flash suddenly appeared in the large model blind test arena Arena.


As soon as developers got started, they noticed something was wrong. The real Gemini 3.8 Flash has been released. Everyone knows what level it should be. But with this "3.8 Flash", the performance of writing code, making SVG, and running Agent has obviously jumped up a notch.



After several rounds of actual testing, a guess quickly spread:

This model is likely to be Google's next-generation flagship hidden behind the label - the Gemini 4 Pro.

Immediately afterwards, an even more exaggerated benchmark comparison picture began to go viral. Coding, Agent, Reasoning, multiple scores ranked GPT-6 Astra and Claude Fable below.


Just a few days ago, several of the most important AI companies were openly discussing whether it was time to slow down. Now, the sound of deceleration has not yet reached the ground, and a new round of competition has begun.

Behind this suspected "god-level model" of Gemini 4 Pro, there is a more dangerous keyword:

RSI.


Gemini 4 Pro suspected to be leaked

On September 2, Google has just officially released Gemini 3.8 Flash. According to officials, this is the latest generation of Flash in the Gemini 3 series, which has significantly improved software engineering, agent tasks and complex multi-step reasoning; from 3.7 Flash to 3.8 Flash, only three weeks passed.

So, when another model with the same name gemini-3.8-flash appeared in Arena, the developers quickly noticed the anomaly.

The real 3.8 Flash is already there, and everyone has a ready reference. And this model looks obviously stronger, and it is clearly the strength that can only be found in the Pro model.


The first batch of actual tests that triggered dissemination focused on SVG, web pages and 3D scenes.

Developer Harshith asked him to draw a sideways house cat using SVG, and put the results together with GPT-6 Astra Max and the official version of Gemini 3.8 Flash. This version in Arena has a visible gap with the official 3.8 Flash in terms of outline, proportion and completeness of details.


From left to right: "4 Pro", Astra, 3.8 Flash

X netizen @thtbee_ continuously tested this suspected Gemini 4 Pro model from multiple angles.

A Voxel style pagoda, the model ran for 8 minutes and directly generated a complete set of interactive 3D scenes.



An Airbus H145 helicopter, it only took about 10 minutes to build a complete 3D display page, which is very detailed. However, this netizen said that the UI elements at the bottom of the page "look a bit bad."



In terms of web design, it took the model about 14 minutes to create a monochrome theme website, which was very clean and complete overall. The design taste issues that Gemini was often complained about in the past seem to have finally been fixed this time.



Another netizen @Mr_Salio directly used this suspected Gemini 4 Pro model to make games. The finished graphics are smooth enough, and the world generation and game design are also very complete.



A more critical clue comes from developer Qwinah.

He claimed that he successfully "ghost routed" to the backend of Gemini 4 Pro and posted a screenshot of suspected internal terminal information.

According to him, G4P-ARGON and VIA_3.8_FLASH appear directly in the internal logo of this model. The maximum output reaches 256K, the context window exceeds 10 million tokens, and it also has cross-session permanent memory, networking without API, back-end automatic compilation and running scripts, and even physical robot control capabilities.


If this information is true, then it is obviously no longer an ordinary Flash upgrade.

At the same time, a Gemini 4 Pro benchmark comparison table also began to circulate in the community.


According to the data in the picture, Gemini 4 Pro scored 88.7% in the DeepSWE v1.1 software engineering test, 95.3% in the Terminal-Bench 2.1 test on the terminal agent, 72.1% in the expert-level knowledge reasoning test HLE-Verified, and 86.8% in the computer operation agent test OSWorld 2.0.

What is even more exaggerated is that on GDPval-AA v2, which measures real knowledge work, it is written as 2064 Elo, directly breaking the 2000 mark.

If these data are true, Gemini 4 Pro can simply "punch Fable and kick Astra", standing on the top of cutting-edge models in one fell swoop.

But the more exaggerated, the more careful you need to be.

It is worth noting that this benchmark table is not completely consistent with the suspected backend screenshot "digged" by Qwinah; the score of Astra, which is used as a comparison item in the benchmark table, does not conform to the official data; both materials are not officially endorsed by Google, Arena or related benchmarks, and are still only information circulated in the community.

The only model that can be confirmed is the model named gemini-3.8-flash, which is completely inconsistent with the performance of Gemini 3.8 Flash.

But there is another very important reason why everyone quickly points to the Gemini 4 Pro: Google’s Pro model has been empty for too long.

Gemini 3.5 Pro has not yet been officially launched, and Gemini 4 has already been confirmed to enter training. In the past few months, Flash has been pushed forward from generation to generation, and there has never been a new model in the Pro line that truly represents the upper limit of capabilities.

Now, a "3.8 Flash" with obviously abnormal abilities suddenly emerged from the Arena.

As a result, the speculation is naturally becoming more and more reasonable - Google may not have any intention of leaving the real big upgrade to 3.5 Pro, but directly to Gemini 4 Pro.


Behind Gemini 4 Pro,

The bigger keyword is RSI

If Gemini 4 Pro is still just a community rumor, then what is more noteworthy is that

Google is significantly accelerating the speed of AI participation in model development.

In this Gemini 4 Pro rumor, the keyword RSI has been mentioned repeatedly.



RSI stands for Recursive Self-Improvement, which means recursive self-improvement. Simply put, AI begins to participate in making stronger AI, and stronger AI further accelerates the development of next-generation models.

Once this cycle is running, it is not just the model ability itself that is accelerated.

Even the speed of model evolution will begin to be accelerated by the model itself.

Reuters disclosed in August this year that Google co-founder Sergey Brin has been pushing Gemini to speed up in recent months and has listed recursive self-improvement as one of the key directions.

Although Google has not publicly announced that it has achieved this closed loop, it is handing over more and more tasks that originally belonged to researchers to Agents, including evaluating models, finding directions for improvement, running experiments, and then feeding the results back to the model development process.

When Gemini 3.8 Flash was released on September 2, Google also mentioned in the official description that a long-running Agent loop is "recursively evaluating and improving the underlying model."


Google DeepMind researcher Yao Shunyu (not Tencent's Yao Shunyu) used the words of Armstrong when he landed on the moon to describe this update after the release of 3.8 Flash:

"That's one small step for the model, one giant step for RSI."


Just recently, Google also included "RSI" in the title of a new paper.

On September 14, Google, GDM and other teams released Dream-RSI, which specifically studies how to enable agents to achieve recursive self-improvement through continuous improvement of exploration strategies. In algorithm engineering, mathematical optimization and GPU kernel tasks, this method shows higher exploration efficiency and lower search cost.

It is precisely because of this that the rumors of Gemini 4 Pro were quickly tied to RSI.

The speculation in the community does not stop at "Gemini 4 Pro is very strong", but also asks to what extent Google has achieved RSI internally.

Google is obviously not the only one taking this path of RSI.

On September 17, US time, Anthropic disclosed a set of internal AI R&D automation data.

Anthropic officials stated that the reason they disclosed these indicators is that they believe that the outside world needs to know how much of the AI ​​research and development in cutting-edge laboratories has begun to be completed by AI itself.


Data shows that as of August this year, Claude has been able to lead 26% of Anthropic AI research and development tasks - that is, humans only need to give high-level goals and supervise, and Claude can basically complete the tasks end-to-end. This proportion was less than 1% in February and March this year.

At the same time, more than 90% of the AI ​​R&D tasks within Anthropic have at least entered the stage of AI collaboration.

In China, Tang Jie, founder and chief scientist of Zhipu, also recently said: "We are still far from reaching recursive self-improvement, but the smallest cycle has appeared: model optimization system, system service model."


Tang Jie’s long article on X, only the beginning is truncated

In the engineering practice disclosed by Zhipu, an Infra Agent driven by GLM-5.3 participated in the design, debugging and optimization of the GLM-5.3-Flash inference infrastructure. This system runs on a cluster of more than 100,000 domestic AI chips. It only took two weeks from the first run-through to taking over all production traffic, increasing the end-to-end throughput to 3.2 times the initial level.


"Slow down" finally leaves only a warning

Just a few days ago, Amodei called on cutting-edge AI companies to join hands in "slowing down."

The first condition he listed to be alert to was RSI.

Model progress in itself is of course a good thing, but the problem is that if AI starts to participate in manufacturing the next generation of AI, the speed of model progress may be faster and faster, but the speed of human understanding, evaluation and control of it may not be accelerated at the same time.

So Amodei has repeatedly emphasized that third-party evaluation, common safety standards and deceleration mechanisms must be established in advance before capabilities cross certain thresholds.

Because once RSI truly forms a positive feedback, the model may have crossed the "safe distance" of braking.

In terms of risk, several cutting-edge laboratories have indeed reached a consensus.

Altman publicly agreed with "slowing down the pace of development of cutting-edge AI" and promised that OpenAI would also introduce independent reviewers with similar employee rights; Musk said that "Dario is right"; Hassabis also said that this is the right direction, but how to do it needs to continue to be discussed.

But the initiative also mentioned that there is no point in slowing down alone. If AI research and development continues to be accelerated by AI, a truly effective slowdown or suspension in the future requires the establishment of a set of common rules: such as the introduction of independent evaluators, and allowing cutting-edge laboratories to jointly set capability thresholds and safety measures, and coordinate to slow down the pace of research and development when necessary.

But until now, for various competitive reasons, common rules have not been established.

Everyone can already say together that "we should be cautious", but when it comes to "specifically how to slow down, who should slow down first, and whether other countries should slow down", the consensus quickly disappears.

There is the most direct model competition between laboratories, and there is greater AI competition between countries. If any country slows down alone, it will bear the risk of giving up the leading window to its opponents; if any country imposes restrictions alone, it will worry about the other side continuing to accelerate.

Several cutting-edge AI laboratories, therefore, behave somewhat inconsistently with their words and deeds.

On Google's side, there is the Gemini 4 Pro mentioned above that is suspected to have been anonymously tested in Arena.

On the Anthropic side, grayscale traces of Opus 5.2 have also begun to appear in the community recently. Some Claude Code users said that they saw the new claude-opus-5-2 model logo through the running status and request routing information, and suspected that some requests have been grayscaled to the next generation Opus.


As for OpenAI, some people said that they saw the model name of gpt-6-sol in the background model list, and some people said that they suspected that they had been grayscaled to a new model. According to the widely circulated news, GPT 6 Sol may be released next week, or it may be on Dev Day on September 29th, US time... In short, it is not far away.


The next round of models from the three cutting-edge AI laboratories have begun to show traces of testing in the community.

The most certain result of this round of “deceleration initiative” so far is not that any company has really slowed down, but that the most important AI companies have publicly talked about the risks in advance.

The model isn't slowing down.

But at least, if something happens one day, they can say: We have warned them.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet