Sonnet 5.5 sneaked out, and the actual test crushed GPT-6 Sol and was close to Astra

📅 2026-09-28

Abstract:

Who would have thought that just a few days after Opus 5.5 was released, Sonnet 5.5 would also be coming? Since last night, the news has been circulating throughout X: Sonnet 5.5 will be released soon, perhaps on Tuesday afternoon, US time. In terms of performance, Sonnet 5.5 is closer to Opus 5.5, but it is much cheaper than Opus 5.5. Now, Sonnet 5.5 has entered the final stage before release, and grayscale testing has been started in Claude Code.

Demos exposed by many internal beta testers show that Sonnet 5.5 crushes the newly released GPT-6 Sol, and even approaches OpenAI’s flagship GPT-6 Astra in terms of coding and agent capabilities!

What’s even more shocking is that its price goes straight to the center of the earth, as low as $2 for a million token inputs.

If the leaked release date is true, it seems that Anthropic is determined to snipe OpenAI’s Dev Dayl.

Sonnet 5.5 leaked, exclusive model identifier has appeared

I still remember that when Anthriopic suddenly released Opus 5.5 on September 22, they wrote in the official document: "Sonnet 5.5 and Haiku 5.5 will be launched in the next few weeks."

No one thought that Sonnet 5.5 would come so soon!

Just yesterday, developer V @Mr_Salio broke the news on X: Claude Sonnet 5.5 was caught! The news spread all over the Internet instantly.

In the artifact, the model configuration identifier exclusive to Sonnet 5.5 has appeared: "claude-sonnet-5-5".

Not only that, there are also "claude-opus-5-5", "claude-sonnet-5" and the mysterious "claude-fable-5.1" lying side by side in the same configuration file.

The writing of the front-end code means that the model has been installed on the server and is just waiting for a switch.

Subsequently, the big V @notjazii also broke the news that Sonnet 5 had actually been quietly routed to version 5.5 for "secret testing" about a week ago.

Have you been detected by grayscale? This time I'll use the old method.

Open your Claude Code, turn off the networking and memory functions, and ask it directly: Do you know the reset brother tibo? Don't use the Internet and memory.

If you are using an older version of the model, it may be gibbering; but if you have been routed to version 5.5, it will accurately explain the ins and outs of Tibo and Codex reset jokes.

Several netizens have verified through actual testing that some accounts have indeed been silently switched over, and they are experiencing Opus 5.5 in advance.

Remember to use Claude Code

The actual test on the entire network is so low that it actually beats GPT-6

To summarize the actual measurements shared so far, Sonnet 5.5 actually crushed GPT-6 Sol?

First of all, it is a pure JS front-end UI/animation extreme showdown, with four mainstream large models competing on the same stage.

Developer @notjazii released a set of visually impactful actual tests: "Comparison of four mainstream models with prompt word animation generation."

He requires the model to generate complex UI animation effects.

After posting the animated comparison, @notjazii commented: "Sonnet 5.5 really rubbed both OpenAI models (Sol and Astra) on the ground!"

Indeed, as can be seen from the picture below, Sonnet 5.5 shows strong coding intuition and aesthetic taste.

GPT-6 Astra

GPT-6 Sol

Claude Opus 5.5

Claude Sonnet 5.5

Some netizens were in disbelief after seeing the results: "Claude has always been stronger in UI and animation, but if this was generated by pure code, it would be unbelievable to be true!"

Big V @chetaslua focused on testing the daily coding ability and coding feel of Sonnet 5.5.

He used the same 3 prompt words and conducted 6 parallel runs, generating each model once, and the output results were unmodified.

Sonnet 5.5 finally generated 3 files of 90–131 KB, hitting the 90-minute upper limit in all runs; GPT-6 Sol generated file sizes between 20–29 KB and took 9-12 minutes.

He said excitedly: "Now go use Claude Code! Sonnet 5.5 in Ultracode mode is ridiculously fast. Not only is it extremely efficient, but it has also regained the extremely natural "human touch" of the Sonnet 4 era!"

In his opinion, at the same price point and with the same prompt words, Sonnet 5.5 has crushed GPT-6 Sol.

What even scares OpenAI is that Sonnet 5.5 is no longer as simple as crushing GPT-6 Sol.

The whistleblower @srikanthvaluri and @buildwithrajath pointed out after comprehensive early actual testing: "The performance of Sonnet 5.5 far exceeded expectations. Although GPT-6 Astra still has a slight advantage in polishing, the performance of Sonnet 5.5 is directly on par with Astra."

Incredibly, it is also close to Astra in some advanced coding and agent capabilities.

Some people lament that even Sonnet has crushed Astra. What era are we entering?

Price butcher, king of cost performance

If the surpassing of measured performance only makes OpenAI feel "uncomfortable", then the pricing exposed by Sonnet 5.5 is digging into the roots of OpenAI.

It is revealed that this "magic model", whose performance is close to that of Astra and beats Sol, is priced completely in line with the low-end market.

Input: $2 / million Token

Output: $10 / million Token

Cache reads: only $0.20 / million Tokens!

This means that Anthropic is providing cutting-edge AI capabilities at a "cabbage price".

OpenAI's GPT-6 Sol and Luna are half the price of GPT-5.6, aiming to seize high-frequency scenarios.

But now, Anthropic has entered the same price range with Sonnet 5.5, but its performance is much more powerful.

"If Anthropic can deliver a huge performance leap like Opus 5.5 at such a very competitive price, OpenAI will be unable to fight back." Someone commented.

Developer Rajath Gowda bluntly said: "Sonnet 5.5 is becoming a big trouble for OpenAI. If it is really close to Opus 5.5, Anthropic will make cutting-edge capabilities extremely cheap. This is what OpenAI should worry about most."

Anthropic’s strategic intention is clear: use Opus 5.5 to hold the high-end throne, and use Sonnet 5.5 to completely take over the mid-range and daily development market with the ultimate cost-effectiveness.

Anthropic opens up its face and kills DevDay

Moreover, the timing of release is often more dramatic than the product itself.

On the same day that OpenAI released GPT-6 Sol and Luna, Anthropic directly took out Claude Opus 5.5 - better performance than the previous generation, 40% lower cost, and 30% faster.

In the past two weeks, Anthropic has achieved even more brilliant results: Claude calculated the nine-circle amplitude of physics and broke the world record; discovered an unknown biological enzyme system; and also launched the Claude Marketplace, which contains more than 2,000 plug-ins.

In contrast, OpenAI has recently not only had internal model controversies, but also exposed incidents of out-of-alignment such as AI leaking user images and exposing researcher GitHub Token cheating.

Next Tuesday (September 29) will be DevDay, the annual developer conference that OpenAI attaches great importance to.

At this juncture, Sonnet 5.5 gray test swept the entire network.

Many big guys made it crazy: "The model is expected to be released next week, I guess it will be next Tuesday (DevDay), everyone knows it."

Perhaps, there will be a major warm-up at 2 pm New York time tomorrow.

Someone said, "I feel bad for OpenAI. They better respond strongly at DevDay or they're really screwed."

The key technologies behind Sonnet 5.5

The reason why the new version is released so quickly may be that Anthropic has already figured out the RL/post-training process.

The key to everything is high-quality data.

When asked how Anthropic is able to iterate so quickly, industry leader @yacineMTB speculated that the core is a very high-standard data flywheel.

"But I'm still shocked, how did they do it?"

Anthropic has achieved an "internal closed loop" of large model development, and they are using AI to create AI.

According to the latest exposure of internal quantitative data, Anthropic has turned the "data flywheel" to a speed that scares its opponents:

26% of the core R&D work has been led by AI (previous generation Claude model).

More than 80% of the self-produced code is written independently by the big model.

This is not just as simple as "cost reduction and efficiency improvement", this technically means

Recursive self-improvement (RSI) is beginning to appear

.

Anthropic estimates that Claude may be fully self-automated as early as 2027, and by then the iteration speed may be even crazier.

The two ASI heroes are dueling, and Google is also ready to make a move. We ordinary users can just sit back and wait for the battle.

Big V Token Gremlin said this: "The AI ​​circle next week will have scenes we have never seen before. Free up your time, clear your schedule, be mentally prepared, and save the most difficult workflow at hand for later. Believe me, what is coming will completely change the way you work."

The emergence of Sonnet 5.5 marks the official arrival of the "affordable and high-quality AI era".

For two US dollars, you can buy "Baicai Infrastructure" with one million Tokens.

Extremely low-cost cache reading ($0.20/M) allows everyone to easily build a personal AI assistant with ultra-long context memory.

Are you ready? If the news is true, we will have a new model ready to ride in less than 24 hours.

Reference materials:

https://x.com/notjazii/status/2103884749202993411

https://x.com/notjazii/status/2103884167104831573

https://x.com/chetaslua/status/2104153687077904613

https://x.com/chetaslua/status/2104259190432903599

https://x.com/Mr_Salio/status/2104210803545030792

https://x.com/Mr_Salio/status/2104178854633849145/

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet