ChatGPT’s new model Bel was revealed to have been pre-trained with parameters as high as 10 trillion

📅 2026-08-31

Abstract:

OpenAI’s most powerful model, Astra, has been revealed to be released next Thursday, and the scope of testing has been expanded. Previously, internal secrets were leaked, and the underlying code of the self-developed chip "Jalapeño" has been officially taken over. The core code written by AI runs 1.8 times faster than top human engineers. There are signs of AI self-recursion within OpenAI!

OpenAI also has a depth charge.

It is rumored that OpenAI has successfully run through the 10 trillion parameter pre-training model code-named "Bel", pointing directly to the ultimate threshold of AGI. Pre-training that even exceeds 10 trillion parameters is just the starting point for Bel. After that, Bel can learn at two speeds.

Ultraman has said: We should have another party for the next generation model release.


What ambitions are hidden in the world after GPT-6?

Bel: The abyssal beast with 10 trillion parameters

This year, OpenAI seems to be stuck in a bottleneck of scale expansion, and even urgently sounded the alarm "Red Code".


In order to concentrate computing power, OpenAI even axed products such as Sora and AI browser Altas.


The next-generation AI model Astra has achieved multiple mathematical breakthroughs in one fell swoop and amazed the world, but it has yet to be released.

In recent weeks, OpenAI has experienced personnel turmoil: chief revenue officer Dennis Dresser, chief operating officer Brad Lightcap, and Fergie Simo, a former deputy to CEO Sam Altman, have resigned.

While the outside world is speculating whether OpenAI has "run out of talent", a number of hard-core technology whistleblowers have revealed that OpenAI has just completed a very large-scale pre-training, codenamed "Bel".


How scary is this "Bel"?

Let’s look at a few core keywords:

1. Breaking through the 10T (10 trillion) parameter mark

The trillion-parameter GPT-4 allows AI to have common sense and logic close to human undergraduates. So what kind of "emergent capabilities" will Bel with 10 trillion parameters show in terms of complex reasoning, long text association, and interdisciplinary multi-modal understanding?

This is a qualitative change.

It is equivalent to packaging the brain capacity of all the top experts in the world today and multiplying it by an exponential amplifier. At this parameter level, AI's understanding of the world model will reach an unprecedented depth.


2. The successor of “Doug”, the ultimate base after GPT-6

The news pointed out that before this, OpenAI had completed pre-training code-named "Doug".

Doug is positioned as the basic model of Project Astra and the rumored GPT-6 (which will undergo extremely high-intensity reinforcement learning alignment later).

As the successor, Bel is the next-generation base for the "post-GPT-6 era" that goes further than Doug.


Bel is directly targeting the Holy Grail of the technology world - AGI (Artificial General Intelligence).

It is reported that the Bel model surpasses Astra in coding, reasoning and long-term agent tasks.

The model can work efficiently for days without human intervention, can self-recover, and coordinate hundreds of parallel subagents, sources said.


3. Claude Fable Killer (The Fable Killer)

@ChrisGPT put it bluntly in his tweet:

Bel is OpenAI’s “monster” model, designed to be a Fable killer.

It should be available before the end of the year, or within a few months of the Astra launch.

He even got the codename six days ago, confirming the reliability of the source.


In theory, Bel may have surpassed GPT-6 and even approached the AGI threshold defined by OpenAI.

In the official release of GPT-5.6, they introduced the "RSI Index", which integrates the results of research and debugging, kernel and training recipe optimization, machine learning experiments, and model self-improvement. In the end, the sol model improved by 16.2 points based on GPT-5.5.


Sol then designed hundreds of architectural experiments for its smaller draft model and started training. And human intervention will only occur in the case of hardware failure and unstable training. Ultimately, the token generation efficiency increased by more than 15%.

Bel has become an existence that continues to evolve in a true sense.


The fast weight layer absorbs lessons learned from validated proofs, code tests, experiments, and tooling trajectories as it works. Slower loops consolidate improvements that survive evaluation into persistent weights and training recipes.

It learns quickly in fast memory, solidifies proven and effective improvements into slow weights, while continuously optimizing the operating mechanism of the next round of learning, and distills the final results into small models that everyone can actually use.

GPT-7 may just be a security snapshot of Bel's status in a certain week.

On Reddit, Leo’s news has triggered a lot of discussion. After all, Leo’s revelations have always been reliable.


Some people speculate that internal models may lead external models by 4.5-6 months.


OpenAI said: "Anthropic can no longer keep up"

If Bel is a technical blow to dimensionality reduction, then computing power is a trick to boost morale within OpenAI.

According to cross-analysis by multiple trackers, OpenAI determines that there is no doubt that it will maintain its lead from the second half of 2026 to 2027.

Why are you so sure? Because of computing power.

The news mentioned that OpenAI’s internal assessment believes that its biggest rival Anthropic has a shortage of computing power and is unable to cope with OpenAI’s next-generation AI.


In the arms race of big models, computing power is ammunition. When the model parameters soar to 10 trillion levels, training once is very expensive.

Although Anthropic is extremely accomplished in model architecture and alignment technology, it is obviously unable to cope with the absolute "aesthetics of violence".

Facing the upcoming public debut of Astra, Anthropic is limited by computing power bottlenecks, and it may be difficult to provide a strong response this year.


While other companies are still struggling to get together 100,000 H100/B200 chips, OpenAI has already used violent aesthetics to smash "Doug" and "Bel".

With the release of OpenAI’s self-developed chip Jalapeño, the moat has not been filled, but has been widened.

The person in charge of OpenAI Codex reveals the “end game”

10 trillion parameters may make you feel a little far away. But within OpenAI, these underlying violent breakthroughs ultimately point to the same AI endgame.

Recently, in an interview with well-known technology podcaster Matthew Berman, OpenAI executive Tibo revealed OpenAI’s future roadmap at the application layer without reservation.


He even said boldly:

The currently extremely powerful Codex model will look like a product of a primitive era in the next 2 to 3 months.


Based on the previous revelations, Tibo pointed to four disruptive changes.

Recursive self-improvement (RSI) happens every day inside OpenAI.

Tibo confirmed that the "internal singularity" not only exists, but has also broken through the business closed loop.

OpenAI has long been using the strongest models to optimize its inference stack, CUDA kernel and even infrastructure.

He gave an example: OpenAI uses the Sol model internally to optimize the Luna model, directly cutting operating costs by 80%!


"Ultra Fast" will reshape human flow.

Tibo revealed that the current Ultra Fast mode internally has achieved up to 14 times acceleration. He predicted that in just 1 to 2 years, this ultra-low latency will become the industry default standard.

When AI's response speed approaches or even exceeds human thinking speed, interaction will become real-time. Workflow will completely return to the "Flow State", and human cognitive load will decrease exponentially.

Computing paradigm shift: Your computer is about to become scrap metal.

With the arrival of next-generation models (such as Astra), Tibo clearly pointed out: the future mainstream of AI will definitely no longer run locally, but will fully shift to large-scale Agent clusters in the cloud.

The ultimate blow: ChatGPT merges with Codex to transform into "Personal AGI".

In the future, "the mechanism will disappear completely." You no longer need to manually write complex prompts, maintain skill files or manage memory, and there is no need to manually schedule sub-agents.

Personal AGI will continuously and passively understand your goals, daily habits and team dynamics, and proactively provide help.

What’s even more amazing is its dynamic UI adaptation. The bottom layer is obviously the same multi-modal, Voice-first technology, but the interface will automatically deform like water based on your identity.

Conclusion: The singularity has arrived, and it will be invisible

No matter how exaggerated the rumors of "Bel" are, or whether Astra's hand-tearing of human expert code is really so silky smooth, this large-scale revelation has sent a signal to the world that cannot be ignored:

The development of artificial intelligence has not stopped, it is just preparing for a storm that can overturn the poker table.

In the second half of 2026, the fun has just begun.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet