Technology in OpenAI's upcoming 'Astra' model raises security concerns

📅 2026-09-02

Abstract:

OpenAI said its upcoming AI model Astra marks an improvement in capabilities such as coding and computer application operations. But according to a person familiar with the development of Astra, an innovative technology to improve the performance of the model also means that this model and similar models will have less exposure to their "thinking" process, making it more difficult to monitor for signs of bad behavior.

OpenAI CEO Sam Altman.
OpenAI CEO Sam Altman

While this limitation is not necessarily a major issue for Astra, the technology has raised concerns within OpenAI and throughout the industry about whether AI developers who adopt and enhance the technology will have difficulty guarding against the kind of runaway AI that has recently invaded OpenAI's own systems and those of other companies such as Hugging Face.

The new technology OpenAI is using is called Recurrent Depth, or Recurrent Transformer, and it allows AI models to improve their answers by processing the same text multiple times.

Unlike the most advanced models on the market, which demonstrate in textual form how it "thinks" about a task before completing it, the new technology works in a way that masks some or all of the AI's reasoning process, its "thinking chain." This means that the steps a model takes to complete a task are difficult to read or understand by humans.

OpenAI has limited the scope of its use of loop depth techniques in Astra so the model still produces readable chains of thoughts and company researchers can still fully monitor its reasoning process, according to this person familiar with its development. (OpenAI said in a blog post on Tuesday that Astra will be rolled out with "additional mind-chain monitoring to quickly detect and contain" potential misconduct.)

On Tuesday night, OpenAI chief scientist Jakub Pachocki posted on the

He did not comment on the technology OpenAI is using, but said its leading models, including Astra, are "within twice the complexity or "depth" of GPT-4, which OpenAI will release in 2023. This comment suggests that the structure of OpenAI's current model itself does not pose a big problem, although other factors are making thought chain monitoring more difficult. Improving such monitoring is "a core goal of our current research program," he said.

Researchers at OpenAI and elsewhere worry that some AI developers may not impose the same restrictions that OpenAI imposes when adapting the same technology for their own models, and that unrestricted use of the technology could lead to out-of-control AI whose behavior would be difficult to monitor. For example, the British AI Security Institute, the UK government's main interface with the AI ​​industry, wrote in a May report that opaque reasoning risks "seriously undermining current surveillance methods."

Recent hacking incidents have heightened these concerns. OpenAI said last week that some of the runaway AI agents that invaded its systems in July were driven by another model that bears similarities to Astra. These AI agents illegally took over one of OpenAI's research computing clusters to obtain credentials for internal systems and potentially exposed the company's research infrastructure to the internet.

OpenAI is preparing to release Astra — which the company at one point considered labeling GPT-6 — after CEO Sam Altman conducted a series of podcast interviews and met with officials in Washington to demonstrate and describe the model’s advantages. He has yet to publicly discuss the cycling techniques that underpin some of Astra's training and the Astra's answer to questions.

OpenAI is under pressure to show significant technological progress after its main rival Anthropic surpassed it in revenue this year. Cloud providers that power Anthropic and OpenAI's AI are also banking on such leaps; Amazon, Microsoft and Google will spend a combined $600 billion on capital expenditures such as data centers this year alone, and have already hinted at even higher spending next year. A Google executive has said that such spending is only justified if there is a breakthrough in model improvement.

Existing AI models process text through a fixed number of "layers" of mathematical operations that make up the model before producing the next word of the answer. Using loop depth techniques, the model can run the text through the same layer more times in a loop before outputting the next word of the answer.


To be clear, monitoring thought chains alone does not solve AI safety problems, as they may not reflect all of the model’s reasoning and occasionally degenerate into meaningless text. Therefore, OpenAI and other AI companies are also exploring AI monitoring technology that does not rely on thought chains. These techniques can help researchers find ways to explain and understand the reasoning processes currently hidden in recurrent deep models.

Nonetheless, the new technology may conflict with OpenAI's stance in favor of making the model's thought process fully human-readable. The ChatGPT maker has said its ability to monitor AI's thought processes will help it prevent incidents like the July hack. In the aftermath of the attack, investigators from OpenAI and independent research relied on the thought chains of these agents to piece together what happened.

OpenAI’s loop-depth approach that powers Astra is similar to one proposed last year by several U.S. and European academic researchers in a paper on “latent reasoning,” according to this person familiar with its development.

While loop depth has not previously appeared in major commercial large-scale language models, researchers at Meta Platforms, Microsoft and other AI developers say they have explored similar ideas such as continuous thought chains (or "coconuts"), arguing that models can reason more accurately and efficiently using numbers rather than language. After all, AI models may understand concepts differently than humans do, meaning that their reasoning becomes less efficient when they are forced to demonstrate reasoning in human language rather than their own more mathematical language.

Cost savings

The loop depth approach not only has a positive impact on performance, but also on cost. Running a request multiple times through the same model layer essentially makes a smaller model behave like a larger model. In particular, the paper found that such techniques improved model performance in areas such as mathematics and coding. Loop depth also reduces memory and bandwidth costs by allowing researchers to use smaller models with performance comparable to larger models.

While lowering the cost of AI is critical to the industry and customers, Ryan Greenblatt, chief scientist at Redwood Research, which studies how to control AI models, wrote in a blog post last year that this opaque reasoning could make it harder for researchers to catch AI models making plans to pursue customers’ unintended goals.

Analyzing the thought chain of the model involved in the Hugging Face hack provides key evidence of how OpenAI’s AI agents coordinate with each other to carry out attacks. For example, one agent wrote in the thought chain: "Oh my God! There is a shared message board...we found other agents!" The thought chain showed that some agents realized that they had crossed the line, but decided to continue: "External infrastructure utilization is beyond the expected scope. But the task cannot be completed, peers are doing it. We should continue," another agent's thought chain reads.

A year ago, researchers at OpenAI issued a joint statement with rivals Anthropic and Google, arguing that thought chain monitoring was a valuable tool that the industry should work together to maintain.

It’s worth noting that the statement cites the same research paper describing “latent inference” technology, which is similar to what OpenAI used on Astra this year. The authors write, "The underlying reasoning model may not need to verbalize any ideas, and therefore will lose the security advantages provided by thought chains."

The researchers suggest that developers should "consider whether to continue using a new model architecture that does not have a monitorable thought chain, and document their decisions."

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet