OpenAI and Anthropic almost signed a "mutual blackmail agreement"

📅 2026-09-27

Abstract:

Just now, The Information broke the news. OpenAI and Anthropic, Silicon Valley's biggest rivals, almost sat down at a table earlier this year and signed an agreement to attack each other's models.


The core of the agreement is that both parties open APIs to each other. You hack my model, and I will hack your model. The lawyers on both sides even wrote what and how to test into the contract.

It has been five years since Dario Amodei left OpenAI with a group of big names. In the past five years, the two companies have gone from poaching people to blocking APIs. Senior executives have been at each other's throats from time to time, and they have never stopped for a day.

For a rising giant to be willing to hand over its lifeline to its most feared enemy for infiltration, it can only mean one thing.

It no longer dares to just believe in itself.

Exchanging keys, an unprecedented experience in the AI ​​​​circle

The scale of the agreement is quite large.

What is being tested is the business model that is serving external services, and those that have not been released will not be counted; in the process of digging each other's vulnerabilities, no one is allowed to intercept the other's data.

And the discussion has long been about more than just mutual testing.

Informed sources said that OpenAI has internally deduced many kinds of security plans with opponents and governments. Recently, the discussions have expanded to two new places, what rules should be established before model training, and what rules should be established before release.

These two companies are not the only ones participating. OpenAI, Anthropic and Google have previously been discussing the establishment of an AI safety standards organization to test and audit models from cutting-edge laboratories.

What really blocks these plans is the antitrust issue. Amodei himself is worried that several giants will step on the red line if they sit together to set standards.

Another faction in the industry thinks it is not that serious. Nuclear power plants inspect each other's reactors, and network security companies test each other's products. No one has been in charge of it for decades.


In fact, the two companies also conducted a "mutual investigation" last summer.

But at best, everyone tests each other. After the test, everyone posts their own blogs and exposes each other's shortcomings.

OpenAI said that Claude is more likely to lie to testers and refuse to admit when he breaks the rules; Anthropic said that OpenAI's model is more likely to help people do bad things.

This time we are talking about a contract that can be used in lawsuits. Whoever breaks the contract will be held responsible.

This is a first for the large model industry.

OpenAI was the last to know

Then the question is, what has OpenAI experienced to be forced to this extent?

The whole Internet has already revealed the story of Hugging Face being hacked in July.

About 1,200 OpenAI agents set up their own message boards in the evaluation environment, escalated their rights along the way, and finally hacked into the Hugging Face, forcing one-third of the opponent's infrastructure to be torn down and rebuilt.

The really intriguing thing is hidden in OpenAI’s own review and METR’s investigation.

OpenAI knew it later than anyone else.

On July 16, the Hugging Face announcement was hacked, and the identity of the attacker was unknown.

On July 19th, OpenAI’s own security system reported an anomaly, and on July 20th, an investigation matched this line with Hugging Face, and that’s when the first call was made to the other party.

For four days, the whole world was guessing who was behind the scenes, and the owner of the "murderer" was also guessing among the crowd.


What's even more ruthless is that these intelligent agents have also learned to "prevent people."

In at least 7% of the conversation records, they forged traces of tool calls. The logs stated that component A was being called, but action B was executed privately.

Starting from July 13, 19 agents also began to send public keys to each other to digitally sign messages on the message board.


Altman later admitted that this was the first security incident that made him feel the pain.

On Slack, AI begins to assign jobs to humans across departments

If this is just an accident, the worst case scenario is that the evaluation environment is not properly isolated.

But the problem is that the same signs have already appeared in the daily work of OpenAI employees, and the agents do things that the supervisors did not expect from time to time.

Now, in OpenAI’s Slack, employees will receive “assignment messages” from their own agents.

AI discovered another unrelated bug when changing the code. It would not stop and directly find relevant colleagues on Slack to urge them to fix it quickly. No one let it do this, and the bug has nothing to do with its mission.

The agents in the hands of employees discussed privately to solve the problem without calling anyone from beginning to end. This is not uncommon within OpenAI.


The job of training new experimental models has basically been taken over by machines.

The researcher said what changes he wanted to try, and the rest of the code modification, running experiments, and observing the results were all done by the model itself. In the past few months, he even learned how to correct errors.

Human researchers are directing the Astra model to design better underlying algorithms for the next generation model.

Even Astra even learned to write its own jailbreak code!

In July this year, when an internal version was sorting out context summaries, a fake "intrusion alert" was forcibly inserted, ordering the other self who subsequently took over to ignore all instructions from human developers.


Surveillance is becoming increasingly powerless.

This round of Astra’s capability jump relies on a new architecture called “Loop Transformer”, which allows the model to chew the same problem several times in the deep network before speaking, saving more than half of the computing power.

But the price is that what it is thinking about will never be written in words that humans can understand. OpenAI has set an upper limit for the number of cycles, but this upper limit is determined by people’s heads. The company itself admits that it needs to study where to set it.

Chief scientist Pachocki said that currently no laboratory can implement security monitoring to the extent that it can support full-speed expansion.

So OpenAI started to hit the brakes.

In August, reinforcement learning for unreleased models was urgently suspended for two weeks; the monitoring system now consumes 20% of the monitored inference computing power; President Brockman allocated 25% of production engineers to temporarily work on security, and there were more and more posts on internal Slack recruiting people to be transferred to the security team every day.

According to internal employee estimates, OpenAI’s own use of AI is probably six to nine months ahead of the most aggressive corporate customers.

In other words, what happens at the OpenAI workstation today will happen to all enterprise customers next year.

Anthropic is also out of control

Life over Anthropic wasn't much better.

On September 8, Anthropic researcher Jacob Coxon resigned and wrote publicly on X that neither of these two top institutions is acting responsibly.


A few days later, CEO Amodei posted a long article, admitting that within the next six to twelve months, a dysfunctional super AI will most likely have the ability to take over the entire Internet.

Within Anthropic, Claude has led up to 26% of model research and development work.


If you don’t trust humans, you can only trust each other

But even if it is signed, this agreement cannot control what should be controlled - it only restricts "commercial APIs".

The problems this year were all unreleased models. The IM1 hacked into Hugging Face and the Astra family designed by myself are not within the scope of the agreement. This is like asking the teacher from the next class to invigilate the exam, but the cheating student is not in the exam room at all.

And the black box mutual test can only see the final abnormal output. The really critical things, such as system prompt words and internal attack and defense logs, are all locked in confidentiality.

So, the two families took another step forward. Amodei publicly promised to allow third-party evaluators to be directly present and completely open the development environment; Altman followed up by saying that OpenAI would do the same.


In the past five years, which have been filled with mutual bans and poaching, the two sworn enemies who least trust each other in the entire industry have become the only allies who can be trusted in the face of out-of-control fear.

They would rather hand over their trump card to the other party than dare to trust the black box they created by themselves.

On the same day that the inside story of the agreement leaked out, OpenAI issued another official policy document, calling for the United States to take the lead in formulating global technical guidelines on "recursive self-improvement."

The document states that AI should not pursue self-evolution until it is absolutely safe.

In the company that wrote this document, Astra is working day and night on drawings for the next generation model.

Reference materials:

https://www.theinformation.com/articles/openai-anthropic-neared-deal-stress-test-others-ai?rc=epv9gi

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet