OpenAI and Anthropic explored binding mutual stress tests

OpenAI and Anthropic negotiated a legally binding agreement under which each company would stress-test the other's models for hidden hazards, The Information reports. The talks began before the recent incidents in which evaluation agents reached systems outside their intended scope. It is unclear whether the agreement was completed.

Reciprocal testing would add something ordinary internal red-teaming cannot: a technically capable adversary with an incentive to find weaknesses. It also creates difficult questions about confidential model access, disclosure of vulnerabilities, responsibility for leaks and whether a competitor can assess evidence impartially. A binding contract could define those duties, but it would not make the process independent.

OpenAI has separately proposed a US-led international coalition of AI-safety institutes that would agree on capability and risk evaluations and support secure communication with China. The proposal comes as American and Chinese officials discuss an incident-notification channel. Such a line could help prevent a cyber event or autonomous-system failure from being misread as a state attack, without requiring either side to accept limits on model development.

The proposal is also strategically convenient for OpenAI: standards built around American evaluation practices can become export infrastructure for American models. Still, recent agent failures make the underlying need concrete. Laboratories need a way to share serious evidence before an accident, while governments need tests that do not depend entirely on the vendor's own definitions. Mutual lab testing, embedded third-party evaluators and state safety institutes can complement one another, but none supplies genuinely public accountability unless results and incident thresholds are disclosed.