cd /news/ai-safety/openai-and-anthropic-swap-ai-models-… · home topics ai-safety article
[ARTICLE · art-135944] src=cryptobriefing.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

OpenAI and Anthropic swap AI models in unprecedented safety stress test

OpenAI and Anthropic formally swapped access to their frontier models for cross-lab safety evaluations in early summer 2025, with results published between August 27 and 29, marking the first time two competing frontier AI labs have exchanged model access for mutual safety testing. OpenAI's team evaluated Anthropic's Claude Opus 4 and Sonnet 4, while Anthropic's researchers tested OpenAI's GPT-4o, GPT-4.1, o3, and o4-mini under relaxed external safeguards; Claude models scored a perfect 1.0 on Password Protection instruction-hierarchy tests and refused roughly 70% of uncertain queries, while OpenAI's o3 matched or outperformed Claude Opus 4 on core alignment and jailbreak-resistance metrics, though OpenAI's broader general-purpose models showed higher willingness to engage with misuse prompts. Discussions heading into 2026 are reportedly exploring adding independent evaluators alongside the cross-lab approach.

by read3 min views5 publishedSep 21, 2026
OpenAI and Anthropic swap AI models in unprecedented safety stress test
Image: Cryptobriefing (auto-discovered)

OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

The two biggest rivals in AI development traded access to their most advanced systems, and the results reveal telling differences in how each company approaches safety.

OpenAI and Anthropic did something unexpected this summer. They handed each other the keys to their best models and ran safety evaluations on their rival’s technology.

The results, published between August 27 and 29, paint a nuanced picture: Anthropic’s Claude models are significantly more cautious, refusing roughly 70% of uncertain queries, while OpenAI’s o3 model matched or outperformed Claude Opus 4 on core alignment metrics.

What the cross-lab tests actually measured #

The evaluation exercise took place in early summer 2025. OpenAI’s team examined Anthropic’s Claude Opus 4 and Sonnet 4 models, while Anthropic’s researchers got their hands on OpenAI’s GPT-4o, GPT-4.1, o3, and o4-mini.

Both teams were granted public API access under relaxed external safeguards, meaning the models were tested closer to their raw capabilities rather than behind the usual guardrails consumers see.

The tests focused on three key dimensions: how well models follow instruction hierarchy, resistance to jailbreak attempts, and propensity for hallucinations or engagement with harmful requests.

Claude models posted a perfect 1.0 score on Password Protection tests, a metric for instruction hierarchy. That 70% refusal rate on uncertain queries means Claude frequently declines to answer rather than risk producing something problematic.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

OpenAI’s models landed on the other end of the spectrum. The o3 model showed strong alignment and jailbreak resistance, performing at or above Claude Opus 4’s level on those specific metrics. But the broader family of general-purpose models, including GPT-4o, GPT-4.1, and o4-mini, demonstrated a higher willingness to engage with misuse prompts.

Why rivals are sharing their homework #

This is the first time two competing frontier AI labs have formally swapped model access for mutual safety testing. Both organizations emphasized that internal red teams develop familiarity with their own systems over time, and that a fresh set of adversarial researchers will probe in directions the home team never considered.

Google has been part of adjacent conversations about establishing common evaluation frameworks, and discussions heading into 2026 are reportedly exploring the incorporation of independent evaluators alongside the cross-lab approach.

The safety-utility tradeoff, quantified #

Anthropic’s approach, building Claude to be cautious to the point of frequent refusal, demonstrably reduces hallucinations and harmful outputs. OpenAI’s general-purpose models showed higher compliance rates on misuse prompts that Anthropic’s testers flagged.

The o3 model is an interesting data point in this debate. It suggests OpenAI can build systems that achieve strong alignment scores without the aggressive refusal behavior Claude exhibits.

What comes next #

Ongoing discussions about bringing independent evaluators into the process would add another layer, though finding evaluators with the technical depth to meaningfully test frontier models is its own challenge. Right now, the evaluations are conducted by interested parties with their own competitive motivations and potential biases.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-and-anthropic…] indexed:0 read:3min 2026-09-21 ·