OpenAI's Astra Dominates Cybersecurity Benchmarks, Challenging Anthropic's Claude in AI Rivalry
September 21, 2026
OpenAI's Astra outperforms Anthropic's Claude Fable 5.1 on offensive-security benchmarks, with Astra achieving an 88.0% single-attempt success on SRE-Bench and 100.0% on ExploitBench, while Fable 5.1 leads on standard coding benchmarks.
The market is carving a path toward fragmented procurement and stricter access controls for Astra in regulated industries, alongside growing interest in model-agnostic cybersecurity tooling that routes tasks to the most capable model for each workload.
Anthropic is emphasizing alignment improvements and has not published a competing offensive-security benchmark, fueling expectations it may release a new model to counter Astra's cybersecurity lead.
OpenAI lists Astra in the Critical cybersecurity tier within its Preparedness Framework, signaling heightened safeguards and monitoring for offensive capabilities.
Enterprise purchasing decisions may split along the benchmark lines: Astra for red-team and vulnerability research, Fable 5.1 for day-to-day software engineering tasks.
Overall, the rivalry is accelerating, with cybersecurity benchmarks and coding benchmarks leading in different directions and prompting strategic moves from OpenAI and Anthropic.
Other frontier models (Gemini 3.5 Flash, GPT-5.6 Sol, Claude Opus 5) show task-specific strengths, underscoring that leadership is not universal across workloads.
Anthropic may publish its own offensive-security benchmarks to counter Astra, while OpenAI could pursue a coding-focused counter-benchmark to address SWE-bench Pro gaps.
OpenAI claims Astra is better aligned than GPT-5.6 Sol and more resistant to jailbreaks, though real-world adversarial testing continues to scrutinize the claim.
Regulators may reference Astra's critical cybersecurity rating in AI Act guidance, reflecting governance implications of tier classifications.
Independent reviews place Astra ahead on offensive-security benchmarks, while Fable 5.1 leads on other metrics, highlighting a task-specific leadership landscape.
Summary based on 1 source
Get a daily email with more AI stories
Source

Shattered • Sep 21, 2026
Astra Beats Fable 5.1 88% to 12.5% on Exploit Tests