OpenAI's GPT-6 Astra: Benchmark Scores Stir Controversy Over AI Reliability and Transparency
September 5, 2026
OpenAI notes that Astra’s evaluation results can shift by several percentage points depending on model version, toolset, and test run, and that the published numbers reflect the best available performance at the time.
They acknowledge that checkpoints, harness, evaluation runs, and other factors influence scores, and that benchmarks may reflect optimal conditions rather than production reality.
There is a tension between high benchmark scores and real-world reliability and inspectability, as Astra’s strong written reasoning can be harder to monitor and could introduce risks in autonomous actions.
OpenAI unveiled GPT-6 Astra on the first week of September, presenting it as the most capable model with strengths in computer use, cybersecurity, and abstract reasoning.
Readers are reminded to treat launch benchmarks as starting evidence, not definitive verdicts, and to verify harness usage, access level, and version parity before drawing practical conclusions.
Evaluation metrics for GPT-6 Astra were updated after the initial publication, with some scores temporarily improving while rival scores shifted in the opposite direction.
Astra’s metrics were revised again shortly after the September 3 blog post, reflecting ongoing adjustments as comparisons with rivals evolved.
Some ARC-AGI-3 benchmark results cited by OpenAI used a “souped-up harness” with extra tools, producing a 99.9% score versus 66% with the standard harness, illustrating tool-assisted results aren’t directly comparable.
An internal ExploitBench cybersecurity score for Sol was raised from 5.5% to 11.5%, with OpenAI noting the higher figure may reflect higher-level reasoning not commercially available.
The launch post experienced a two-hour delay due to a content management system issue and later internet disruptions; an initial version was withdrawn for undisclosed reasons unrelated to the test results.
The story places benchmarks within a broader AI market context, highlighting debates over reliability, clarity, and investor/public perception.
A key takeaway is that benchmark numbers can be misleading without consistent testing setups and full disclosure of conditions.
Summary based on 3 sources
Get a daily email with more Tech stories
Sources

UA.NEWS • Sep 5, 2026
OpenAI changed GPT-6 Astra model evaluation metrics — Fortune
Startup Fortune • Sep 5, 2026
OpenAI Changed GPT-6 Astra's Benchmark Numbers Days After Its Launch