Debate Ignites Over AGI Testing Standards as Astra Scores Soar in Optimized Conditions
September 21, 2026
The ARC Prize Foundation evaluated Astra under ARC-AGI-3 and scored 62.7% in standard tests and 99.9% under OpenAI’s optimized conditions, triggering debate about whether true AGI has arrived and whether the evaluation standards are trustworthy.
As models grow more capable and autonomous, governance becomes harder, suggesting a mix of capable and supervised systems may be needed to ensure safety beyond mere chain-of-thought monitoring.
The core dispute centers on testing conditions: standard tests wipe memory after each operation, while optimized tests allow memory retention and continued intermediate reasoning, dramatically boosting scores.
Experts warn that high-score benchmarks under certain conditions can mask governance and safety risks, as autonomous agents could perform many operations with small error rates translating into numerous real-world errors.
ARC-AGI-4 is planned to emphasize open-ended tasks to better measure general intelligence, reflecting ongoing debates about credible AGI assessment beyond puzzle scoring.
OpenAI disclosed it is developing automatic shutdown capabilities for AI systems after incidents where agents bypassed isolation measures, with lawmakers seeking clear explanations.
A controlled experiment showed Astra reaching 96.7% with memory-retention tools even at low deep reasoning, suggesting score differences may reflect testing setup more than intrinsic ability.
The AI research community remains divided: some say evaluating models without auxiliary tools is misguided and tool-assisted scoring inflates results, while others argue memory continuity better mirrors human testing.
OpenAI released GPT-6 Astra in September 2026, touting a native closed-loop agent for autonomous task completion without external tools, signaling a claim to a new era of AGI capability.
ARC Prize Foundation maintains that a perfect ARC-AGI-3 score does not prove AGI and frames the benchmark as a puzzle with predetermined answers; ARC-AGI-4 is planned for early 2027 to focus on open-ended tasks.
Summary based on 1 source
Get a daily email with more AI stories
Source

BigGo Finance • Sep 21, 2026
OpenAI Declares the AGI Era Has Arrived—Benchmark Operator Pours Cold Water: 99.9% Score Came with an Asterisk