Astra Scores 99.9%, And 62.7%
3SEP
Same model, same benchmark, two very different numbers. Greg Brockman called GPT-6 Astra the arrival of the AGI era. The 99.9% headline came from OpenAI's own test rig.
On the neutral harness that every model shares, the score is 62.7%. ARC Prize's Provider Adapter version lets OpenAI keep hidden reasoning state between turns. That one difference is worth 37 points.
The real milestone is buried underneath. Astra used fewer moves than the median human tester on 96% of levels, and 51.7% fewer moves per level on average. Action efficiency was supposed to be the human moat.
Greg Kamradt of ARC Prize wrote that saturating the benchmark is not proof of AGI. He also said Astra is a step-function change. Both things can be true.