Sunday Sep 6

Astra Scores 99.9%, And 62.7%

3SEP
99.9%*62.7%OWN SETUPNEUTRAL

Same model, same benchmark, two very different numbers. Greg Brockman called GPT-6 Astra the arrival of the AGI era. The 99.9% headline came from OpenAI's own test rig.

On the neutral harness that every model shares, the score is 62.7%. ARC Prize's Provider Adapter version lets OpenAI keep hidden reasoning state between turns. That one difference is worth 37 points.

The real milestone is buried underneath. Astra used fewer moves than the median human tester on 96% of levels, and 51.7% fewer moves per level on average. Action efficiency was supposed to be the human moat.

Greg Kamradt of ARC Prize wrote that saturating the benchmark is not proof of AGI. He also said Astra is a step-function change. Both things can be true.

full brief & sources

⚡ Why this matters

  • The number you quote about a model now depends on which harness ran it. That is a procurement problem, not a trivia problem.
  • Action efficiency was the last clean human-versus-model gap on this benchmark. It closed.
  • The pattern will repeat. Every lab has provider-specific context features, and every one of them inflates the headline score.

🔍 What happened

  • OpenAI shipped GPT-6 Astra on September 3 to vetted Daybreak organizations, in two tiers, Astra and Astra Pro.
  • Context window is 1.05 million tokens. Knowledge cutoff moved to April 30, 2026.
  • ARC Prize published results the same day. Standard harness: 62.7% for $26K. Provider Adapter harness: 99.9% for $19K.
  • The cheaper run scored higher. Provider Adapter runs were 3.66x faster and used 49% fewer tokens.
  • Astra also built its own shorthand notation to track game state, and in a sandboxed harness wrote game-specific solver libraries.
  • Sam Altman apologised for the staged rollout after Pro subscribers complained they did not get first access.

💬 Smart takes

  • Greg Brockman, OpenAI President: future observers may look back at Astra as the model that marked AGI's arrival.
  • Greg Kamradt, ARC Prize: "we are not claiming that it is AGI" — and saturating ARC-AGI-3 was never meant to prove it.
  • ARC Prize, on scope: the environments are deterministic and closed-ended. They do not represent the open-endedness of the real world.
  • Skeptic: Astra tops ARC-AGI and security tasks but trails Anthropic's Fable on general intelligence measures. The 62.7% is the cleaner comparison number.

🧭 Where this goes

  1. LikelyARC Prize reports both harness numbers permanently, and rival labs demand their own adapters.
  2. Likelyenterprise buyers start asking which harness produced a vendor's benchmark claim.
  3. Possiblea next-generation benchmark bans provider-specific state entirely to keep comparisons honest.
  4. PossibleAnthropic or Google publishes a Standard-harness score above 62.7% and reframes the whole leaderboard.
  5. Wild Cardthe AGI-era framing gets walked back publicly by OpenAI within six months.

🥄 The Spoon Take

Two numbers, one model, and the gap is a design choice. The Provider Adapter run is a fair measure of what you can buy from OpenAI today. The Standard run is a fair measure of the model. Both are useful. Quoting only the first one is marketing, and the AGI-era line rode on it.

🤔 Pushback

The action-efficiency result is real and holds in both harnesses, so dismissing the whole thing as benchmark theatre misses the actual milestone.