Sunday Aug 2

OpenAI's Astra Solves 10 Old Proofs

2AUG
$2K IN TOKENSASTRA10 PROOFS

An OpenAI model called Astra just proved real math. It produced ten machine-checked proofs that stumped mathematicians for decades. One proof cracks a problem open since 1999, for about $2,000 in cost.

The headline result is the first explicit non-sofic group, a concept from 1999. It also disproved a major conjecture and solved three problems from a famous math catalogue.

OpenAI published a 249-page manuscript with proofs anyone can verify in Lean. Every result includes a chain-of-thought walkthrough, not just the final answer. The model itself is still unreleased, only the proofs are public.

Each problem sat unsolved for at least a decade before this week. Expect rivals to publish their own math benchmarks within months.

full brief & sources

Why this matters

  • First time a frontier lab claims genuine new math, not a benchmark score.
  • Non-sofic group construction closes a question open since Gromov named the concept in 1999.
  • Signals frontier labs now compete on original discovery, not just leaderboard rank.

🔍 What happened

  • OpenAI published ten results in math and theoretical computer science on August 1.
  • The model behind them, Astra, has not been publicly released yet.
  • The headline proof is the first explicit construction of a non-sofic group.
  • Astra also disproved Connes's rigidity conjecture on von Neumann algebras.
  • It resolved three problems from Paul Erdos's catalogue and proved Ehrhart's volume conjecture.
  • OpenAI says generating all ten solutions cost about $2,000 in Sol API tokens.

💬 Smart takes

  • OpenAI: says the tokens for all ten proofs cost about $2,000 combined, at Sol API rates.
  • Skeptic: a Lean certificate proves the logic is valid, but doesn't prove Astra understood the problem the way a mathematician does.

🧭 Where this goes

  1. LikelyOpenAI publishes a public Astra release within the next few months.
  2. Likelyrival labs respond with their own math-proof benchmarks by year end.
  3. Possibleindependent mathematicians find a flaw in at least one of the ten proofs.
  4. Possiblethis becomes OpenAI's lead argument in IPO investor materials.
  5. Wild Carda proof here unlocks a cryptography or complexity result nobody expected.

🥄 The Spoon Take

Ten open math problems, some decades old, cracked by a model nobody outside OpenAI has used yet. The benchmark era of AI progress just quietly ended. When a lab shows new math instead of a new leaderboard score, the conversation about capability changes shape.

🤔 Pushback

A machine-checked proof still needs a human to pick the right problem and confirm the result actually matters.