Monday Jul 13

GPT-5.6 Proves A 50-Year-Old Conjecture

13JUL
UNDER 1 HOUR64 AGENTS1 PROOF

An AI just solved a decades-old math problem alone. OpenAI says GPT-5.6 Sol Ultra proved a 1973 conjecture in under an hour. Nobody has peer-reviewed it, and this exact problem has fooled experts before.

The math: cover every edge of a graph with cycles, each edge counted exactly twice. Mathematicians Szekeres and Seymour posed it decades apart, and nobody had cracked it.

GPT-5.6 ran 64 subagents at once, each testing a different angle. Most agents were told to explore, not converge, early on. The model leaned on an old theorem, then closed the proof with linear algebra.

Mathematician Thomas Bloom called it clean, even elementary. Nobody has independently verified it, and this exact conjecture has swallowed flawed proofs before.

full brief & sources

Why this matters

  • First time a model produced a genuinely new proof of an open problem, not just a known one restated.
  • It shipped the same week GPT-5.6 went fully public, doubling as a capability demo.
  • If it holds up, it's evidence models can now do original math research, not just verify it.

🔍 What happened

  • The Cycle Double Cover Conjecture was posed by George Szekeres in 1973 and independently by Paul Seymour in 1979.
  • It claims any bridgeless graph has a set of cycles that together cover each edge exactly twice.
  • OpenAI had GPT-5.6 Sol Ultra run up to 64 subagents in parallel, managed 'aggressively and dynamically.'
  • Early rounds pushed the agents toward diverse approaches before converging on one proof strategy.
  • The proof reduces the problem to cubic graphs and leans on the 8-flow theorem plus a linear-algebra argument.
  • OpenAI published the full prompt and proof publicly the next day.

💬 Smart takes

  • Ethan Knight, OpenAI: the model produced the proof using 64 subagents in just under an hour.
  • Thomas Bloom, mathematician: called the proof 'very nice' and 'elementary' - the kind of result that could have been found in the 1980s.
  • Skeptic: this conjecture has attracted multiple flawed proofs over the decades, and this one hasn't passed peer review yet.

🧭 Where this goes

  1. LikelyOpenAI keeps publishing math results as flagship proof points for GPT-5.6's capability.
  2. Possiblea mathematician finds a subtle gap in the proof within weeks, given the conjecture's track record.
  3. Possiblerival labs race to show their own models solving open problems, turning math into a benchmark war.
  4. Wild Cardthis becomes the first AI-generated proof formally accepted into a peer-reviewed math journal.

🥄 The Spoon Take

A model didn't just answer a question, it picked a fight nobody had won in fifty years and walked away with a proof. That's a different kind of milestone than a benchmark score. But math has a brutal review process, and this conjecture has burned confident people before.

🤔 Pushback

If a flaw turns up in peer review, this becomes a cautionary tale about confident-sounding AI math, not a landmark.