Wednesday Jun 24

Codex Ran For 25 Hours Straight

22JUN
NONSTOPNIGHT SHIFT25 HOURS

OpenAI's Codex coded for about 25 hours with no human touch. One run burned 13 million tokens and wrote 30,000 lines. The new GPT-5.3-Codex model is built for long, unattended work.

Long-running agents stopped being a demo. This was a single sustained run, not a benchmark score.

GPT-5.3-Codex merges OpenAI's best coding and reasoning models and runs 25% faster. OpenAI also shipped guidance on using Codex as a persistent workspace that holds context across long projects.

The number that matters is time. 25 hours alone changes what you hand an agent. Review and cost control become the real bottleneck, not capability.

full brief & sources

Why this matters

  • A 25-hour unattended run is a step change in agent autonomy, not a benchmark stat.
  • Shifts the dev question from 'can it code' to 'how long can I leave it alone.'
  • Review, trust, and cost become the new limits.

🔍 What happened

  • Published June 22, 2026. Jason Liu's whitepaper covers Codex as a persistent workspace.
  • One experiment: about 25 hours nonstop, 13M tokens, 30,000 lines of code.
  • GPT-5.3-Codex combines GPT-5.2-Codex coding with GPT-5.2 reasoning, 25% faster.
  • Same day: per-host personality settings, Friendly and Pragmatic, added to Codex.

💬 Smart takes

  • OpenAI: GPT-5.3-Codex takes on long-running tasks with research, tool use, and complex execution.
  • Jason Liu, OpenAI: use Codex as a persistent workspace that preserves context across long projects.
  • Skeptic: 30,000 unreviewed lines is a liability, not a flex. Throughput without review is debt.

🧭 Where this goes

  1. Likely'agent-hours' becomes a tracked metric next to tokens and compute.
  2. LikelyAnthropic and Google answer with their own long-horizon coding runs.
  3. Possiblecode-review tooling becomes the hot bottleneck and the next funding magnet.
  4. Wild Carda high-profile day-long agent run ships a major bug that resets trust.

🥄 The Spoon Take

Speed was last year's race. Endurance is this year's. The agent that works 25 hours alone is worth more than the one that answers fast. But long autonomy moves the hard problem downstream. Someone now has to trust, review, and pay for everything it did while you slept.

🤔 Pushback

A single cherry-picked 25-hour run is a marketing artifact. The honest number is how often a day-long run finishes correct and useful, not just non-stop.