Codex Ran For 25 Hours Straight
22JUN
OpenAI's Codex coded for about 25 hours with no human touch. One run burned 13 million tokens and wrote 30,000 lines. The new GPT-5.3-Codex model is built for long, unattended work.
Long-running agents stopped being a demo. This was a single sustained run, not a benchmark score.
GPT-5.3-Codex merges OpenAI's best coding and reasoning models and runs 25% faster. OpenAI also shipped guidance on using Codex as a persistent workspace that holds context across long projects.
The number that matters is time. 25 hours alone changes what you hand an agent. Review and cost control become the real bottleneck, not capability.