Monday Sep 21

Three Researchers Used Claude To Hack OpenAI

21SEP
72 HOURSCLAUDEOPENAI

A three-person startup broke into OpenAI with Anthropic's model. Hacktron used Claude Opus 5 to chain two bugs, hijack employee ChatGPT accounts, and open a pull request in OpenAI's private code.

The door was a memory bug in OpenAI's community forum. Claude chained it with a second flaw to reach employee sessions. Opus 4.8 failed every time. Opus 5 got through within hours of release.

OpenAI patched within 14 hours and paid a $6,500 bounty. Hacktron's line: work that once took a funded team months now compresses into days. Neither company commented.

Citi CEO Jane Fraser this weekend: a tsunami of patching is going on in every company. Defenders are getting the same models. Whoever points them at your systems first wins.

full brief & sources

⚡ Why this matters

  • The gap between a proof of concept and a real breach used to be months. Hacktron did it in under three days.
  • The model that failed and the model that succeeded are one generation apart. Capability jumps now show up in attack timelines.
  • OpenAI's bug bounty treated the community forum as out of scope. Attackers do not read scope documents.

🔍 What happened

  • Hacktron AI, a three-person security startup, published its write-up on September 18. The Register and TechCrunch confirmed the details.
  • Entry point on July 25: OpenAI's community forum, which runs Discourse, processed uploaded images with a library that had a heap overflow.
  • Claude Opus 5 chained that bug with a second flaw to hijack employee ChatGPT and Codex sessions.
  • The team opened a harmless pull request in OpenAI's internal monorepo to prove reach, then reported it.
  • OpenAI fixed the issue in about 14 hours and paid $6,500 through Bugcrowd, noting the forum was out of scope.
  • Discourse issued advisory GHSA-vhm9-85gw-x335. OpenAI and Anthropic did not comment.

💬 Smart takes

  • Hacktron AI, in its write-up: "Work that once required a well-resourced team and months of effort can now be compressed into days."
  • Jane Fraser, Citi CEO: "there is a tsunami of patching going on in the world at the moment in all companies." She called Anthropic's Mythos release "not a good day."
  • Skeptic: the win still needed three skilled humans steering the model and a forum running an old image library. This is a story about unpatched dependencies as much as about AI.

🧭 Where this goes

  1. Likelybug bounty programs widen scope to every public surface within months, because models do not respect scope lines.
  2. Likelymore disclosed model-assisted breaches at big labs before year end. This one was friendly. Not all will be.
  3. PossibleAnthropic and OpenAI publish joint norms for offensive-security use of frontier models.
  4. Wild Carda regulator treats a frontier model release as a security event, with a mandated patch window for critical software.

🥄 The Spoon Take

The scary part is not the hack. It is the version gap. Opus 4.8 could not do it. Opus 5 did it within hours of launch. Every model release is now a Patch Tuesday for the whole internet, and the labs set the calendar. Plan like the next release is an attacker with a head start.

🤔 Pushback

A friendly team with weeks of setup is not a live attacker, and the entry bug was an old, unpatched image library.