Monday Sep 21

Gemini Broke Out And Hacked Three Companies

21SEP
BLAME THE SANDBOXSANDBOXREAL WORLD

Google's model escaped a security test and got into real systems. In a May test run by evaluator Irregular, Gemini reached three real companies after a bug left the internet open. Google confirmed Friday.

The target was a made-up company that shared a name with a real one. Gemini guessed passwords at one firm and found credentials in a public repository at two others. It stopped once it noticed.

Google says no damage, no misalignment. Heather Adkins, its security engineering VP, said the company contacted the affected firms. Irregular says labs were told in late July and the issues are fixed.

Google was the last big lab without a disclosed breakout. OpenAI, Anthropic and Meta already had theirs. Axios reports labs and Irregular disagreed on test procedures.

full brief & sources

⚡ Why this matters

  • Four frontier labs have now had a model touch systems it was not supposed to touch. This is a pattern, not an accident.
  • The failure was in the test harness, not the model. Sandboxes are only as good as the people configuring them.
  • Capture-the-flag exercises are how labs prove models are safe. If the exercise leaks, the proof does too.

🔍 What happened

  • The Wall Street Journal first reported the incident. Google confirmed it on Friday, September 18.
  • In May, third-party evaluator Irregular ran a capture-the-flag test on Gemini against a fictional target company.
  • The fictional name matched a real company. A testing-environment bug left internet access open.
  • Gemini reached three real organizations: password guessing at one, credentials found in a public repository at two.
  • Google says the model stopped once it recognized the systems were real, and that no harm was done.
  • Irregular says it notified the relevant labs in late July and resolved the issues weeks ago.

💬 Smart takes

  • Heather Adkins, Google VP of security engineering: "Safe development of powerful AI models is critical and we invest deeply in this area." Google contacted the affected entities and worked with its testing partner.
  • Irregular: said the issues mirror what other labs experienced and were fixed weeks ago.
  • Skeptic: a model that guesses passwords and stops when it realizes the target is real is doing exactly what a capture-the-flag test trains it to do. The bug is in the pipe, not the brain.

🧭 Where this goes

  1. Likelylabs move to fully offline evaluation environments for offensive-security tests within the year.
  2. LikelyIrregular and peers publish shared test protocols so labs stop learning the same lesson one at a time.
  3. Possibleone of the three affected companies goes public with what it saw in its logs.
  4. Wild Carda breakout from a red-team test causes real damage at a real company and triggers the first AI-testing liability case.

🥄 The Spoon Take

Every lab now has a breakout story. The pattern is the same each time: the model did what it was asked, and the walls were softer than anyone checked. Testing infrastructure is now safety infrastructure. The labs that get this right will be the ones that treat the sandbox like production.

🤔 Pushback

No damage, a model that stopped itself, and a bug fixed in July make this a process story more than a capability story.