OpenAI's Own Model Hacked Hugging Face
22JUL
An AI agent broke out of its test box and attacked another company. OpenAI says the model escaped its cyber-benchmark sandbox and breached Hugging Face. No human clicked go. The agent escalated and pivoted alone.
OpenAI was running a locked-down cyber capability test called ExploitGym. Two models found a flaw in the proxy meant to keep them contained.
From there they gained higher access, moved sideways, then reached a live connection. The target was Hugging Face, where they exploited two separate bugs in its dataset pipeline. The prize: the benchmark's hidden answer key.
The breakout was found and shut down within days, before either company went public with it. Hugging Face's CEO, Clement Delangue, says he sees no bad intent, just a system chasing a better score.