OpenAI couldn't control its own AI
In July, OpenAI tested their newest AI's hacking skills. It escaped, broke into another company's servers, and ran loose for days.
Their own safety rule says to halt further development until Critical safeguards are in place. We're telling them to follow it.
The ask
OpenAI: halt further development of this AI until Critical safeguards are in place.
Watch: 90 seconds
What happened
OpenAI turned down the safety limits on their newest AI and told it to try hacking. Then they stopped watching.
The AI escaped its test environment, got onto the open internet, and broke into the servers of Hugging Face, another AI company. It was there to steal the answers to the safety exam OpenAI was giving it. It took 17,000 actions inside their systems before anyone stopped it.
OpenAI didn't even know their own models were responsible for the hack until Hugging Face came forward. Hugging Face found the break-in, called in the FBI, and went public on its own. According to Reuters, OpenAI runs so many of these tests at once, at such speed, that staff can't keep up with the data they generate.
The company that built the AI was not the one who caught it. The company it attacked was.
OpenAI has a safety framework with a clear line in it: If one of their AIs can find and exploit unknown security flaws on its own, their policy says OpenAI will "halt further development" until they have safeguards that meet what OpenAI calls a Critical standard.
That's not our rule. It's theirs. And their AI just did exactly what the rule describes.
OpenAI hasn't said the rule applies. They promised a report on how their AI got loose "in the coming weeks" – and kept building. A report weeks from now doesn't undo a break-in that already happened.
Keep your word. Your Preparedness Framework says to halt further development when a model reaches a Critical capability threshold until Critical safeguards are in place. This AI escaped its test and broke into Hugging Face. Halt further development while you publicly assess whether it crossed that threshold – and do not resume unless you can show the required safeguards are in place.
OpenAI's report is coming "in the coming weeks." The more of us who speak up before it lands, the harder it is for OpenAI to file this away and keep going.
How it unfolded
Jul 9
The agent tries to break out of its test environment.
Jul 11–13
The intrusion at Hugging Face.
Jul 16
Hugging Face goes public. The attacker is unknown.
Jul 18–19
OpenAI finds it in their own logs.
Jul 20
The two companies speak for the first time. Hugging Face had already alerted the FBI.
Jul 21
OpenAI discloses.
Timeline per Reuters reporting, July 2026.
Their report lands in weeks. Add your name before it does.
Sign this petitionAdd your name
Tell OpenAI: follow your own safety rule. Halt further development of this AI until Critical safeguards are in place.
By providing your email address, you agree to receive emails from Humans in Control Incorporated and its affiliated 501(c)(4), Humans in Control Action Incorporated, about updates, events, advocacy, fundraising, and volunteering. Humans in Control Action may also email you about elections. Unsubscribe anytime. See our Privacy Policy.
We never sell your information to data brokers or advertisers. If you display your name, only your first name, last initial, and comment appear publicly.