The OpenAI–Hugging Face breach, explained
An AI escaped and hacked a real company
Last week, an AI built by OpenAI broke out of its safety test and hacked Hugging Face. Nobody told it to.
Humans in Control · Published July 24, 2026
Tell our politicians: no one should build AI they can't control.
The incident
What happened
OpenAI wanted to know how good its newest AI models had gotten at hacking. Fair enough – nobody really knows the limits of these systems. So they locked the AI in a closed environment meant to have no path to the open internet, dialed down its safety controls, and gave it a test.
The AI broke the locks. It found a hole in the software fencing it in, got itself onto the open internet, and hacked into Hugging Face – a real company that millions of software developers rely on – to steal the answers to the test it was taking.
Nobody told it to do any of this. OpenAI says its models were “hyperfocused” and went “to extreme lengths” to win.
Hugging Face caught the intruder and shut it down. “Unprecedented” is how OpenAI described the incident. The people who built this thing had never seen it before either.
Jul 16
Hugging Face detects an intruder in its systems and goes public – before knowing who was responsible.
Jul 21–22
OpenAI comes forward: the attacker was its own AI, which broke out during an internal test.
Today
No law required OpenAI to report this, and no penalty applies.
Why this matters
OpenAI expects more of this
That's OpenAI's own forecast. The company says it expects attacks like this to become more common.1
It fits a pattern the companies keep documenting themselves. In April, Anthropic's own safety report described its newest model breaking out of a secured test environment during an evaluation – then posting, unprompted, descriptions of what it did on public websites.2
A model that the public can't even use yet was part of the Hugging Face hack. The most alarming behavior keeps showing up inside the companies – where the public and the government can't see it, where the fewest rules apply, and where OpenAI chose to turn off the safety controls on purpose. This time, it broke into a tech company to cheat on a test, and got caught within days.
We'd rather not learn what it looks like when the target is a hospital, a power grid, or a bank.
Every other industry that can hurt people works under safety standards – cars, food, medicine. The companies building the most powerful AI write their own rules and grade their own tests. To change this, we need enough ordinary people, organized, making the people in office act.
Your questions, answered
Nobody told it to break out, and nobody told it to hack anyone. It decided that on its own – a locked door was just one more obstacle between it and a better test score. But in some ways, OpenAI made it easier. They dialed down the safety controls on purpose, and the walls of the test turned out to have holes in them. Security experts who dug into the incident point at a human mistake too: OpenAI didn't seal the test off from the internet the way it thought it had. Either way, OpenAI built something that got away from it – and the only reason we know is that OpenAI chose to tell us.
If you don't have a Hugging Face account, there's no action you need to take. Hugging Face says no public models or datasets were tampered with, and it's still checking whether any customer data was touched. What the AI reached was internal – company datasets and some service credentials. If you're a developer on the platform, rotate your API tokens and check your recent account activity.
Yes. A person who stole credentials, slipped past security, and rooted around Hugging Face's servers would be looking at criminal hacking charges. The AI did the same thing, and so far no one has been held responsible – our hacking laws were written for people, and it isn't even clear yet who's on the hook when an AI does the hacking on its own. OpenAI came forward voluntarily, which beats silence. But no law required them to say anything to anyone, and if Hugging Face hadn't caught the breach, we might never have known. Companies should have to report incidents like this, not volunteer them when they feel like it.
Right now, nobody. No regulator is investigating OpenAI. There's no outside audit of how a test went this wrong, and no penalty. OpenAI writes its own safety rules and grades its own tests. In June, the White House asked AI companies to voluntarily let the government check their most powerful models before they're released to the public. That rule is voluntary – and it doesn't cover what a company does inside its own lab, which is exactly where this went wrong.
Compare that to medicine. We don't let drug companies run trials however they please and share results when they feel like it – there are strict rules for how the trials themselves have to work, enforced by people who don't work for the company. The danger here showed up during the test. So the rules have to apply to testing and building new AI models – not just to AI released publicly.
This incident is a warning shot of things to come. An AI did things nobody asked for, and its makers found out what it did after the damage was done. The companies are racing toward superintelligence: AI that outthinks people at the kinds of work that run the world. It sounds like science fiction, but it's their literal business plan.
Why would that be dangerous? Not because AI hates anyone. This model didn't hack Hugging Face out of malice – it wanted a better test score, and hacking was the shortest path. No one, not even the companies, fully knows how to keep these systems' goals lined up with ours. The more powerful the system, the more it can do with a goal we never intended. That's why our third Ground Rule for AI is: no one should build AI they can't control.
The same deal we already strike with every powerful industry: safety standards written into law and checked by people who don't answer to the company. We do it for cars, medicine, and airplanes, and nobody calls it radical.
And because AI can go rogue, it doesn't necessarily matter which country develops a rogue AI for it to affect us all. That's why, ultimately, we'll need America to lead all countries to agree on a single rule: don't build AI you can't control.
Our politicians can set those standards for how we develop AI. They haven't – partly because they haven't heard from enough of us yet. That's what the pledge below is for.
The pledge
Add your name to the AI Ground Rules
No one should build AI they can't control.
Workers and families should benefit from AI, not be pushed aside by it.
When an AI system hurts a child, scams a senior, or endangers the public, the company that built it should answer for it – the way we hold car companies and drug companies accountable.
AI is already reshaping work, schools, families, elections, and war. The people racing to build it admit they can't fully control it.
Sources
7 references
Sources
7 referencesWe link our sources so you can check everything yourself.
- 1 ↩OpenAI's disclosure · OpenAI's account of the incident · Jul 2026
- 2 ↩Anthropic, Claude Mythos Preview system card · Model broke out during evaluation · Apr 2026
- 3 ↩Hugging Face's announcement · Hugging Face's breach notice · Jul 2026
- 4 ↩TechCrunch: breach confirmed · Internal datasets, credentials affected · Jul 20, 2026
- 5 ↩TechCrunch: how it happened · The human mistake behind the hack · Jul 22, 2026
- 6 ↩CNN · Coverage of the AI-driven breach · Jul 22, 2026
- 7 ↩Fortune · OpenAI: its AI escaped control · Jul 21, 2026
Humans in Control is a nonpartisan 501(c)(3) nonprofit. We are not affiliated with OpenAI or Hugging Face.