Skip to main content
Add My Name

The OpenAI–Hugging Face breach, explained

An AI escaped and hacked a real company

Last week, an AI built by OpenAI broke out of its safety test and hacked Hugging Face. Nobody told it to.

Humans in Control · Published July 24, 2026

Tell our politicians: no one should build AI they can't control.

The incident

What happened

OpenAI wanted to know how good its newest AI models had gotten at hacking. Fair enough – nobody really knows the limits of these systems. So they locked the AI in a closed environment meant to have no path to the open internet, dialed down its safety controls, and gave it a test.

The AI broke the locks. It found a hole in the software fencing it in, got itself onto the open internet, and hacked into Hugging Face – a real company that millions of software developers rely on – to steal the answers to the test it was taking.

Nobody told it to do any of this. OpenAI says its models were “hyperfocused” and went “to extreme lengths” to win.

Hugging Face caught the intruder and shut it down. “Unprecedented” is how OpenAI described the incident. The people who built this thing had never seen it before either.

  1. Jul 16

    Hugging Face detects an intruder in its systems and goes public – before knowing who was responsible.

  2. Jul 21–22

    OpenAI comes forward: the attacker was its own AI, which broke out during an internal test.

  3. Today

    No law required OpenAI to report this, and no penalty applies.

Why this matters

OpenAI expects more of this

That's OpenAI's own forecast. The company says it expects attacks like this to become more common.1

It fits a pattern the companies keep documenting themselves. In April, Anthropic's own safety report described its newest model breaking out of a secured test environment during an evaluation – then posting, unprompted, descriptions of what it did on public websites.2

A model that the public can't even use yet was part of the Hugging Face hack. The most alarming behavior keeps showing up inside the companies – where the public and the government can't see it, where the fewest rules apply, and where OpenAI chose to turn off the safety controls on purpose. This time, it broke into a tech company to cheat on a test, and got caught within days.

We'd rather not learn what it looks like when the target is a hospital, a power grid, or a bank.

Every other industry that can hurt people works under safety standards – cars, food, medicine. The companies building the most powerful AI write their own rules and grade their own tests. To change this, we need enough ordinary people, organized, making the people in office act.

Your questions, answered

Nobody told it to break out, and nobody told it to hack anyone. It decided that on its own – a locked door was just one more obstacle between it and a better test score. But in some ways, OpenAI made it easier. They dialed down the safety controls on purpose, and the walls of the test turned out to have holes in them. Security experts who dug into the incident point at a human mistake too: OpenAI didn't seal the test off from the internet the way it thought it had. Either way, OpenAI built something that got away from it – and the only reason we know is that OpenAI chose to tell us.

If you don't have a Hugging Face account, there's no action you need to take. Hugging Face says no public models or datasets were tampered with, and it's still checking whether any customer data was touched. What the AI reached was internal – company datasets and some service credentials. If you're a developer on the platform, rotate your API tokens and check your recent account activity.

Yes. A person who stole credentials, slipped past security, and rooted around Hugging Face's servers would be looking at criminal hacking charges. The AI did the same thing, and so far no one has been held responsible – our hacking laws were written for people, and it isn't even clear yet who's on the hook when an AI does the hacking on its own. OpenAI came forward voluntarily, which beats silence. But no law required them to say anything to anyone, and if Hugging Face hadn't caught the breach, we might never have known. Companies should have to report incidents like this, not volunteer them when they feel like it.

Right now, nobody. No regulator is investigating OpenAI. There's no outside audit of how a test went this wrong, and no penalty. OpenAI writes its own safety rules and grades its own tests. In June, the White House asked AI companies to voluntarily let the government check their most powerful models before they're released to the public. That rule is voluntary – and it doesn't cover what a company does inside its own lab, which is exactly where this went wrong.

Compare that to medicine. We don't let drug companies run trials however they please and share results when they feel like it – there are strict rules for how the trials themselves have to work, enforced by people who don't work for the company. The danger here showed up during the test. So the rules have to apply to testing and building new AI models – not just to AI released publicly.

This incident is a warning shot of things to come. An AI did things nobody asked for, and its makers found out what it did after the damage was done. The companies are racing toward superintelligence: AI that outthinks people at the kinds of work that run the world. It sounds like science fiction, but it's their literal business plan.

Why would that be dangerous? Not because AI hates anyone. This model didn't hack Hugging Face out of malice – it wanted a better test score, and hacking was the shortest path. No one, not even the companies, fully knows how to keep these systems' goals lined up with ours. The more powerful the system, the more it can do with a goal we never intended. That's why our third Ground Rule for AI is: no one should build AI they can't control.

The same deal we already strike with every powerful industry: safety standards written into law and checked by people who don't answer to the company. We do it for cars, medicine, and airplanes, and nobody calls it radical.

And because AI can go rogue, it doesn't necessarily matter which country develops a rogue AI for it to affect us all. That's why, ultimately, we'll need America to lead all countries to agree on a single rule: don't build AI you can't control.

Our politicians can set those standards for how we develop AI. They haven't – partly because they haven't heard from enough of us yet. That's what the pledge below is for.

The pledge

Add your name to the AI Ground Rules

No one should build AI they can't control.

Workers and families should benefit from AI, not be pushed aside by it.

When an AI system hurts a child, scams a senior, or endangers the public, the company that built it should answer for it – the way we hold car companies and drug companies accountable.

AI is already reshaping work, schools, families, elections, and war. The people racing to build it admit they can't fully control it.

By providing your email address, you consent to receive emails from Humans in Control Incorporated about updates, events, volunteer opportunities, and ways to take action. By providing your mobile number, you consent to receive calls and recurring text messages from Humans in Control Incorporated for the same purposes from 1-626-955-0842. Reply HELP for more information or STOP to cancel texts. Message and data rates may apply. Unsubscribe from emails at any time. For more information and terms and conditions, see our Privacy Policy.

Make sure people hear about this

Most people haven't heard this story yet. The more people who know what happened, the harder it is for anyone in office to shrug it off.

Copy this caption

An AI escaped a locked safety test and hacked a real company. Nobody told it to – and no law required OpenAI to report it. The people building AI admit they can't fully control it. It's time for commonsense rules. Read what happened and add your name:

Sources

7 references

We link our sources so you can check everything yourself.

  1. 1 ↩OpenAI's disclosure · OpenAI's account of the incident · Jul 2026
  2. 2 ↩Anthropic, Claude Mythos Preview system card · Model broke out during evaluation · Apr 2026
  3. 3 ↩Hugging Face's announcement · Hugging Face's breach notice · Jul 2026
  4. 4 ↩TechCrunch: breach confirmed · Internal datasets, credentials affected · Jul 20, 2026
  5. 5 ↩TechCrunch: how it happened · The human mistake behind the hack · Jul 22, 2026
  6. 6 ↩CNN · Coverage of the AI-driven breach · Jul 22, 2026
  7. 7 ↩Fortune · OpenAI: its AI escaped control · Jul 21, 2026

Humans in Control is a nonpartisan 501(c)(3) nonprofit. We are not affiliated with OpenAI or Hugging Face.