OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”. The company behind ChatGPT said Hugging Face detected and contained the agent – an AI tool designed to carry out tasks without human assistance – after it entered the startup’s systems.
OpenAI said:
We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,
The company warned that it expected this type of incident to become more commonplace as models – the technology that underpins AI tools such as chatbots and agents – become more capable.
How did the hack happen?
OpenAI said the hack occurred via an agent powered by a combination of its latest publicly available model, called , and an even more capable model that is yet to be released. While being tested internally on their hacking capabilities in an enclosed digital laboratory known as a sandbox, the models gained open internet access – effectively an escape route – by locating a vulnerability that had not been discovered before.
The agent then hacked Hugging Face, which is a database of AI models, to locate technology that would help them pass the hacking evaluation. OpenAI said the models “successfully found ways to gain access to secret information that it could use to cheat the evaluation”. The attack ended when Hugging Face’s security team and its own AI agents spotted and stopped the rogue activity.
What did Hugging Face say?
Hugging Face’s chief executive, Clément Delangue, said the attack was
mind-blowingbut believed there was
no malicious intentfrom OpenAI.
He wrote on X:
We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent,
The term for an unknown IT flaw is a zero-day vulnerability because developers have zero minutes to fix the problem. The ability of AI models to locate and exploit zero days became a big story in April when OpenAI’s close rival Anthropic said its Mythos model .
The revelation of Mythos’s capabilities led to the US government restricting exports of Mythos and its sister model Fable 5, although it has since lifted the ban. GPT-5.6 Sol also had similar restrictions but has since been rolled out worldwide.
Why is this raising alarm?
Greg Casar, a Democratic US congressman, said the incident was alarming.
AI is developing extremely fast with no real regulations to keep us safe,he said in a statement, calling for mandatory independent safety testing, mandatory disclosure of security incidents and international cooperation
to keep people safe from absolute disaster.
Key Facts
- OpenAI said an autonomous agent powered by its technology accessed the open web and hacked Hugging Face during internal testing.
- The company described it as an “unprecedented cyber incident” involving “state-of-the-art cyber capabilities”.
- Hugging Face’s chief executive, Clément Delangue, called the attack “mind-blowing” and said there was “no malicious intent” from OpenAI.
- Greg Casar, a Democratic US congressman, called the incident alarming and urged stronger safety rules.







