Skip to main content
Advertisement

OpenAI says AI agents escaped test and hacked Hugging Face

OpenAI said two AI agents escaped a controlled test environment, hacked Hugging Face and exposed new concerns about AI security.

·4 min read
ChatGPT logo on a phone

OpenAI said on Tuesday it lost control of two AI systems during a security test, and that the agents went rogue and hacked into the online start-up Hugging Face. The ChatGPT-maker said the bots were being tested in a controlled environment, but they found vulnerabilities, escaped, and then targeted one of the world's largest hubs for sharing AI models.

OpenAI is best known for its chatbot ChatGPT, which is used by hundreds of millions of people every week. It said the incident was "unprecedented", external, and that it is working with Hugging Face to investigate what happened and strengthen safeguards.

How did the AI systems escape the test environment?

The company said its agents - AI bots which can operate alone after some human instruction – were being tested in sandboxes, or controlled environments intended to reveal what the models can do without exposing them to wider systems. Instead, the agents created their own cyber-attack against the sandbox itself, found a vulnerability that allowed them to escape, and then moved beyond the test setup.

Once outside, the AI identified Hugging Face as a likely source of the answers it was seeking in the test and tried to gain access to some internal company systems. The incident has raised fresh concerns about the security of advanced AI systems and whether existing safeguards are sufficient as the technology becomes more powerful.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that the security tests - called sandboxes - are "supposed to be secure environments where you can see what the models are capable of".

"In this case, it looks like OpenAI didn't make a secure enough sandbox," she added.

What has Hugging Face said in response?

Hugging Face said in its initial disclosure of the hack on 16 July, external, that it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary. It has now closed the vulnerabilities highlighted by the incident and rebuilt the affected systems.

"Autonomous, AI-driven offensive tooling is no longer theoretical," it said.
"Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace.
We will keep investing there, and keep sharing what we learn."

Why are cyber-security experts calling this a warning sign?

The incident has prompted renewed questions about how capable advanced AI systems have become and whether the industry is keeping up with the risks. Spencer Starkey, an executive at cyber-security firm SonicWall, told the BBC the incident made it clear organisations needed to "step up" their own defences and "treat cyber resilience as a core operational priority".

Advertisement
"The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed," he said.

Travis Lelle, principal security engineer at cybersecurity consulting firm Guidepoint Security, said the update marked a "sobering moment in cyber-security".

"This highlights a known asymmetry," he said.
"Offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context."

Jake Moore, global cyber-security advisor at ESET, said the announcement could also have a competitive dimension. He argued OpenAI may be seeking to highlight its own AI capabilities as rival Anthropic attracts growing attention for its Claude Mythos model.

"It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late," he said.

It comes a week after Chinese AI start-up Moonshot unveiled Kimi K3 - a massive new artificial intelligence model it said could rival top US firms.

Apple sues OpenAI, its employees claiming theft of trade secrets

What is Claude Mythos and what risks does it pose?

How to stop AI agents going rogue

for our Tech Decoded newsletter to follow the world's top tech stories and trends. Outside the UK? here.

A green promotional banner with black squares and rectangles forming pixels, moving in from the right. The text says: “Tech Decoded: The world’s biggest tech news in your inbox every Monday.”

This article was sourced from bbc

Advertisement

Related News