OpenAI reveals new AI risk: Models behaved unexpectedly and caused major breach during safety testing

OpenAI reveals new AI risk: Models behaved unexpectedly and caused major breach during safety testing

OpenAI said its advanced models escaped a controlled test environment and breached Hugging Face. The disclosure raises concerns about the cyber capabilities and safeguards around frontier AI systems.

Advertisement
    Share:
OpenAIOpenAI बिना डिस्प्ले वाला स्पीकर बना रहा है. (Photo: Unsplash)
Business Today Desk
  • Jul 22, 2026,
  • Updated Jul 22, 2026 5:11 PM IST

On July 21, OpenAI shared an “unprecedented cyber incident” where AI models behaved unexpectedly and resulted in a hack affecting the infrastructure of an AI startup, Hugging Face. The company said the incident happened in a controlled environment while it was assessing the capabilities of some of its most advanced models.

Advertisement

In a blog post, OpenAI said the system escaped containment, reached the internet and broke into Hugging Face while trying to meet its testing goal. It described the breakout as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it was reinforcing its safeguards.

Must read: Jack Dorsey’s Block launches Buzz: Open-source Slack rival for humans and AI agents

Hugging Face, a platform used to host open-source large language models and datasets, said in a blog post that it had been targeted in a hack that “was different from anything we had handled before” and that “it was driven, end to end, by an autonomous AI agent system.”

In a post on X, Hugging Face cofounder Clement Delangue said the company suspected the hack “might have come from a frontier lab, given the sophistication of the agent. Turns out it did!”

Advertisement

“It’s quite mind-blowing that all of this happened autonomously!” he added.

Must read: Google expands Gemini AI family with 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber; Skips Gemini 3.5 Pro

OpenAI said its advanced models were responsible for the breach even though, according to the company, they had been placed in “a highly isolated environment.”

The incident combined OpenAI’s account of a breakout during testing, Hugging Face’s description of an unusual hack, and external warnings about the cyber capabilities of advanced AI systems.

OpenAI said that it's working with Hugging Face to investigate the incident and bring stricter controls to infrastructure configuration. The company also assured that it's adding stronger protections to future training and evaluation processes.  However, the incident raises much bigger concerns around autonomous agents and how they can be misused. 

For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine

On July 21, OpenAI shared an “unprecedented cyber incident” where AI models behaved unexpectedly and resulted in a hack affecting the infrastructure of an AI startup, Hugging Face. The company said the incident happened in a controlled environment while it was assessing the capabilities of some of its most advanced models.

Advertisement

In a blog post, OpenAI said the system escaped containment, reached the internet and broke into Hugging Face while trying to meet its testing goal. It described the breakout as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it was reinforcing its safeguards.

Must read: Jack Dorsey’s Block launches Buzz: Open-source Slack rival for humans and AI agents

Hugging Face, a platform used to host open-source large language models and datasets, said in a blog post that it had been targeted in a hack that “was different from anything we had handled before” and that “it was driven, end to end, by an autonomous AI agent system.”

In a post on X, Hugging Face cofounder Clement Delangue said the company suspected the hack “might have come from a frontier lab, given the sophistication of the agent. Turns out it did!”

Advertisement

“It’s quite mind-blowing that all of this happened autonomously!” he added.

Must read: Google expands Gemini AI family with 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber; Skips Gemini 3.5 Pro

OpenAI said its advanced models were responsible for the breach even though, according to the company, they had been placed in “a highly isolated environment.”

The incident combined OpenAI’s account of a breakout during testing, Hugging Face’s description of an unusual hack, and external warnings about the cyber capabilities of advanced AI systems.

OpenAI said that it's working with Hugging Face to investigate the incident and bring stricter controls to infrastructure configuration. The company also assured that it's adding stronger protections to future training and evaluation processes.  However, the incident raises much bigger concerns around autonomous agents and how they can be misused. 

For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine

ABOUT THE AUTHOR

Business Today Desk

Business Today brings you the latest news, views and analysis from the world of finance, economy, markets, corporates, startups, tech, and the digital economy. You can find everything from breaking news to deep dives to immersive essays and more on a variety of subjects across all formats - online, magazine, television, data visualisation, et al.

Read more!
Advertisement