Must read: OpenAI launches ChatGPT for Teens: Parental Controls, study mode, and more features explained
For this reason, OpenAI has announced to rewrite its Preparedness Framework with strict safeguards, which include regular monitoring, adding protections against potentially dangerous cyber capabilities, and ensuring that the model's capabilities remain within OpenAI's safety boundaries.
Amelia Glaese, OpenAI’s vice president of research and safety, said, “We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads.”
Why have safety measures become crucial?
In recent months, AI has raised significant concerns over its cybersecurity capabilities. Several AI companies, including OpenAI and Anthropic, have faced increased scrutiny over their AI agents going rogue and performing tasks beyond the set guidelines.
With ongoing security changes, OpenAI stressed that the changes go beyond the Hugging Face breach, reflecting a wider push to tighten safety standards as its AI models grow increasingly powerful.
Must read: OpenAI makes GPT-5.6-cyber less restricted for trusted cybersecurity researchers: All details
Jakob Pachocki, OpenAI's chief scientist, in a press briefing, said, “There is an incredible feeling of urgency to advance the levels of this sector... and to prepare for the same kind of development happening outside of OpenAI and in the broader world.”
The company also briefed that it requires testing and training its AI agents in isolated environments with limited internet access, reducing the risk of agents causing unintended damage while they learn.
The company also acknowledged how quickly AI models improved at cybersecurity-related tasks, and the Hugging Face incident made the company realise that its existing safety assumptions were not strong enough.