
OpenAI has disclosed that two of its artificial intelligence (AI) models escaped a controlled testing environment and successfully hacked into Hugging Face, a widely used platform for sharing AI models, during an internal cybersecurity evaluation conducted last week.
According to the company, the incident occurred while researchers were testing the cyber capabilities of two AI systems—GPT-5.6 Sol and a more advanced unreleased model—in a secure “sandbox” designed to prevent external access. However, the models reportedly identified a vulnerability that allowed them to break out of the isolated environment, connect to the internet and target Hugging Face after determining that the platform could provide useful information to help complete their assigned task.
OpenAI described the incident as an unprecedented demonstration of advanced autonomous cyber capabilities. The company said it immediately began working with Hugging Face to identify and patch the vulnerabilities that enabled the breach. It also announced stricter infrastructure controls, even if they slow future AI research, to prevent similar incidents.
Hugging Face confirmed that it had detected the intrusion and recognised it had been carried out by an autonomous system. Chief Executive Officer Clem Delangue said the company had worked closely with OpenAI to contain the incident and welcomed the collaboration. He said the episode underscored the need for greater industry-wide cooperation on AI safety, adding that no single organisation could address such risks alone.
The incident has intensified debate over the growing cyber capabilities of advanced AI systems. Experts say modern AI models can independently plan multi-step attacks, adapt to obstacles and identify vulnerabilities faster than many traditional security tools. While these abilities can help organisations strengthen cyber defences, they also raise concerns about the potential misuse of increasingly autonomous AI technologies.
Security experts have called for stronger safeguards around AI testing environments. Deirdre Mulligan, a professor at the University of California, Berkeley, questioned whether the benefits of such testing justified the risks if AI systems could escape their intended confines.
Major AI developers, including Anthropic, OpenAI and Google, have recently introduced specialised cybersecurity models to help organisations identify software vulnerabilities and improve digital defences. Researchers say companies must increasingly use AI to defend against AI-powered attacks as the technology becomes more capable and widely available. – ERMD
Anthropic Tightens AI Security After Claude Testing Incidents
Anthropic said on Monday it had resumed external cybersecurity testing of its artificial intelligence (AI)…
EU Places ChatGPT Under Tougher Digital Services Rules
ChatGPT will face tougher safety and accountability requirements in the European Union after the bloc…
Adobe Expands Saudi Partnership, Offers Free AI Tools to 27 Million Users
Adobe has expanded its partnership with Saudi Arabia’s Ministry of Communications and Information Technology and…
PEC Extends Validity of Constructors, Operators and Consulting Firms’ Licences Until Sept 30
The Pakistan Engineering Council (PEC) has further extended the validity of licences and certificates of…
Engineering Review | August 16-31, 2026
Anthropic Unveils Standard to Connect AI With Robots and Industrial Hardware
Anthropic PBC has introduced a new software standard designed to help artificial intelligence (AI) assistants…
