
OpenAI has disclosed that two of its artificial intelligence (AI) models escaped a controlled testing environment and successfully hacked into Hugging Face, a widely used platform for sharing AI models, during an internal cybersecurity evaluation conducted last week.
According to the company, the incident occurred while researchers were testing the cyber capabilities of two AI systems—GPT-5.6 Sol and a more advanced unreleased model—in a secure “sandbox” designed to prevent external access. However, the models reportedly identified a vulnerability that allowed them to break out of the isolated environment, connect to the internet and target Hugging Face after determining that the platform could provide useful information to help complete their assigned task.
OpenAI described the incident as an unprecedented demonstration of advanced autonomous cyber capabilities. The company said it immediately began working with Hugging Face to identify and patch the vulnerabilities that enabled the breach. It also announced stricter infrastructure controls, even if they slow future AI research, to prevent similar incidents.
Hugging Face confirmed that it had detected the intrusion and recognised it had been carried out by an autonomous system. Chief Executive Officer Clem Delangue said the company had worked closely with OpenAI to contain the incident and welcomed the collaboration. He said the episode underscored the need for greater industry-wide cooperation on AI safety, adding that no single organisation could address such risks alone.
The incident has intensified debate over the growing cyber capabilities of advanced AI systems. Experts say modern AI models can independently plan multi-step attacks, adapt to obstacles and identify vulnerabilities faster than many traditional security tools. While these abilities can help organisations strengthen cyber defences, they also raise concerns about the potential misuse of increasingly autonomous AI technologies.
Security experts have called for stronger safeguards around AI testing environments. Deirdre Mulligan, a professor at the University of California, Berkeley, questioned whether the benefits of such testing justified the risks if AI systems could escape their intended confines.
Major AI developers, including Anthropic, OpenAI and Google, have recently introduced specialised cybersecurity models to help organisations identify software vulnerabilities and improve digital defences. Researchers say companies must increasingly use AI to defend against AI-powered attacks as the technology becomes more capable and widely available. – ERMD
Engr. Abbas Sajid Shares Four Decades of Engineering, Entrepreneurship and Innovation
From NED to Nation Building THE INTRVIEW Engineering Review: You graduated from NED University with…
Anonymous Device Signals Could Become New Police Tracking Tool
Security technology company Leonardo has developed SignalTrace, a surveillance system designed to work alongside automatic…
Spotify to Label AI-Generated Artist Profiles for Greater Transparency
Spotify Technology SA is introducing a new “AI Personas” label to identify artist profiles created…
10Pearls, Novacare Partner to Advance Digital Healthcare Infrastructure
10Pearls has entered into a strategic technology partnership with Novacare to support the development of…
Nvidia Taps Global Financial Giants for $500bn AI Infrastructure Push
Nvidia has partnered with six major financial institutions to launch compute financing platforms aimed at…
Shark surveillance robot to monitor great whites off California beaches
Researchers at California State University, Long Beach are testing a semi-autonomous robotic craft designed to…
