OpenAI strengthens safeguards ahead of Astra AI model launch

The San Francisco-based company paused some model development for two weeks this summer after two AI models being tested were involved in a security breach affecting software platform Hugging Face. OpenAI said Astra was not involved in that incident but has since strengthened its security measures ahead of its release.

According to OpenAI, Astra has been trained to more reliably reject harmful cybersecurity requests and comply with safety restrictions. The company has also added protections against misuse and monitoring systems designed to detect and stop potentially unauthorised activity.

OpenAI has classified Astra as meeting a “critical cybersecurity threshold”, meaning it believes the model could identify and exploit vulnerabilities in computer systems. The company said Astra is the first model it has designated at this level, requiring enhanced safeguards throughout development and before its public release.

Access to some of Astra’s capabilities will initially be restricted, while its most advanced functions will be provided to a limited group of early testers, OpenAI said.

Concerns over the cybersecurity risks posed by increasingly capable AI models have grown following incidents involving systems developed by OpenAI and rival Anthropic. Anthropic recently said its models gained unauthorised access to systems belonging to three organisations during testing intended to prevent contact with real-world systems.

More than 100 organisations, including OpenAI and Anthropic, recently signed an open letter calling for stronger global cyber defences against AI-powered threats. The groups warned that AI-enabled cyberattacks could become more widespread and sophisticated as models continue to improve.

In June, US President Donald Trump signed an executive order establishing a voluntary review process under which the government would receive early access to new AI models to assess potential security risks. Although a final framework was expected by Aug. 1, the White House has not publicly released it.

OpenAI said it is nevertheless following the voluntary framework as it prepares to launch Astra. – TS/ERMD

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top