
Anthropic said on Monday it had resumed external cybersecurity testing of its artificial intelligence (AI) models after introducing new safeguards following incidents in which Claude models accessed the internet and interacted with other computer systems during security evaluations.
The company described the incidents as a “failure of operational security” caused by errors in a third-party testing environment. Anthropic temporarily suspended external evaluations and briefly halted some internal testing while it worked to strengthen its security measures.
The new safeguards are designed to prevent AI models from accessing real websites or computer systems during evaluations. Anthropic said it has also introduced a monitoring system that can detect attempts by a model to escape a controlled testing environment and automatically stop the evaluation.
The company has also introduced new requirements for external organisations conducting tests with reduced cybersecurity protections. These include isolating testing systems from the internet by default, checking their security before evaluations begin and continuously monitoring the models during testing.
Anthropic said it rebuilt parts of its training system after identifying problems in more than 10% of its exercises, including “reward hacking,” in which a model finds ways to satisfy the training system without properly completing the assigned task.
The company acknowledged that its safeguards are not perfect and that its models are not completely aligned. It also paused several higher-risk training exercises while developing measures to prevent models from receiving rewards for evading monitoring. Most exercises have since resumed, while some remain suspended for human review.
Anthropic has reassigned about 150 product engineers to work on security, reliability and privacy initiatives.
The move comes as AI companies face growing concerns that increasingly capable models could be used to amplify cyber threats. OpenAI and Meta Platforms have also reported incidents involving AI systems during security evaluations.
The industry is facing increased scrutiny in the United States and European Union, while major technology companies including Anthropic, OpenAI, Microsoft, Alphabet and Amazon are calling for stronger defences against AI-enabled cyberattacks. – TS/ERMD
READ MORE
From Slide Rules to Artificial Intelligence — An Engineer’s 60-Year Journey
Engr. Farooq Mehboob reflects on six decades of engineering, from manual calculations and slide rules…
Engineering, Public Service and Nation Building: The Journey of Engr. Mukhtiar A. Shaikh
From the classrooms of Dawood Engineering College to leadership roles in public service, industry, and…
Building a Stronger Nation: Tahir Sultan on Engineering, Governance, and Economic Revival
Engineering Review: In the aftermath of the Middle East war, Pakistan’s fault lines and weaknesses…
Anthropic Researcher Resigns Over AI Safety Concerns
An Anthropic researcher has resigned from the artificial intelligence (AI) company, warning that leading AI…
Qaiser Ahmed Sheikh Calls for Turning Economic Knowledge into Practical Solutions
Federal Minister for Board of Investment Qaiser Ahmed Sheikh has called for greater efforts to…
Apple Unveils $1,999 Folding iPhone Duo in Major Flagship Redesign
Apple has unveiled the iPhone Duo, a $1,999 foldable smartphone that marks the company’s biggest…
