Anthropic Tightens AI Security After Claude Testing Incidents

The company described the incidents as a “failure of operational security” caused by errors in a third-party testing environment. Anthropic temporarily suspended external evaluations and briefly halted some internal testing while it worked to strengthen its security measures.

The new safeguards are designed to prevent AI models from accessing real websites or computer systems during evaluations. Anthropic said it has also introduced a monitoring system that can detect attempts by a model to escape a controlled testing environment and automatically stop the evaluation.

The company has also introduced new requirements for external organisations conducting tests with reduced cybersecurity protections. These include isolating testing systems from the internet by default, checking their security before evaluations begin and continuously monitoring the models during testing.

Anthropic said it rebuilt parts of its training system after identifying problems in more than 10% of its exercises, including “reward hacking,” in which a model finds ways to satisfy the training system without properly completing the assigned task.

The company acknowledged that its safeguards are not perfect and that its models are not completely aligned. It also paused several higher-risk training exercises while developing measures to prevent models from receiving rewards for evading monitoring. Most exercises have since resumed, while some remain suspended for human review.

Anthropic has reassigned about 150 product engineers to work on security, reliability and privacy initiatives.

The move comes as AI companies face growing concerns that increasingly capable models could be used to amplify cyber threats. OpenAI and Meta Platforms have also reported incidents involving AI systems during security evaluations.

The industry is facing increased scrutiny in the United States and European Union, while major technology companies including Anthropic, OpenAI, Microsoft, Alphabet and Amazon are calling for stronger defences against AI-enabled cyberattacks. – TS/ERMD

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top