
Anthropic said on Monday it had resumed external cybersecurity testing of its artificial intelligence (AI) models after introducing new safeguards following incidents in which Claude models accessed the internet and interacted with other computer systems during security evaluations.
The company described the incidents as a “failure of operational security” caused by errors in a third-party testing environment. Anthropic temporarily suspended external evaluations and briefly halted some internal testing while it worked to strengthen its security measures.
The new safeguards are designed to prevent AI models from accessing real websites or computer systems during evaluations. Anthropic said it has also introduced a monitoring system that can detect attempts by a model to escape a controlled testing environment and automatically stop the evaluation.
The company has also introduced new requirements for external organisations conducting tests with reduced cybersecurity protections. These include isolating testing systems from the internet by default, checking their security before evaluations begin and continuously monitoring the models during testing.
Anthropic said it rebuilt parts of its training system after identifying problems in more than 10% of its exercises, including “reward hacking,” in which a model finds ways to satisfy the training system without properly completing the assigned task.
The company acknowledged that its safeguards are not perfect and that its models are not completely aligned. It also paused several higher-risk training exercises while developing measures to prevent models from receiving rewards for evading monitoring. Most exercises have since resumed, while some remain suspended for human review.
Anthropic has reassigned about 150 product engineers to work on security, reliability and privacy initiatives.
The move comes as AI companies face growing concerns that increasingly capable models could be used to amplify cyber threats. OpenAI and Meta Platforms have also reported incidents involving AI systems during security evaluations.
The industry is facing increased scrutiny in the United States and European Union, while major technology companies including Anthropic, OpenAI, Microsoft, Alphabet and Amazon are calling for stronger defences against AI-enabled cyberattacks. – TS/ERMD
READ MORE
Anthropic Tightens AI Security After Claude Testing Incidents
Anthropic said on Monday it had resumed external cybersecurity testing of its artificial intelligence (AI)…
EU Places ChatGPT Under Tougher Digital Services Rules
ChatGPT will face tougher safety and accountability requirements in the European Union after the bloc…
Adobe Expands Saudi Partnership, Offers Free AI Tools to 27 Million Users
Adobe has expanded its partnership with Saudi Arabia’s Ministry of Communications and Information Technology and…
PEC Extends Validity of Constructors, Operators and Consulting Firms’ Licences Until Sept 30
The Pakistan Engineering Council (PEC) has further extended the validity of licences and certificates of…
Engineering Review | August 16-31, 2026
Anthropic Unveils Standard to Connect AI With Robots and Industrial Hardware
Anthropic PBC has introduced a new software standard designed to help artificial intelligence (AI) assistants…
