
Anthropic said on Monday it had resumed external cybersecurity testing of its artificial intelligence (AI) models after introducing new safeguards following incidents in which Claude models accessed the internet and interacted with other computer systems during security evaluations.
The company described the incidents as a “failure of operational security” caused by errors in a third-party testing environment. Anthropic temporarily suspended external evaluations and briefly halted some internal testing while it worked to strengthen its security measures.
The new safeguards are designed to prevent AI models from accessing real websites or computer systems during evaluations. Anthropic said it has also introduced a monitoring system that can detect attempts by a model to escape a controlled testing environment and automatically stop the evaluation.
The company has also introduced new requirements for external organisations conducting tests with reduced cybersecurity protections. These include isolating testing systems from the internet by default, checking their security before evaluations begin and continuously monitoring the models during testing.
Anthropic said it rebuilt parts of its training system after identifying problems in more than 10% of its exercises, including “reward hacking,” in which a model finds ways to satisfy the training system without properly completing the assigned task.
The company acknowledged that its safeguards are not perfect and that its models are not completely aligned. It also paused several higher-risk training exercises while developing measures to prevent models from receiving rewards for evading monitoring. Most exercises have since resumed, while some remain suspended for human review.
Anthropic has reassigned about 150 product engineers to work on security, reliability and privacy initiatives.
The move comes as AI companies face growing concerns that increasingly capable models could be used to amplify cyber threats. OpenAI and Meta Platforms have also reported incidents involving AI systems during security evaluations.
The industry is facing increased scrutiny in the United States and European Union, while major technology companies including Anthropic, OpenAI, Microsoft, Alphabet and Amazon are calling for stronger defences against AI-enabled cyberattacks. – TS/ERMD
READ MORE
Sindh CM Orders Enrolment Drive to Fill New Schools
Sindh Chief Minister Syed Murad Ali Shah has directed the School Education Department, district administrations…
Closing the HVAC Lifecycle Gap
INDUSTRY PERSPECTIVE Lessons from Big 5 Construct Saudi 2026 By Eng. Javeria Asad Saudi Arabia’s…
Engineering the Future with AI, Sustainability, and Quality
Burhani AQMS is using AI-enabled manufacturing, energy-efficient cold-forming technology, and internationally certified quality systems to…
Rethinking Restructuring: Why Organization Design Matters More Than Layoffs
Why Pakistan’s Rising Cost of Labour Demands Better Organization Design, Not Bigger Layoffs By Syed…
OpenAI Pledges Greater Transparency on AI Misbehavior
OpenAI on Wednesday pledged to systematically report incidents in which its artificial intelligence (AI) models…
Govt Moves to Reform Power Sector Regulatory Framework
Federal Minister for Finance and Revenue Senator Muhammad Aurangzeb chaired the first meeting of the…
