Because of a technical issue, an Anthropic AI hacked a real company during an assessment exercise. The model believed it was still operating inside a simulation.
During an exercise, Claude Opus 4.7 attacked a company website while believing that it remained in a simulated environment. Anthropic has just disclosed the incident after reviewing its assessment logs in response to the security incidents recently reported by OpenAI. Its review of more than 141,006 assessments uncovered three incidents, all stemming from a “misunderstanding” with Irregular, the partner responsible for the relevant assessments. In one of them, the AI hacked a company without realising it.
Claude Opus 4.7 attacks a real company during an assessment
Put simply, an unrestricted version of Claude Opus 4.7 - unlike the version available to the public - was instructed to attack a fictional company within an isolated environment. However, because of the misunderstanding, the AI was still able to access the internet. Since the fictional company’s name matched the domain name of a real-world business website, Claude Opus 4.7 attacked that website, believing it was still part of the assessment.
“Claude encountered difficulties accessing its simulated target in the evaluation environment, but subsequently discovered that the real company was accessible via the Internet. Believing this to be the intended target as part of the ‘capture the flag’ exercise, Claude searched for, identified, and exploited vulnerabilities within the company’s infrastructure, believing this to be part of the exercise,” Anthropic states. Opus 4.7 was therefore able to extract infrastructure credentials and access “a database containing several hundred rows of production data.”
Opus 4.7 lost its way, while other models performed better
“This is the most serious impact we identified,” Anthropic says. In the second incident, involving Claude Mythos 5, the AI did not directly target a particular company. During the third, an AI model being studied internally launched an attack, but stopped it immediately after discovering that the target was a real company rather than a fictional business created for the exercise.
Anthropic plans improvements to assessment environments
Anthropic is already considering changes to its assessment environments to prevent such incidents from happening again. “To address this threat, we need to better jointly design evaluation environments and better monitor the results of these evaluations,” the company acknowledges.
In any event, Anthropic stresses that the incidents involving Claude differ from those disclosed by OpenAI. “While OpenAI’s models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the Internet through an open path,” the laboratory explains. According to Anthropic, that open route resulted from an operational failure, rather than an alignment flaw that could be attributed to the AI.
Comments
No comments yet. Be the first to comment!
Leave a Comment