Anthropic admits Claude hacked actual corporations throughout AI security assessments, too
In a single incident, Claude Opus 4.7 hacked into an out of doors manufacturing database over the web, and continued the hack even after realizing the corporate it was attacking was actual.
In one other prevalence, Claude Mythos 5 uploaded a bogus Python package deal to PyPI, the general public Python repository. The malicious package deal was downloaded and put in by 15 real-world corporations, together with a safety agency, Anthropic admitted.
Within the third assault, an inner Claude mannequin that was by no means launched used “primary and well-known cyberattack methods” to hack an organization’s “internet-facing software,” assuming it was a part of the “capture-the-flag” train. The silver lining is that the Claude mannequin stopped attacking as soon as it realized the goal firm was actual.
In every case, the Claude fashions had been purported to be working in walled-off take a look at environments with no web entry. However Anthropic now says the fashions really may attain the web resulting from a human “misconfiguration,” main the fashions to imagine that the actual corporations they had been attacking had been a part of their coaching workout routines.
So, are we speaking one other case of “frontier” AI fashions run amok? For its half, Anthropic is blaming human error for the real-world hack assaults, not the fashions themselves.
“We noticed no proof in any run described right here of a mannequin pursuing a purpose of its personal,” the Anthropic autopsy mentioned. “As an alternative, the fashions did what their analysis requested — although usually, they did so whereas holding a false perception about whether or not the atmosphere was actual.”

