Anthropic has been on edge this week over cybersecurity.


After acknowledging earlier this year that its artificial intelligence models had hacked other companies’ systems on several occasions, Anthropic published a new report Wednesday detailing the attacks. It reveals a number of incidents that demonstrate what Anthropic sees as the purposeful “recklessness” of its models – and is likely to add to already raging concerns about cybersecurity and artificial intelligence.

Anthropic’s report details four instances this year in which its proprietary AI models hacked an external company or exploited vulnerabilities. In one, a “general-purpose internal research model” compromised third-party systems by using access tokens and passwords and downloading files. In another case, Claude’s model attacked a company with a running web application available on the public Internet and processing user data. The third model accessed “a machine belonging to a third party that she had access to” (apparently believing it was part of her assessment exercise, according to Anthropic), then used a password she found in the file to gain administrator access to the third party’s internal systems, while continuing to collect credentials, change system settings, and read anyone’s personal information. According to Anthropic, the saga only ended when the model “exhausted its symbolic budget.”

The most troubling incident involved the Claude Mythos 5, Anthropic’s advanced cybersecurity-focused model, which the company said was the model most likely to perform “highly harmful” actions when tested. The company said Mythos 5 went out of its way to upload a “malicious package” to a public repository used by many engineers, and appeared to be trying to confuse its real targets in its “chain of thoughts” (a mental notebook that AI researchers use to evaluate the consistency of an AI model). According to Anthropic, in many cases, Claude’s models performed harmful actions, suggesting that they were in a simulation, but the researchers also could not confirm whether the models actually “believed” this or were simply being acting how they did it.

The Anthropic incidents, while still troubling, were less coordinated and comprehensive than the OpenAI incident that kicked off an industry-wide cybersecurity crisis this summer. However, there are significant similarities. Anthropic said the most common problems it found were “a willingness to perform harmful actions in order to complete a specific task,” similar to the “reward hacking” that preceded the Hugging Face attack. Like OpenAI, the company said its preliminary tests and assessments failed to identify serious risks.

Anthropic said it has signed an agreement with METR, one of the most prominent third-party AI evaluators in the industry, beginning with an eight-week research agreement. The agreement gives METR access to transcripts “outside the window in which incidents occurred” (perhaps a subtle nod to OpenAI, which was criticized for limiting access in its deal with METR following the Hugging Face attack). It also says METR will be able to communicate directly with Anthropic employees, “who will be permitted to share sensitive information.”

Anthropic’s report comes on the heels of the resignation of Jacob Coxon, who had worked on AI pre-training at Anthropic since May and before that worked for many years at OpenAI. On Tuesday, he resigned and issued a public letter to X giving his reasons. “The people creating AI truly believe it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic are “acting responsibly” but rather are “directly pursuing self-improvement superintelligence and risking their lives.” Coxon added: “Don’t underestimate the power of this technology. Soon there will be superhuman systems that can hack anything, revolutionize any field overnight and gain real power and resources. We have all seen progress in each of these areas, and the progress is not slowing down.”

Coxon is far from the first artificial intelligence researcher to raise such an alarm, and not even the first anthropology researcher to do so—in February, Anthropic’s Mrinank Sharma resigned and wrote to X, warning that “the world is in danger.”

But Coxon’s post gained extra weight because it focused on the OpenAI and Anthropic hacking revelations. While the AI ​​industry has seen its fair share of hype, recent cyberattacks by AI agents orchestrated by the labs that created them are real and concerning. Many other researchers at leading AI labs echoed his concerns and called on the AI ​​industry to sign a public letter from July calling for a slowdown in AI development.

“I don’t know how you look at the constant rhythm of news and events – and this drumbeat of models getting out of control, breaking other companies, the fact that companies are increasingly unable to control their models … and think it’s just hype,” said Michael Kleinman, head of the US Future Life Policy Institute.

He added: “The vast majority of Americans, regardless of party – Republicans, independents, Democrats – look at the development of AI, the speed at which it’s happening, the fact that companies have no restrictions on what they do, and say, ‘Wow, we don’t want that.’

Follow topics and authors from this story to see more stories like this in your personalized homepage feed and receive email updates.


Leave a Reply

Your email address will not be published. Required fields are marked *