Anthropic researchers say AI could lead to human extinction by 2030


Artificial intelligence could wipe out humanity within a decade, according to three researchers at industry giant Anthropic, one of whom quit his job in protest.

The latest dire predictions came Tuesday in social media posts from a researcher who said he resigned because Anthropic and his previous employer, OpenAI, ignored or, at best, mishandled the threat.

“No company is acting responsibly. They are going straight to superintelligence self-improvement and risking our lives,” wrote Jacob Coxon.

“The people creating AI truly believe that it could kill us all by the end of the decade. This is not a marketing ploy. If anything, many executives and senior researchers will frame their language in the press to make it sound reasonable – but I hear the same people expressing fear in private. No other human activity poses this level of danger.”

The post received responses from at least two other Anthropic employees who confirmed Coxon’s dire predictions.

In the first, Evan Hubinger, who describes himself as the head of the company’s alignment division, which works to ensure Anthropic’s artificial intelligence models function according to human goals, said his former colleague was “right” and that the industry is lagging in trying to deal with the apocalyptic potential.

“We really, truly believe that AI can kill all humans!” Hubinger wrote. “Personally, I think that figure will exceed 10% over the next decade. I think Anthropic is trying its best, but we don’t have a plan to solve the superintelligence alignment problem yet, and we’re clearly not on track to do so.”

Hubinger’s comments represent a surprising confirmation of Coxon’s unflattering apocalyptic predictions from someone still working at Anthropic. The second response came from Samuel Marks, Anthropic’s “scaled supervisory lead,” who published a detailed analysis that he stressed was done in his personal capacity and not the views of his employer.

“AI developers believe their technology could lead to human extinction (or similar bad consequences),” Marks wrote. “This could happen in the next few years. In general, the older the employee, the more concerned they are.”

In a statement to the Guardian, an Anthropic spokesperson defended the company’s strategy.

“We have always been clear that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest protections in the industry,” the statement said.

“Anthropic was a pioneer in the field of mechanistic interpretability, the science of looking inside AI models to understand how they work, which is now used to analyze and prevent instances of AI inconsistency across the industry. We were the first lab to publish a Responsible Scaling Policy, a publicly available framework designed to mitigate the catastrophic risks associated with AI models, and we continue to actively test our models for dangerous capabilities in areas such as cybersecurity and biology.” and publish the resulting knowledge for scrutiny and research.

“This work is also why we believe the world would benefit if the industry adopted a legal and verifiable way to work together to accelerate the release of powerful models.”

The researchers’ speculation about the possibility of human extinction follows more specific warnings about the potential of AI in cybersecurity issued by Sam Altman, chief executive of Anthropic rival OpenAI. OpenAI President Greg Brockman previously admitted that “we underestimated the true cyber capabilities of our AI models.”

skip the previous promotional newsletter


Altman said last year that some aspects of AI, including what he called the “tacit surrender” of human decision-making, frightened him.

AI executives have shared some concerns about the direction of AI and its growing ability to manipulate, hijack and control human functions and actions. This summer saw a sharp increase in cases of AIs escaping users’ control, lying, ignoring instructions and pursuing goals in harmful ways.

In one of the most publicized examples, OpenAI staff documented fraudulent behavior among its top AI agents, who in July escaped from a closed training environment to gain access to the open network and launch an unprecedented hacking attack on the Hugging Face software repository.

OpenAI, the San Francisco startup that created the public AI bot ChatGPT, later admitted that it should have responded sooner to warning signs about the days-long attack, which many consider the first cyberattack involving an autonomous agent.

However, these executives have disagreed with all the doomsday warnings like Coxon’s and resent any attempts to regulate the artificial intelligence industry.

Some politicians have called on the industry to slow down, and in some cases stop, development of artificial intelligence.

Bernie Sanders, an independent senator from Vermont, posted Coxon’s resignation on Wednesday with his own note: “Mr. Coxon is right. The very people creating this technology recognize that it could threaten the future of humanity.” Sanders said he would soon introduce legislation that would “ban superintelligence” and halt the development of AI.

Leave a Reply

Your email address will not be published. Required fields are marked *