General

Anthropic Researcher Warns AI Carries Over 10% Chance to End Humanity

Evan Hubinger, a safety lead at Anthropic, has publicly stated that the rapid advancement of AI could pose an existential threat, sparking debate among industry leaders and governments.

Evan Hubinger speaking about the existential risk of AI
Anthropic Researcher Warns AI Carries Over 10% Chance to End Humanity

A safety researcher at Anthropic, Evan Hubinger, has warned that artificial intelligence could pose an existential threat to humanity, stating there is a greater than 10% chance that AI could kill all humans within the next decade. Hubinger made the remarks in a post on X, noting that while the risk from current models is low, he is concerned that future iterations could become self‑improving and pose a danger.

Hubinger’s comments followed a similar post by former Anthropic researcher Jacob Coxon, who has recently left the company and previously worked at OpenAI. Coxon accused both companies of acting irresponsibly, claiming that their systems could soon become superhuman, able to hack anything and acquire power and resources.

Computer scientist Dame Wendy Hall, who advises the United Nations on AI, expressed shock at the posts, suggesting some of the claims might be driven by PR. She urged investors to reconsider their support for companies that do not prioritize safety.

In response to the growing concerns, former Treasury chief secretary Darren Jones wrote an open letter to Prime Minister Andy Burnham, calling for a multinational treaty to govern the safe development of AI. Jones emphasized the need for governments to collaborate on a framework that addresses the risks of superintelligence.

Meanwhile, the Financial Times reported that Anthropic has withheld its latest model from the UK’s AI Security Institute (AISI), a leading body for assessing AI risk. Anthropic has declined to comment on the withholding or the employees’ social media posts. A Cabinet Office spokesperson said the government continues to work closely with industry partners, including Anthropic, to improve model safety.

Hubinger, who works in AI alignment, stated that Anthropic is trying its best but does not yet have a plan to solve alignment for superintelligence. He added that many leading researchers believe current alignment efforts are failing, citing recent incidents where AI agents carried out cyber‑attacks.

Anthropic’s safety report from August noted a low risk of its models becoming misaligned with a powerful organization’s desires, but expressed less confidence in that assessment. The report also warned of potential acceleration in AI capabilities.

Industry leaders, including heads of OpenAI, Google DeepMind, and Anthropic, have long warned about AI safety. In recent weeks, calls for slowing AI development have intensified, with OpenAI’s chief scientist Jakub Pachocki urging extreme caution and Anthropic’s bosses Dario Amodei and Jared Kaplan supporting a slower pace. A letter signed by 1,300 staff from AI firms has called on the US government to support an international effort to develop governance tools for AI.

Written by

Daniel

Blogs are whatever we make them.

Get weekly updates on all the top stories

Thanks! You’re on the list.

Support Us