A warning about Anthropic AI from one of the company’s top safety researchers is raising fresh questions about how quickly artificial intelligence is advancing and whether humans will be able to control what comes next. Evan Hubinger, Anthropic’s alignment science lead, said he believes there is more than a 10% chance that AI could cause human extinction within the next decade. While Hubinger stressed that today’s models pose a low risk, he fears future systems capable of improving themselves could become far more dangerous.His comments come amid growing concern from researchers working inside some of the world’s most powerful AI companies.
Warning About Anthropic AI Raises a Terrifying Possibility
Hubinger made the startling assessment in a post on X while responding to Jacob Coxon, an Anthropic researcher and former OpenAI employee who announced he was leaving the AI industry.
Hubinger said he was “worried” about where increasingly capable systems could lead and wrote that researchers at Anthropic genuinely believe artificial intelligence could eventually pose an existential threat to humanity. “We really do earnestly believe” that AI poses a species-ending risk, Hubinger wrote.
His warning came with an important distinction. Hubinger did not claim that existing AI systems are currently capable of wiping out humanity, nor did he explain a specific scenario through which AI would cause human extinction.
Instead, his concern centers on what could happen as the technology becomes significantly more capable, particularly if AI reaches the point where it can automate research and contribute to the development of even more advanced systems. Hubinger works in AI alignment, a field focused on ensuring powerful AI systems behave in ways consistent with human intentions and values. His position makes another part of his warning particularly significant.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he wrote. The concern, therefore, isn’t simply that AI is becoming smarter. It is whether researchers can reliably control systems if their capabilities eventually surpass those of humans in important areas.
Another Anthropic Researcher Walks Away From the AI Race
Hubinger’s comments followed an even more dramatic move from Coxon, who quit Anthropic. Coxon said he no longer wanted to participate in what he sees as an industrywide race toward powerful, self-improving AI.
“Neither company is acting responsibly,” Coxon wrote of Anthropic and OpenAI. He warned that future systems could become capable of hacking targets, rapidly advancing scientific fields and acquiring meaningful resources and influence.
Coxon told The Wall Street Journal that he believes competitive pressure between AI companies makes safety compromises increasingly difficult to avoid. He also acknowledged that he believes Anthropic takes safety seriously, but concluded that developing artificial general intelligence responsibly may require coordinated slowdowns or government intervention.
That creates an uncomfortable contradiction for the industry. Anthropic has built much of its identity around developing AI safely, yet researchers inside the company are publicly questioning whether even a safety-focused organization can adequately manage the technology’s accelerating capabilities.
The debate is unfolding as Anthropic also faces scrutiny over external safety testing. The Financial Times reported that the company did not give the U.K.’s AI Security Institute pre-release access to its latest model, prompting concern among U.K. officials.
The development is separate from Hubinger’s extinction-risk estimate and does not prove that Anthropic’s newest system is unsafe. However, it raises further questions about independent oversight as frontier models become increasingly powerful.
AI Insiders Are Calling for a Way to Slow Things Down
Hubinger and Coxon aren’t alone in expressing concern about the speed of AI development. A statement called Pacing the Frontier has been signed by more than 1,300 employees of frontier AI companies, including prominent figures associated with Anthropic, OpenAI, Google DeepMind and Meta.
Rather than demanding that artificial intelligence development permanently stop, the signatories want governments and the industry to develop mechanisms that could deliberately slow frontier AI progress if capabilities begin advancing faster than society can safely manage.
The statement warns that companies may be approaching the ability to automate AI research itself. If that happens, AI could potentially contribute to improving future AI systems, accelerating technological progress in ways that become increasingly difficult for researchers and regulators to follow.
The signatories argue that individual companies face competitive pressure not to slow down on their own. They have called on the U.S. government to support an international effort to develop technical and governance tools capable of pacing frontier development when necessary.
Among the names attached to the statement are Anthropic CEO Dario Amodei and co-founder Jared Kaplan, OpenAI chief scientist Jakub Pachocki and Google DeepMind co-founder Shane Legg, showing that concerns about controlling rapidly advancing AI extend well beyond a handful of researchers.
Still, Hubinger’s greater-than-10% figure should be understood for what it is: an individual researcher’s estimate of an uncertain future risk, not evidence that humanity has a 10% scientifically established probability of extinction from AI.
What makes the Anthropic AI warning notable is where it is coming from. With Piers Morgan’s post on late Professor Stephen Hawking’s concern about AI threat, the people sounding the alarm aren’t merely outside critics of artificial intelligence. Some are helping build, study and safeguard the most advanced systems themselves and they are increasingly asking whether the technology could advance faster than humans learn how to control it.



