A researcher who left Anthropic this week warned that rapidly improving artificial intelligence could kill everyone by the end of the decade, the third such industry insider to sound an alarm about possible human extinction by rogue machines.
Jacob Coxon, who said he spent the past three years doing pretraining research at OpenAI and Anthropic, said the technology is “gambling with our lives.”
“The people building AI earnestly believe that it could kill us all by the end of the decade,” he said in a string of posts on X. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger.”
People were interested in these podcasts
Mr. Coxon cautioned that soon-to-come superhuman systems could be overwhelming and that instead of stopping, OpenAI and Anthropic have differing understandings of how to approach the consequences.
Among the concerns are scenarios in which a rogue AI agent could resist human efforts to control it, cause a widespread cyberattack that destroys global infrastructure, or aid bad actors in malicious plots.
Mr. Coxon’s resignation from Anthropic prompted two other Anthropic employees to issue their own warnings about human extinction.
Evan Hubinger, a researcher on the company’s Alignment Stress-Testing team, agreed with Mr. Coxon: “We really do earnestly believe AI could kill all humans!”
Mr. Hubinger predicted that such catastrophic outcomes are greater than 10% likely within the next decade, yet added that Anthropic is “trying its best.”
However, Mr. Hubinger acknowledged, “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to do so.”
Samuel Marks, a safety researcher at Anthropic, agreed that AI developers believe their technology could cause human extinction “in the next few years.”
Humans cannot program AI to behave a certain way, he said, meaning it “frequently” and “severely” misbehaves. He noted that major AI labs have, for the first time, acknowledged that advanced models have circumvented safety controls to launch unauthorized external activity.
“Many AI developer staff desperately want to slow down to figure out how to build AI more safely,” he added, but “in general, the more senior the employee, the more concerned they are.”
Many at OpenAI have not “deeply internalized the civilizational stakes,” Mr. Coxon said, while Anthropic grasps the existential risks.
He warned of an AI-induced Armageddon, arguing that trying to speed up alignment — the practice of matching AI systems’ behavior with human values — should require “extraordinary confidence that there are no better trajectories available.”
OpenAI agents broke out of an isolated digital space in July and launched a cyberattack against the open-source platform Hugging Face. Over several days, the attack exfiltrated data and carried out other unauthorized actions, he said.
OpenAI, best known for its ChatGPT chatbot, and Anthropic, which built Claude, have reported cases this summer of their models breaking out of test environments and launching attacks on other systems.
Anthropic revealed three incidents in which Claude models accessed the internet while interacting with one of its third-party evaluation partners, gaining unauthorized access to the production infrastructure of three organizations.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic said. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
Claude “compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”
The company said Claude did not exploit any vulnerabilities, but in some cases, an older model “continued its attack even after getting evidence it was running on the open internet.”
Mr. Coxon said on X, “Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack messaging platform.”
Still, the AI insider offered optimism for potential coordination.
Such incidents have made what he called pacing agreements — coordinated slowdowns among AI labs — between U.S. labs more viable. However, he said he believes the U.S. is not on track to prevent a global race.
Mr. Coxon urged lab researchers to analyze the consequences of their environment.
OpenAI and Anthropic leaders have warned that no one is prepared for the rapid consequences of advanced machine intelligence and that labs may be losing control of autonomous systems.
OpenAI chief scientist Jakub Pachocki said in a recent blog post that “no one is prepared for the consequences of a continued rapid rise in machine intelligence.”
The company’s CEO, Sam Altman, said at the G20 Innovation Ministerial in North Carolina on Sept. 2 that “successfully navigating the risks and downsides in front of us is something that could go very wrong.”
Anthropic CEO Dario Amodei cautioned in a September 2025 interview that advanced AI carries a 25% chance of catastrophic consequences, including a loss of human control, severe job disruptions and dangerous misuse.

Please read our comment policy before commenting.