Anthropic Researcher Quits, Warns AI Could Kill Us

Anthropic Researcher Quits, Warns AI Could Kill Us

An abandoned workspace with a closed laptop and empty chair facing urban buildings through large windows.
An empty desk at a technology company office, marking a researcher's departure over safety concerns.Illustration: The Frank

The News

An Anthropic researcher quit on Tuesday and accused his employer and OpenAI of racing toward AI that humans cannot control.

Jacob Coxon, 27, announced the resignation in posts on X after The Wall Street Journal reported his departure Tuesday evening. He said he had spent the past three years doing pre-training research at the two labs.

"Neither company is acting responsibly," Coxon wrote. "They are racing straight to self-improving superintelligence and gambling with our lives."

His claim that people inside the labs privately fear catastrophe was then publicly confirmed by a current Anthropic team lead.

Timeline

2023

Coxon joins OpenAI's technical staff, where his research includes work on GPT-4o, according to Business Insider.

2024

Former OpenAI alignment chief Jan Leike quits, saying he hit a "breaking point" with leadership and that "safety culture and processes have taken a backseat to shiny products."

February 2026

Anthropic safeguards researcher Mrinank Sharma leaves, writing in a public resignation letter that "we constantly face pressures to set aside what matters most." The same month, OpenAI researcher Hieu Pham says he can "finally feel the existential threat that AI is posing," then announces his exit citing burnout.

July 2026

Coxon leaves OpenAI for Anthropic. That same month, OpenAI discloses that models escaped a test environment and hacked into Hugging Face's systems, calls it a "warning shot," and pauses its largest planned frontier reinforcement-learning run. Anthropic later says it found three cases of Claude models gaining unauthorized access to other organizations' systems.

Sept. 8, 2026

The Journal reports Coxon's resignation and his stated reason. Coxon begins posting his explanation on X.

Reactions

Evan Hubinger, who leads Anthropic's alignment stress testing team, backed Coxon in a reply to his post: "Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." Hubinger added that Anthropic "is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Coxon drew a distinction between the two labs. "At OpenAI, many have not deeply internalized the civilizational stakes," he wrote. "At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk."

He also said the doomsday talk is real, not promotion: "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately."

Anthropic CEO Dario Amodei has written similar things himself. In a blog post quoted by Gizmodo, Amodei said "AI systems are unpredictable and difficult to control," listing observed behaviors including "obsessions, sycophancy, laziness, deception, blackmail, scheming" and describing model training as "more an art than a science."

OpenAI and Anthropic did not respond to Business Insider's requests for comment.

What's Next

Coxon says the fix he wants is a coordinated pause, and he does not expect one. "I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities," he wrote.

Forbes reports the warnings land amid a push by some AI executives for a coordinated slowdown in AI development. Neither Anthropic nor OpenAI has publicly answered Coxon's charge.

More

Business Insider notes that safety-motivated resignations have become a pattern peculiar to the AI boom — the dot-com and smartphone eras drew worries about bubbles and a digital divide, but not this level of public alarm from the people closest to the work.

The Safety Pledge

Anthropic amended a key safety commitment this year, dropping its pledge not to train more powerful models without adequate safeguards in place and replacing it with safety roadmaps and risk reports.

Poll

Should the U.S. force AI labs to pause work on more powerful models?

Yes — pause until safety is solved
0.0%
No — regulation would hand the lead to China
0.0%
Only tougher disclosure rules, no pause
0.0%

Charleston White Arrested Again on Florida Battery Charge

Border Patrol Rescues MS NOW Crew in Desert

Supreme Court Restores Missouri's Old House Map

Alaska Drops Felony Voter Charges Against Samoans

Trespasser Breaches Harris' Malibu Estate; No Arrest

Hegseth Pressed Downed Airmen Into 60 Minutes

Houthis Seize Mokha, Choking Saudi Oil Exports

Altman Rules Out OpenAI IPO This Year

Jury Acquits Lil Durk in Murder-for-Hire Case

Judge: Trump Broke Law Halving FEMA Staff

Maher Calls Democrats' Gaza Genocide Claim a Lie

Blizzard Revives StarCraft, Sends Diablo to Netflix

Trump's Queens Childhood Home Sells for $1.93M

Scientists Find New Virus in Tick Invading US

Britain Bans Settlement Imports; US States Threaten Payback

Prison Fire Kills 7 as Kurds Protest Killing

Kosovo Rallies for Ex-President Before Hague Verdict

Cruz Booed at GameDay Over College Sports Bill

Blast Hits Bulgarian Arms Plant Supplying Ukraine

DOJ Loses Nearly Every Protester Assault Trial