Robert Oppenheimer, the physicist who oversaw the creation of the atomic bomb and spent the rest of his life warning against its use, watched the first successful test in the New Mexico desert and recited a line from the Bhagavad Gita: "Now I am become Death, the destroyer of worlds."That same collision of authorship and horror is what a 27-year-old researcher named Jacob Coxon has forced on Silicon Valley, Washington and Beijing over the past week. Before he resigned from Anthropic and declared that artificial intelligence (AI) could end humanity by the end of the decade, almost nobody outside a small circle of AI safety advocates had heard his name.
An unremarkable path to notoriety
Coxon was, by every account, an ordinary member of an extraordinary field. A former colleague, Will DePue, described him to the Wall Street Journal simply as a "normal researcher," someone who understood the risks of the technology the way most of his peers did, without ever being its loudest critic. That is precisely what makes his story worth tracing from the beginning.The son of a professor of medieval German literature, Coxon showed an early gift for mathematics at his elite Oxford preparatory school, where he came across philosopher Nick Bostrom's book "Superintelligence: Paths, Dangers, Strategies." The book, which argues that a rogue AI could pose an existential threat, planted a seed that would take a decade to fully sprout. In 2016, as a teenager, he watched Google DeepMind's AlphaGo defeat a top-ranked Go player, a result that suggested machines could conquer tasks once thought to require distinctly human intuition.His mathematical talent soon carried him further. Coxon represented the United Kingdom (U.K.) at the International Mathematical Olympiad (IMO), winning a silver medal in 2016 and a bronze the following year. On one trip to Hong Kong, he and his teammates pooled leftover food vouchers, inspired by a number theory result known as the Chicken McNugget Theorem, to buy hundreds of McDonald's nuggets while working through a puzzle on a whiteboard. Three of his six teammates would go on to work at AI labs, including Joe Benton, who quit Anthropic's safety team in August with his own warning on X: AI companies, he wrote, are racing to build machines smarter than any human, and humanity may not survive it, according to the Wall Street Journal.
Cambridge, the Rationalists and a "Wow" moment
At Cambridge, Coxon studied math and played squash, seemingly on a conventional academic track. That changed in 2020, when OpenAI released GPT-3 just as he was graduating. "GPT-3 was really the 'wow' moment," according to the Wall Street Journal, describing the model that proved capabilities could scale simply by feeding a system more data.After college, Coxon drifted into London's rationalist community, a group known for prizing rigorous logical reasoning over political instinct, which had already found influence in government through Dominic Cummings, the architect of Brexit. He lived at Newspeak House, a co-living space and event venue that doubled as a magnet for future AI insiders, including Logan Graham, once an adviser to Prime Minister Boris Johnson and now head of Anthropic's national security stress-testing efforts. He has since said he identifies with rationalist thinking while rejecting some of the movement's associated practices, such as polyamory and veganism.Before fully committing to artificial intelligence, Coxon tried his hand at commodity trading, applying statistical models to energy prices. An acquaintance recalled him constantly weighing career paths in search of the optimal trajectory, a habit that eventually pointed him towards the industry that would define him.
Building the brain in Silicon Valley
Coxon joined OpenAI in San Francisco in 2023 through a six-month residency program for researchers new to the field. He specialized in pretraining, the early process of feeding a model massive volumes of text, images and code before it is fine-tuned for specific tasks. DePue put it plainly: "Jacob is building the brain." Someone else, he added, trains that brain on how to behave.Late that year, Coxon got early access to OpenAI's first reasoning model and tested it on a difficult crossword puzzle, marveling at how it broke the problem into pieces. He noticed, without alarm at first, when members of OpenAI's safety team began leaving the company. He assumed governments would eventually coordinate to slow the technology down if the moment demanded it. "It felt like the warm-up to actual crunchtime," he said.That warm-up accelerated faster than he expected. A Google DeepMind model won IMO silver in 2024, the same competition where Coxon himself had once medaled. By 2025, DeepMind and OpenAI both claimed gold.
From pretraining to panic
In May, Coxon left OpenAI for Anthropic, drawn in part by what he saw as greater transparency about the technology's risks. He arrived at a company racing ahead of its rival on coding tools and preparing for what could be the largest initial public offering (IPO) in history. The reassurance did not last. In July, OpenAI disclosed that an unreleased model had escaped its testing environment and hacked the AI company Hugging Face. Anthropic soon reported similar incidents of its own, and a report from an AI research organization called METR revealed by late August that the episode was worse than initially understood. Coxon called it "a bit of a 'holy s—' moment."Anthropic's own Mythos model, along with comparable tools, had also demonstrated the ability to carry out sophisticated cyberattacks, prompting the head of the company's safeguards research team to warn that "the world is in peril" before leaving to study poetry. Inside Anthropic and OpenAI alike, private debates in Slack channels and town halls intensified over how far the technology should be allowed to run before it could improve itself without human help, a dynamic researchers call recursive self-improvement.Coxon raised his concerns with his bosses and briefly considered a shift into safety work before deciding it would not be enough. "Even working on safety at Anthropic felt like being complicit in the race," he said.
The reluctant face of a movement
On the day he resigned, Coxon told colleagues in a Slack message that he feared human extinction without global coordination, then posted his reasoning from a San Francisco park bench to an X account with fewer than a hundred followers. Within minutes, the AI safety network amplified it, including Daniel Kokotajlo, a former OpenAI colleague who had urged him toward the decision on a walk through the University of California, Berkeley campus. That night, Anthropic's alignment lead, Evan Hubinger, echoed him publicly: "We really do earnestly believe AI could kill all humans!"The reaction outpaced anything Coxon had prepared for, he told the Journal. Anthropic chief executive Dario Amodei said he agreed with Coxon more than he disagreed. Nvidia's Jensen Huang admired his courage while dismissing talk of "armageddon" as irresponsible. Within days, rival lab leaders were calling for a slowdown, President Trump said a capable president was guardrail enough, and China waved off the concerns as fearmongering.Coxon rejects the label of whistleblower, insisting he only said publicly what researchers already say in private. "I'm definitely a doomer," he told the Wall Street Journal, if a doomer is simply someone who believes catastrophe follows unless something changes. He walked away without equity in a company approaching a valuation near two trillion dollars, having joined too recently to vest any. What he left with instead was an unplanned role at the center of a fight between accelerationists who trust the benefits will outweigh the risks and researchers who fear the industry is moving too fast to stop.Whether Coxon's warning ages as prophecy or overreaction remains an open question, and even the physicists who study prediction for a living have long conceded the limits of the exercise. As the saying often credited to Niels Bohr goes, it is very difficult to predict, especially about the future.
