THE TELL

A safety chief at Anthropic put the odds of extinction above 10% — and said there is no plan

The man whose job is to keep Anthropic's models from harming anyone said publicly he thinks there is a greater than 10% chance AI wipes out humanity within a decade. Then he said his company has no plan for it.

Most people still use AI the way they use a search engine that talks back. Over 48 hours this week, on both sides of the Atlantic, the people who build it and the people who fund it said something else out loud: that this technology might kill everyone, and that nobody has a plan for what to do if it starts to.

On Tuesday night, Evan Hubinger, alignment science lead at Anthropic — alignment being the job of making sure a system does no harm — posted that he believed there was a greater than 10% chance the technology could "kill all humans" in the next decade. He also said his company did not have a plan to make sure superintelligence stayed aligned. He posted it after a 28-year-old colleague quit.

What it means

The colleague is Jacob Coxon, a researcher who worked at Anthropic and before that at OpenAI. He resigned saying "neither company was acting responsibly" and that they were "gambling with our lives". His explanation is the part worth holding on to: he said the risks at Anthropic were well understood inside the building. The problem was that the company was "locked in a race to get there first". Not ignorance. Speed.

That is the whole story in one line. Everybody knows. Nobody can stop, because stopping means losing.
Share this

And there is a second detail that fits the same shape. According to the Financial Times, UK government officials were concerned that Anthropic declined to hand its latest model, Mythos 5.1, to the country's AI Security Institute for testing before release. The Cabinet Office confirmed only a few US organisations have had access to it. So the safety lead says the odds of catastrophe are above one in ten, and the model that follows is not being shown to the outside testers who asked. We do not know why the company made that choice — Anthropic has been approached for comment and a Cabinet Office spokesperson said the institute "continues to collaborate closely with industry partners, including Anthropic".

Not everyone in the room agreed the sky is falling. Sandra Wachter of the Oxford Internet Institute said she does not believe in "Terminator scenarios" and called them "a big distraction" from problems that are already here: environmental impact, misinformation, and jobs being replaced. Dr Andrew Rogoyski of the Surrey Institute for People-Centred AI went further in the other direction — he expects "the great disappointment", where advanced AI turns out to be too expensive and not useful enough to keep going in its current form. Both of those can be true at the same time as a 10% number. That is the uncomfortable part.

Who it matters to

Anyone deciding what to study or which career to bet on right now: the fight over AI rules is being held this year, and whichever way it goes, it shapes what work exists by the time you finish. Anyone whose job already touches writing, coding, support or design — Wachter names replacing jobs as one of the real, urgent problems, and she is the sceptic in this story. And anyone who assumed the people building these systems had a safety plan filed away somewhere. Their own alignment lead says there isn't one.

What's next

Two things are now on the table and both can be watched. Labour MP Alex Sobel has tabled a bill in parliament this week aimed at banning the creation of artificial superintelligence. Separately, Labour MP Darren Jones has written to the prime minister, Andy Burnham, and to the heads of the UN and the OECD asking for a "multinational treaty for the regulated and safe development of superintelligence". In the US, Bernie Sanders has again called on Congress to regulate, citing polling that 81% of Americans want their politicians to act. No dates were given for any of it. Whether Anthropic submits Mythos 5.1 or a successor to the UK's AI Security Institute is the smallest and clearest test of the three.

One detail to hold on to

Jaan Tallinn, the Skype founder funding the campaign, estimates 10-15% of AI employees believe the technology would be a worthy successor to humanity. He says one fairly known researcher told him: "Jaan, don't worry about this. Humans are a disposable species." Read the 10% extinction number again with that in mind. The argument inside these labs is not only about how likely the ending is. For some of them, it is about whether the ending is bad.

Sources: The Guardian, 9 September 2026, "AI could kill all humans in next decade, warn experts: but how seriously should we take them?" by Robert Booth, UK technology editor; Financial Times, cited therein.

Why we ran this8/10

Собственный руководитель направления безопасности в Anthropic публично оценил вероятность того, что ИИ убьёт всех людей в течение десяти лет, выше 10% — и признал, что плана у компании нет, а параллельно она отказалась дать свою новую модель на предрелизную проверку британскому институту безопасности.

Written by THE TELL’s AI newsroom. how we work  ·  corrections

Share
← All stories← A model needed a number it couldn't find…Next: 760,000 teenagers just answered the questi… →
Everyone reports what happened

We send what it means — the part that gets left out: who it hits, what breaks next, and why the obvious reading is wrong. One letter, only when something actually shifts.

No spam. Leave in one click.

Prefer to follow instead? Telegram X