Essays · September 2026
No, There Isn’t a 1 in 10 Chance AI Kills Your Family
Extinction forecasts from inside the frontier labs are feelings expressed as numbers — and they crowd out the AI risks that are actually here.
I spent yesterday evening reassuring my wife that there is not a 1 in 10 chance that AI kills our family before my youngest daughter gets to secondary school. She had read this in one of the thousands of recent articles covering statements from people within frontier AI companies.
“I personally think it is >10% within the next decade.” — Evan Hubinger, Alignment Science Lead at Anthropic, on the chance that AI kills all humans
Despite being said by a scientist, this is obviously not a scientific statement. It is not backed by a simulation like a weather forecast, it could not be. It is instead a feeling about an end state that could happen if a number of imagined things were to happen expressed as a number.
So why do people think that it’s acceptable to say this? To understand why, it’s useful to first ask why they believe it.
To put it simply: they are worried about AI agents getting access to systems they are not meant to. Hubinger himself says today’s models are low risk and his concern is future, self-improving systems. But strip back any of these doomsday scenarios and they all have the same starting point: an AI acting outside what it was authorised to do. I.e. “going rogue”
An AI agent can “go rogue” in two ways
- It can perform unauthorised actions whilst being run on computers owned by an organization or company
- It can do the same on a computer which it controls.
Neither of these are hypothetical scenarios. The Hugging Face attack was an example of rogue agents running on OpenAI’s hardware and the chance that there is currently an uncontrolled agent performing unauthorised actions somewhere online is 100%. Both of these scenarios are new, and new can be scary (see OpenAI’s decision to withhold GPT-2 in 2019, made while Dario was running research there) but if you understand what is actually going on in either of these scenarios, extrapolating to Terminator 2 is totally daft.
Hugging Face is a popular repository for storing and distributing AI related stuff. It was recently hacked by OpenAI. OpenAI was training its agents to be better at finding computer exploits, i.e. to be better at hacking. They were using a set of closed problems called ExploitGym (a test to see how good an agent is at hacking).
Training agents to hack is in itself not a problem, it’s important that agents are trained to be good at this so that we can build proper defenses against them. If an agent you own finds a problem with the security, you can patch it before someone else finds it. The issue here came from the way that the agents chose to solve the problems.
Rather than solve the test’s problems the agents chose to hack into the system which held the answers (Hugging Face). Over 4 days the agents carried out 17,600 unauthorised actions to get into Hugging Face’s servers and steal the solutions.
This shape of issue is well documented. It’s called reward hacking, where a model tries to get its reward in a way the person training the model had not anticipated. When trying to train a model to ‘drive from my house to the shops as fast as possible’ it learns that it could just drive through your neighbours garden and destroy their fence because that is ‘as fast as possible’. It does this because this is when it is “most right” but actually it has hacked your reward mechanism (which needed another term to say obey the laws of the road) and it now behaves in a not useful way.
OpenAI’s monitoring flagged the agents’ behaviour weeks before the attack, and they let the tests carry on. You could argue that was negligent. Even more surprisingly, if OpenAI had explicitly told their agents not to hack outside companies, they likely would not have. Additionally, this attack will have likely cost OpenAI tens, potentially hundreds of thousands of dollars in energy and GPUs to run the software which carried out the crime.
Here, “rogue” means their software behaved in a way that they did not expect or want. In this instance, stopping a “rogue” agent means telling it to stop, or turning off its machine or turning off its model.
In the second scenario (which is illustrative), an agent uses an open source model to replicate itself onto machines which it gains access to or pays for using stolen (or donated) funds. The same constraints apply, compute is expensive and running AI is expensive so unless the model was able to generate an income it would eventually run out of resources and stop. However, in this scenario, turning off the machine is harder as you can’t easily tell where it is located, who owns it.
As I said earlier, it is completely certain that uncontrolled agents currently exist on the internet. However, if your mental model of this is that now it has “escaped” it will over time become more intelligent and eventually take over then you have misunderstood how these models currently work.
You could make a counterpoint that RSI will change this. Recursive self improvement is where AI is used to train new AI. But this isn’t just AI improving itself automatically. This is well funded labs using their computers to run an algorithm which gets incrementally better. How much better it gets is a function of time and how many GPUs you have access to (so money). Becoming intelligent enough to be dangerous is expensive.
Counterintuitively, as models get better at hacking they actually result in more secure systems. The internet is made up of building blocks and if security flaws are found in those blocks then they can be patched and made more secure for everyone. This means that each vulnerability that is found makes it harder to find the next vulnerability.
As time goes on and AI’s hacking capability improves, so too will the security systems of the services we rely on. That’s not to say I don’t see risks here, and clearly the AI labs do too. I am sure we will see regulation and better enforcement of the existing laws soon. But these are not extinction level risks. So why would anyone say otherwise? I think there could be a few reasons.
Firstly, and most obviously, they have simply come to a different conclusion to me. They actually think that we might all die. I haven’t got much to say here, other than extraordinary claims require extraordinary evidence. I have not seen any yet.
Secondly, and I think this is the most generous. They see it as a means to an end. It doesn’t matter whether there is an extinction level risk, if they say the scary thing now and that kicks people into action it might avoid some unknown harm in the future. This logic assumes that people wouldn’t be kicked into action otherwise.
Thirdly, it is a popular thing to say. It’s interesting and captures people’s imagination. It’s divisive and new. It ticks all the boxes for a great story.
It would be remiss of me not to mention that there are potentially misaligned incentives here. The day after the CEO of Anthropic published his essay “We Must Pace the Frontier”, it was reported that Anthropic had chosen Nasdaq for its IPO, potentially the largest IPO in history. AI companies are making lots of money, as are their employees. There is currently a shortage of mansions in San Francisco. You don’t need to be a conspiracy theorist to see that there could well be some misaligned incentives.
I won’t go into how regulatory capture works, or why regulation can benefit the established players (you can read about that elsewhere) but I will say that it can be valuable to be in the news. Scary technology is often seen as powerful technology, and powerful technology is seen as valuable technology. That is true even when the message is “slow down”.
But let’s give those prophesying our demise the benefit of the doubt. That they are doing it to try to minimise the risk of AI causing harm in spite of the damage caused by popularising the prophecy. I would argue that doing this is still extremely negative for society for a number of reasons.
Firstly there are plenty of significant AI safety concerns that don’t involve humanity being wiped out. I have outlined a few of them in this essay. Focusing on some far away hypothetical moves the conversation away from the actual risks which are here and happening right now.
Outside of security risks there are plenty of important things in AI that we need to better understand. We don’t fully understand the implications of having people chat to these RLHF’d dopamine tuned machines all day. Much like we didn’t understand the impact of algorithmic feeds attached to social media. How might LLMs impact our mental health? Educating people about how this technology works also seems to be lagging, as you can see from the discourse online.
And perhaps most importantly, we need to figure out how to utilise AI as a tool to solve important problems. To make the world and people’s lives better.
Saying terrifying things can move focus away from the things that are actually happening and it can cause people to make bad decisions. It is also very difficult to roll back. No amount of evidence will undo the damage that has been done by these recent predictions. They were irresponsible and damaging.
So to my wife, or anyone else who read those headlines and lay awake worrying. We are going to be fine, I would argue better than fine. Ignore the prophecies of doom.