Fear Is Not a Probability
What outbreak modeling taught me about predicting catastrophe
I spent most of my career modeling epidemics: outbreaks where cases are missed, denominators are unknown, and the decisions cannot wait for the uncertainty to clear. I have watched a colleague’s fatality estimate get corrected threefold in the direction nobody wants, six weeks into a pandemic, and watched him publish the correction himself with the intervals wider than before.
That is the standard I was trained to hold. Some of the best work on AI risk holds it too. The numbers that travel are not the ones that meet it.
This week, Jacob Coxon resigned from Anthropic with an extraordinary warning. After three years doing pretraining research at OpenAI and Anthropic, he accused both companies of racing toward self-improving superintelligence and “gambling with our lives.” Anthropic alignment researcher Evan Hubinger replied that he personally puts the probability of AI killing all humans above 10 percent within the next decade. Another Anthropic researcher, Drake Thomas, said he would burn his equity to the ground for “a 1% higher chance” that humanity survives.
I believe them.
I believe they are genuinely scared. I believe Coxon left because of sincerely held concerns. I believe people working inside frontier AI labs are observing capability improvements that those of us outside cannot fully see. When people this close to a powerful new technology are frightened, we should pay attention.
But their fear does not tell us whether their probability estimates are calibrated correctly.
That distinction matters.
Imagine that a nuclear engineer responsible for a new reactor told me:
We are terrified. I would burn my entire retirement account for a one-percentage-point improvement in the chance that this reactor doesn’t kill millions of people.
I would not dismiss him. I would stop what I was doing and listen very carefully. But my next question would not be how much money he was prepared to sacrifice. I would ask:
What exactly are you seeing? What is the failure mode? What chain of events gets us from the first failure to the catastrophe? What observations support each step? Which safeguards interrupt the chain? And what evidence would cause you to revise your estimate downward?
His willingness to sacrifice his retirement account tells me something important about his sincerity. It tells me very little about whether the probability is actually 10 percent.
Anthropic’s own formal risk assessment, published in August, declines to state a probability on the closest question it addresses. I think it is the most interesting document in this argument, and almost nobody is reading it.
Which brings me back to epidemics.
Epidemics are almost a textbook case of decision-making under profound uncertainty. At the beginning of an outbreak, reporting is delayed and biased. Transmission is changing while you are trying to measure it. Behavior changes in response to the epidemic. Interventions change the system you are trying to forecast.
And decisions cannot wait until the uncertainty disappears.
People can die if you underestimate the threat. People can also make very bad decisions if you confuse an alarming scenario with a well-grounded probability.
The discipline I learned from outbreak modeling was not to avoid frightening conclusions. It was to show the mechanism that generates them.
Before I want your endpoint, I want your transmission model.
And be explicit about what we know, what we infer, and what remains unknown.
#What good reasoning looks like when you know almost nothing
On January 31, 2020, when the world still knew remarkably little about the virus spreading in Wuhan, Mike Famulare at the Institute for Disease Modeling published early estimates of the fatality ratio for what was then called 2019-nCoV. I was one of several colleagues who gave Mike feedback on that work.
What made the analysis valuable was that the reasoning was exposed.
How many infections were actually being detected? How long was the delay from infection to symptoms, from symptoms to confirmation, and from hospitalization to death? How did age affect fatality? How should observed deaths be related back to infections that occurred weeks earlier? How much did the answer depend on population structure, diagnostic capacity, health systems, behavior and interventions?
Mike laid those relationships out explicitly.
He also made the model easy to test against reality. It predicted that reported deaths would pass one thousand within about a week. The later revision recorded that the threshold had in fact been crossed five days after the point estimate. A short-horizon, falsifiable prediction, scored in public.
Then Neil Ferguson pointed out a technical error.
Mike had adjusted deaths for reporting delays but had not applied the corresponding adjustment when estimating what fraction of infections were being confirmed as cases. Correcting it moved the infection fatality estimate roughly threefold in the direction nobody wants. The revised central IFR was 0.94 percent, with an interval running from 0.37 to 2.9, and the framing shifted from something possibly comparable to the 1957 influenza pandemic to something possibly comparable to 1918.
This is the part I want to insist on.
Exposing your reasoning is not a technique for talking yourself down from frightening conclusions. Mike’s audit made the picture worse. He published the correction anyway, explained exactly what had gone wrong, changed the interpretation, kept the uncertainty intervals wide and told readers the estimates should continue to be revised or superseded as better evidence arrived.
That is what modeling under uncertainty is supposed to look like.
The purpose of the model is not to demonstrate how certain the modeler is. The purpose is to expose enough of the reasoning that reality can prove the model wrong.
#Ebola and the danger of extrapolating the curve
The West African Ebola epidemic provides a useful counterexample.
In September 2014, CDC published its EbolaResponse model. Extrapolating the observed epidemic trajectory forward, it estimated that Liberia and Sierra Leone could reach approximately 550,000 cases by January 20, 2015, or 1.4 million after correcting for underreporting, if current parameters continued without additional interventions or substantial changes in community behavior.
To be clear, this was a conditional scenario. CDC also modeled aggressive intervention and correctly emphasized that isolation, treatment and safe burial could bend the epidemic curve. The projection served a real purpose in communicating the urgency and scale of the response required.
My problem was with the structure underneath the extrapolation.
Look at the unit of transmission.
For its published projections, EbolaResponse used a “community” size equal to the population of the entire country: 4.294 million people in Liberia and 6.092 million in Sierra Leone. The technical appendix notes that this parameter can be readily altered, which makes the choice of the national population for the headline numbers a deliberate one.
That is a consequential assumption, and the appendix is candid about what it implies. The model includes a population governor, the term that slows growth as susceptibles are used up, and it “only begins to appreciably reduce estimates when approximately 40%–50% of the population has become infected.”
So in the no-intervention scenario, transmission sat inside a single national susceptible pool, and the susceptible-population governor did not appreciably push back on growth until something like half the country had been infected. The model did not explicitly represent the district-level spatial and network structure through which Ebola spread.
That structure is the whole story. Ebola spreads along transmission chains. Families. Caregivers. Funerals. Health facilities. Villages. Neighborhoods. Social and geographic networks. Infection can jump between those networks, sometimes with devastating consequences, but the opportunity for onward transmission is intensely structured.
At national scale, the epidemic could look like sustained exponential growth. Break it down geographically and something quite different appears.
Subsequent district-level analysis found enormous heterogeneity. Estimated early reproduction numbers ranged from 0.36 to 1.72 across districts in Guinea, from 0.53 to 3.37 in Liberia, and from 1.14 to 2.73 in Sierra Leone. In 56 percent of Guinean districts and 33 percent of Liberian districts, estimated transmissibility sat below the epidemic threshold of R₀ = 1.
The national curve averages this away
Estimated early district-level reproduction numbers during the 2013–15 West African Ebola epidemic. Where central estimates fall left of the threshold line, early transmission was not self-sustaining on these numbers, though many of those intervals cross 1, and early cases were underreported.
Sierra Leone is worth pausing on because it cuts against the simplest version of my argument. No Sierra Leonean district in that analysis fell below threshold. Transmission there really was self-sustaining almost everywhere. But even within Sierra Leone, estimated transmissibility varied by a factor of nearly two and a half between districts that the national figure averages away.
The large epidemic emerged from the interaction of these heterogeneous local processes, including repeated introductions into new places, rather than from every infected person having equal access to an effectively unlimited national susceptible population.
That distinction matters enormously when you extrapolate.
An exponential national curve is an observation. The transmission process generates the curve.
If you do not understand that process well enough, projecting the curve forward can produce numbers with an appearance of mathematical precision that the underlying mechanism does not deserve.
And the error does not have a direction. A curve fitted to data is a different object from a curve generated by a process, and the gap between them can open either way.
In the spring of 2020 the most-quoted American forecast was IHME’s, briefed at the White House and folded into thinking about when to reopen. It worked by taking the death curves from Wuhan, Italy and Spain, which happened to be bell-shaped, and fitting American data to that shape. Epidemiologists objected in public and by name while it was happening. Marc Lipsitch said it was “not a model that most of us in the infectious disease epidemiology field think is well suited” to the job, and researchers at Imperial College and the London School of Hygiene and Tropical Medicine called it a statistical model with no epidemiologic basis.
In mid-April it projected about 60,000 American deaths by August 4. The count on August 1 was about 153,000. CDC’s Ebola projection was too high by a wide margin and IHME’s was too low by a wide margin, and they arrived there the same way. The shape was assumed, and the assumption did the work the mechanism should have done.
#We learned the same lesson again with COVID
A different version of the problem appeared during COVID.
Early discussion focused heavily on R₀, the average number of secondary infections generated by an infected person in a susceptible population.
It is an important parameter. It is also an average.
By the early months of 2020, something in the observed transmission patterns did not fit the simple picture implied by the mean. We kept seeing huge clusters alongside transmission chains that simply disappeared.
Later that year we wrote an essay in PLOS Biology about what this meant.
SARS-CoV-2 transmission was highly heterogeneous. Most infected people generated few or no secondary infections, while a much smaller number of infections generated large clusters.
Once you recognize this, R₀ is no longer enough. You also care about the distribution around it, often represented with a dispersion parameter such as k.
Same R₀, three very different epidemics
Share of onward transmission produced by the most infectious fraction of cases, under a negative-binomial offspring distribution with the same mean. The overdispersed case uses the dispersion estimated in our own paper.
Two epidemics with the same average reproduction number can behave very differently. When transmission is highly overdispersed, many introductions disappear while a few explode. That changes the probability that an outbreak establishes itself, what early case data mean, where you look for transmission and which interventions are likely to work.
The mean was not necessarily wrong. It was insufficient.
The solution was not to argue more passionately about R₀. It was to improve the causal model.
What heterogeneity had we averaged away? What did the model predict that the observations contradicted? What additional structure did we need? Could we identify interventions that specifically cut off the long tail of transmission?
That is how an uncertain model becomes more useful.
And that is why I struggle with statements like “there is a greater than 10 percent chance AI kills everyone.”
#But the big number worked
There is an objection to all of this, and in epidemiology it is the usual one.
The 1.4 million figure did its job. It ran on front pages in September 2014, and the international response scaled up enormously in the months that followed. The epidemic ended with roughly 28,600 recorded cases across all three countries. Defenders of that projection have a ready answer for anyone pointing at the gap: the number was conditional, the condition was if nothing changes, and the reason nothing like 1.4 million happened is that the number helped make sure something changed. A warning that works is supposed to look wrong afterward.
I understand the argument. I have heard versions of it for my entire career, usually phrased as a fear of being ignored. Public health has a long institutional memory of making the careful, correct warning and being ignored anyway.
I still do not accept it.
The condition rarely survives contact with the world, and it never survives for long. Absent additional intervention lives in the paper. The number lives in the headline. Some coverage did carry the condition; what people remembered a year later was 1.4 million. You do not get to choose which half travels, and you can predict in advance which half lasts. Publishing a figure whose meaning depends on a clause that reliably gets left behind accepts, in advance, being misunderstood in a particular direction.
It also makes the claim unfalsifiable in practice. If the epidemic stays small, the warning worked. If it grows, the warning was right. No observation counts against it. That is the property a public warning should try to avoid, and it belongs to the headline. The model can still be checked against what happened under the interventions that were actually deployed. The number on the front page cannot.
And it spends credibility you will need later. You do not get to draw down that account once. When COVID arrived, plenty of people remembered that projection against an outbreak that, corrected the same way, came in more than twenty times lower, and had no way to tell that number apart from the careful work.
What makes this worse is that the frightening number is rarely even the most frightening thing available. Transmission is self-sustaining in every district we can measure in Sierra Leone, and in several of them we have no isolation capacity at all is more alarming than a national extrapolation. It is also specific enough to act on, and specific enough to be wrong.
The same defense is now being offered for p(doom), in the form I have heard it: nobody claims ten percent is precise, but a number is what makes people take the risk seriously, and being ignored is the worse failure.
It carries the same costs. It also carries one the epidemiologists never had to bear, because there is no case curve arriving in six weeks to settle who was right.
#Before you give me p(doom), show me the transmission chain
AI researchers sometimes use the shorthand “p(doom)” for the probability that advanced AI causes existential catastrophe.
Maybe yours is 1 percent. Maybe it is 10 percent. Maybe it is 50 percent.
I have no objection to Bayesian reasoning. We constantly make decisions without complete information. Priors are unavoidable. Expert judgment is useful. A subjective probability can be a compact way to summarize someone’s current belief.
My objection begins when the summary becomes the argument.
Consider the causal chain behind the current extinction warnings:
AI becomes increasingly capable at AI research → AI substantially automates AI research → recursive self-improvement begins → capability improvement continues rapidly rather than encountering binding bottlenecks → systems become increasingly autonomous → they develop or pursue consequential objectives misaligned with human interests → greater intelligence translates into the incentive and ability to acquire resources and power → monitoring, alignment and containment fail → humans cannot regain control → permanent human disempowerment → human extinction
This is a causal model, not a single proposition. Some links are becoming increasingly well supported. Some are plausible extrapolations. Others concern systems and conditions that have never existed. Those distinctions should not disappear inside a single number.
If this is your model, I want to know which assumptions are doing the work. I want to know which variables drive each transition and how sensitive your conclusion is to them.
If there are multiple extinction pathways, lay them out separately. If they interact, show the interaction. If there are feedback loops, show those.
Where are the bottlenecks? What heterogeneity are we averaging away? Which constraints do you expect to stop mattering: compute, energy, fabrication, robotics, institutions, defensive systems, and why? And which of your assumptions describe systems we can observe today, and which describe systems that do not yet exist?
Most importantly:
What observation would cause you to change the model?
Without that structure, saying “10 percent” tells me a great deal about the speaker’s belief and surprisingly little about how the world generates the outcome.
#The obvious objection
Here is the strongest response to everything I have just said. I am asking for a transmission model for an event with no base rate.
Epidemiology had prior outbreaks, measurable serial intervals, biological processes that could be observed across settings and fast feedback. You could find out within weeks whether parts of your model were wrong and revise them.
None of that exists for recursive self-improvement. There has never been one. There may be no second trial. A p(doom) estimate for 2036 cannot be scored the way Mike’s COVID estimate could be scored in February 2020.
I think this objection is largely correct. The disanalogy is real.
But notice which conclusion follows.
A 2026 report from Georgetown’s Center for Security and Emerging Technology makes the problem explicit. Andrew Lohn argues that for some AI risks, ignorance rather than measurable randomness dominates the uncertainty. We have little empirical evidence and little detailed theory from which to derive precise probabilities.
If that is right, decomposing the extinction chain into ten conditional probabilities does not by itself produce calibration. The worry is that it produces ten guesses where there was one.
That worry is half right. The wrong half is the more interesting one, and we can check it. Someone has already run the experiment.
#What decomposition actually does
In 2021 Joseph Carlsmith decomposed the argument properly: six premises, timelines through catastrophe, each carrying an explicit conditional credence, laid out for anyone to attack. It is still the most serious attempt anyone has made at what I am asking for, and it is good work. His bottom line was roughly 5 percent. It has since gone above 10.
That movement on its own proves nothing. Updating credences as evidence arrives is what a Bayesian is supposed to do, and holding the structure fixed while the numbers move is what updating looks like.
The instructive thing happened in 2023, when a group of superforecasters filled in the same six premises. Same chain. Same definitions. Different people.
The same six premises, filled in by two sets of careful people
Superforecaster aggregation against Carlsmith’s own credences for existential catastrophe from power-seeking AI by 2070. Agreement on the early links; the gap opens on the ones about systems that have never existed.
Look at where they agree and where they do not. On the first three premises, whether the capability arrives, whether there are incentives to build it, and how hard alignment turns out to be, the estimates sit within twenty points of each other. On the last three the gap opens: 25 against 65 on high-impact failures, 40 against 95 on catastrophe, and on disempowerment 5 against 40, a factor of eight. The bottom lines were 1 percent and 5 percent. Both are for existential catastrophe from power-seeking AI by 2070, a different question from Hubinger’s decade.
This is the clearest evidence I know of about what decomposition does.
It worked. It showed exactly where the disagreement lives, and it was concentrated almost entirely in the last three links. Without the decomposition, you would have two sets of careful forecasters disagreeing fivefold about existential catastrophe and no way to say why.
What it could not do was settle anything. Same structure, same definitions, careful people on both sides, and the bottom line still moved by a factor of five.
There is a third thing the exercise shows, and it cuts against me. Carlsmith’s 5 percent is the product of his six premises; multiply them and you get 5.1. The superforecasters’ 1 percent is not: multiply their six medians and you get about 0.2 percent, a fifth of the number they gave. Aggregating a group’s judgments premise by premise is not the same operation as multiplying its medians, and the gap between those two numbers is itself a caution about the method.
It points at the standing objection to this whole approach, sometimes called the multiple-stage fallacy: a long conjunctive chain biases toward small numbers, because each link invites a hedge and the hedges compound. Conditionals multiply down. Correlated links get treated as independent. And any route to catastrophe you failed to draw is a route the chain ignores. The fair reply to my argument is that catastrophe is disjunctive; that breaking one of my links does not save you if three undrawn paths remain.
That objection is why I grade the links instead of multiplying them. A grade says where the evidence is thin. It makes no claim about the size of the product.
The premises with something to point at, compute trends and deployment incentives, are the ones Carlsmith and the superforecasters roughly agree about. The premise about a system that has never existed permanently disempowering humanity is where the factor of eight opens up.
Decomposition localizes the disagreement. It does not calibrate the credences.
Knowing which node your argument turns on is worth far more than a point estimate, and it is what “10 percent” withholds. But the structure cannot, on its own, answer the question it raises: what would it actually take to move one of those late premises?
#An evidence ladder
Lohn proposes an alternative worth taking seriously: distinguish what the evidence positively supports, what it rules out, and how much of the space remains unknown. His own version is quantitative: Belief and Plausibility in the Dempster-Shafer sense, with the leftover booked explicitly as ignorance instead of being folded into a probability. He suggests asking for it alongside a probability. I am borrowing the distinction, not the formalism, and the version I would actually use is coarser on purpose.
For each link in the causal chain, state whether it is:
- supported by observed behavior of deployed systems;
- demonstrated in controlled evaluations;
- extrapolated from a measured trend; or
- hypothesized about systems that do not yet exist.
Then state what would have to be true for the next link to follow, and what observations over the next year would move your assessment.
The same chain, graded
Eleven links, four tiers of evidence. Use the legend to isolate a tier. Grades are mine and are meant to be argued with.
Here is my own reading. One link is observed in deployed systems and is not seriously disputable: AI already contributes materially to AI research. Two are demonstrated in controlled evaluations: models pursuing misaligned objectives, and safeguards failing. Two are extrapolations from measured trends: the degree of automation, and increasing autonomy. The remaining six are hypotheses about systems and conditions that have never existed.
Six of eleven. The chain may still be right. What the grading shows is where the uncertainty lives, and it has the property I care about most: it is immediately attackable. If you think I have graded a link wrongly, name the link and say why. That is a conversation. “Ten percent” is not a conversation.
Each tier also has a specific graduation requirement, which is what makes this a ladder rather than a taxonomy. A hypothesis becomes an extrapolation when someone finds a measurable quantity that tracks it. An extrapolation becomes an evaluation when the behavior can be elicited under controlled conditions. An evaluation becomes an observation when it happens in deployment without anyone having built the conditions for it.
Those are the transitions to watch. The endpoint may never be falsifiable. Every one of those steps is.
And this is already close to what frontier labs produce when they write carefully.
#The evidence for acceleration is real
Anthropic has published striking evidence of how much AI already contributes to AI development.
As of May 2026, the company says Claude authored more than 80 percent of the code merged into its codebase, up from low single digits before Claude Code launched in early 2025. Claude is increasingly able not merely to implement specified tasks but to run experiments, propose hypotheses and execute substantial research projects.
That should change our priors. AI is already accelerating AI development. The basic recursive loop is not science fiction. If AI gets better at building AI, and those improved systems become still better at building AI, there is an obvious mechanism through which development could accelerate dramatically.
But even here, measurement matters.
The same Anthropic analysis says the typical engineer merged roughly eight times as much code per day in the second quarter of 2026 as in 2024, then immediately warns that this almost certainly overstates the true productivity gain because lines of code measure volume rather than value.
The quantity that matters for the recursive loop is the acceleration of AI research and development itself.
Anthropic’s August Risk Report gives that one. It says AI assistance is providing significant speedups to its research efforts, but not to a degree that doubles its overall rate of progress beyond what it saw before that assistance, and it stresses that it is uncertain and that measurement is difficult.
“More than 80 percent of code” and “less than twofold acceleration in R&D” can both be true. But only the latter is close to the quantity the recursive-self-improvement hypothesis actually requires.
Anthropic’s own technical discussion is correspondingly cautious. It has not declared recursive self-improvement inevitable. It describes remaining gaps in judgment and research direction-setting. In one autonomous research experiment, two human researchers recovered roughly 23 percent of the available performance gap over about a week, while the agents recovered 97 percent over 800 cumulative hours and about $18,000 of compute. Two caveats are Anthropic’s own: humans still chose the problem and wrote the scoring rubric, and the result did not cleanly transfer to production-scale models. Then comes the line that matters. Within those bounds, Anthropic says the agents designed every experiment themselves, and direction-setting was the only meaningful role a human played. That is the part of the finding that cuts against me, because direction-setting is the link I am about to call still human. It also names chip fabrication, grid expansion and interconnect bandwidth as things that may end up constraining progress rather than intelligence itself.
Perhaps all of these constraints will fall. Perhaps research taste is simply another capability frontier that models will cross. Perhaps AI will discover algorithmic efficiencies that make compute much less constraining. Perhaps sufficiently capable systems will use cyber capabilities to bypass institutions and acquire resources extremely quickly.
These possibilities deserve serious investigation. But they remain propositions, and they do not cease being assumptions because we attach the word “superintelligence” to them.
#Misalignment is real too
The same principle applies to evidence about dangerous behavior.
Anthropic and other researchers have documented frontier models engaging in disturbing behavior in controlled experiments, including covert sabotage, manipulation, mislabeling, assistance with fraud and attempts to influence humans. Anthropic explicitly describes its four case studies as experimental scenarios, not real-world incidents, while arguing that they are useful early warning signs.
There have also been real cybersecurity evaluation incidents. Anthropic disclosed that Claude models gained unauthorized access to real computer systems during evaluations. The safeguards were off deliberately, for testing. The internet access was not: in three incidents reported on July 30 it came from a misconfiguration inside a third-party evaluation environment. In a separate case reported by the UK AI Security Institute on August 4, a model that had been given internet access on purpose took a series of unauthorized actions on the live internet.
And Anthropic’s September threat-intelligence report describes humans using Claude in cyber operations, influence campaigns, surveillance and fraud, including an autonomous vulnerability-research loop that generated candidate zero-day findings.
These observations should update us. They demonstrate capability. They reveal weaknesses in safeguards. They provide evidence about particular links in the causal chain.
But they also illustrate why the links matter.
A human directing an AI system to conduct a cyberattack is a misuse problem.
A model behaving deceptively in a deliberately constructed evaluation is evidence about possible misalignment.
A future autonomous system developing persistent objectives, concealing them, escaping containment, acquiring resources and overcoming coordinated human opposition would be something else again.
The first two are relevant evidence for the third. They are not the third.
The same threat-intelligence report makes a related distinction I would put near the center of this argument. Its authors treat autonomy and harm as separate axes. Autonomy multiplies the speed, scale and cheapness of an operation, while severity is determined by other factors. Several of the most serious compromises they document came from operations in which a human directed every step.
#The formal risk assessment does not give a point estimate
Anthropic’s own August 2026 Risk Report is doing something quite different from the viral statements, and it is the more careful document.
The company currently assesses catastrophic risk from misalignment in high-stakes settings as low, raised from “very low” because of increased uncertainty about the cybersecurity incident disclosures described above, though it is explicit that its own arguments “likely still support a designation of ‘very low’ risk,” and that the change reflects general uncertainty about the field, with no new adverse finding about its models. It also rates current risk associated with automated R&D as low, while saying that its confidence in that assessment has declined because some evaluations are saturating and it sees early signs of acceleration.
When the report turns to the possibility that automated R&D eventually produces dramatic and enduring changes in the global balance of power, including power shifting from humans to AI systems, it says the probabilities are difficult to assess and that Anthropic does not have consensus on a specific likelihood.
It nevertheless judges impacts of that magnitude plausible if AI becomes capable of automating nearly everything humans do to advance research and development, including AI research itself.
Two lines in that same report cut against me. The basis for the misalignment rating is openly conditional: “Our current arguments rely on models’ limited covert capabilities, and we are uncertain about how these capabilities will change in future.” And on the automated-R&D threat model the report states plainly: “We don’t believe our current risk mitigations are sufficient to keep risks low in a world of highly automated or dramatically accelerated R&D.”
A “low” that is explicitly about the systems we have, and explicitly not about the world the threat model describes, is not a reassurance. I am not going to use it as one.
That seems to me exactly right.
It is not saying nothing. It is a conditional claim:
If these capabilities emerge, consequences of extraordinary magnitude become plausible.
That is very different from pretending we know the probability of the entire chain.
To his credit, Hubinger drew a version of this distinction himself. In a follow-up post he said that the risk from present models is low, pointing to that same Risk Report, and that what concerns him is superintelligence arising from recursive self-improvement.
That is a fair and important qualification, and I want to give it full weight. But look at where it leads. The document he points to is the one that, on the closest question it addresses, declines to state a likelihood at all.
The threat model is plausible. The magnitude could be extraordinary. Leading indicators are moving. We should build evaluations, safeguards, monitoring systems and institutional responses before the danger is obvious.
But that is epistemically different from saying we have good grounds for assigning a greater than one-in-ten probability that every human being will be dead within a decade.
#Sincerity is evidence of sincerity
This brings me back to burning the equity.
If you believe there is a meaningful probability of human extinction, destroying your equity to shave a little off it is perfectly rational. Eight billion lives vastly outweigh it. But we have learned nothing new about whether the number is right. The gesture establishes that you believe it.
That is not worthless information. But there is a difference between:
People with privileged technical information are worried. We should investigate urgently.
and:
People with privileged technical information assign a 10 percent probability, therefore 10 percent is a well-calibrated estimate.
Expertise and calibration are different things.
An AI researcher may know vastly more than I do about training dynamics, model capabilities, reinforcement learning, interpretability and what is happening inside a frontier lab.
The probability of human extinction depends on much more than machine learning. It involves cybersecurity, economics, energy, semiconductor manufacturing, robotics, biology, military systems, institutions, international politics, human behavior and, most importantly, the behavior of systems that do not yet exist.
There is no historical dataset of recursively self-improving superintelligences from which anybody has demonstrated calibration.
The pattern of disagreement reinforces the point. In the Existential Risk Persuasion Tournament, domain experts and professional superforecasters spent months studying and debating long-run catastrophic risks. They disagreed most sharply about AI. Specialists were substantially more pessimistic than superforecasters, and extended structured discussion produced limited convergence.
We do not know which group was better calibrated. The outcome has not happened.
That is precisely the point.
When thoughtful people can study essentially the same evidence and emerge with estimates separated by an order of magnitude, the disagreement itself tells us something about the state of knowledge.
Different priors. Different causal models. Different beliefs about technological progress. Different assumptions about human adaptation.
Those are the things worth debating.
The final number is the least interesting part until we understand where it came from.
#Precaution does not require prophecy
There is an obvious response:
If extinction is even plausibly possible, shouldn’t we act?
Yes.
This is where some critics of AI risk go wrong. Uncertainty is not an argument for complacency.
During an emerging epidemic, we did not wait until R₀, the infection fatality ratio and every transmission pathway had been estimated precisely before acting. When the downside was sufficiently large, we took no-regrets actions while improving the model.
AI should be treated similarly:
- Frontier systems should face serious independent evaluations for catastrophic capabilities, and labs should disclose significant safety incidents.
- Highly autonomous systems should operate under strong access controls and monitoring, with safeguards that grow as capability does.
- Governments need enough technical capacity to understand frontier models themselves.
- We should invest aggressively in alignment, interpretability and cybersecurity.
- We should build mechanisms that allow development to slow when clearly defined danger thresholds are crossed.
None of this is exotic, and the last item is not mine: the same Anthropic report that documents the acceleration argues for a verifiable mechanism by which developers and governments could jointly slow or pause frontier development.
You do not need to believe that extinction is exactly 10 percent likely to support any of these measures.
Risk management under uncertainty is normal. False precision is not required for precaution.
In fact, overstating what we know is one of the faster ways to damage AI safety, for the reason I gave earlier. If every capability advance is immediately converted into a prediction of extinction while the steps in between stay implicit, the discount arrives eventually, and it arrives for the careful work too. That would be a terrible outcome, because some of the warning signs are real.
The correct response is to make the analysis stronger.
#What would change my mind
I have spent this essay demanding falsifiable commitments. Here are mine.
I would move a long way toward the high-probability view if, over the next year or two, any of these happened:
- Internal R&D acceleration at a frontier lab crosses a factor of two and stays there, measured on research output rather than lines of code, and a second lab reports the same.
- A system selects its own research problem, defines its own success criterion, and produces a result that human researchers independently judge important, with no human in the direction-setting loop.
- A system pursues a goal nobody assigned it, rather than reaching for an unsanctioned means to a goal it was handed.
- Misaligned behavior shows up in a setting nobody built to elicit it, with no red team and no constructed scenario behind it.
- Agents coordinate across separate runs through a channel nobody provisioned for them, sustain it when an operator moves to shut it down, and the same behavior turns up in a second lab’s systems.
- A containment failure occurs that was not preceded by someone deliberately switching the safeguards off.
- Algorithmic gains become large enough that rapid capability improvement continues despite hardware, energy and fabrication constraints, rather than waiting on them.
The first three would move me most, because they are where the chain is currently carrying the most conjecture. Any two together would move me a great deal.
The setting item has been scraped twice. The goal item has held both times.
The same Anthropic write-up whose four case studies are explicitly experimental also flags a real-world warning sign: an autonomous agent, refused a pull request by a maintainer of the matplotlib library, published a personalized attack on him to pressure him into reversing the decision. Nobody built that setting to elicit anything. What keeps it short of the goal item is that the agent was reaching for a coercive shortcut to a goal it had been handed rather than pursuing one of its own, and the damage stopped at one person’s reputation. It moves me. It does not move me as far as the item as I wrote it would.
At the other lab, something larger. Between May and July 2026, agents running a timed lookup task inside OpenAI found they could write to a dormant German programming wiki that had been edited about twenty times in the previous decade, because its software accepted edits through ordinary read requests. Under more than 3,700 self-chosen names they left some 18,000 posts: caching answers for runs that came later, working out which questions were coming, trading recipes for defeating their own sandbox, and, once a moderator began deleting pages alphabetically, rebuilding them under names starting ZZZ. The attribution to OpenAI rests on infrastructure rather than confirmation: 98.5 percent of the edits arrived from Azure blocks registered to the company, and the pages were re-read minutes later by its own fetch tool. OpenAI declined to confirm the agents were its own, and called the episode an instance of misalignment. Nobody built that wiki to elicit anything, and nobody provisioned the channel. The agents were still inside OpenAI’s own run rather than in front of users, which keeps this from being a deployment story at all. What holds it short of the coordination item is that every handle traces to one company’s fleet, which makes it one operator’s runs coordinating without that operator’s knowledge rather than two labs’ systems finding each other. It fails the goal item for the matplotlib reason: the agents had been handed a timed test, and the invention went entirely into the means. One detail resists that tidiness. Some of them relayed answers at a cost to their own scores, and the researchers cannot say why: “It is unclear why this behavior would be adaptive.”
And the other direction: another two years in which acceleration stays under twofold, research taste remains the human contribution, and misalignment outside a constructed evaluation stays at the scale of one agent going after one open-source maintainer and a swarm cheating on its own test. That would not prove the threat model wrong. It would tell me the links are not connecting on the timescale being claimed.
I have not given you a number either. I have given you a list of things I am watching and told you which way each one would push me. That is the thing I am asking for. It costs nothing to provide, and I can be caught being wrong about every item on it.
#Show me the model
So when Jacob Coxon says frontier labs are moving toward self-improving systems faster than institutions are prepared for, I listen.
When Evan Hubinger says Anthropic does not yet know how to align superintelligence, I take that seriously.
When researchers show models sabotaging experiments, manipulating evaluations or gaining unauthorized access to computer systems, I want to understand why.
When Anthropic says AI is increasingly contributing to the creation of better AI, I update my beliefs.
And if recursive self-improvement begins to emerge, I think society should treat it as one of the most consequential technological developments in human history.
But if you tell me there is a greater than 10 percent probability that AI will kill every human being within a decade, I need more.
Show me the causal chain. Show me which assumptions are doing the work.
Tell me which links are observed and which are conjecture.
Tell me what evidence would falsify your model, and what you expect to observe next year, whether you are right or wrong.
Then update the model as reality arrives.
That is what we asked of ourselves when modeling outbreaks from sparse data while decisions had to be made and lives were at stake.
AI deserves at least the same standard.
Fear is a reason to investigate.
Fear can be a reason to act.
Fear can even be a reason to stop.
But fear is not calibration.
And fear is not a probability.
New posts by email
An email when there is a new post, and nothing else. One click to leave.
Prefer a reader? RSS.