walkthrough · 02

How bad, and how it happens, are two separate questions

Public argument fuses them constantly. One person says "AI risk" meaning a chatbot giving bad medical advice. Another means humanity losing control of its future forever. They use the same words and talk past each other.

the idea to hold on to

Severity and cause are independent. Misuse can produce a small harm or a permanent one. Misalignment can produce a trivial bug or a catastrophe. You cannot read the severity off the mechanism, or the mechanism off the severity. So this is a grid, not a ladder.

Rows are how bad. Columns are how it arrives. Click any square.

Misuse
Misalignment
Structural
Loss of control
Global catastrophic
Existential
Suffering (s-risk)

Pick a square

Each one is a different combination, and most public arguments are two people standing in different squares.

The word that causes the most trouble

Existential does not mean everyone dies. In the framing used in this literature it means the permanent destruction of humanity's long term potential. Extinction is one way that happens. A permanent unrecoverable dystopia is another. So is permanent stagnation.

This matters because people hear "existential risk," picture extinction, decide it sounds like science fiction, and dismiss the whole category. The claim being made is usually broader and, to its proponents, more plausible than the one being dismissed.

The disagreement that is actually about priorities

There is a real and serious argument that attention to speculative long term risk pulls resources and political will away from harms happening now: discrimination in deployed systems, labour effects, surveillance, environmental cost, and the concentration of power in a handful of companies. Researchers who make this case are not denying that future risk exists. They are making a claim about attention.

Researchers on the other side argue the two are complementary, that the same governance capacity serves both, and that waiting for a harm to be measurable is a poor strategy for harms that are irreversible.

Both of those are arguments about what to do, not findings about what is true. Drawing them as a settled question in either direction would be dishonest.

where this framework comes from

The grid is not mine. Drawing catastrophic risk on two axes rather than one ladder is Bostrom and Cirkovic's move, from the introduction to Global Catastrophic Risks in 2008, where risks are plotted on scope against intensity. The severity axis here, including the insistence that existential is broader than extinction, follows Bostrom's 2013 four class typology. The causal axis is assembled from others: misuse and misalignment are common currency in the field, and the structural category is specifically Zwetsloot and Dafoe, 2019.

What is new here is only the pairing of these two particular taxonomies, the twelve cells, and the wording inside them. It is a recombination of existing work, and the MIT AI Risk Repository, which catalogues over 1,700 risks drawn from 65 identified frameworks, is a good reminder that no taxonomy in this area is canonical, including this one.

Belle looking unimpressed
before you quote it

Someone is going to tell you the probability. Ask them what outcome they mean, by when, and whether they are counting the chance it never happens. Most of the time the number falls apart in your hands.

The number you will see quoted

Sooner or later someone tells you the probability. It has a nickname, p(doom), no agreed definition, no resolution date, and no way for anyone to be calibrated on it. That does not make it meaningless. It does make it a weak instrument, and it is worth seeing why by building one yourself.

Build the sentence. Same number, three choices, and the meaning changes completely.

what outcome counts
by when
counting what

Twenty four combinations, all defensible readings of the same two words. This is why quoting somebody's number without their definition tells you very little.

What the surveys actually found

Individual quotes travel further than survey data, which is unfortunate, because surveys of published researchers are the closest thing to a measurement in this area. Two findings are worth holding on to.

widespread across experts
unstableanswers shift with wording
splitforecasters vs domain experts

The first is that the range among people who work on this professionally is enormous, spanning several orders of magnitude. That spread is itself the most robust finding. The second is that answers move substantially when the question is rephrased, which is a documented result and a serious problem for treating any single figure as a measurement.

There is also a persistent gap between superforecasters with strong general track records and domain specialists, with the forecasters consistently lower. Neither group has feedback on this particular question, so neither can claim calibration.

Why the number is a weak instrument, stated fairly

Against quoting it. There is no agreed operationalisation, so two figures are usually not comparable. There is no resolution date and no feedback loop, so nobody has a track record on it. There is no defensible base rate to anchor against. And a single scalar destroys the structure of the underlying argument, which is where all the actual content lives.

For quoting it. Refusing to put a number on anything hides real disagreement behind vague language, makes positions impossible to compare, and lets people avoid committing to a view they are in fact acting on. Decisions get made either way, and an unstated probability is still a probability.

Both of these are good arguments. The honest position is that the number is a conversation starter and a terrible conclusion.

three ways it gets misused

Quoting a figure without the definition it came with. Treating a person's estimate as though it were a measurement. And using someone's number as a badge of which side they are on, which turns a question about the world into a question about tribe.

Sources

Origin of the global catastrophic risk definition and of the scope against intensity grid this page borrows its form from.
The canonical definition, plus the four class typology: extinction, permanent stagnation, flawed realisation, subsequent ruination.
Origin of the structural risk category. Their point: technology can cause harm even when no single actor misuses it and it behaves as intended.
The most authoritative institutional source used here, backed by more than thirty governments. Source of the three way causal taxonomy and of the statement that the evidence base is uneven.
A living database of over 1,700 risks, drawn from 65 identified classifications and frameworks. Figures change as it is updated; checked August 2026. Evidence that no single taxonomy is canonical.
A four category academic alternative: malicious use, AI race, organisational risks, rogue AIs.
The 1 in 6 figure, the distinction between existential risk and existential catastrophe, and Ord's own caveats about his numbers.
Philosophical critique of every current definition, including the ones this page uses. A working paper, so not itself peer reviewed.
The distinction from specification gaming, stated by the authors rather than inferred.
The clearest statement of the systemic route to an existential outcome with no single misaligned system. A preprint, not a peer reviewed journal article.
A six premise decomposition with a probability attached to each premise, which is the opposite of a single headline number.
The originating definition of s-risk, including the authors' own description of it as speculative.
Even handed reconstruction of the arguments against, used here so the sceptical side is quoted rather than characterised.
A preregistered experiment, and the only direct empirical test of the crowding out claim. Limited to individual attitudes.
The clearest worked example of one person giving several different numbers for several precisely defined outcomes, with his own half a significant figure caveat.
An early argument for taking the term apart rather than quoting it, and the ancestor of the axes this page's builder is made from.
The high estimate camp rejecting the metric itself, which is worth knowing before treating the number as a scoreboard.
The strongest public statement of the extinction case, including the conditional clause usually stripped out when it is quoted.
The best documented CEO figure. Amodei's own wording was that things go really, really badly; the definition attached to it in print is the reporter's, not his.
His reason: a number would imply a level of precision that is not there.
The sceptical argument in his own words, and confirmation that he gave no number.
A concerned researcher setting out a framework, and proposing the probabilities be obtained by polling experts rather than supplying his own.
The survey data behind the spread and the framing effect. Three question wordings produced different medians from the same population.
Exact definitions of catastrophe and extinction, the persistent superforecaster and domain expert gap, and the failure to converge after extended debate.
The reference class argument, and the observation that these numbers function partly as identity signals.
Included so the case for quoting a number is made by someone who believes it, not paraphrased by someone who does not.
The most constructive alternative: belief and plausibility as a pair, with the gap between them representing ignorance.
The conditionals critique, with worked examples of what a useful statement would look like.
The exact twenty two words and the signatory list, which is worth reading before anyone tells you what it said.
Withholds a number rather than adjudicating between the published estimates, which is itself a position worth noting.