The Shape of Harm
Thirteen drugs. Five kinds of harm. You decide what counts.
Some drugs do more damage than others — but "damage" isn't one thing. A drug can wreck your health slowly, kill you in a single night, or mostly hurt the people around you. Rank them one way and alcohol comes first. Rank them another way and it doesn't.
This page scores thirteen drugs on five kinds of harm, using published research. Then it hands you the sliders. There is no neutral setting. Deciding how much each kind of harm counts is a judgment about values, not a fact you can look up — so you make it.
Build the ranking yourself
Move a slider to say how much that kind of harm matters to you, and the table regroups as you go. There is no correct setting — that is the point.
The result is grouped into bands, not ranked one to thirteen. A band break falls where one drug outranks another with probability 90% or higher under the model's normal approximation, and a band is held together only where every pair inside it stays below that. The bands are a reading aid, not a statistical stratum: the grouping is one of several the model permits, and it moves when you move the sliders. For any specific pair, use Compare two substances below the chart rather than reading the horizontal rules. That probability is a closed-form approximation rather than a count of simulation runs, and it does not apply the simulation's clamping at 0 and 100.
All five columns are in play, so this is the combined model — four columns about the person using and one about everyone else, added together. It is not a per-person risk figure.
Everything matters the same. Simple, but it's a real stance — it says a drug's damage to society counts exactly as much as its damage to your liver.
This is the model's own uncertainty under a normal approximation — not a significance test, and not a count of simulation runs. The interval widths it propagates are assigned by evidence grade rather than estimated from data, so the percentage describes this model's confidence in itself, not the strength of the underlying evidence.
How to read the numbers
Open any drug and every cell carries an evidence grade. A 55 backed by a cohort study and a 55 backed by nothing look identical on a bar chart, so each is now marked Measured Indirect or Judgment. Of the 65 cells in this table, 21 rest on a drug-specific measured figure, 19 on a class proxy or extrapolation, 24 on judgment alone, and one has been suspended. That ratio is the single most useful thing on this page for deciding how much weight to give any particular number.
Every bar now carries its interval, and the numbers are ranges rather than points. Each cell's width comes from its evidence grade — ±8 where a drug-specific measurement exists, ±15 for a proxy, ±20 for judgment — propagated to the composite as though errors within a drug move together, which is the conservative direction. The result is uncomfortable and correct: of the 78 pairwise comparisons between these thirteen drugs, 41 have overlapping intervals. Only about half the ordering on this page is distinguishable from noise. The bands hold; the neighbours do not.
One drug is drawn as a hatched band rather than a bar. Methamphetamine's damage to others cell has been suspended — the only value available for it came from a formula that its neighbours in the same column are not held to, in a country where the drug was rare (section 07). Rather than print a number nobody should trust, the band shows every score meth could take, from that cell being 0 to it being 100. It spans four rank positions. That is the honest width of what is not known, and no other cell in this table has been examined closely enough to be sure it deserves a point rather than a band.
Every drug on this list is dangerous; the ranking is relative, not a verdict on safety. Grouped into bands under your current weights, worst first. Ordering within a band is not supported by the evidence and is not claimed. Tap any drug for its full profile — all five scores, how many Americans use it,39 its federal status, deaths per year, dependence risk, what is least certain about it, and every source behind it.
Go deeper
Everything above is the short version. Each of these is a full section.
Safety information
The part of this page that matters if the subject is more than academic.
The numbers on this page describe populations, not you. Individual risk swings hugely with dose, frequency, mixing, mental-health history, and setting.
Interactions are where people actually get hurt. Depressant stacking — alcohol plus benzodiazepines plus opioids — multiplies overdose risk. For serotonergic drugs the risks are not equivalent: interaction risk differs substantially by drug, dose and mechanism. MAOIs and lithium warrant particular caution. SSRIs and SNRIs may alter or blunt effects rather than amplify them, but the evidence is incomplete and does not support treating every antidepressant as equivalent. If blunting leads someone to redose, that is a plausible hazard rather than a demonstrated one. Check a combination chart before mixing anything, and don't treat "antidepressant" as one category.
A personal or family history of psychosis or bipolar disorder meaningfully raises the acute psychiatric risk of psychedelics and stimulants. Research protocols exclude people with that history outright, and treat preparation, a safe setting and the presence of a sober companion as the real safeguards rather than optional extras.44
For dependence or withdrawal — especially alcohol or benzodiazepines, where quitting unsupervised can be fatal — talk to a clinician. Medically supervised detox exists for exactly this reason.
Every substance on this page can seriously harm you. The ranking says which is worse. It does not say that anything here is safe, and rank 12 of 13 is still a psychoactive drug with a documented way of killing or damaging people.
This is a comparative scale, and comparative scales are easy to misread. Alcohol scoring 81 and psilocybin scoring 9 does not mean psilocybin is harmless — it means the two do different amounts of a thing that both do. Not one of these thirteen has a score of zero, and not one of them is without a documented serious harm.
And the supply has largely collapsed the low end of this scale. In the current US market, illicitly obtained pills and powders routinely contain something other than what they are sold as. Counterfeit benzodiazepines and prescription opioids frequently contain illicit fentanyl; substances sold as LSD have turned out to be NBOMe compounds, which — unlike LSD — have killed people at ordinary doses; material sold as MDMA is often adulterated. A drug's position in this table describes the molecule, not the thing in your hand. Anything not from a pharmacy should be treated as an unknown, which is why testing and naloxone appear in the notice above.
If this is more than academic
The short form of what is above, worth repeating:
These numbers describe populations, not you. Interactions are where people actually get hurt — depressants stack, and MAOIs are the antidepressant class that kills. A personal or family history of psychosis or bipolar disorder raises the psychiatric risk of psychedelics and stimulants. Alcohol and benzodiazepine withdrawal can be fatal unsupervised; supervised detox exists for that reason.
Carry naloxone, test the supply, don't use alone. Those three do more to change outcomes than anything else here for opioids and for the counterfeit-pill supply, which is where most overdose deaths now are. They are not universal: naloxone reverses opioid overdose and nothing else — not alcohol poisoning, not a stimulant cardiac event, not severe withdrawal, not a psychiatric emergency. For those, the answer is emergency care.
Reliable sources
PsychonautWiki — per-drug pharmacology, dosing, interactions.
TripSit combination chart — what's dangerous to mix.
DrugBank / Drugs.com — prescription interactions.
US support: SAMHSA National Helpline, 1-800-662-4357 — free, confidential, 24/7, English and Spanish, for families as well as users. Treatment search at findtreatment.gov. Mental-health crisis: call or text 988.
What the five harms mean
What each column measures, what each of the thirteen rows actually covers, and how well sourced every drug is.
What the five kinds of harm mean
Plain-language names for the five, with what each one actually covers. Later sections use the formal terms from the research — "harm to others," "acute crisis" — for the same things.
The whole point of keeping them separate is that a single score hides the shape. Alcohol and benzodiazepines are both depressants whose withdrawal can kill — but alcohol is bad at everything, while benzodiazepines are mild on most axes and terrifying on one.
And what the thirteen rows mean
Where the categories are too broadA score has to attach to a defined substance, route and pattern — not to a drug name. Several rows here are categories doing the work of several distinct things, and in at least one case the underlying sources disagree about which thing they meant. Stating the intended referent is the minimum fix; splitting the rows would be the real one.
That is one row of this table changing the reinforcing value of another row, between two of the most ordinary substances here — coffee and cigarettes. And it runs in both directions: in a pilot from the same laboratory, psilocybin inside a structured cessation programme left 12 of 15 smokers abstinent at six months.45 Small and uncontrolled, so it settles nothing clinically — but a table where psilocybin sits at rank 11 and nicotine at rank 5 cannot express the sentence "one of these is being tested as a treatment for the other" at all.
The same logic runs through the rest: depressants stack, stimulants and opioids are now routinely co-used, and most cocaine and methamphetamine deaths in section 04 also involve an opioid. A thirteen-row table has no place to put any of it. Real use is combinations, and this instrument can only score elements — one at a time, each assumed to be acting alone.
That 18-point gap is combustion. Not nicotine, which both rows share, and which is why both score near the top on getting hooked. Smokeless nicotine still carries roughly 28% higher all-cause mortality than using no tobacco at all,48 so this is emphatically not a story about a safe product. It is a story about which part of the cigarette does the killing — and it is the only place in this table where a single variable has been isolated cleanly enough to answer that.
Three of thirteen rows are categories, not substances, and two of those — cocaine and nicotine — silently mix sources that meant different things. That is a design defect, not a data defect: no amount of extra citation fixes it. It is fixed by splitting the rows, which would mean rescoring.
The evidence, column by column
Where the numbers come from, what they actually say, and where the published sources disagree with each other. Every figure below is traceable to a citation at the foot of the page.
Addiction
Strongest evidenceThe hardest data in the table, and now covering eight of the thirteen drugs rather than four. NESARC followed a nationally representative US sample across two waves — 15,918 nicotine users, 28,907 alcohol, 7,389 cannabis, 2,259 cocaine — and estimated lifetime probability of moving from first use to dependence.2
Lifetime transition probability, and the point at which half of all dependence cases had appeared.2
- At the ten-year mark the ordering compresses sharply: 15.6% nicotine, 14.8% cocaine, 11.0% alcohol, 5.9% cannabis — closely matching earlier National Comorbidity Survey estimates.23
- Speed of progression to a use disorder is fastest for heroin (median ~0 months), then cocaine (0–4 yrs), cannabis (1–6 yrs), tobacco (1–27 yrs), and alcohol (3–15 yrs) — a synthesis across several cohorts, which is why it's given as ranges rather than points.3
- Psychedelics sit at the lower end for mechanistic reasons, not just statistical ones — lower relative risk, which is not the same as low risk. Psilocybin has limited reinforcing effects and only marginal, transient self-administration in animals — the standard abuse-liability model.4 It can produce tolerance, but there's no evidence of a withdrawal syndrome after chronic administration.5 NIDA does not consider it addictive, as it doesn't drive uncontrollable drug-seeking.5
- A telling detail: on the ARCI euphoria scale used to predict abuse liability, psilocybin doesn't reliably score — it instead raises a dysphoria measure historically used to predict the absence of abuse potential.5
All thirteen drugs, and what is actually measured
Two survey familiesThe four NESARC figures cover only a third of the table. Adding the National Comorbidity Survey family — Anthony, Warner & Kessler's comparative epidemiology, summarised across drug classes — extends coverage to eight. Four drugs have no comparable figure at all, and saying so is more useful than inventing one.
Why the two columns disagree. NESARC and NCS were run two decades apart on different diagnostic criteria, and they measure subtly different things — NESARC estimates lifetime cumulative probability of transition using survival analysis, while the NCS figures are lifetime dependence histories among users.232 For nicotine the gap is large (67.5% vs 32%); for cannabis the two nearly agree (8.9% vs 9%). The ordering is stable across both, which is why the column's ranking is trustworthy while its exact digits are not.
One finding sits awkwardly across this whole column. Nicotine has the highest population dependence rate ever measured — and in the laboratory it is a weak and inconsistent reinforcer. Griffiths gave oral nicotine against placebo to eighteen people who had never smoked; some chose it reliably, others avoided it, and the drug is unreliable as a reinforcer in non-humans altogether.41 Population capture rates and individual reinforcement are measuring different things, and only one of them is what "getting hooked" sounds like it means. This column reports the first.
The four with no percentage still have evidence — just not in this form. Benzodiazepines: roughly four million daily US users, many meeting dependence criteria, with withdrawal possible after about a month of daily use.28 Caffeine: use disorder estimated at 8–20% of consumers, with a DSM-5-recognised withdrawal syndrome.2223 MDMA and ketamine: NESARC-III assesses "club drug" and hallucinogen use disorders, but publishes no use-to-dependence transition probability comparable to the figures above.33
Withdrawal danger
Sources disagree — range shownThe distinction that matters is unpleasant versus lethal. Only the GABA depressants can kill by cessation.
Alcohol. Severe withdrawal can progress to delirium tremens, a medical emergency. Reported frequency is roughly 3–5% of hospitalized withdrawal patients.6 Mortality figures vary meaningfully by source and era, and it would be false precision to pick one:
- StatPearls: historically as high as 20%; now around 1% with advances in critical care and prompt treatment.6
- Medscape: 5–15% even with appropriate treatment, closer to 5% under modern ICU management, versus as high as 35% before intensive care existed.7
- Clinical reviews commonly cite untreated mortality in the 15–37% band.78
The honest reading: treated DTs is now a low-single-digit-mortality event; untreated it is a substantial fraction. That gap — not any single percentage — is the argument for supervised detox, and it's why alcohol scores 95 on this column.
Benzodiazepines. Withdrawal parallels alcohol's through the shared GABA-A mechanism, spanning mild anxiety to life-threatening delirium or seizures, with documented cases progressing to convulsive seizures and nonconvulsive status epilepticus.9 Hence the 90 — the one extreme column in an otherwise moderate profile.
Opioids and stimulants produce severe withdrawal — and for stimulants a psychologically dangerous crash — but not the classically lethal physical syndrome of GABA-depressant cessation. Scored high, below 95.
Organ and body harm
Mixed: strong + estimated- Tobacco's high organ score is really combustion and tar rather than the molecule — which is why its acute column is near zero while its chronic column is severe.
- Opioids score lower here (~55) than their composite implies: the danger concentrates in addiction and overdose rather than in progressive tissue damage. That is a comparison rather than an absence — chronic use still carries endocrine effects, severe constipation, hypoxic injury from repeated overdose, and infection risk where injected. A genuinely different harm shape from alcohol.
- Ketamine's 40 is the most under-appreciated number in the table. Chronic heavy use is strongly linked to ketamine-induced uropathy — ulcerative cystitis, urothelial ulceration, reduced bladder capacity and fibrosis — first described as a clinical entity in 2007.1011 It can progress to hydronephrosis and renal failure, and case reports also link chronic use to cholangiopathy affecting the biliary tract.1112 In the UK, ketamine misuse and related uropathy among 16–24 year-olds doubled between 2010 and 2020.12
- Psychedelics have low physiological toxicity and no established fatal overdose at recreational doses; their real risk is acute and psychological, not bodily.4
Harm to others
Published figuresAlcohol is the clear outlier, and this single column is what drives its overall #1 ranking. In the Lancet MCDA, alcohol's harm-to-others part-score was 46 — more than double the next drug, with heroin at 21 and crack cocaine at 17.1 Meanwhile the most harmful drugs to individuals were crack (37), heroin (34), and methamphetamine (32).1 That split is the entire per-user-versus-societal crossover, in the source's own numbers.
Acute crisis
Mixed: strong + estimated- Opioids (95): highest single-episode death risk — respiratory depression, compounded by a fentanyl-adulterated illicit supply.
- Cocaine (75) and meth (70): acute cardiovascular catastrophe — MI, stroke, hyperthermia — possible even in young or first-time users; cocaine raised myocardial-infarction risk 23.7-fold in the first hour in the original case-crossover study — on a very wide interval (95% CI 8.5–66.3) resting on nine exposed cases.31
- Psychedelics (35–45): the one column where they leave the floor. Not toxicity — the acute psychological event, plus dangerous behavior in unprepared or unsupervised users and exacerbation of illness in those with or predisposed to psychotic disorders.4 LSD scores above psilocybin largely because the experience runs roughly twice as long, widening the window.
- Tobacco (2): essentially no acute-crisis risk. Its harm is entirely chronic — the mirror image of the psychedelic profile.
What backs each drug
Every drug in the table now rests on at least three independent sources, and most on four or more. This is the audit: what each one is actually built from, and where the evidence is thinnest.
Dose and lethality
How much is a normal dose, how much would kill you, and why that ratio is less clean than it looks.
The missing dimension: how much
Everything above scores drugs as categories — as if "cocaine" were one fixed thing. It isn't. The same substance is a different risk at a different dose, and the gap between a normal dose and a fatal one varies more than a hundredfold across this table — from about six doses at one end to over a thousand at the other. That gap has a name.
Robert Gable estimated, for 20 common substances, how many typical recreational doses it would take to kill a healthy 70 kg adult with no tolerance.13 A ratio of 10 means ten doses could be fatal. A ratio of 1000 means the lethal dose is effectively unreachable by accident.
Note the scale is logarithmic — each gridline is a tenfold jump. Gable is emphatic that these are ordinal, not arithmetic: you can say nitrous oxide is safer than GHB, but not that it is "20 times" safer.13 Two of the main thirteen are absent: Gable's benzodiazepine figure is for flunitrazepam specifically rather than the class, and nicotine's acute lethal dose is disputed — neither belongs on a chart of comparable ratios. Two substances shown here — GHB and nitrous oxide — sit outside the thirteen, included only to place the others on Gable's full scale.
Why a small ratio is so dangerous in practice
So the most common way people actually die on this list — combining drugs — is precisely what these numbers exclude by design. Gable also notes the ratios ignore tolerance, and that accidents, aggression and addiction were deliberately left out.
A ratio of 6 doesn't mean you're safe until dose six. It means the distance between the effect you want and the effect that kills you is small enough that ordinary variation can close it. Four things close it faster than people expect:
- Purity is unknown in an illicit supply. The ratio assumes you know what you took. If a bag is twice as strong as the last one, your "one dose" was two — and against a ratio of 6, you just spent a third of your margin without deciding to.
- Mixing collapses the margin. Safety ratios are single-substance figures. Depressants stack: alcohol with benzodiazepines with opioids share a respiratory-depression mechanism, so the combined margin is far narrower than any individual number implies.
- Tolerance moves the wanted dose, not always the lethal one. Gable notes explicitly that his ratios don't reflect tolerance, and that where tolerance to the desired effect grows faster than tolerance to the toxic effect, the ratio narrows.13 This is why a long-time user chasing the old feeling is in more danger than a beginner, not less — and why returning to a previous dose after a break is a classic fatal error.
- Route changes everything. Gable found intravenous heroin carried the greatest combined risk of dependence and acute lethality of the 20 substances he assessed; oral psilocybin the least.15 Same molecule, different route, different risk.
Dose-response: the curve inside a single drug
Alcohol as the worked exampleAlcohol is the best-studied case of "the amount is the whole story," and the honest answer has moved. A pooled analysis of 599,912 current drinkers across 83 prospective studies found all-cause mortality rose with consumption in a curved relationship, with minimum risk at or below roughly 100 g of alcohol per week — about five to six UK glasses or pints as the paper puts it, or roughly seven US standard drinks at 14 g each.16
The older claim that light drinking is protective — the J-shaped curve — is now contested rather than settled. A meta-analysis of 87 studies reproduced the classic J-shape without adjustment, showing reduced mortality among low-volume drinkers, but the shape substantially reflects study-design bias, including comparing drinkers against ex-drinkers who quit because they were already unwell.17 The GBD 2016 analysis concluded that the consumption level minimising an individual's risk is zero.18
What the 2020 analysis actually varies by. It estimated two thresholds for every region, five-year age group, sex and year: the TMREL, the intake that minimises health loss for a population, and the non-drinker equivalence, the intake at which a drinker's risk matches a non-drinker's.19 Age moved them a lot. Region moved them some. Sex did not move them at all — and that null result is one of the paper's headline conclusions, not a footnote.
One standard drink here is 10 g of ethanol, the GBD convention — not the 14 g US standard drink used in the lethal-dose table further down. The two are not interchangeable.19
Deaths per user, not deaths per drug
Cohort dataYour other question — deaths per person — has real data behind it. A Danish register study followed 20,581 people in treatment for substance use disorders, recording 1,441 deaths across 111,445 person-years, and computed standardised mortality ratios: how many times more likely each group was to die than the general population.20
Standardised mortality ratios versus the general population. MDMA's crude mortality was 1.7 per 1,000 person-years and its SMR was not significantly elevated.20
Read this one carefully — it is the most misinterpretable number on the page. These are people in treatment, which selects for the most severe cases, and the SMR captures all-cause death, not just the drug. That cannabis appears at 4.9× does not mean cannabis kills at half the rate of heroin; it largely reflects who ends up in treatment and what else is going on in their lives. The ordering is informative. The absolute values do not transfer to a casual user.
For scale at population level: US overdose deaths peaked at 107,941 in 2022, fell to 105,007 in 2023, and then dropped to 79,384 in 2024 — an age-adjusted rate of 23.1 per 100,000, down 26.2% in a single year and the largest fall in the decade.37 The counts in the table below are 2024 finals for that reason; earlier drafts of this page used 2023, which now understates how fast this is moving.21
All thirteen: dose, lethal dose, and deaths
Audited cell by cellFirst, what these three things actually mean, because they are easy to confuse and they answer different questions:
- Typical dose — the amount someone actually takes to get the effect they want. Varies with tolerance, purity, and route.
- Lethal dose — roughly how much would kill a healthy adult with no tolerance. For several drugs this has never been established in humans, because there aren't enough deaths to study.
- Safety ratio — lethal dose divided by typical dose. It describes how much room a substance leaves for error, not how many doses any particular person could survive. A ratio of 6 means the fatal amount sits close enough to the ordinary amount that normal variation can close the gap; a ratio of 1000 means it is effectively unreachable by accident. These are ordinal population estimates, not a personal count. And they describe a pure substance in controlled conditions: under illicit-market conditions there is no stable lethal dose, because purity, adulteration, route, tolerance, individual sensitivity, co-use and how fast help arrives all move it. The ratio is a property of the molecule, not of the situation anyone is actually in.
- Deaths per year — how many people actually die. This is a completely different question from lethal dose, and the two often disagree, which is the most important thing on this page.
Nicotine has a wide safety ratio — you cannot realistically smoke yourself to death tonight — yet it kills roughly 480,000 Americans a year. Those deaths come from decades of smoking, not from overdose. Its acute score is near zero and its chronic score is near maximum, and both are correct.
Psilocybin is the mirror image: essentially no deaths ever recorded, one documented case in a patient with a transplanted heart. Cannabis has no confirmed overdose deaths at all.
Alcohol sits in the worst possible position — a narrow safety ratio of about 10 and a death toll of roughly 178,000 a year, more than double the entire US overdose count. (Smoked tobacco kills more, at about 480,000, but it has no comparable acute danger: nobody dies of an overnight cigarette overdose. Alcohol is the only substance here that is near the top on both counts.) It is dangerous both ways at once, which no other drug in this table manages.
2. Deaths per year is not risk per user. A drug used by fifty million people will produce more deaths than one used by fifty thousand, even if it is far safer per person. Alcohol's enormous total partly reflects how many people drink.
3. "No recorded deaths" is not "safe." It means low acute toxicity. Psychedelics still carry psychiatric risk, and accidents while impaired are real. Cannabis has no overdose deaths but a measurable psychosis association.24
4. The lethal dose assumes conditions almost nobody meets — no tolerance, no other substances, known purity, healthy adult.13 Real overdoses usually involve a mixture, an unknown dose, or a lost tolerance after a break.
The audit, stated plainly. Of the thirteen rows above: four rest on human toxicity evidence plus category-level mortality surveillance (opioids, alcohol, meth, cocaine) — and the qualifier carries weight: the CDC figures are for headings such as psychostimulants, mostly meth and synthetic opioids other than methadone, most of those deaths involve more than one substance, and a death is counted under every drug it involved. Two have a solid dose figure but no separate death tracking (MDMA, caffeine). Four have lethal doses extrapolated from animal studies because human deaths are too rare to study — Gable marks these with a question mark, and so does this table (ketamine, cannabis, LSD, psilocybin). One is split between two different measures entirely (nicotine). One is dangerous mainly in combination, so its solo number understates it badly (benzodiazepines). Only a third of this table rests on direct human measurement, and not one of the death columns is a clean per-substance count. The honest presentation is to say which third, and to say what the counts actually count.
Who gets harmed
Overdose deaths by age, sex and race — and the eleven drugs with no equivalent data.
The other missing dimension: who
The alcohol model above varies by age and region because someone built one. No equivalent exists for the other eleven drugs — there is no published minimum-risk intake for cannabis or cocaine by age and region, and inventing one would be the exact failure this page keeps correcting. But that does not mean nothing is known. It means the evidence takes a different shape for each drug, and the honest move is to say which shape.
What is resolved by demographics, and resolved well, is who dies of an overdose. US vital statistics break that down by age, sex, and race and ethnicity — and the 2024 figures record the largest single-year fall in a decade.
All thirteen: what is actually known about who
Four tiersApplying the same audit to the demographic question that section 07 applies to the harm columns. Only one drug has the full treatment, and five have essentially nothing.
The audit, stated plainly. One drug of thirteen has a demographic model. Three have demographic mortality data that cannot be crossed with the drug itself. Three have an age gradient measuring a different quantity. Five have nothing. An age-and-region control for all thirteen would be about ninety per cent fabrication, which is why this section is a map of the gap rather than a twelfth copy of the alcohol widget.
How reliable is this?
Agreement with two expert panels, an audit of every column, and what happens when the uncertainty is propagated.
How closely does this agree with Nutt 2010?
Nutt, King & Phillips (2010, The Lancet) had an expert panel score 20 drugs on 16 weighted criteria. Plotting their published overall-harm score against this table's composite at equal weighting, the two track closely: Spearman's ρ = 0.93 across the eleven overlapping drugs, and tightest at the extremes. (This was 0.94 before methamphetamine's damage-to-others cell was suspended — suspending a cell moves the composite, and every figure downstream of it.)
ρ = 0.94 is therefore an upper bound, inflated by an unknown amount. Some of it is real — most columns were not built from Nutt — but no honest reading treats this as validation. The fix is not to delete the plot. It is to stop calling it something it isn't, and to go find a panel that had nothing to do with this table.
Each point is one drug. Both axes run 0–100, but the two instruments were built differently, so what matters is that points rise together, not that they sit on the dashed line of exact agreement. Psilocybin and LSD anchor the floor in both; alcohol tops both. This plot is pinned to equal weighting — moving the sliders above does not move it, so it stays a fixed calibration reference rather than something the weights can tune. The two informative disagreements — meth and benzodiazepines — are explained below.
A fresh panel not used to construct this table
Canada 2026 · agreement dropsThe proper test is a panel that never informed these scores. In January 2026 a Canadian MCDA did exactly that: 20 experts from six provinces scored 16 drugs on 16 harm dimensions at a two-day decision conference.38 Alcohol came first at 79, then tobacco 45, nonprescription opioids 33, cocaine 19, methamphetamine 19, cannabis 15.
One component can be isolated, and it cuts the other way. Restricting the Nutt comparison to the same six drugs Canada scored gives ρ = 0.94, not a lower figure — so the smaller drug set does not explain the gap, and controlling for it makes the contrast sharper rather than softer. Everything else remains confounded, and no arithmetic here separates leakage from the genuine methodological differences below.
The disagreement is not random either. It is almost entirely tobacco — 2nd in Canada, 5th here. That is a real methodological difference rather than an error on either side: the Canadian panel scored population-level harm, explicitly weighting how many people use each drug, so a legal product used by millions rises. This table scores harm per user and keeps prevalence out. Both are defensible; they answer different questions, and the gap between them is the same per-user-versus-societal crossover the model at the top of this page is built around.
One caveat on independence, which cuts against this page's own framing. David Nutt and Lawrence Phillips — two of the three authors of the 2010 paper — are co-authors on the Canadian study, which also uses the same MCDA method and decision-conference format. So this is a fresh panel, a different country, a different decade and different data, but it is not a methodologically independent tradition. A genuinely independent check would have to come from outside the MCDA family altogether. That check does not currently exist, and this page should not pretend otherwise.
Could these numbers come from a formula?
A fair challenge to everything above: if the digits are judgment, why not replace the judgment with arithmetic? This section tries. The answer splits cleanly by column, and the exercise surfaced one real error in the table.
Harm to others
This column needs the least judgment of any. Nutt's panel already published weighted harm-to-others part-scores; rescaling them so alcohol anchors at 100 yields a value directly, grounded in an expert panel rather than in a single author's intuition.
score = 100 × (nutt_others ÷ 46)
Nutt published part-scores for six of the thirteen drugs here.1 Running the formula on all six is the test the earlier version of this section skipped — it agrees closely on four and diverges sharply on two:
The chart used 2 for a time. It no longer uses any value. Correcting 70 to 2 moved methamphetamine from 2nd to 3rd under equal weighting and from 2nd to 4th under the societal preset — but a value obtained by a method the rest of the column is not held to, in a country where the drug was then rare, is not a value this table can defend. The cell is now suspended, and methamphetamine is drawn as a band spanning every score it could take. It still ranks at or near the top for harm to the person using, which the evidence does strongly support.
There is a reason not to simply rewrite them, but it undercuts the column rather than rescuing it. Nutt's figures are UK-2010 and drug-specific in ways that don't transfer: his 17 is crack, not powder cocaine, and his heroin 21 is not the same object as this table's "opioids" in a fentanyl era. Rewriting those two cells would import that mismatch rather than resolve it — but the identical objection applies to meth, whose near-zero score reflects how little methamphetamine circulated in Britain in 2010, not what it does to families in the US now.
So the honest label for this column is not "derivable" but "checked against a formula, and corrected in one place." For the record, applying it to all six would drop cocaine from 65 to 37 and opioids from 65 to 46, which under the everyone else preset narrows cocaine's lead over meth from 12 points to 3. The ordering survives. The margins do not.
Acute crisis
Gable's safety ratios give a principled toxic component — take a log, because the ratios span three orders of magnitude:
toxic = 1 − log(safety_ratio) ÷ log(1000)
That works well for the poisons: alcohol lands within 7 points of the table's value, meth within 3. Then it collapses. It scores LSD's acute risk at zero, because LSD has no reachable lethal dose — while the table says 45.
Both are correct about different things. Gable measures lethality; the acute column also contains psychological crisis, which has no lethal-dose analogue and no published scale. Adding a second component and taking whichever dominates cut the error from 23.5 to 14.3 — a genuine structural finding: acute risk is two independent things wearing one label. But the psychological values are this page's own, so that half remains judgment, and the improvement is partly circular.
Withdrawal danger
Derivable as tiers, not digits. The published evidence supports a four-step ladder — lethal (alcohol, benzodiazepines), severe but not classically lethal (opioids, stimulants), mild, none — because real mortality data exists only for the top tier.67 The gap between 90 and 95 for benzos versus alcohol is not measuring anything; the gap between tier 1 and tier 2 is.
Addiction
The surprise. NESARC gives hard lifetime-dependence figures for only four drugs; adding the NCS/Anthony family reaches eight, but four still have no comparable percentage at all, and three of the eight rely on a drug-class proxy.2 Worse, no scaling of that data reproduces the table. Rescaling linearly puts cocaine at 31; the table says 70.
Adding a speed-of-onset term fixes cocaine and breaks nicotine, because nicotine has the highest lifetime risk but the slowest capture — a median of 27 years. The best fit available still missed by about 22 points on average, well outside the ±10 the document claims.
The honest conclusion: the addiction column is not a formula in disguise. It is judgment that blends lifetime probability, onset speed, and compulsion intensity in a ratio never made explicit — and only five of its thirteen values rest on a drug-specific measured percentage.
Organ harm · Harm to self
Organ harm would need disability-adjusted life-years per user-year, standardised across substances. That data doesn't exist in comparable form — alcohol's liver and cancer burden is heavily studied, ketamine's bladder toxicity is known mostly from case series.11
Harm to self has a worse problem: it isn't independent. It overlaps addiction, organ harm and acute crisis, so a composite that includes all four double-counts. Nutt avoided this by defining sixteen criteria — nine harms to the user, seven to others — designed to minimise overlap, and weighting them explicitly. This table did not, which is a structural flaw in its column design rather than a bad number within it. It was six columns when that criticism was first written; it is five now, because one of them was deleted for exactly this reason.
If the digits are shaky, how much does the ranking move?
A stated error bar is worth nothing until you push it through the model. So: take each cell's interval — ±8 where a drug-specific measurement exists, ±15 for a proxy or extrapolation, ±20 for judgment, as marked on every cell in the drug panels — treat those as 95% intervals, resample all 65 cells 20,000 times, and re-rank each draw. An earlier version used one flat width per column, which assumed a measured cell and a guessed cell were equally trustworthy. They are not, and the grades now say so.
Run one, independent errors. Each cell wobbles on its own. The ranking barely moves: alcohol takes first place in 97% of draws under equal weights and 100% under the societal preset. But this is the flattering assumption, and it flatters for a structural reason — five independent errors partly cancel when you average them, so the composite comes out steadier than its inputs.
Run two, correlated errors. Real mistakes are not independent. If this table has misjudged methamphetamine, it has probably misjudged it across several columns at once. Adding a shared per-drug bias — one author, one blind spot — gives the honest picture:
Where the answer holds, it holds firmly. Where it doesn't, it doesn't at all: under the per-user view the top spot is a genuine three-way contest between methamphetamine, alcohol and opioids, and this page's habit of naming a single winner there is not supported by its own error bars.
Rank intervals, equal weighting — the bar is the 90% interval, the mark is the median:
Simulation method — parameters, and how to reproduce it
These figures are computed in your browser when this section loads, not looked up from a stored table. Press re-run to resample; the numbers should move by a point or two and no more.
- Distribution
- Gaussian, via Marsaglia polar transform
- Cell standard deviation
- Derived from each cell's evidence grade, not assumed flat: measured 4, indirect 7.5, judgment 10 — i.e. the ±8 / ±15 / ±20 intervals shown in the drug panels, read as 95% intervals. A suspended cell is excluded from the draw and handled as a range instead.
- Shared per-drug bias
- sd 5, drawn once per drug per iteration and added to all five of its cells (correlated mode only)
- Implied column correlation
- Not a single number — it depends on the two cells' grades, because the shared bias is fixed at sd 5 while cell noise varies. Between two measured cells it is 0.61; two indirect, 0.31; two judgment, 0.20; measured against judgment, 0.35. Range 0.20 to 0.61. An earlier version of this panel quoted 0.50 and 0.32, which were the values from the previous flat-width model and were left behind when the grades took over.
- Truncation
- Clamped to [0, 100] after perturbation
- Iterations
- 20,000 per run
- Seed
- 7 on first load (mulberry32); the re-run button uses a fresh random seed
- Weight vectors
- equal 5/5/5/5/5 · per-user 8/6/8/1/7 · societal 4/4/4/10/5 · one-bad-night 2/4/1/1/10, in column order addiction, withdrawal, organ, others, acute
- Weights perturbed?
- No. Weights are the reader's values, not an uncertain quantity — only the evidence cells are resampled
- Reference implementation
- Published alongside this page as
stability-simulation.js. The full analysis is also reproducible from data files — see below.
The correlated model is itself an assumption, and a consequential one — it is what moves alcohol from 97% to 79%. A shared bias of sd 5 says roughly that a systematic misjudgment about one drug is about as large as the noise on any single cell. That is a plausible guess about how a single author goes wrong, not a measured quantity. Both modes are exposed above so you can see how much rests on it.
The broad bands survive: nothing at the top ever falls to the bottom, and no psychedelic ever climbs into the top half. But adjacent ranks are not distinguishable. Opioids spans ranks 1–4. Cocaine and methamphetamine overlap almost completely. Ketamine, MDMA and cannabis are one undifferentiated block, and so are LSD, psilocybin and caffeine. Reporting these as an ordered list of thirteen implies a precision the model does not have — five or six broad bands, depending on the weighting, is what the evidence supports.
That was presented as reassurance. It is the opposite. A column that changes almost nothing when removed is a column carrying almost no information — while still letting the same harm be counted twice by anyone who moved two sliders. It has been deleted, and every score, preset and simulation on this page now runs on five columns. The reader-facing consequence is small; the honest description of the instrument is meaningfully different.
Known weaknesses
Everything wrong with this page that its author is aware of, collected in one place so nobody has to reconstruct it from ten sections. The first group describes what is missing from the evidence itself; the three after it describe what is wrong with the instrument built on top of that evidence, ordered roughly by how much damage each does to the conclusions. None of these are hidden elsewhere on the page — this is a summary, not a confession.
That means the model cannot become valid by replacing a few scores. Outcome rates must be separated from causal and contextual modifiers first. The full review is published as
INSTRUMENT_AUDIT_2026-07-23.md, every cell is classified in evidence-audit.csv, and the unscored v6 specification is in instrument-v6.json.
The evidence base — what has never been measured
Not faults in this instrument but absences in the literature it is built from. Most of the structural problems in the next group are downstream of these, and no amount of care in scoring would remove them.
Structural — the design of the instrument
Problems with how the thing is built, which no amount of extra evidence would fix.
Evidential — what is actually behind the numbers
Problems with the inputs rather than the architecture.
And it is load-bearing. Alcohol leads this table only while that column is weighted at 3 or above. Below that, methamphetamine takes the top. Remove the column entirely and score on the four genuinely individual-level criteria and the order is methamphetamine 78, alcohol 75, opioids 74 — alcohol is no longer first. So the headline result of this page depends on the one column whose denominator differs from the other four.
That is not a reason to discard the finding: alcohol's societal harm is real, well evidenced, and the reason most serious analyses put it near the top. It is a reason to stop describing the composite as harm per user without qualification. The honest description is four per-user columns plus one societal column, combined by a weighting the reader chooses. Try it: set damage to others to 1 and watch the top of the table change
Validation — what has not been checked
Where agreement has been tested, and why the tests are weaker than they look.
What a stronger source can — and cannot — fix
The audit changed the earlier claim that every weak cell only needed a better literature search. Some do. Others are attached to the wrong question, so stronger evidence would only make the mismatch more confidently wrong.
Future work
In rough order of how much each would improve the result per unit of effort. The first three are the ones that would change what this is, rather than how good it is.
References
Forty-nine sources. Tags mark what kind of evidence each one is — primary for original peer-reviewed studies, clinical for practice references, review for syntheses of a literature.
Legacy model v5 — 13 substances × 5 criteria = 65 cells. Criteria: getting hooked, danger in quitting, damage to your body, damage to others, one bad night. A sixth criterion, damage to your life, existed in earlier versions and was deleted for overlapping the other four. Exactly one cell is suspended — methamphetamine's damage to others — and that drug is drawn as a band rather than a bar in consequence. Cell intervals derive from evidence grades at ±8 measured, ±15 indirect, ±20 judgment. Band breaks are drawn at P ≥ 0.90 under complete linkage. Grade tally: 21 measured, 19 indirect, 24 judgment, 1 suspended.
Anything on this page that contradicts the paragraph above is a bug, and the data files are the tiebreak. The legacy score files are cross-checked against the page; the audit files are separate and deliberately do not overwrite v5.
Re-checking found real errors. Two examples. Gable's own abstract states that most published acute-toxicity fatalities involve a co-intoxicant — a caveat this page had omitted entirely from a section built on his ratios, and which materially changes what a safety ratio means. And two of the values plotted from him, caffeine at 100 and nitrous oxide at 150, are not in the 2004 paper the page cites; they appear to come from a later book chapter, and are now marked estimated. Neither error was invented — both were inherited by not looking closely enough.
The checks that came back clean are worth as much. NESARC's four dependence probabilities and all four sample sizes are exact. Gable's every plotted value falls inside its published band. The remaining 33 references are the obvious next piece of work, and the honest position until then is that this page's numbers are as good as one careful reading, not as good as an audit.
scores.csv — all 65 cells with their evidence grade, interval and the specific basis for each;
references.json — the 49 sources structured, with the note explaining what each one supports;
weighting.json — the criteria, the weight presets, and every simulation parameter;
reproduce.py — standard library only, no dependencies;
stability-simulation.js — the rank-stability analysis as a standalone script;
evidence-audit.csv — all 65 cells classified by evidential and construct problem;
instrument-v6.json — the unscored redesign specification;
INSTRUMENT_AUDIT_2026-07-23.md — the reasoning and development protocol.
Running
python3 reproduce.py rebuilds the ranking under each preset, the composite intervals, the pairwise-overlap count, the correlations with Nutt 2010 and Canada 2026, the rank-stability analysis, and a reference audit. Anything it prints that disagrees with this page means the page is wrong. That check has already earned its place: suspending methamphetamine's damage-to-others cell moved the Nutt correlation from 0.94 to 0.93 and the Canada correlation from 0.64 to 0.55, and the script caught that the page was still quoting the old figures.
Reproducibility is not validity. It means someone else can get these numbers from these inputs — not that the inputs are right. Twenty-four of sixty cells are author judgment, and no amount of reproducible arithmetic turns a judgment into a measurement. The script makes the analysis checkable, not correct, and it says so when you run it.
A validated version would need, at minimum: a preregistered protocol fixing the intended use, denominator and time horizon before any scoring; criteria redesigned to remove overlap, with an operational definition and a stated endpoint for each — "getting hooked" currently blends transition probability, onset speed, compulsion severity and relapse into one number; a systematic evidence review with predefined search strategies and risk-of-bias assessment for every drug-by-criterion cell, of which there would be well over a hundred; a multidisciplinary panel — addiction medicine, toxicology, epidemiology, psychiatry, health economics, criminology, lived experience — scoring independently from identical dossiers, then reconciling through structured elicitation; swing weighting for any default weights, since setting every slider to 5 asserts that a 0→100 move in acute mortality matters exactly as much as a 0→100 move in public cost, which is a claim and not a neutrality; distributions instead of digits, propagated through the ranking; reliability testing — inter-rater, test–retest, and replication by a second panel working from the same protocol; validation against sources deliberately held out from construction; transportability testing across regions, eras and supply conditions; and publication of the protocol, data and code for independent replication.
Section 07 does a small piece of this — it propagates the stated uncertainty and finds the ordering robust in bands but not in ranks. Section 02 does another small piece, badly at first and then honestly. Everything else on that list is undone, and most of it is a research programme rather than a revision. The sliders should stay exactly where they are, though: separating what the evidence says from how much each harm should count is the one structural thing this page already gets right.
Section 07 tests whether a formula could replace that judgment. Short version: harm-to-others can be derived wherever Nutt published a part-score — doing so revealed a nearly seventy-point error in methamphetamine and two further gaps, at opioids and cocaine, that remain uncorrected in the table; acute crisis is half derivable; withdrawal is ordinal only; and addiction, organ harm and harm-to-self cannot currently be derived at all. The stated ±10 accuracy is too confident — for the addiction column the real spread is closer to ±20, and harm-to-self double-counts other columns by construction.
About this project
What this instrument claims to measure, what it does not, and how to check it yourself.
The chart on this page remains the transparent legacy model. The scientist-facing framework defines the future evidence process, and four exact pilot estimands are now provisionally frozen for feasibility and content review.
Open the Research Framework · Inspect the four defined pilot questions · See the first evidence pilot
A structured, single-author comparison of the relative harm associated with recreational use of thirteen substance categories among US adults, as of 2026 — scored per person using, not per population.
- The denominator is mostly harm per user — but not entirely. Prevalence is excluded from four of the five columns: a drug is not scored as worse simply because more people take it. Damage to others is the exception, since it descends from expert-panel judgments of societal footprint, and that column is what puts alcohol first. Weight it below 3 and methamphetamine leads instead. US prevalence is shown in each drug's profile39 as context, never as an input. This is why nicotine sits mid-table here and second in a population-level analysis (section 02). Only the damage to others column carries societal footprint, and only because its source did.
- It is an exploratory MCDA, not a validated instrument. The scores come from one author reading published evidence, not from a panel, a systematic review, or a preregistered protocol. Section 07 sets out precisely what would have to change, and how much of it is missing.
- It is not personalised risk. No number here applies to a person. Individual risk turns on dose, route, frequency, tolerance, supply purity, co-use, and health history — none of which are inputs to this model.
- It is not stable across place or time. It is anchored to US 2026 conditions in a fentanyl-adulterated supply. Several columns would move substantially in another country or another decade.
- The exposure denominator is undefined, and that is a real limitation rather than a quibble. "Per user" does not distinguish someone who drinks twice a year from someone who drinks daily, and those two cannot share a meaningful harm score. A rigorous version would state a target such as expected harm per 1,000 user-years at a specified pattern of use, then score against it. This page cannot yet, so every number should be read as an implicit average over a mixed and unstated population.
- Four things that move real-world risk have no axis here at all. Route — smoking, injecting, insufflating and swallowing produce very different risks from the same molecule. Dose–response — each cell collapses a nonlinear curve into one point, so a drug that is modest at one pattern and severe at another gets a single number. Co-use — most real harm involves combinations, and the rows here are scored as though each were taken alone (section 05). Supply — for opioids and counterfeit pills, much of the present danger is uncertain potency and adulteration, which is a property of the market rather than the pharmacology, and a properly built instrument would separate intrinsic harm, route harm and current-market supply harm into different columns.
- The scores assume an unstated context. Supervised settings, sterile equipment, a known dose, a companion present, and access to treatment all substantially change outcomes, and none of them are inputs. Griffiths' hallucinogen safety work is the clearest case: under screening, preparation and monitoring, persisting adverse reactions are rare,44 which means a large share of what the acute column measures is circumstance rather than chemistry.
- It scores harm only, and harm is not the whole of a drug's effect. There is no benefit column, no therapeutic column, and no column for why anyone takes any of these in the first place. That omission is deliberate — a harm ranking is a coherent thing to build and a harm-minus-benefit ranking is not, since the two are not measured in comparable units — but it means the table is silent on a substantial clinical literature. Psilocybin sits near the floor here while being studied as a treatment for tobacco addiction,45 depression and end-of-life anxiety; ketamine is a licensed anaesthetic and an approved antidepressant; benzodiazepines and opioids are prescribed daily for good reasons. A drug scoring low here is not endorsed, and a drug scoring high is not without use.
The big-picture reads
Legal status is poorly aligned with this model's estimated harm
The two legal, culturally normalized drugs — alcohol and tobacco — do much of the real damage, while several substances people moralize about most sit at the floor. Nutt found the same: LSD, mushrooms and MDMA all ranked near the bottom despite Class A status. Note this is misalignment, not an inverse relationship — opioids, cocaine and meth are illegal and genuinely near the top. The claim is that legal status is a poor predictor of harm, not that it runs backwards.
"Most harmful" has two different answers
Per user, opioids and meth are arguably worst. For society, alcohol wins on its enormous harm-to-others. That's not a contradiction — it's the crossover. Toggle the presets above and you can watch the top of the list swap.
Dopamine explains a lot of the pattern — but not addiction
How directly a drug acts on the dopamine reward system tracks the addiction column reasonably well here: psychedelics (5-HT2A, little dopamine reward) sit at the floor, stimulants and opioids at the ceiling, the rest strung between. That is a useful intuition, not an explanation. Modern accounts of addiction also involve stress systems, learned cues and habit formation, relief of withdrawal, executive control, how fast a drug reaches the brain, and social context — which is why nicotine, a weak and inconsistent reinforcer in the laboratory,41 nonetheless produces the highest dependence rate ever measured. One axis orders this table; it does not account for it.