Pick a country, an age, and a sex, and I give you a ranked list of what people like you eventually die of, with a rough likelihood next to each. The numbers come from the WHO, cover 185 countries, and describe a population rather than you. Everything is computed once at build time and written to one small JSON file per country. The build recomputes those files from the committed intermediate data and fails if a value differs, which once caught tied probabilities sorting in a different order on the Linux runner.
Eventually, not this year
The snapshot version of this question, what did people who died at your age die of, misleads. A 25-year-old is far more likely to die in a car crash this year than of heart disease, but most 25-year-olds survive the year, and over a lifetime the slow diseases win. So I answer the forward question. Of everyone alive at your age today, what fraction will eventually die of each cause? Everyone dies of something, so the shares sum to one. Dying of one cause removes you from the pool for all the others, which makes this a competing-risks calculation.
Two sources for the cause split
Survival, the chance of making it from one age band to the next, comes from the WHO Global Health Observatory life tables, which exist for every member state on five-year bands from 2000 to 2021. They decide when a cohort dies. What it dies of needs a second source, and I use one of two depending on the country.
Where the WHO Mortality Database has any registered deaths coded in ICD-10 between 2000 and 2019, I use it. That's 101 countries. Fifty of them have all 20 years, 24 have fewer than 15, and Lebanon has one. The other 84 countries fall back to the Global Health Estimates 2021 files, which are modeled rather than registered, come on seven coarse age bands, and carry less cause detail. The site badges those as modeled and gives them wider intervals.
Both sources end up on the GHE cause list. Each ICD-10 code maps to the most specific GHE range it falls in. The Mortality Database keys countries by WHO numeric code, so there's a crosswalk to ISO 3166-1 alpha-3 to line the three sources up.
Death certificates are noisy. A real share of them carry garbage codes: unspecified heart failure, cancer of unknown primary site, injury of undetermined intent, symptoms with no cause named. Left alone, "unspecified" outranks real causes. Following the GHE method, I redistribute each garbage code's deaths onto plausible real causes within the same country, year, sex, and age cell, in proportion to the deaths already there. Unspecified cardiovascular deaths go to cardiovascular causes, unknown-site cancers to the malignancies, undetermined-intent injuries to the injury causes.
Carrying the rates forward
Today's rates frozen for 80 years would be wrong, so both halves of the model are fit on 2000 to 2019 and projected. I leave the pandemic years out of the fit.
The all-cause level is Lee-Carter (Lee and Carter, 1992), per country and sex. It factors the log death rate into an age profile and one time index,
\[\ln m(x,t) = a(x) + b(x)\,k(t),\]taken from the first singular vector of the centered log-rate matrix. I forecast \(k\) as a random walk with drift and damp the drift by 0.98 a year, so an 80-year projection saturates instead of running off in a straight line. Projected rates are floored at half the lowest rate ever observed for that age.
The cause mix within an age band is a composition, so I forecast it in centered-log-ratio space. Fit a linear trend per cause, shrink the slope toward zero for rare causes, damp it by 0.9 a year so it saturates within about a decade, and transform back. Every forecast stays positive and sums to one, and stays near the recent mix instead of chasing a 20-year slope across a lifetime.
The cohort life table
A cohort alive at age \(x\) today reaches age \(x+k\) in year \(2025+k\), so each future band is evaluated on that year's projected rates. Within a band of width \(w\) I treat the hazard as constant, so the chance of dying in it given survival to it is \(q = 1 - e^{-w\,m}\). The WHO tables use their own within-band adjustments, and my \(q\) differs from their published values by up to 0.004 in infancy and by more in the oldest bands. With \(S_b\) for survival to band \(b\) and \(f_c(b)\) for the projected share of that band's deaths due to cause \(c\),
\[\pi_c = \sum_b S_b\,q_b\,f_c(b).\]The open 85+ band takes everyone left, so the shares sum to one by construction. For men aged 55 to 59 in Japan cancer comes first at about one in three, with lower respiratory infections, stroke, and other circulatory disease clustered near one in twelve. For men aged 40 to 44 in the United States cancer leads at about a fifth, with dementia and ischemic heart disease just behind.
Uncertainty
Each cohort is run 200 times with the Lee-Carter drift and the cause trends perturbed by their estimated standard errors, and the interval is the 5th to 95th percentile of the draws. That band is narrower than a proper random walk would give. I apply the drift error linearly in the horizon and don't simulate the walk's own year-to-year noise, so the spread grows in proportion to the horizon and comes from the drift estimate alone. It's then clamped by hand-set constants, wider for modeled countries, and no half-width exceeds 20 percentage points. The edges of the band are mine, which is why the site leads with rank order and rounds to "about 1 in 3".
The composition forecast starts from each country's 2019 row, and ten countries in the registered tier (the Bahamas, Iran, Iraq, Morocco, Norway, Puerto Rico, Saudi Arabia, El Salvador, Tunisia, and Venezuela) have no 2019 data. The fit fills missing years with zero deaths, so those ten start from a near-uniform cause mix and their rankings aren't reliable.