Calculator methodology · Trust & method
How Do Actuaries Validate a Gompertz-Makeham Fit Against Observed Population Death Rates?
Actuaries validate a Gompertz–Makeham fit with three sequential checks: a chi-square goodness-of-fit test against observed death rates, an AIC comparison against rival models, and Cox–Snell residual plots.
On this page
What validation means The chi-square test The A/E ratio AIC model comparison Cox–Snell residuals Where the model fails How this reaches your number FAQsWhat does ‘validating a mortality model’ actually mean in practice?
Validation means comparing the model’s predicted death rate at each age against the observed rate from a reference life table, then applying formal statistical tests to decide whether the gap is random noise or systematic misfit.
The quantity under test is the age-specific mortality rate, written q(x) — the chance of dying within a year at age x. The model produces a predicted q(x) at every age; the reference life table supplies the observed q(x). Validation puts the two side by side across all ages and asks one question: does the difference look like ordinary sampling scatter, or a consistent pattern? A consistent pattern means the model is wrong somewhere and gets rejected; random scatter means it is accepted. The reference tables here are country-sex period life tables — which mortality dataset the calculator validates against covers where they come from.
The decision rule is directional: a systematic deviation between predicted and observed q(x) rejects the model, while random scatter around the observed rates accepts it. The benchmark is the observed q(x) column of WHO Global Health Observatory and UN World Population Prospects country-sex period life tables. WHO Global Health Observatory; UN World Population Prospects
Every test below is a different lens on that one comparison. If you want the model those tests judge, see how the Gompertz-Makeham hazard function is constructed. The first and most common test measures the fit across all ages at once.
How does the chi-square goodness-of-fit test detect a poor Gompertz-Makeham fit?
The chi-square test compares observed deaths in each age band to the deaths the model predicts. A statistically significant result — p below 0.05 — signals that the model’s hazard curve deviates from the empirical life table by more than chance variation would explain.
The test needs two counts per age band: observed deaths, O, and expected deaths, E, where E comes from the fitted model applied to that band’s population. It sums the squared gaps, scaled by the expected counts, into a single chi-square statistic. The larger the accumulated gap, the larger the statistic, and the smaller the p-value. Its degrees of freedom equal the number of age bands minus the number of estimated parameters — typically three for Gompertz–Makeham (the baseline level, the ageing slope, and the Makeham background constant). A significant result does not just say “misfit”; actuaries then inspect which bands contributed most to the statistic to see where the model is straining.
The chi-square goodness-of-fit test rejects a mortality model when p < 0.05, using observed-versus-expected deaths per age band with degrees of freedom equal to the number of bands minus the three fitted Gompertz–Makeham parameters. Pearson chi-square goodness-of-fit; standard actuarial model-validation method
The chi-square gives one verdict for the whole curve, but it cannot tell you the direction of a misfit. That is the A/E ratio’s job.
What is the Actual-to-Expected (A/E) ratio and why do actuaries use it alongside chi-square?
The A/E ratio divides observed deaths by model-predicted deaths in each age band. A ratio of 1.00 is a perfect fit; ratios consistently above or below 1.00 across adjacent bands reveal directional bias that a single chi-square p-value cannot locate.
Where chi-square returns one number for the whole curve, the A/E ratio returns one number per age band, so it maps the misfit by age. A run of A/E values above 1.00 means the model is under-predicting deaths there; a run below 1.00 means it is over-predicting. That directional, age-located reading is exactly what the aggregate chi-square statistic hides. The Gompertz–Makeham model has a well-known A/E signature at the oldest ages: in high-longevity populations it tends to over-predict deaths above 85, pushing A/E below 1.00 there.
A/E equals 1.00 at a perfect fit; a practical working rule treats A/E drifting outside a roughly 0.95–1.05 band across three or more consecutive age groups as a directional bias serious enough to re-estimate the parameters. Actual-to-Expected mortality experience analysis; standard actuarial practice
Both chi-square and A/E answer “does this model fit?” Neither answers “is this the best model available?” — which is where AIC comes in.
How does AIC comparison tell actuaries whether Gompertz-Makeham is the right model to use?
The Akaike Information Criterion (AIC) rewards fit but penalises complexity. Actuaries fit Gompertz–Makeham alongside simpler models (plain Gompertz) and more complex ones (Kannisto, Beard); the model with the lowest AIC wins. A ΔAIC above 2 over the next-best model is treated as meaningful.
AIC is defined as 2k − 2·ln(L), where k is the number of parameters and L is the maximised likelihood — how well the model reproduces the data. Adding parameters can only improve raw fit, so AIC docks 2 points per parameter to stop actuaries rewarding needless complexity. The model that balances fit against parsimony best has the lowest AIC. This is a different question from chi-square: chi-square asks whether a model fits adequately on its own; AIC asks which model fits best relative to its rivals. Adding the Makeham background term to a plain Gompertz model typically lowers AIC most in younger-adult populations, where age-independent background deaths are a meaningful share of the total.
AIC = 2k − 2·ln(L), with k the parameter count and L the maximised likelihood; the conventional actuarial threshold is that a ΔAIC greater than 2 between two models justifies preferring the lower-AIC one. Akaike Information Criterion (Akaike, 1974); standard model-selection practice
Chi-square, A/E, and AIC all work on grouped, population-level counts. To catch misfit hiding inside individual records — especially in the tails — actuaries turn to residuals.
What do Cox-Snell residuals reveal that aggregate tests miss?
Cox–Snell residuals transform individual survival times so that, if the model is correct, they follow a unit-exponential distribution. A quantile plot of these residuals exposes age-specific misfit — particularly at the tails — that stays invisible in chi-square or AIC summaries.
The idea is a clever change of scale. Under a correctly specified model, running each person’s survival time through the model’s own cumulative hazard should produce residuals with a known, standard shape: a unit-exponential distribution. Plot those residuals against exponential quantiles and, if the model is right, they land on a 45-degree line. Systematic departures from that line reveal where the model misfits — and because this works at the individual-record level, it catches problems the grouped aggregate tests smooth over. The Gompertz–Makeham model’s residuals characteristically bend above the 45-degree line at the oldest ages, the visual signature of it underestimating survival at extreme old age.
Under a correct model, Cox–Snell residuals follow a unit-exponential distribution and fall on a 45-degree line against exponential quantiles; a curve above that line at the oldest ages signals the Gompertz–Makeham model underestimating survival there. Cox & Snell, J. R. Statist. Soc. B, 1968
That tail behaviour is not an accident of one dataset — it is a structural property of the model, with a specific age where it starts to break.
At what ages does the Gompertz-Makeham model typically fail validation, and why?
The model reliably fits ages 30–85. Above 85, its exponential hazard assumption over-predicts observed mortality — a known structural limitation called mortality deceleration — which causes systematic validation failures in the very oldest, supercentenarian data.
Across the working adult range, roughly 30 to 85, the exponential Gompertz term tracks real death rates closely and the fit passes the tests above. Beyond about 85, observed death rates stop accelerating and begin to plateau, while the model’s exponential term keeps climbing — so the model predicts more deaths than actually occur. This gap is mortality deceleration, and it is the model’s most-documented weakness; it is why period tables are used as the validation benchmark rather than sparse extreme-age data. Actuaries handle it two ways: cap the hazard at a maximum yearly value, or switch above about age 90 to a logistic-type (Kannisto) model that lets the curve plateau.
The Gompertz–Makeham model fits well across ages 30–85 but fails above 85 because observed mortality decelerates (the hazard plateaus) while the exponential term keeps rising — corrected in practice by capping the hazard or switching to a Kannisto logistic model above about age 90. Mortality-deceleration / late-life plateau; gamma-Gompertz and Kannisto extensions, demographic literature
Knowing exactly where a model stops being trustworthy is itself a validation result — and it is what connects all this testing to the number the calculator actually shows you.
How does this validation process connect to the numbers the calculator shows you?
A parameter set only earns its place after passing these checks — a non-significant chi-square, an A/E near 1.00, and the lowest AIC among the models tested — against the reference life tables. When a fit fails those checks, the response is to re-estimate the parameters, not to ship the misfit.
This is what “trust the number” means concretely: the estimate you see rests on a parameter set that reproduced real observed death rates closely enough to pass standard actuarial tests, within the age range where the model is known to hold. Outside that range — at the extreme oldest ages — the honest move is the correction described above, not a silent over-estimate. That is the whole reason a validation node exists: to show the checks behind the output rather than ask you to take the output on faith. To see how those validated parameters become a personal figure, read how validated parameters become your personalised estimate, and for the wider picture, how actuarial accuracy compares across calculator types. The practical takeaway is simple: treat a life-expectancy number as trustworthy only when its model names the tests it passed, the data it was checked against, and the ages where it stops working — then read your own result as a validated range, not a verdict. Get your survivorship-adjusted estimate.
Frequently asked questions about validating a Gompertz-Makeham fit
How do actuaries test if a mortality model fits real data?
They compare the model’s predicted death rate at each age with the observed rate from a reference life table, then apply formal tests — chi-square goodness-of-fit, the Actual-to-Expected ratio, AIC comparison, and Cox–Snell residual plots — to judge whether the gaps are random noise or systematic misfit.
What is a goodness-of-fit test in actuarial science?
A goodness-of-fit test measures how closely a model’s predicted values match observed data. In mortality work, it compares predicted deaths to observed deaths across age bands and returns a statistic — such as chi-square — whose p-value indicates whether any mismatch is beyond chance.
What is the chi-square test used for in life table validation?
It tests whether a fitted model’s expected deaths match the observed deaths in a life table across age bands. A p-value below 0.05 signals systematic misfit; degrees of freedom equal the number of age bands minus the number of fitted parameters (three for Gompertz–Makeham).
How is AIC used to compare mortality models?
AIC = 2k − 2·ln(L) scores each model by fit while penalising extra parameters. Actuaries fit several models to the same data and prefer the one with the lowest AIC; a difference greater than 2 over the next-best model is treated as a meaningful improvement.
What are Cox-Snell residuals in survival analysis?
They are transformed survival times that follow a unit-exponential distribution if the model is correct. Plotted against exponential quantiles, they should fall on a 45-degree line; departures — especially at the tails — reveal where the model misfits at the individual-record level.
How do actuaries know when a Gompertz model fails?
Failure shows up as a significant chi-square, an A/E ratio drifting away from 1.00 across consecutive age bands, a higher AIC than a rival model, or Cox–Snell residuals bending off the 45-degree line — most often at ages above 85, where mortality decelerates.
What does A/E ratio mean in mortality experience studies?
A/E is observed deaths divided by model-expected deaths in an age band. A value of 1.00 is a perfect match; consistently above 1.00 means the model under-predicts deaths, consistently below means it over-predicts — locating directional bias the aggregate chi-square cannot.

Leave a Reply