Signature
X2 = sum_(j=1)^m [(O_j - E_j)^2 / E_j]; df = m - c - 1
| Inputs | Definition | Unit |
|---|---|---|
O_j | Observed count for target j, such as new cases in an age band | cases |
E_j | Count predicted by the model for target j with the candidate parameter set | cases |
m | Number of targets, equal to the length of the O_j and E_j lists | count |
c | Number of model parameters estimated or tuned against the same counts | count |
X2 | Pearson statistic summed over the calibration targets | none |
|---|---|---|
df | Degrees of freedom under the NIST/SEMATECH rule | none |
Function
Chi-square distribution of a sum of squared standard normal deviates
Maps the degrees of freedom k to the distribution of the sum of k independent squared standard normal variables. Health economic models meet it as the reference for test statistics, each read from the right-hand tail: a likelihood ratio between nested survival models, Cochran's Q for heterogeneity in a meta-analysis and Pearson's statistic for calibration targets. It is the gamma distribution with shape k/2 and scale 2. The formulae for Cochran's Q, I-squared and the between-study variance are on the Aggregate Data Meta-Analysis page (HE-FM-ADMA-002 to HE-FM-ADMA-004). AIC and BIC, which TSD 14 prefers for choosing between survival curves, are on the Akaike Information Criterion page (HE-FM-AIC-001 and HE-FM-AIC-004) and the Bayesian Information Criterion page (HE-FM-BIC-001).
Try this function
Implementations
Excel
Pearson calibration score and degrees of freedom from two ranges
With observed counts in a range named Observed, predicted counts in a range named Predicted and the number of calibrated parameters in CalibParams, the formulas return the score and its degrees of freedom.
=SUMPRODUCT((Observed-Predicted)^2/Predicted); =COUNT(Observed)-CalibParams-1
Assumptions
Expected counts of at least 5 for a Pearson chi-square reference
The chi-square approximation needs adequate expected counts; the NIST/SEMATECH handbook asks for at least 5 in each bin and advises combining bins in the tails when counts are lower.
Parameters tuned to the same targets make a Pearson p-value descriptive
When parameters are searched over many candidate sets against the same targets, the number of degrees of freedom used up is unclear, so a p-value from X2 is descriptive and the score is better used to rank candidate sets. The one subtracted in m minus c minus 1 applies only when the predicted counts are constrained to the observed total; without that constraint it is not subtracted.
Worked examples
Pearson score for four age bands with the first parameter set
Predictions of 45, 88, 112 and 105 against observed counts of 52, 81, 120 and 97 give terms of about 1.089, 0.557, 0.571 and 0.610, totalling 2.827. With two calibrated parameters the rule leaves one degree of freedom and a p-value of about 0.09, against about 0.59 on four.
O_j = [52,81,120,97]; E_j = [45,88,112,105]; m = 4; c = 2; X2 = 2.827; df = 1
Pearson score for four age bands with the second parameter set
Predictions of 50, 85, 125 and 90 score about 1.013, so this set ranks ahead of the first.
O_j = [52,81,120,97]; E_j = [50,85,125,90]; m = 4; c = 2; X2 = 1.013; df = 1
Common errors
Pearson gaps divided by the observed instead of the predicted counts
Dividing each squared gap by the observed count gives about 2.740 for the first parameter set instead of 2.827. It is a different statistic, and mixing the two across candidate sets distorts the ranking.
Pearson p-value quoted without its degrees of freedom rule
The first set's score of 2.827 gives a p-value of about 0.59 when the four counts are treated as independent with expectations fixed in advance, and about 0.09 under the NIST/SEMATECH rule with two calibrated parameters. Both are defensible counts of degrees of freedom, so a p-value quoted without its rule is misleading; the score is better used to rank candidate parameter sets than as a formal test.
Sources
Pearson statistic and degrees of freedom in the NIST goodness-of-fit test
NIST/SEMATECH. e-Handbook of Statistical Methods. Section 1.3.5.15, Chi-Square Goodness-of-Fit Test. National Institute of Standards and Technology; accessed 2 October 2026. Test statistic: the sum of (O minus E) squared over E across bins; critical region: chi-square with degrees of freedom equal to the number of non-empty bins minus the number of estimated parameters plus one, written in this package's notation as m minus (c + 1); and the requirement that the expected frequency be at least 5.
Chi-squared among goodness-of-fit measures for model calibration
Enns EA, Cipriano LE, Simons CT, Kong CY. Identifying best-fitting inputs in health-economic model calibration: a Pareto frontier approach. Medical Decision Making. 2015;35(2):170-182. Table 1, which lists chi-squared with likelihood and log-likelihood as probabilistic measures of fit to calibration targets.
Chi-squared and likelihood measures compared in probabilistic calibration
Karnon J, Vanni T. Calibrating models in economic evaluation: a comparison of alternative measures of goodness of fit, parameter search strategies and convergence criteria. PharmacoEconomics. 2011;29(1):51-62. Abstract, which reports that the chi-squared measure differentiated between parameter sets by accuracy to a far greater degree than the likelihood measure.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0