Pearson chi-square score for model calibration targets with its degrees of freedom

Scores a candidate parameter set by scaling each squared gap between an observed count and the model's predicted count by the predicted count, so a gap of 7 cases matters more for a target of 45 than for one of 450. When the score is read against chi-square, the NIST/SEMATECH rule subtracts the number of parameters estimated from the same counts, plus one, from the number of targets, which gives df. The one is subtracted because the predicted counts are constrained to add to the observed total, as in the article's example where both parameter sets total 350, the same as the observed counts. Treating the counts as independent with expectations fixed in advance gives m degrees of freedom instead, and the article's example shows how far the two readings can differ.

Signature

X2 = sum_(j=1)^m [(O_j - E_j)^2 / E_j]; df = m - c - 1
Inputs
InputsDefinitionUnit
O_jObserved count for target j, such as new cases in an age bandcases
E_jCount predicted by the model for target j with the candidate parameter setcases
mNumber of targets, equal to the length of the O_j and E_j listscount
cNumber of model parameters estimated or tuned against the same countscount
Output
X2Pearson statistic summed over the calibration targetsnone
dfDegrees of freedom under the NIST/SEMATECH rulenone

Function

Chi-square distribution of a sum of squared standard normal deviates

Maps the degrees of freedom k to the distribution of the sum of k independent squared standard normal variables. Health economic models meet it as the reference for test statistics, each read from the right-hand tail: a likelihood ratio between nested survival models, Cochran's Q for heterogeneity in a meta-analysis and Pearson's statistic for calibration targets. It is the gamma distribution with shape k/2 and scale 2. The formulae for Cochran's Q, I-squared and the between-study variance are on the Aggregate Data Meta-Analysis page (HE-FM-ADMA-002 to HE-FM-ADMA-004). AIC and BIC, which TSD 14 prefers for choosing between survival curves, are on the Akaike Information Criterion page (HE-FM-AIC-001 and HE-FM-AIC-004) and the Bayesian Information Criterion page (HE-FM-BIC-001).

Try this function

Implementations

  • Excel

    Pearson calibration score and degrees of freedom from two ranges

    With observed counts in a range named Observed, predicted counts in a range named Predicted and the number of calibrated parameters in CalibParams, the formulas return the score and its degrees of freedom.

    =SUMPRODUCT((Observed-Predicted)^2/Predicted); =COUNT(Observed)-CalibParams-1

Assumptions

  • Expected counts of at least 5 for a Pearson chi-square reference

    The chi-square approximation needs adequate expected counts; the NIST/SEMATECH handbook asks for at least 5 in each bin and advises combining bins in the tails when counts are lower.

  • Parameters tuned to the same targets make a Pearson p-value descriptive

    When parameters are searched over many candidate sets against the same targets, the number of degrees of freedom used up is unclear, so a p-value from X2 is descriptive and the score is better used to rank candidate sets. The one subtracted in m minus c minus 1 applies only when the predicted counts are constrained to the observed total; without that constraint it is not subtracted.

Worked examples

  • Pearson score for four age bands with the first parameter set

    Predictions of 45, 88, 112 and 105 against observed counts of 52, 81, 120 and 97 give terms of about 1.089, 0.557, 0.571 and 0.610, totalling 2.827. With two calibrated parameters the rule leaves one degree of freedom and a p-value of about 0.09, against about 0.59 on four.

    O_j = [52,81,120,97]; E_j = [45,88,112,105]; m = 4; c = 2; X2 = 2.827; df = 1
  • Pearson score for four age bands with the second parameter set

    Predictions of 50, 85, 125 and 90 score about 1.013, so this set ranks ahead of the first.

    O_j = [52,81,120,97]; E_j = [50,85,125,90]; m = 4; c = 2; X2 = 1.013; df = 1

Common errors

  • Pearson gaps divided by the observed instead of the predicted counts

    Dividing each squared gap by the observed count gives about 2.740 for the first parameter set instead of 2.827. It is a different statistic, and mixing the two across candidate sets distorts the ranking.

  • Pearson p-value quoted without its degrees of freedom rule

    The first set's score of 2.827 gives a p-value of about 0.59 when the four counts are treated as independent with expectations fixed in advance, and about 0.09 under the NIST/SEMATECH rule with two calibrated parameters. Both are defensible counts of degrees of freedom, so a p-value quoted without its rule is misleading; the score is better used to rank candidate parameter sets than as a formal test.

Sources

  • Pearson statistic and degrees of freedom in the NIST goodness-of-fit test

    NIST/SEMATECH. e-Handbook of Statistical Methods. Section 1.3.5.15, Chi-Square Goodness-of-Fit Test. National Institute of Standards and Technology; accessed 2 October 2026. Test statistic: the sum of (O minus E) squared over E across bins; critical region: chi-square with degrees of freedom equal to the number of non-empty bins minus the number of estimated parameters plus one, written in this package's notation as m minus (c + 1); and the requirement that the expected frequency be at least 5.

    View source →

  • Chi-squared among goodness-of-fit measures for model calibration

    Enns EA, Cipriano LE, Simons CT, Kong CY. Identifying best-fitting inputs in health-economic model calibration: a Pareto frontier approach. Medical Decision Making. 2015;35(2):170-182. Table 1, which lists chi-squared with likelihood and log-likelihood as probabilistic measures of fit to calibration targets.

    View source →

  • Chi-squared and likelihood measures compared in probabilistic calibration

    Karnon J, Vanni T. Calibrating models in economic evaluation: a comparison of alternative measures of goodness of fit, parameter search strategies and convergence criteria. PharmacoEconomics. 2011;29(1):51-62. Abstract, which reports that the chi-squared measure differentiated between parameter sets by accuracy to a far greater degree than the likelihood measure.

    View source →

Canonical Identity