Dirichlet posterior parameter, mean and beta marginal for one destination from transition counts and a Dirichlet prior

A Dirichlet prior combined with multinomial counts gives a Dirichlet posterior in which each parameter is its prior parameter plus its count. The posterior mean of one destination is that parameter over the prior total plus the number of patients, and the destination's probability on its own follows Beta(alpha_j, alpha_0 minus alpha_j), whose second parameter is written b_j here. A zero prior reproduces the observed proportions; a prior of one per destination is uniform over all valid rows.

Signature

alpha_j = a_j + n_j; b_j = A_0 + N - alpha_j; E_j = alpha_j / (A_0 + N)
Inputs
InputsDefinitionUnit
a_j0 for the improper zero prior, 1 for the uniform priornone
n_jPatients who ended the interval in destination jpatients
A_0K times the prior per destination when the prior is the same for eachnone
NSum of the counts over all destinationspatients
Output
alpha_jUnit: none, must be above zero—
b_jPosterior parameters of all other destinations added togethernone
E_jUnit: probability—

Function

Dirichlet distribution for one row of a transition matrix: moments, conjugate update from transition counts and beta marginals

Treats the probabilities of moving from one health state to each of K destinations as a single uncertain vector that always adds to one, with one positive parameter per destination. The ratios of the parameters fix the means and their sum fixes the precision; observed transition counts add to a Dirichlet prior to give a Dirichlet posterior, and each component on its own follows a beta distribution. Sampling a row from normalised gamma draws is HE-FM-BETA-005 on the beta distribution page. Notation follows the Dirichlet Distribution article.

Computational function

  • Computational function: Dirichlet row from transition counts and a prior with beta-marginal intervals, correlations and a sampled row

    Fits a Dirichlet row from a vector of transition counts and a prior, returns every component's mean, standard deviation and central interval from its beta marginal with the full correlation matrix, and, given uniform numbers, returns one sampled row from normalised gamma quantiles. Beta quantiles and gamma quantiles have no closed form, so this procedure cannot run in the site calculator; Excel BETA.INV and GAMMA.INV, R qbeta and qgamma and Python scipy.stats return them. The inputs differ from the formulas': a count vector of any length, a prior, a coverage level and optional uniform numbers.

    Inputs and outputs: n: Transition counts to each destination; required. Unit: patients.; a: Prior parameter, one value for every destination or one per destination; default 0. Unit: none.; u: Uniform numbers, one per destination, for one sampled row; optional. Unit: probability.; level: Coverage of the central interval; default 0.95. Unit: proportion.; alpha: Posterior parameters (HE-FM-DIRI-002). Unit: none.; mean, sd: Mean and standard deviation of each component (HE-FM-DIRI-001). Unit: probability.; corr: Correlation matrix of the components. Unit: none.; lower, upper: Central interval of each beta marginal. Unit: probability.; gamma, row: Gamma quantiles with shape alpha and scale 1, and the sampled row they give. Unit: none; probability.

    Assumption: The counts describe one starting state over the model cycle, structural zeros have been left out, and the uniform numbers are independent, one per destination, in each iteration.

    Worked example (Dirichlet(70, 20, 10) with fixed uniform numbers): The means are 0.70, 0.20 and 0.10 with standard deviations of about 0.0456, 0.0398 and 0.0299, the correlation of the first two components is about minus 0.76 and the progression interval runs from about 0.128 to 0.283; uniform numbers of 0.30, 0.80 and 0.50 give gamma quantiles of about 65.38, 23.63 and 9.67 and a sampled row of 0.6625, 0.2395 and 0.0980, which the article shows as 0.663, 0.239 and 0.098. n = 70, 20, 10; a = 0; u = 0.30, 0.80, 0.50; lower_2 = 0.1280; upper_2 = 0.2834; row = 0.6625, 0.2395, 0.0980

    Worked example (Uniform prior with a zero count): Counts of 0, 45 and 15 with a prior of 1 give Dirichlet(1, 46, 16), a return mean of 0.0159 and an interval of about 0.0004 to 0.0578, as in the article; the same counts with the zero prior stop with a message, because a parameter of zero gives no valid distribution. n = 0, 45, 15; a = 1; alpha = 1, 46, 16; mean_1 = 0.0159; lower_1 = 0.0004; upper_1 = 0.0578

    Excel: With the posterior parameters in AlphaVec and the uniform numbers in UVec, =GAMMA.INV(UVec,AlphaVec,1) returns the gamma quantiles into GammaVec and =GammaVec/SUM(GammaVec) the sampled row into RowVec; =AlphaVec/SUM(AlphaVec) gives MeanVec, and =BETA.INV(0.025,AlphaVec,SUM(AlphaVec)-AlphaVec) and the same with 0.975 give LowerVec and UpperVec. In a probabilistic analysis UVec holds RAND() in each cell.

    R: dir_row <- function(n, a = 0, u = NULL, level = 0.95) { alpha <- n+a; if (any(alpha <= 0)) stop("every parameter must be above zero: add a prior or leave out a structural zero"); a0 <- sum(alpha); v <- alpha*(a0-alpha)/(a0^2*(a0+1)); cv <- -outer(alpha, alpha)/(a0^2*(a0+1)); diag(cv) <- v; q <- (1-level)/2; out <- list(alpha = alpha, mean = alpha/a0, sd = sqrt(v), corr = cv/sqrt(outer(v, v)), lower = qbeta(q, alpha, a0-alpha), upper = qbeta(1-q, alpha, a0-alpha)); if (!is.null(u)) { g <- qgamma(u, alpha, 1); out$gamma <- g; out$row <- g/sum(g) }; out } Base R only; dir_row(c(70, 20, 10), 0, c(0.30, 0.80, 0.50)) returns the first example.

    Python: def dir_row(n, a=0, u=None, level=0.95): alpha = [x+y for x, y in zip(n, a if isinstance(a, list) else [a]*len(n))]; assert min(alpha) > 0, "every parameter must be above zero"; a0 = sum(alpha); v = [x*(a0-x)/(a0**2*(a0+1)) for x in alpha]; k = range(len(alpha)); cov = [[v[i] if i == j else -alpha[i]*alpha[j]/(a0**2*(a0+1)) for j in k] for i in k]; q = (1-level)/2; g = [stats.gamma.ppf(x, s) for x, s in zip(u, alpha)] if u else None; return {"alpha": alpha, "mean": [x/a0 for x in alpha], "sd": [math.sqrt(x) for x in v], "corr": [[cov[i][j]/math.sqrt(v[i]*v[j]) for j in k] for i in k], "lower": [stats.beta.ppf(q, x, a0-x) for x in alpha], "upper": [stats.beta.ppf(1-q, x, a0-x) for x in alpha], "gamma": g, "row": [x/sum(g) for x in g] if g else None} Needs import math and from scipy import stats; returns the same values as the R function.

    Test (Sampled row adds to one and every mean lies inside its interval): The sampled row adds to 1 and every component's mean lies between its lower and upper limits. Expected result: TRUE. FALSE shows the gamma quantiles divided by the parameter total of 100 in place of their own total of 98.69, which gives a row adding to 0.987. Excel check: =AND(ABS(SUM(RowVec)-1)<1E-12,SUMPRODUCT(--(MeanVec>=LowerVec),--(MeanVec<=UpperVec))=COUNT(MeanVec))

    Common error (One uniform number shared by every destination): Feeding the same RAND() cell to all three GAMMA.INV calls makes the gamma quantiles move together instead of independently, so the row no longer follows the Dirichlet; with u of 0.5 for every destination the progression probability would be close to its mean in every iteration.

    Source: Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Appendix A, Dirichlet sampling: draw x_1, ..., x_k from independent gamma distributions with common scale and shape parameters alpha_1, ..., alpha_k and let theta_j = x_j over their sum; Microsoft. GAMMA.INV function. Microsoft Support; accessed 2 October 2026 (full text read). If beta = 1, GAMMA.INV returns the standard gamma distribution, and if alpha is 0 or less it returns the #NUM! error value; Microsoft. BETA.INV function. Microsoft Support; accessed 2 October 2026 (full text read). Returns the inverse of the cumulative beta probability density function.

    alpha_j = a_j + n_j; alpha_0 = sum_(j=1)^K alpha_j; mean_j = alpha_j / alpha_0; sd_j = sqrt(alpha_j * (alpha_0 - alpha_j) / (alpha_0^2 * (alpha_0 + 1))); lower_j = F^-1((1 - level) / 2 | alpha_j, alpha_0 - alpha_j); upper_j = F^-1((1 + level) / 2 | alpha_j, alpha_0 - alpha_j); g_j = G^-1(u_j | alpha_j, 1); row_j = g_j / sum_(k=1)^K g_k

Try this function

Implementations

  • Excel

    Posterior Dirichlet parameter, beta marginal and mean from named cells

    With PriorJ, CountJ, PriorSum and CountSum named, the formulas return the posterior parameter, the second parameter of the beta marginal and the posterior mean, held in AlphaJ, BetaJ and MeanJ. With the posterior parameters of the whole row in a range named AlphaRow, the row checks below apply.

    =PriorJ+CountJ; =PriorSum+CountSum-AlphaJ; =AlphaJ/(PriorSum+CountSum)

Assumptions

  • Counts from the population, starting state and interval of the model

    The counts come from patients in the same starting state, followed over an interval equal to the model cycle without irregular or censored follow-up, each counted in exactly one destination.

  • Zero prior valid only when every destination was observed

    Setting every a_j to zero gives an improper prior whose posterior is proper only if each destination has at least one observation; a clinically impossible transition, such as a return from death, is left out of the row as a structural zero rather than given prior weight.

Worked examples

  • Progression count of 20 out of 100 with the zero prior

    The posterior parameter is 20, the beta marginal is Beta(20, 80) and the mean is the observed proportion of 0.20, as in the article; the central 95 per cent of that beta runs from about 0.128 to 0.283.

    a_j = 0; n_j = 20; A_0 = 0; N = 100; alpha_j = 20; b_j = 80; E_j = 0.2
  • Progression count of 20 out of 100 with a uniform prior over three destinations

    One is added to each of the three counts, so the parameter is 21 and the mean 21 / 103 = 0.204, as in the article's Dirichlet(71, 21, 11); with 100 patients the prior changes little.

    a_j = 1; n_j = 20; A_0 = 3; N = 100; alpha_j = 21; b_j = 82; E_j = 0.2039
  • Return from the progressed state never observed in 60 patients

    With 0 returning, 45 remaining progressed and 15 dying, the zero prior gives no valid distribution, while the uniform prior gives Dirichlet(1, 46, 16), a mean return probability of 1 / 63 = 0.016 and a beta marginal Beta(1, 62) whose central 95 per cent runs from about 0.0004 to 0.058, as in the article.

    a_j = 1; n_j = 0; A_0 = 3; N = 60; alpha_j = 1; b_j = 62; E_j = 0.0159

Common errors

  • Replacing a zero count with a small arbitrary number

    Substituting, say, 0.01 for a zero count makes the spreadsheet run but hides the prior it implies; the article asks for the prior to be reported, and a uniform prior adds one to every destination, not to the zero cell alone.

  • Building a row from a percentage and an assumed sample size

    Multiplying a published percentage by an assumed number of patients sets the precision alpha_0 by assumption, which sits uneasily with NICE's requirement that distributions are not chosen arbitrarily but represent the available evidence.

  • Giving prior weight to a transition that cannot happen

    A uniform prior over a row that includes an impossible move, such as a return from death, makes that probability positive in every draw; structural zeros are left out of the row instead.

Sources

  • Conjugate Dirichlet update and noninformative priors in Bayesian Data Analysis

    Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Section 3.4: the posterior distribution for the multinomial probabilities is Dirichlet with parameters alpha_j + y_j; a uniform density is obtained by setting alpha_j = 1 for all j; setting alpha_j = 0 for all j gives an improper prior, and the posterior is proper if there is at least one observation in each of the k categories.

    View source →

  • Beta marginal of a single Dirichlet component in Bayesian Data Analysis

    Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Appendix A: the Dirichlet is the conjugate prior distribution for the parameters of the multinomial distribution and a multivariate generalisation of the beta; the marginal distribution of a single theta_j is Beta(alpha_j, alpha_0 minus alpha_j).

    View source →

  • Briggs, Ades and Price on zero transition counts

    Briggs AH, Ades AE, Price MJ. Probabilistic sensitivity analysis for decision trees with multiple branches: use of the Dirichlet distribution in a Bayesian framework. Medical Decision Making. 2003;23(4):341-350. doi:10.1177/0272989X03255922 (abstract read). Abstract: by adopting a Bayesian approach, the problem of observing zero counts for transitions of interest can be overcome.

    View source →

  • NICE PMG36 on distributions that represent the available evidence

    National Institute for Health and Care Excellence. NICE technology appraisal and highly specialised technologies guidance: the manual (PMG36). London: NICE; published 31 January 2022, last updated 31 March 2026 (full text of chapter 4 read). Section 4.7.10: the distributions chosen for probabilistic sensitivity analysis should not be chosen arbitrarily but chosen to represent the available evidence on the parameter of interest, and their use should be justified.

    View source →

Canonical Identity