Signature
alpha_j = a_j + n_j; b_j = A_0 + N - alpha_j; E_j = alpha_j / (A_0 + N)
| Inputs | Definition | Unit |
|---|---|---|
a_j | 0 for the improper zero prior, 1 for the uniform prior | none |
n_j | Patients who ended the interval in destination j | patients |
A_0 | K times the prior per destination when the prior is the same for each | none |
N | Sum of the counts over all destinations | patients |
alpha_j | Unit: none, must be above zero | — |
|---|---|---|
b_j | Posterior parameters of all other destinations added together | none |
E_j | Unit: probability | — |
Function
Dirichlet distribution for one row of a transition matrix: moments, conjugate update from transition counts and beta marginals
Treats the probabilities of moving from one health state to each of K destinations as a single uncertain vector that always adds to one, with one positive parameter per destination. The ratios of the parameters fix the means and their sum fixes the precision; observed transition counts add to a Dirichlet prior to give a Dirichlet posterior, and each component on its own follows a beta distribution. Sampling a row from normalised gamma draws is HE-FM-BETA-005 on the beta distribution page. Notation follows the Dirichlet Distribution article.
Computational function
Computational function: Dirichlet row from transition counts and a prior with beta-marginal intervals, correlations and a sampled row
Fits a Dirichlet row from a vector of transition counts and a prior, returns every component's mean, standard deviation and central interval from its beta marginal with the full correlation matrix, and, given uniform numbers, returns one sampled row from normalised gamma quantiles. Beta quantiles and gamma quantiles have no closed form, so this procedure cannot run in the site calculator; Excel BETA.INV and GAMMA.INV, R qbeta and qgamma and Python scipy.stats return them. The inputs differ from the formulas': a count vector of any length, a prior, a coverage level and optional uniform numbers.
Inputs and outputs:
n: Transition counts to each destination; required. Unit: patients.;a: Prior parameter, one value for every destination or one per destination; default 0. Unit: none.;u: Uniform numbers, one per destination, for one sampled row; optional. Unit: probability.;level: Coverage of the central interval; default 0.95. Unit: proportion.;alpha: Posterior parameters (HE-FM-DIRI-002). Unit: none.;mean,sd: Mean and standard deviation of each component (HE-FM-DIRI-001). Unit: probability.;corr: Correlation matrix of the components. Unit: none.;lower,upper: Central interval of each beta marginal. Unit: probability.;gamma,row: Gamma quantiles with shape alpha and scale 1, and the sampled row they give. Unit: none; probability.Assumption: The counts describe one starting state over the model cycle, structural zeros have been left out, and the uniform numbers are independent, one per destination, in each iteration.
Worked example (Dirichlet(70, 20, 10) with fixed uniform numbers): The means are 0.70, 0.20 and 0.10 with standard deviations of about 0.0456, 0.0398 and 0.0299, the correlation of the first two components is about minus 0.76 and the progression interval runs from about 0.128 to 0.283; uniform numbers of 0.30, 0.80 and 0.50 give gamma quantiles of about 65.38, 23.63 and 9.67 and a sampled row of 0.6625, 0.2395 and 0.0980, which the article shows as 0.663, 0.239 and 0.098.
n = 70, 20, 10; a = 0; u = 0.30, 0.80, 0.50; lower_2 = 0.1280; upper_2 = 0.2834; row = 0.6625, 0.2395, 0.0980Worked example (Uniform prior with a zero count): Counts of 0, 45 and 15 with a prior of 1 give Dirichlet(1, 46, 16), a return mean of 0.0159 and an interval of about 0.0004 to 0.0578, as in the article; the same counts with the zero prior stop with a message, because a parameter of zero gives no valid distribution.
n = 0, 45, 15; a = 1; alpha = 1, 46, 16; mean_1 = 0.0159; lower_1 = 0.0004; upper_1 = 0.0578Excel: With the posterior parameters in AlphaVec and the uniform numbers in UVec,
=GAMMA.INV(UVec,AlphaVec,1)returns the gamma quantiles into GammaVec and=GammaVec/SUM(GammaVec)the sampled row into RowVec;=AlphaVec/SUM(AlphaVec)gives MeanVec, and=BETA.INV(0.025,AlphaVec,SUM(AlphaVec)-AlphaVec)and the same with 0.975 give LowerVec and UpperVec. In a probabilistic analysis UVec holds RAND() in each cell.R:
dir_row <- function(n, a = 0, u = NULL, level = 0.95) { alpha <- n+a; if (any(alpha <= 0)) stop("every parameter must be above zero: add a prior or leave out a structural zero"); a0 <- sum(alpha); v <- alpha*(a0-alpha)/(a0^2*(a0+1)); cv <- -outer(alpha, alpha)/(a0^2*(a0+1)); diag(cv) <- v; q <- (1-level)/2; out <- list(alpha = alpha, mean = alpha/a0, sd = sqrt(v), corr = cv/sqrt(outer(v, v)), lower = qbeta(q, alpha, a0-alpha), upper = qbeta(1-q, alpha, a0-alpha)); if (!is.null(u)) { g <- qgamma(u, alpha, 1); out$gamma <- g; out$row <- g/sum(g) }; out }Base R only;dir_row(c(70, 20, 10), 0, c(0.30, 0.80, 0.50))returns the first example.Python:
def dir_row(n, a=0, u=None, level=0.95): alpha = [x+y for x, y in zip(n, a if isinstance(a, list) else [a]*len(n))]; assert min(alpha) > 0, "every parameter must be above zero"; a0 = sum(alpha); v = [x*(a0-x)/(a0**2*(a0+1)) for x in alpha]; k = range(len(alpha)); cov = [[v[i] if i == j else -alpha[i]*alpha[j]/(a0**2*(a0+1)) for j in k] for i in k]; q = (1-level)/2; g = [stats.gamma.ppf(x, s) for x, s in zip(u, alpha)] if u else None; return {"alpha": alpha, "mean": [x/a0 for x in alpha], "sd": [math.sqrt(x) for x in v], "corr": [[cov[i][j]/math.sqrt(v[i]*v[j]) for j in k] for i in k], "lower": [stats.beta.ppf(q, x, a0-x) for x in alpha], "upper": [stats.beta.ppf(1-q, x, a0-x) for x in alpha], "gamma": g, "row": [x/sum(g) for x in g] if g else None}Needsimport mathandfrom scipy import stats; returns the same values as the R function.Test (Sampled row adds to one and every mean lies inside its interval): The sampled row adds to 1 and every component's mean lies between its lower and upper limits. Expected result: TRUE. FALSE shows the gamma quantiles divided by the parameter total of 100 in place of their own total of 98.69, which gives a row adding to 0.987. Excel check:
=AND(ABS(SUM(RowVec)-1)<1E-12,SUMPRODUCT(--(MeanVec>=LowerVec),--(MeanVec<=UpperVec))=COUNT(MeanVec))Common error (One uniform number shared by every destination): Feeding the same RAND() cell to all three GAMMA.INV calls makes the gamma quantiles move together instead of independently, so the row no longer follows the Dirichlet; with u of 0.5 for every destination the progression probability would be close to its mean in every iteration.
Source: Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Appendix A, Dirichlet sampling: draw x_1, ..., x_k from independent gamma distributions with common scale and shape parameters alpha_1, ..., alpha_k and let theta_j = x_j over their sum; Microsoft. GAMMA.INV function. Microsoft Support; accessed 2 October 2026 (full text read). If beta = 1, GAMMA.INV returns the standard gamma distribution, and if alpha is 0 or less it returns the #NUM! error value; Microsoft. BETA.INV function. Microsoft Support; accessed 2 October 2026 (full text read). Returns the inverse of the cumulative beta probability density function.
alpha_j = a_j + n_j; alpha_0 = sum_(j=1)^K alpha_j; mean_j = alpha_j / alpha_0; sd_j = sqrt(alpha_j * (alpha_0 - alpha_j) / (alpha_0^2 * (alpha_0 + 1))); lower_j = F^-1((1 - level) / 2 | alpha_j, alpha_0 - alpha_j); upper_j = F^-1((1 + level) / 2 | alpha_j, alpha_0 - alpha_j); g_j = G^-1(u_j | alpha_j, 1); row_j = g_j / sum_(k=1)^K g_k
Try this function
Implementations
Excel
Posterior Dirichlet parameter, beta marginal and mean from named cells
With PriorJ, CountJ, PriorSum and CountSum named, the formulas return the posterior parameter, the second parameter of the beta marginal and the posterior mean, held in AlphaJ, BetaJ and MeanJ. With the posterior parameters of the whole row in a range named AlphaRow, the row checks below apply.
=PriorJ+CountJ; =PriorSum+CountSum-AlphaJ; =AlphaJ/(PriorSum+CountSum)
Assumptions
Counts from the population, starting state and interval of the model
The counts come from patients in the same starting state, followed over an interval equal to the model cycle without irregular or censored follow-up, each counted in exactly one destination.
Zero prior valid only when every destination was observed
Setting every a_j to zero gives an improper prior whose posterior is proper only if each destination has at least one observation; a clinically impossible transition, such as a return from death, is left out of the row as a structural zero rather than given prior weight.
Worked examples
Progression count of 20 out of 100 with the zero prior
The posterior parameter is 20, the beta marginal is Beta(20, 80) and the mean is the observed proportion of 0.20, as in the article; the central 95 per cent of that beta runs from about 0.128 to 0.283.
a_j = 0; n_j = 20; A_0 = 0; N = 100; alpha_j = 20; b_j = 80; E_j = 0.2
Progression count of 20 out of 100 with a uniform prior over three destinations
One is added to each of the three counts, so the parameter is 21 and the mean 21 / 103 = 0.204, as in the article's Dirichlet(71, 21, 11); with 100 patients the prior changes little.
a_j = 1; n_j = 20; A_0 = 3; N = 100; alpha_j = 21; b_j = 82; E_j = 0.2039
Return from the progressed state never observed in 60 patients
With 0 returning, 45 remaining progressed and 15 dying, the zero prior gives no valid distribution, while the uniform prior gives Dirichlet(1, 46, 16), a mean return probability of 1 / 63 = 0.016 and a beta marginal Beta(1, 62) whose central 95 per cent runs from about 0.0004 to 0.058, as in the article.
a_j = 1; n_j = 0; A_0 = 3; N = 60; alpha_j = 1; b_j = 62; E_j = 0.0159
Common errors
Replacing a zero count with a small arbitrary number
Substituting, say, 0.01 for a zero count makes the spreadsheet run but hides the prior it implies; the article asks for the prior to be reported, and a uniform prior adds one to every destination, not to the zero cell alone.
Building a row from a percentage and an assumed sample size
Multiplying a published percentage by an assumed number of patients sets the precision alpha_0 by assumption, which sits uneasily with NICE's requirement that distributions are not chosen arbitrarily but represent the available evidence.
Giving prior weight to a transition that cannot happen
A uniform prior over a row that includes an impossible move, such as a return from death, makes that probability positive in every draw; structural zeros are left out of the row instead.
Sources
Conjugate Dirichlet update and noninformative priors in Bayesian Data Analysis
Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Section 3.4: the posterior distribution for the multinomial probabilities is Dirichlet with parameters alpha_j + y_j; a uniform density is obtained by setting alpha_j = 1 for all j; setting alpha_j = 0 for all j gives an improper prior, and the posterior is proper if there is at least one observation in each of the k categories.
Beta marginal of a single Dirichlet component in Bayesian Data Analysis
Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2013 (full text of the electronic edition read). Appendix A: the Dirichlet is the conjugate prior distribution for the parameters of the multinomial distribution and a multivariate generalisation of the beta; the marginal distribution of a single theta_j is Beta(alpha_j, alpha_0 minus alpha_j).
Briggs, Ades and Price on zero transition counts
Briggs AH, Ades AE, Price MJ. Probabilistic sensitivity analysis for decision trees with multiple branches: use of the Dirichlet distribution in a Bayesian framework. Medical Decision Making. 2003;23(4):341-350. doi:10.1177/0272989X03255922 (abstract read). Abstract: by adopting a Bayesian approach, the problem of observing zero counts for transitions of interest can be overcome.
NICE PMG36 on distributions that represent the available evidence
National Institute for Health and Care Excellence. NICE technology appraisal and highly specialised technologies guidance: the manual (PMG36). London: NICE; published 31 January 2022, last updated 31 March 2026 (full text of chapter 4 read). Section 4.7.10: the distributions chosen for probabilistic sensitivity analysis should not be chosen arbitrarily but chosen to represent the available evidence on the parameter of interest, and their use should be justified.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0