Functions & Formulae

Each applied formula has its own function page, with a signature, implementations, and tests.

Binomial distribution of an event count among patients at risk

P(R = r) = C(n, r) * p^r * (1 - p)^(n - r), r = 0, 1, ..., n

Gives the probability that exactly r of n patients have an event within a fixed period when each patient has the same probability p of the event and outcomes are independent. Read as a function of p for an observed count, the same expression is the binomial likelihood, proportional to p^r (1 minus p)^(n minus r), which peaks at the observed proportion r/n. The formulae below give the probability function, the chance of a zero count, the mean and variance of the count, the estimated probability with its standard error, the Wald and Wilson intervals and the logit model for binomial data in NICE DSU evidence synthesis. The conjugate beta update of a binomial probability is set out on the Beta Distribution page (HE-FM-BETA-002), and the general Wald interval for any estimate on the Asymptotic Normality page (HE-FM-AN-001).

  • Binomial probability of exactly r patients with an event

    P_r = C(n, r) * p^r * (1 - p)^(n - r)

    Gives the probability that exactly r of n patients at risk have the event within the period. C(n, r) is the binomial coefficient, n! divided by r! (n minus r)!, the number of ways of choosing which r of the n patients have the event; Excel returns it with COMBIN. A single patient, the case n = 1, is a Bernoulli trial.

  • Binomial probability of no events among n patients

    P_0 = (1 - p)^n

    The binomial probability function with r = 0. The chance of a zero count depends on the expected count n p: it stays high when n p is small, so a rare event in a small study often gives no events at all, and a zero count does not show that the probability is zero.

  • Mean, variance and standard deviation of a binomial event count

    E_R = n * p; V_R = n * p * (1 - p); SD_R = sqrt(V_R)

    Gives the expected number of patients with the event and the variance and standard deviation of that count. In a cohort model the number moving is the mean n p, with no chance variation. In a microsimulation of n patients with a fixed probability, the count varies around that mean with variance n p (1 minus p), the stochastic or first-order uncertainty of the ISPOR-SMDM task force. The variance is largest when p is one half. The function sqrt is the square root.

  • Maximum likelihood estimate of a binomial probability and its standard error

    p_hat = r / n; SE_p = sqrt(p_hat * (1 - p_hat) / n)

    The binomial likelihood, proportional to p^r (1 minus p)^(n minus r), peaks at the observed proportion r/n, the maximum likelihood estimate. Its standard error is estimated by putting p_hat in place of p in the sampling variance p (1 minus p)/n, so larger studies give more precise estimates.

  • Wald confidence interval for a binomial proportion

    p_hat = r / n; L_Wald = p_hat - z * sqrt(p_hat * (1 - p_hat) / n); U_Wald = p_hat + z * sqrt(p_hat * (1 - p_hat) / n)

    The normal-approximation interval for a binomial proportion, the form most often quoted. It is symmetric about p_hat, has zero width when r is 0 or n, and its lower limit can fall below zero. Brown, Cai and DasGupta showed that its coverage behaves erratically and that the problem is far more persistent than had been appreciated. The general Wald interval for any estimate is on the Asymptotic Normality page (HE-FM-AN-001).

  • Wilson score interval for a binomial proportion

    p_hat = r / n; m_W = p_hat + z^2 / (2 * n); h_W = z * sqrt(p_hat * (1 - p_hat) / n + z^2 / (4 * n^2)); d_W = 1 + z^2 / n; L_Wilson = (m_W - h_W) / d_W; U_Wilson = (m_W + h_W) / d_W

    Inverts the large-sample test for a proportion instead of plugging p_hat into the standard error. Its centre is pulled from p_hat towards one half, it is asymmetric near the bounds and its limits cannot fall outside the possible range for a probability. Brown, Cai and DasGupta recommend it, or the equal-tailed Jeffreys interval, for small n.

  • Binomial arm event probability from the logit model in NICE DSU TSD 2

    p_ik = exp(mu_i + delta_ik) / (1 + exp(mu_i + delta_ik))

    In the TSD 2 framework the number of events r_ik in arm k of trial i is binomial with probability p_ik among n_ik patients, and the logit of p_ik is the trial's baseline log odds mu_i plus, for arms other than the control, the trial-specific log odds ratio delta_ik. Inverting the logit gives the arm probability below; for the control arm delta_ik is zero. The function exp is the exponential function.