Cohen's d from two group means and the pooled standard deviation

Divides the difference between two observed group means by the pooled standard deviation of the two groups, the usual sample estimate of Cohen's d. The pooled standard deviation weights each group's variance by its degrees of freedom. The sign carries the direction of the effect, so instruments that run in opposite directions have to be aligned first. The function sqrt is the square root.

Signature

s_p = sqrt(((n_1 - 1) * s_1^2 + (n_2 - 1) * s_2^2) / (n_1 + n_2 - 2)); d = (xbar_1 - xbar_2) / s_p
Inputs
InputsDefinitionUnit
n_1Number of participants in group 1people
s_1Sample standard deviation of the outcome in group 1instrument points
n_2Number of participants in group 2people
s_2Sample standard deviation of the outcome in group 2instrument points
xbar_1Observed mean outcome in group 1, usually the intervention groupinstrument points
xbar_2Observed mean outcome in group 2, usually the comparator groupinstrument points
Output
s_pPooled within-group standard deviation of the two groupsinstrument points
dDifference in group means divided by the pooled standard deviationstandard deviation units

Function

Cohen's d standardisation and re-expression

Maps a difference between two group means on a continuous outcome to standard deviation units, so that trials which measured the same construct on different instruments can be compared and pooled, and maps a standardised difference back to a quantity a decision model can use, here a log odds ratio for responding. A confidence interval for the standardised difference is the large-sample interval HE-FM-AN-001 applied with its standard error, and a comparator response probability combined with an odds ratio gives the intervention probability by HE-FM-NMA-004; neither is restated here. Hedges' small-sample correction and its standard error belong on the Hedges' g page.

Computational function

  • Computational function: trial summary data to an illustrative QALY gain through Cohen's d

    Takes the summary data a trial report gives, group means, standard deviations and sizes, together with the external inputs a cost-utility model needs, a comparator response probability, utility values for responders and non-responders and a duration, and returns an illustrative QALY gain per person. It chains four steps: Cohen's d from HE-FM-CD-001, the logistic conversion to an odds ratio from HE-FM-CD-002, the intervention response probability from the comparator probability and the odds ratio as in HE-FM-NMA-004, and the extra proportion responding multiplied by the utility difference and the duration. The inputs therefore differ from either formula's variables.

    Inputs and outputs: xbar_1, xbar_2: Group means, intervention and comparator, with higher scores better; required. Unit: instrument points.; s_1, s_2: Group standard deviations; required, above zero. Unit: instrument points.; n_1, n_2: Group sizes; required, at least 2 each. Unit: people.; p_0: Comparator probability of responding; required, above 0 and below 1. Unit: probability.; u_R, u_N: Utility values for responders and non-responders; required. Unit: utility on the QALY scale.; L: Duration over which response status and the utility difference hold; optional, default 1. Unit: years.; s_p, d, lnOR, OR, p_1: intermediate outputs as defined in HE-FM-CD-001, HE-FM-CD-002 and HE-FM-NMA-004.; Delta_Q: Illustrative QALY gain per person. Unit: QALYs.

    Assumption: The function uses the uncorrected d. The article applies Hedges' small-sample correction first and rounds the result to 0.54, which gives 0.0186 QALYs instead of the 0.0188 below; with a pooled estimate from a meta-analysis, the pooled standardised difference would replace d. The logistic conversion is an approximation, nothing is discounted within the duration, and each external input carries its own uncertainty.

    Worked example (Article fatigue trial with uncorrected d): Computed here for illustration from the article's inputs: d of about 0.5432 gives an odds ratio of about 2.6789, an intervention response of about 0.5345 against 0.30, and a QALY gain of about 0.0188 over one year. xbar_1 = 58; xbar_2 = 52; s_1 = 12; s_2 = 10; n_1 = 30; n_2 = 30; p_0 = 0.30; u_R = 0.78; u_N = 0.70; L = 1; s_p = 11.0454; d = 0.5432; lnOR = 0.9854; OR = 2.6789; p_1 = 0.5345; Delta_Q = 0.0188

    Worked example (Equal group means give no QALY gain): With both means at 52 the standardised difference is 0, the odds ratio is 1, the intervention response equals the comparator's 0.30 and the QALY gain is 0, a limiting case. xbar_1 = 52; xbar_2 = 52; s_1 = 12; s_2 = 10; n_1 = 30; n_2 = 30; p_0 = 0.30; u_R = 0.78; u_N = 0.70; L = 1; s_p = 11.0454; d = 0; lnOR = 0; OR = 1; p_1 = 0.30; Delta_Q = 0

    Excel: =1.814*CohensD in a cell named LogOR; =ComparatorResp*EXP(LogOR)/(1-ComparatorResp+ComparatorResp*EXP(LogOR)) in a cell named InterventionResp; =(InterventionResp-ComparatorResp)*(UtilResp-UtilNon)*Duration in a cell named QALYGain. CohensD and PooledSD come from the Excel cells of HE-FM-CD-001, and ComparatorResp, UtilResp, UtilNon and Duration are named input cells.

    R: cd_qaly_gain <- function(m1, m2, sd1, sd2, n1, n2, p0, u_resp, u_non, years=1) { sp <- sqrt(((n1-1)*sd1^2+(n2-1)*sd2^2)/(n1+n2-2)); d <- (m1-m2)/sp; or <- exp(1.814*d); p1 <- p0*or/(1-p0+p0*or); list(sp=sp, d=d, or=or, p1=p1, qaly=(p1-p0)*(u_resp-u_non)*years) } Returns a list of the intermediate values and the QALY gain, read with [["qaly"]].

    Python: def cd_qaly_gain(m1, m2, sd1, sd2, n1, n2, p0, u_resp, u_non, years=1.0): sp = math.sqrt(((n1-1)*sd1**2+(n2-1)*sd2**2)/(n1+n2-2)); d = (m1-m2)/sp; odds_ratio = math.exp(1.814*d); p1 = p0*odds_ratio/(1-p0+p0*odds_ratio); return {"sp": sp, "d": d, "or": odds_ratio, "p1": p1, "qaly": (p1-p0)*(u_resp-u_non)*years} Requires the math module; returns a dictionary of the intermediate values and the QALY gain.

    Test (Article response probability from a standardised effect of 0.54): With the CohensD cell replaced by the article's rounded Hedges' g of 0.54 and ComparatorResp set to 0.30, the model's InterventionResp cell rounds to the article's 0.533. Expected result: TRUE. FALSE shows the odds ratio applied to the probability instead of the odds. Excel check: =ROUND(InterventionResp,3)=0.533

    Test (Article QALY gain from a standardised effect of 0.54): With the same inputs and UtilResp 0.78, UtilNon 0.70 and Duration 1, the model's QALYGain cell rounds to the article's 0.0186. Expected result: TRUE. FALSE shows the intervention response used in place of the extra proportion responding. Excel check: =ROUND(QALYGain,4)=0.0186

    Common error (Carrying the standardised effect into the model as a probability or utility): Reading a d of 0.54 as 54 percentage points more responders, or adding it to a utility value, overstates the effect many times over. The chain gives about 0.23 more responders and a QALY gain of about 0.019, and only through a comparator response rate and utility values that need their own sensitivity analysis.

    Source: Schünemann HJ, Vist GE, Higgins JPT, Santesso N, Deeks JJ, Glasziou P, Akl EA, Guyatt GH. Chapter 15: Interpreting results and drawing conclusions. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA (editors). Cochrane Handbook for Systematic Reviews of Interventions version 6.5. Cochrane; 2024. Section 15.5.3.3 on re-expressing a standardised mean difference as an odds ratio and combining it with a comparator risk. Chinn S. A simple method for converting an odds ratio to effect size for use in meta-analysis. Statistics in Medicine. 2000;19(22):3127-3131.

    s_p = sqrt(((n_1 - 1) * s_1^2 + (n_2 - 1) * s_2^2) / (n_1 + n_2 - 2)); d = (xbar_1 - xbar_2) / s_p; lnOR = 1.814 * d; OR = exp(lnOR); p_1 = p_0 * OR / (1 - p_0 + p_0 * OR); Delta_Q = (p_1 - p_0) * (u_R - u_N) * L

Try this function

Implementations

  • Excel

    Pooled standard deviation and Cohen's d in two cells

    With named cells Mean1, Mean2, SD1, SD2, N1 and N2, the first formula returns the pooled standard deviation in a cell named PooledSD and the second returns d in a cell named CohensD.

    =SQRT(((N1-1)*SD1^2+(N2-1)*SD2^2)/(N1+N2-2)); =(Mean1-Mean2)/PooledSD

Assumptions

  • Equal spread in the two populations

    Cohen defined d for two populations whose standard deviations are assumed equal, and the pooled standard deviation estimates that common value. Where a treatment changes the spread of outcomes, Glass's delta, which divides by the comparator group's standard deviation alone, answers a different question.

  • Same direction of scale in every study

    The method does not correct for instruments on which higher scores mean opposite things. Before standardising, the means from one set of studies are multiplied by minus one so that the sign of d means the same everywhere.

  • Uncorrected for small-sample bias

    The sample d tends to overstate the population effect in small samples. Cochrane reviews report Hedges' adjusted g, which multiplies d by a correction factor that approaches one as the sample grows; in the article's trial of 60 participants it changes 0.543 to 0.536. The correction and its standard error are covered on the Hedges' g page.

Worked examples

  • Article fatigue trial with 30 patients per arm

    Illustrative trial from the article: intervention mean 58 with standard deviation 12, usual care mean 52 with standard deviation 10, 30 patients in each arm. The pooled standard deviation is the square root of 122, about 11.0454, and d is about 0.5432, shown as 11.05 and 0.543 in the article.

    xbar_1 = 58; xbar_2 = 52; s_1 = 12; s_2 = 10; n_1 = 30; n_2 = 30; s_p = 11.0454; d = 0.5432
  • Equal group standard deviations return that standard deviation

    Computed here for illustration: with both standard deviations equal to 10, the pooled value is 10 whatever the group sizes, and the same six-point difference gives a d of 0.6, a limiting case.

    xbar_1 = 58; xbar_2 = 52; s_1 = 10; s_2 = 10; n_1 = 30; n_2 = 30; s_p = 10; d = 0.6
  • Unequal group sizes weight the larger group

    Computed here for illustration: with 20 patients in the intervention arm and 40 in usual care, the pooled standard deviation moves towards the larger group's 10, to about 10.6964, and d rises to about 0.5609.

    xbar_1 = 58; xbar_2 = 52; s_1 = 12; s_2 = 10; n_1 = 20; n_2 = 40; s_p = 10.6964; d = 0.5609

Common errors

  • Comparing Cohen's d across populations with different spread

    The same six-point difference gives a d of 0.75 in a narrow trial population with a standard deviation of 8 and 0.40 in a broad pragmatic population with a standard deviation of 15, computed here for illustration. A larger d does not by itself mean a larger benefit.

  • Standard errors entered as standard deviations

    Dividing the difference by a pooled standard error instead of the pooled standard deviation inflates d by about the square root of the group size. In the article's trial, 11.0454 divided by the square root of 30 is about 2.0166, which would give a d of about 2.98 instead of 0.54, computed here for illustration.

  • Cohen's d and Hedges' g used interchangeably

    Reporting a standardised effect without saying whether it is the uncorrected d or Hedges' adjusted g, the Cochrane standardised mean difference, leaves the reader unable to reproduce it. The difference is small above about 20 participants but should still be stated.

Sources

  • Pooled standard deviation and Cohen's d for independent groups

    Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Frontiers in Psychology. 2013;4:863. Equation 1, which gives Cohen's d for two independent groups with the pooled standard deviation as the denominator, and the discussion of Hedges' correction.

    View source →

  • Pooled standard deviation in the Review Manager algorithms

    Deeks JJ, Higgins JPT, on behalf of the Cochrane Statistical Methods Group. Statistical algorithms in Review Manager. Cochrane; May 2022. Section on individual study estimates for continuous outcomes, which defines the pooled standard deviation across the two groups.

    View source →

  • Cohen's definition of d for two populations

    Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988 (Routledge digital edition 2013). Section 2.2, equations 2.2.1 and 2.2.2, which define d as the difference in population means divided by the common within-population standard deviation.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0

Cohen's d from two group means and the pooled standard deviation | HealthEconomics.wiki