VerifiedEvidence: highv1.0.0

Effective Sample Size

The sample size a simple random sample would need to match the precision of an actual sample from a more complex design, such as clustered.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, the Effective Sample Size (ESS) is the number of statistically independent observations that would provide the same precision as a correlated, weighted or simulated dataset. It is founded on sampling theory, Bayesian computation and survey methodology, where dependence among observations reduces the amount of independent information contained within the data. In health economics, effective sample size is used in cluster randomised trials, complex survey analyses, Bayesian estimation and probabilistic simulation to quantify the information available for statistical inference.

Mathematically, effective sample size is derived by adjusting the observed sample size to account for correlation, unequal weighting or autocorrelation. For clustered sampling, ESS is obtained by dividing the observed sample size by the design effect. In Bayesian Markov chain Monte Carlo analyses, ESS estimates the equivalent number of independent posterior samples after accounting for serial correlation. Larger effective sample sizes indicate greater estimation precision.

In practice, effective sample size is calculated during study design, statistical analysis and Bayesian computation. Health economists use ESS to adjust standard errors, determine required sample sizes, evaluate convergence of simulation algorithms and quantify the precision of parameter estimates used in decision models and health technology assessments.


Purpose

Used to quantify the amount of independent statistical information contained within correlated or weighted data, adjust sample size calculations, assess estimation precision and evaluate Bayesian simulation efficiency.


Mathematical Formulae

Primary Formula

For clustered sampling:

n?ff = n / DEFF

where:

n?ff = effective sample size

n = observed sample size

DEFF = design effect

Supporting Formulae

Design Effect:

DEFF = 1 + (m ? 1)?

For Markov chain Monte Carlo:

ESS = N / (1 + 2 ? ???)

where:

N = total posterior samples

?? = autocorrelation at lag k

Related Mathematical Methods

Design Effect

Intraclass Correlation Coefficient

Cluster Randomised Trials

Survey Sampling

Markov Chain Monte Carlo

Bayesian Inference

Autocorrelation Analysis


Example

A cluster randomised trial recruits 1,200 patients.

Average cluster size = 20

Intraclass correlation coefficient = 0.05

Design effect:

DEFF = 1 + (20 ? 1) ? 0.05

DEFF = 1.95

Effective sample size:

n?ff = 1,200 / 1.95

n?ff = 615

Although 1,200 patients were recruited, the statistical information is equivalent to approximately 615 independently sampled individuals.


Excel Implementation

FunctionExample FormulaHealth Economics Application
Formula=B2/C2Calculate effective sample size (n � DEFF)
Formula=1+((D2-1)*E2)Calculate design effect from cluster size and ICC
ROUND=ROUND(B2/C2,0)Report effective sample size
IF=IF(C2>1,"Cluster adjustment required","No adjustment required")Flag studies requiring variance adjustment

VBA (Optional)

Automate calculation of effective sample size for clustered studies, weighted survey data and Bayesian simulation output.


Sources

Kish L. Survey Sampling.

Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis.

Donner A, Klar N. Design and Analysis of Cluster Randomization Trials in Health Research.

Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.

Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes.

Library

Publications

1
  • Book

    Statistical Analysis of Cost-Effectiveness Data — Willan & Briggs, 1st Edition ed., 2006 (John Wiley & Sons)

    A synthesis of statistical methods for analysing cost-effectiveness data, including net-benefit regression, confidence intervals for the ICER, cost-effectiveness acceptability curves, and covariate adjustment. Part of the Wiley Statistics in Practice series.

Frequently Asked Questions (6)

  • What is effective sample size?

    The sample size a simple random sample would need to match the precision of an actual sample from a more complex design, such as clustered.

    Source: Kish 1965

  • What does effective sample size tell us about a complex design?

    Effective sample size is the number of participants a simple random sample would need to carry the same precision as an actual sample drawn from a more complex design, such as a clustered survey. It tells us how much usable information the design really holds, which is often less than the raw headcount because clustering makes observations partly redundant. A study of a thousand people spread across a few clusters may carry the precision of far fewer independent ones. The true informational size of a sample is what it captures. Kirkwood and Sterne (2003) describe this.

    Source: Kirkwood & Sterne 2003

  • How is effective sample size calculated?

    Effective sample size is calculated by dividing the actual sample size by the design effect, the ratio of the design's variance to that of simple random sampling. A design effect above one, common with clustering, therefore gives an effective sample size smaller than the actual number. So effective sample size is calculated as the actual size divided by the design effect, which translates the complex design's reduced precision into the size of an equivalent independent sample, showing how much usable information the complex sample provides, and it depends on factors such as the intraclass correlation and cluster size that determine the design effect.

    Source: Kish 1965

  • Why does effective sample size matter?

    Effective sample size matters because it reflects the true precision of estimates from a complex design, which can be considerably less than the actual number of observations suggests, so using the nominal sample size would overstate the precision. It is important for both analysis and planning. So effective sample size matters for correctly gauging the information in a complex sample, since clustering and weighting reduce precision, and recognising the smaller effective sample prevents overconfidence in the estimates and informs how large a study must be, which is why it is considered when designing and interpreting surveys and clustered studies.

    Source: Kish 1965

  • How does effective sample size relate to the design effect?

    Effective sample size relates to the design effect as its inverse scaling of the actual sample size: the effective sample size equals the actual sample size divided by the design effect, so a larger design effect, indicating greater loss of precision, gives a smaller effective sample size. So the design effect and effective sample size are directly linked, with the design effect quantifying how much the design inflates variance relative to simple random sampling and the effective sample size expressing the equivalent independent sample, and together they convey the impact of the design on precision, which is important for the analysis and planning of complex or clustered studies.

    Source: Kish 1965

  • How is effective sample size used in study design?

    Effective sample size is used in study design to ensure that a complex or clustered study collects enough observations to achieve the desired precision, since the design effect reduces the information per observation; the actual sample size is set larger than the target effective size by the design effect. So effective sample size is used, together with the design effect, to plan adequately powered complex studies, because achieving a required precision demands more observations under clustering than under simple random sampling, and calculating the effective sample size ensures the study is large enough to yield estimates as precise as intended despite the loss of efficiency from the design.

    Source: Kish 1965

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 15 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-057

Stable URI · Machine-readable · Resolvable · CC BY 4.0