VerifiedEvidence: highv1.0.0

Cluster Randomization

A randomisation technique in which entire groups, such as clinics, rather than individuals, are randomly assigned to different treatment conditions.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept


Theoretically, Cluster Randomization is a randomised study design in which groups of participants, rather than individual participants, are allocated to intervention or control arms. Clusters may comprise hospitals, general practices, schools, communities or geographical regions. The concept is founded on experimental design, probability theory and hierarchical statistical modelling. It exists to prevent contamination between participants, facilitate delivery of group-level interventions and accommodate interventions that are naturally implemented at an organisational or community level.

Mathematically, Cluster Randomization is characterised by correlation among observations within the same cluster. This intracluster dependence is quantified by the intraclass correlation coefficient (ICC), which inflates the variance of treatment effect estimates relative to individually randomised trials. The impact of clustering is incorporated through the design effect, which adjusts both sample size calculations and statistical analyses using mixed-effects models, generalised estimating equations or other hierarchical modelling approaches.

In practice, clusters are randomly allocated to treatment groups before participant recruitment or intervention delivery. Sample size calculations account for the expected ICC and average cluster size to ensure adequate statistical power. Analysis explicitly models the clustered structure of the data to obtain valid standard errors, confidence intervals and hypothesis tests. Cluster randomised trials are widely used in public health, primary care, implementation science and health services research, where they frequently generate evidence for health technology assessment and health economic evaluation.


Purpose


Used to evaluate interventions delivered at the group level or where individual randomisation is impractical, while accounting for correlation between participants within the same cluster.


Mathematical Formulae

Primary Formula

Design Effect:

DE = 1 + (m ? 1)?

where:

  • DE = design effect
  • m = average cluster size
  • ? = intraclass correlation coefficient (ICC)

Supporting Formulae

Effective Sample Size:

n?ff = n / DE

Intraclass Correlation Coefficient:

? = ��? / (��? + ��?)

where:

  • ��? = between-cluster variance
  • ��? = within-cluster variance

Related Mathematical Methods

  • Intraclass Correlation Coefficient
  • Design Effect
  • Mixed-Effects Models
  • Generalised Estimating Equations
  • Hierarchical Linear Models
  • Multilevel Modelling
  • Variance Components Analysis
  • Sample Size Determination

Example


A primary care trial randomises 40 general practices to either a new prescribing intervention or usual care. Each practice contributes an average of 50 patients, and the estimated ICC is 0.02.

Design Effect:

DE = 1 + (50 ? 1) ? 0.02

DE = 1 + 0.98 = 1.98

If the planned sample size without clustering is 1,000 participants, the effective sample size becomes:

n?ff = 1,000 / 1.98 = 505

The increased sample size required to compensate for clustering is incorporated into the trial design before recruitment.


Excel Implementation

FunctionExample FormulaHealth Economics Application
POWER=1+((B2-1)*C2)Calculate the design effect using average cluster size and ICC.
QUOTIENT=B3/B4Estimate the effective sample size after adjusting for clustering.
SUM=SUM(D2:D41)Calculate the total number of participants across clusters.
AVERAGE=AVERAGE(E2:E41)Calculate the mean cluster size.
VAR.S=VAR.S(F2:F41)Estimate between-cluster variability for hierarchical analyses.

VBA (Optional)


VBA can automate cluster randomisation schedules, calculate design effects and generate sample size adjustments for clustered trial designs.


Sources

  • Donner A, Klar N. Design and Analysis of Cluster Randomization Trials in Health Research.
  • Hayes RJ, Moulton LH. Cluster Randomised Trials.
  • Campbell MK, Piaggio G, Elbourne DR, Altman DG. CONSORT Extension for Cluster Randomised Trials.
  • Eldridge SM, Kerry S. A Practical Guide to Cluster Randomised Trials in Health Services Research.
  • NICE. Health Technology Evaluation Manual.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.

Frequently Asked Questions (6)

  • What is cluster randomisation?

    A randomisation technique in which entire groups, such as clinics, rather than individuals, are randomly assigned to different treatment conditions.

    Source: Donner & Klar 2000

  • When is cluster randomisation necessary?

    Some interventions are delivered to whole groups rather than to individuals, such as a change to how a clinic operates or a public health campaign in a community, and cannot sensibly be assigned patient by patient. Cluster randomisation suits these by allocating entire units, such as clinics or villages, to the treatments compared. It is also used to prevent contamination, where patients in the same setting would otherwise influence one another. The intervention's group-level nature is what makes it necessary. Campbell and colleagues (2004) describe such trials.

    Source: Campbell et al. 2004

  • Why is cluster randomisation used?

    Cluster randomisation is used when an intervention is delivered at the level of a group rather than an individual, such as a change in how a clinic operates, when randomising individuals is impractical, or when there is a risk of contamination, where individuals in the same setting assigned to different treatments might influence each other, diluting the comparison. Randomising whole clusters avoids this by giving everyone in a cluster the same condition. So cluster randomisation suits group-level interventions and situations where individual randomisation would be infeasible or lead to contamination.

    Source: Donner & Klar 2000

  • What are the statistical implications of cluster randomisation?

    The statistical implications of cluster randomisation arise because individuals within a cluster tend to be more similar to each other than to those in other clusters, so their responses are correlated, measured by the intracluster correlation. This correlation reduces the effective sample size, so a cluster-randomised trial needs more participants than an individually randomised one for the same power, and the analysis must account for the clustering. Ignoring the correlation would understate uncertainty and give misleading results. So cluster randomisation requires larger samples and analysis methods that handle the clustered structure.

    Source: Donner & Klar 2000

  • What are the challenges of cluster-randomised trials?

    The challenges of cluster-randomised trials include the need for larger sample sizes to offset the reduced efficiency from within-cluster correlation; the requirement for analysis methods that account for clustering; the risk of imbalance between arms when few clusters are randomised; and issues of consent and recruitment at both cluster and individual levels. Selection bias can arise if individuals are recruited after clusters are assigned. These challenges mean cluster-randomised trials are designed and analysed carefully, accounting for the clustering, with attention to the number of clusters and to recruitment and consent procedures.

    Source: Donner & Klar 2000

  • How does cluster randomisation differ from individual randomisation?

    Cluster randomisation differs from individual randomisation in that whole groups are assigned to treatments rather than single participants, so everyone in a cluster receives the same condition, whereas individual randomisation assigns each participant separately. Cluster randomisation suits group-level interventions and avoids contamination but is less statistically efficient, needing larger samples because of within-cluster correlation, and requires analysis that accounts for clustering. Individual randomisation is more efficient and balances groups better but is unsuitable when interventions act at the group level or contamination is a concern. The choice depends on the intervention and setting.

    Source: Donner & Klar 2000

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 13 Nov 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-CTM-015

Stable URI · Machine-readable · Resolvable · CC BY 4.0