Concept Architecture
How a quasi-experimental design estimates a causal effect
A quasi-experimental design estimates the effect of an intervention without random assignment by using an external rule, timing change, threshold, comparison trend, or other design feature to construct a credible counterfactual. Its strength comes from the assignment mechanism and explicit identifying assumptions, not from the size of the dataset or the complexity of adjustment. This page explains the main designs, the assumptions that make them causal, and the tests and sensitivity analyses needed for credible interpretation.
The counterfactual is the missing outcome
For each unit, the causal effect compares the outcome under intervention with the outcome that would have occurred without it. Only one potential outcome is observed, so a design must use other units, times, or rules to estimate the missing counterfactual. Quasi-experimental credibility depends on whether that comparison reproduces the untreated potential outcome.
For unit (i), the individual treatment effect is:
$$ \tau_i=Y_i(1)-Y_i(0) $$
The average treatment effect is:
$$ ATE=E[Y(1)-Y(0)] $$
Because both potential outcomes cannot be observed for the same unit at the same time, the estimand and identifying assumptions must be stated before analysis.
Quasi-experimental and adjusted observational studies differ
An ordinary observational analysis often relies on measuring and adjusting for all important confounders. A quasi-experimental design exploits a feature that creates plausibly exogenous variation in treatment, such as a policy boundary, implementation date, eligibility threshold, or external instrument. Regression adjustment or matching can support either design but does not by itself create a natural experiment.
| Design feature | Source of counterfactual | Central assumption |
|---|---|---|
| Difference-in-differences | Change over time in a comparison group | Untreated outcome trends would have been parallel |
| Interrupted time series | Pre-intervention level and trend | No concurrent event explains the post-intervention change |
| Regression discontinuity | Units just above and below an assignment threshold | Potential outcomes change smoothly at the cutoff |
| Instrumental variables | Treatment variation induced by an instrument | Instrument affects outcome only through treatment and satisfies other IV assumptions |
| Synthetic control | Weighted combination of untreated units | The weighted control reproduces the treated unit's counterfactual trajectory |
| Controlled before-after | Change in a comparison group | Groups would have changed similarly absent intervention |
Target-trial thinking clarifies the question
Even when randomisation is impossible, specifying the hypothetical pragmatic trial helps align eligibility, treatment strategies, assignment, time zero, follow-up, outcomes, and analysis. Misalignment can create immortal-time, selection, and prevalent-user biases. The emulated trial also clarifies which causal effect the quasi-experiment can and cannot identify.
A target-trial specification includes:
- Eligibility criteria and recruitment period.
- Intervention and comparator strategies.
- Assignment mechanism or natural experiment.
- Time zero and start of follow-up.
- Outcome and follow-up horizon.
- Intercurrent events and censoring.
- Causal estimand and analysis plan.
Natural experiments use external variation
A natural experiment arises when policy, administration, geography, timing, or another external process changes exposure in a way that approximates random or conditionally random assignment. Calling an event natural does not establish exogeneity. The mechanism should be reconstructed and tested for manipulation, anticipation, selective migration, and correlated changes.
Examples can include phased policy roll-out, administrative lotteries, geographic boundaries, eligibility rules, supply shocks, or provider practice patterns. Each requires a design-specific argument showing why treated and comparison outcomes would otherwise have evolved appropriately.
Difference-in-differences compares changes
Difference-in-differences, or DiD, compares the before-after change in a treated group with the corresponding change in an untreated comparison group. Time-invariant differences between groups are removed, but time-varying confounding can remain. The key assumption is that untreated outcome trends would have been parallel.
For treated group (T), control group (C), post period 1, and pre period 0:
$$ \widehat{\tau}{DiD}=(\overline{Y}{T1}-\overline{Y}{T0})-(\overline{Y}{C1}-\overline{Y}_{C0}) $$
A regression form is:
$$ Y_{it}=\alpha+\gamma Treat_i+\delta Post_t+\tau(Treat_i\times Post_t)+\varepsilon_{it} $$
The coefficient (\tau) estimates the DiD effect under the assumptions and correct variance estimation.
Parallel trends cannot be proven from one pre-period
Parallel trends means that, without treatment, the treated and comparison groups would have experienced the same average change. Similar baseline levels are not required, and equal levels do not guarantee parallel trends. Several pre-intervention periods help assess trend compatibility but cannot prove the unobserved post-treatment counterfactual.
Review should examine:
- Pre-intervention trends on the outcome and related outcomes.
- Changes in population composition and data capture.
- Anticipation before formal implementation.
- Concurrent policies or shocks affecting groups differently.
- Differential timing and intensity of exposure.
- Functional form and sensitivity to comparison periods.
Staggered adoption requires modern DiD methods
When units adopt treatment at different times, conventional two-way fixed-effects models can combine comparisons with negative or unintuitive weights and use already-treated units as controls. Heterogeneous effects across cohorts or time can produce misleading averages. Group-time or cohort-specific estimators should match the adoption pattern.
The analysis should report which units serve as controls at each time, event-time effects, cohort effects, anticipation windows, and aggregation weights. A single regression coefficient can hide important variation.
Event studies show timing but need discipline
An event-study model estimates effects before and after treatment relative to an omitted period. Pre-treatment coefficients can reveal incompatible trends or anticipation, while post-treatment coefficients describe dynamics. Multiple noisy pre-period estimates should not be interpreted as formal proof that parallel trends holds merely because none is statistically significant.
Confidence intervals, joint tests, functional-form sensitivity, and an appropriate comparison group are needed. Binning distant periods can improve precision but changes interpretation.
Interrupted time series estimates level and trend changes
Interrupted time series, or ITS, uses repeated observations before and after an intervention to estimate whether the outcome changed beyond its prior level and trend. It is strongest with many time points, a clearly timed intervention, stable measurement, and no concurrent shock that could explain the change. A controlled ITS adds an unaffected series to strengthen the counterfactual.
A segmented regression can be written as:
$$ Y_t=\beta_0+\beta_1Time_t+\beta_2Intervention_t+\beta_3TimeAfter_t+\varepsilon_t $$
Here, (\beta_2) estimates an immediate level change and (\beta_3) estimates a change in trend. Autocorrelation, seasonality, nonlinearity, implementation lag, and changing variance should be addressed.
A simple before-after study is usually weak
A single outcome before and after implementation cannot distinguish the intervention from secular trend, regression to the mean, maturation, seasonality, or other events. Adding a comparison group or a longer time series improves causal leverage. Statistical significance of the pre-post difference does not solve the missing counterfactual.
Before-after evidence can still describe implementation and generate hypotheses. Causal language should match the design's limitations.
Regression discontinuity uses a treatment threshold
Regression discontinuity, or RD, applies when treatment assignment changes at a known cutoff in a continuously measured running variable. Units just above and below the threshold can be comparable if potential outcomes and sorting behave smoothly. The estimate is generally local to the cutoff.
For cutoff (c), the sharp RD effect is:
$$ \tau_{RD}=\lim_{x\downarrow c}E[Y|X=x]-\lim_{x\uparrow c}E[Y|X=x] $$
Local polynomial methods, appropriate bandwidths, and robust inference are usually preferred to high-order global polynomials. Graphical presentation is an essential diagnostic.
Manipulation threatens regression discontinuity
If individuals or administrators can precisely manipulate the running variable, units near the cutoff may differ systematically. Density tests, heaping checks, institutional knowledge, and covariate continuity can detect concerns. A discontinuity in another policy or measurement rule at the same threshold also threatens attribution.
The running variable and assignment rule should be recorded as implemented, not only as written. Exceptions and imperfect compliance may convert a sharp RD into a fuzzy RD.
Fuzzy regression discontinuity uses the cutoff as an instrument
In fuzzy RD, the probability of treatment changes discontinuously at the threshold but not from zero to one. The outcome discontinuity is divided by the treatment-probability discontinuity to estimate a local effect for compliers under instrumental-variable assumptions.
$$ \tau_{FRD}=\frac{\lim_{x\downarrow c}E[Y|X=x]-\lim_{x\uparrow c}E[Y|X=x]}{\lim_{x\downarrow c}E[D|X=x]-\lim_{x\uparrow c}E[D|X=x]} $$
Weak first-stage changes create imprecision, and the result should not be generalised automatically away from the cutoff.
Instrumental variables use externally induced treatment variation
An instrumental variable, or IV, affects treatment receipt and is used to isolate variation not confounded with the outcome. A valid instrument must be relevant, independent of unmeasured causes of the outcome, and affect the outcome only through the treatment under the exclusion restriction. Monotonicity is also needed for the usual local average treatment effect interpretation.
Potential instruments include policy eligibility, distance, provider preference, or random encouragement, but none is valid from its label alone. Direct effects, selective access, and correlated provider quality commonly threaten validity.
Two-stage least squares estimates an IV effect
For a continuous treatment or linear approximation, two-stage least squares first predicts treatment from the instrument and covariates, then relates predicted treatment to outcome:
$$ D_i=\pi_0+\pi_1Z_i+\boldsymbol{\pi}^{\top}\mathbf{X}_i+u_i $$
$$ Y_i=\beta_0+\tau\widehat{D}_i+\boldsymbol{\beta}^{\top}\mathbf{X}_i+\varepsilon_i $$
Standard errors must account for the two-stage estimation. Using predicted treatment in an ordinary second-stage regression without correct IV procedures produces invalid inference.
Weak instruments produce unstable estimates
An instrument that only weakly changes treatment can generate imprecise, biased, and highly sensitive estimates. First-stage strength, confidence intervals robust to weak instruments, and the institutional mechanism should be reported. A conventional threshold statistic should not replace substantive assessment.
The IV estimate often applies to people whose treatment changes because of the instrument. It may not equal the average effect for all patients or policy settings.
Synthetic control constructs a weighted counterfactual
Synthetic control methods create a weighted combination of untreated units that reproduces the treated unit's pre-intervention outcomes and predictors. The post-intervention gap estimates the effect under the assumption that the synthetic control continues to represent the untreated path. The method is useful for one or a few treated aggregate units with a sufficiently rich donor pool.
Weights (w_j) are commonly constrained so that:
$$ w_j\geq0,\quad \sum_{j=1}^{J}w_j=1 $$
The donor pool should exclude units affected by the intervention or related spillovers. Poor pre-treatment fit weakens credibility regardless of post-treatment gap size.
Placebo tests support synthetic-control inference
Placebo or permutation analyses apply the method to untreated units or false intervention dates to assess whether the treated gap is unusual relative to available controls. Results depend on donor-pool size, fit, and exchangeability. A visually large gap is not sufficient without context.
Sensitivity analyses should test donor exclusions, predictor periods, alternative specifications, and in-time placebos. Inference should acknowledge the small number of aggregate units.
Matching creates comparability on observed variables
Matching pairs or weights treated and untreated units with similar observed covariates. It can improve balance and reduce model dependence, but it cannot remove confounding from unmeasured variables. Matching is best viewed as a design step rather than proof of quasi-random assignment.
The propensity score is:
$$ e(\mathbf{X})=P(D=1|\mathbf{X}) $$
Balance should be assessed on covariates and their distributions after matching, not by the propensity model's predictive accuracy. Lack of overlap may require restricting the target population.
Weighting changes the target estimand
Inverse probability weights can estimate average effects in different target populations depending on their form. Extreme weights increase variance and sensitivity to model misspecification. Trimming, truncation, or overlap weighting changes interpretation and should be pre-specified.
For an average treatment effect weight:
$$ w_i=\frac{D_i}{e(\mathbf{X}_i)}+\frac{1-D_i}{1-e(\mathbf{X}_i)} $$
Covariates should be selected from causal knowledge, not only from their association with treatment or outcome.
Confounding remains a central threat
Quasi-experimental designs reduce confounding only through their specific assignment mechanism. Covariate adjustment may improve precision or address conditional assumptions, but it cannot rescue an invalid instrument, non-parallel trends, manipulated cutoff, or contaminated comparison group. The causal diagram and institutional process should agree.
Sensitivity analysis can assess how strong unmeasured confounding or violations would need to be to change the conclusion. It should not be presented as proof that no unmeasured confounding exists.
Spillovers can contaminate the comparison
An intervention may affect untreated units through referral, information, prices, workforce movement, infection transmission, or shared providers. This violates simple no-interference assumptions and can attenuate or exaggerate estimates. Geographic buffers or network-aware models may help, but they change the estimand.
The analysis should define whether it estimates direct, indirect, total, or population effects. Excluding neighbouring units without reporting the decision can conceal important system consequences.
Composition changes can mimic an intervention effect
Policy implementation can change who appears in the dataset, seeks care, is diagnosed, or remains enrolled. An apparent outcome improvement can reflect a healthier post-intervention population. Stable eligibility rules do not guarantee stable observed composition.
Analysts should compare demographics, baseline risk, entry, exit, missingness, coding, and denominator construction over time and between groups. Balanced panels and repeated cross-sections answer different questions.
Measurement changes threaten identification
New coding systems, screening intensity, electronic records, reporting incentives, or outcome definitions can coincide with an intervention. A discontinuity in measurement can be mistaken for a real effect. Outcomes not expected to respond but measured by the same system can serve as falsification checks.
Data-generation processes should be mapped before and after implementation. Consistent labels do not guarantee consistent measurement.
Anticipation and implementation lag affect timing
People may change behaviour before a policy begins if it is announced in advance. Implementation may also phase in after the official date. Using the legal start date as the only exposure point can dilute or mis-time effects.
Event-time plots, exclusion windows, alternative dates, and measures of actual uptake can examine timing. The main estimand should specify intention to implement, availability, receipt, or treatment on the treated.
Standard errors must follow the assignment process
Observations can be correlated within people, providers, regions, or time. Standard errors clustered at too low a level can overstate precision. With few clusters, conventional cluster-robust methods can also perform poorly.
The inference method should reflect the level at which treatment varies and the source of residual correlation. Randomisation inference, wild-cluster bootstrap, or small-sample corrections may be appropriate in some designs.
Placebo and falsification tests probe assumptions
A placebo test applies the design to a time, outcome, group, or treatment that should not be affected. A detected effect can reveal confounding, measurement change, or model misspecification. A null placebo result supports but does not prove the identifying assumption.
Useful tests include:
- False intervention dates before implementation.
- Outcomes unrelated to the intervention mechanism.
- Populations ineligible for treatment.
- Alternative geographic or policy boundaries.
- Covariate discontinuities at an RD cutoff.
- Pre-treatment effects in event studies.
Robustness should target plausible failures
Running many arbitrary models and selecting the most favourable is not robustness. Sensitivity analyses should correspond to threats identified in the causal design. The main result and all planned alternatives should remain visible.
Relevant analyses can include:
- Alternative comparison groups and windows.
- Alternative trends, bandwidths, and functional forms.
- Exclusion of concurrent policies or contaminated units.
- Adjustment for measured time-varying confounders.
- Bounds for unmeasured confounding or misclassification.
- Alternative definitions of treatment timing and intensity.
- Negative controls and leave-one-out analyses.
Heterogeneous effects need design support
Policy effects can differ by baseline risk, service capacity, geography, socioeconomic position, and implementation intensity. Subgroup analysis is credible only when the design's identifying assumptions remain plausible within subgroups. Small samples and multiple testing can create unstable findings.
Distributional analysis should distinguish true effect heterogeneity from differential exposure, uptake, and measurement. Equity conclusions require both effect and access pathways.
Economic evaluation can use quasi-experimental effects
Quasi-experimental estimates can provide real-world treatment effects, utilisation changes, costs, uptake, and implementation consequences for economic models. The model should preserve the estimate's population, intervention contrast, follow-up, estimand, and local nature. A local RD or IV effect should not be applied automatically to the whole eligible population.
Uncertainty should include sampling error and design uncertainty. Scenarios can test alternative estimates from credible specifications rather than selecting one preferred coefficient.
Cost analyses need the same causal design
Healthcare costs are skewed, correlated over time, and affected by survival and enrolment. Applying a causal design to the clinical outcome but comparing raw costs can produce inconsistent evidence. The cost estimand, censoring, price year, and inference should align with the quasi-experiment.
Policy costs, implementation resources, spillovers, and costs shifted across sectors should be included when relevant. A reduction in recorded payer expenditure may represent unmet need or transfer rather than efficiency.
Worked difference-in-differences example
Suppose hospital admissions fall from 120 to 90 per 10,000 in a region implementing a policy and from 100 to 95 in a comparison region. The treated decline is 30 and the comparison decline is 5. The DiD estimate is therefore a reduction of 25 admissions per 10,000.
$$ \widehat{\tau}_{DiD}=(90-120)-(95-100)=-25 $$
The estimate is causal only if the comparison trend represents what would have happened in the treated region and no differential concurrent change explains the gap. Uncertainty, clustering, population composition, and pre-trends must still be assessed.
Common mistakes
Quasi-experimental language can make an observational comparison sound causal without demonstrating an assignment mechanism. The following errors should be checked before the design supports an effect claim. Each threatens a different part of identification.
- Calling matching or multivariable adjustment quasi-random assignment.
- Treating similar baseline levels as proof of parallel trends.
- Using one pre-period to claim a trend assumption is satisfied.
- Applying conventional two-way fixed effects without examining staggered adoption.
- Fitting high-order RD polynomials across the full data range.
- Assuming an instrument is valid because it strongly predicts treatment.
- Ignoring spillovers, anticipation, or concurrent policies.
- Clustering standard errors below the treatment-assignment level.
- Generalising a local RD or IV effect to the entire population.
- Reporting many specifications while hiding those that alter the conclusion.
Reporting a quasi-experimental design
Transparent reporting should make the source of treatment variation and every identifying assumption visible. Readers should be able to reconstruct why the comparison estimates the intended counterfactual. Graphs and institutional detail are often as important as the regression table.
- Define the causal question, estimand, population, treatment, comparator, outcome, and time zero.
- Describe the assignment mechanism and why it is plausibly exogenous.
- State every design-specific assumption and the evidence supporting it.
- Report pre-trends, balance, first stage, cutoff diagnostics, or pre-fit as applicable.
- Report timing, anticipation, implementation, spillovers, and concurrent events.
- Use inference aligned with clustering, serial correlation, and the number of assignment units.
- Present placebo, falsification, robustness, and sensitivity analyses.
- State the population and policy contexts to which the estimate can be generalised.
The decision standard
A quasi-experimental estimate is credible when the intervention-assignment mechanism creates a defensible counterfactual and the observed data behave consistently with the required assumptions. Statistical adjustment and large samples cannot replace that design argument. The conclusion should remain tied to the identified population, timing, and policy contrast, with plausible violations tested rather than hidden behind a causal label.
Related Concepts (2)
Library
Publications
1
Economic Evaluation in Clinical Trials — Glick, Doshi, Sonnad & Polsky, 2nd Edition ed., 2015 (Oxford University Press)
Practical guidance on conducting cost-effectiveness analyses alongside controlled trials, covering trial design, measurement of costs and quality-adjusted life years, handling censored and missing data, and reporting stochastic uncertainty. Volume 4 in the Handbooks in Health Economic Evaluation series.
BookView source →
Frequently Asked Questions (7)
What is a quasi experimental design?
A research design estimating an intervention's causal effect without random assignment, instead using features such as a natural experiment or matched comparison group.
Source: Shadish, Cook & Campbell 2002
What distinguishes a quasi-experimental design from a true experiment?
A quasi-experimental design studies the effect of an intervention without randomly assigning who receives it, relying instead on naturally occurring or administratively determined groups. What distinguishes it from a true experiment is the absence of randomisation, so the compared groups may differ systematically at the outset. It is used when randomising is impossible or unethical, as with a policy applied to a whole region, and it leans on design features and analysis to approximate the missing randomisation. Comparison without random allocation is its mark. Shadish and colleagues (2002) describe these designs.
Source: Shadish et al. 2002
What is a quasi-experimental design?
A quasi-experimental design is a research design that estimates an intervention's causal effect without random assignment, instead using features such as a natural experiment, a matched comparison group, or timing to approximate the conditions of an experiment. Described by Shadish, Cook, and Campbell, quasi-experimental designs are used when randomisation is impractical or unethical but a causal question remains, employing structures and analyses that strengthen causal inference. Because they lack randomisation, they are more vulnerable to confounding than randomised experiments, so their designs and analyses aim to reduce this threat and support credible causal conclusions.
Source: Shadish, Cook & Campbell 2002
What are examples of quasi-experimental designs?
Examples of quasi-experimental designs include the difference-in-differences approach, comparing changes over time between groups exposed and not exposed to an intervention; interrupted time series, examining changes in a trend after an intervention; regression discontinuity, comparing units just above and below a threshold determining exposure; and controlled before-and-after studies using matched comparison groups. Each uses structure or timing to approximate an experiment without randomising. These designs exploit natural variation or thresholds to estimate causal effects, offering stronger inference than simple observational comparisons where randomisation is not possible.
Source: Angrist & Pischke 2009
Why are quasi-experimental designs used?
Quasi-experimental designs are used when randomisation is impractical, unethical, or impossible, but a causal question about an intervention's effect must still be addressed, such as evaluating policies, programmes, or interventions applied to whole groups. They exploit natural experiments, thresholds, or timing to approximate experimental conditions, providing stronger causal inference than simple observational comparisons. This makes them valuable for evaluating real-world interventions where trials cannot be conducted. By using design features to reduce confounding, quasi-experimental designs offer a means of estimating causal effects outside the randomised experiment.
Source: Shadish, Cook & Campbell 2002
How do quasi-experimental designs strengthen causal inference?
Quasi-experimental designs strengthen causal inference by using structural features that reduce confounding without randomisation: difference-in-differences controls for stable differences between groups and common time trends; regression discontinuity uses the near-random assignment around a threshold; interrupted time series uses the pre-intervention trend as a comparison; and matching creates comparable groups. These features approximate the balance randomisation provides, isolating the intervention's effect from other influences. By exploiting such structure, quasi-experimental designs can support credible causal claims, though more assumptions are required than in randomised experiments, so their validity depends on those assumptions holding.
Source: Angrist & Pischke 2009
What are the limitations of quasi-experimental designs?
The limitations of quasi-experimental designs stem from the lack of randomisation, so they rely on assumptions to rule out confounding, such as parallel trends in difference-in-differences or continuity in regression discontinuity, and if these assumptions fail, the estimates can be biased. Unmeasured confounding remains a threat, and the designs may apply only to particular groups, such as those near a threshold, limiting generalisability. These limitations mean quasi-experimental estimates are held with more caution than randomised results, and their validity depends on the plausibility of the identifying assumptions, which must be examined.
Source: Shadish, Cook & Campbell 2002
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 22 Sep 2026
Content version: 1.0.0
Canonical Identity
- Term code
- HE-ES-CTM-074
Stable URI · Machine-readable · Resolvable · CC BY 4.0