VerifiedEvidence: highv1.0.0

Missing Data

Unobserved or unavailable values required for a specified analysis, whose consequences depend on the pattern, mechanism and target estimand.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Missing Data

Missing data are values needed for an analysis that were not observed or are unavailable in the dataset being used. They can affect a clinical outcome, cost, covariate, date or even whether a person enters the recorded population. This page identifies what is missing, explains assumptions about why, and shows how the assumptions change an estimate and its uncertainty.

Find the missingness before choosing a method

Begin with a clear analysis population and a variable-by-time inventory. A blank value may mean a test was not ordered, a participant stopped answering, care occurred elsewhere, a code was not captured or the value genuinely does not apply. Those possibilities have different meanings and should not all be recoded as zero.

PatternExampleWhy it matters
Missing outcomeA participant's follow-up quality-of-life score is unavailable.The treatment effect may differ among people without follow-up.
Missing covariateDisease severity was not recorded for some people.Adjustment and eligibility definitions may change.
Intermittent missingnessA middle visit is absent but later visits exist.Longitudinal methods can use observed neighbouring information under assumptions.
DropoutMeasurements cease after a particular visit.Reasons for stopping may predict later outcomes.
Unobserved careA patient attends another provider outside the dataset.An apparent absence of treatment or events may be incomplete coverage.

Report counts and percentages by treatment group, variable and time, along with known reasons and the characteristics of people with and without data. Prevent avoidable losses through careful collection and follow-up before relying on statistical repair. Define the intended treatment effect or other estimand, including how events such as treatment discontinuation are handled, before calling a value “missing.”

What MCAR, MAR and MNAR actually assume

The mechanism concerns the probability that a value is missing, conditional on information available in the study. Missing completely at random (MCAR) means missingness is unrelated to observed and unobserved values relevant to the analysis. Missing at random (MAR) permits dependence on observed information, but after conditioning on suitable observed variables, missingness does not additionally depend on the missing value itself. Missing not at random (MNAR) describes remaining dependence on the unobserved value or another inadequately accounted-for factor.

AssumptionIllustrative storyImplication
MCARA random equipment failure deletes some measurements independently of patient status.Complete cases may be unbiased for certain estimands but lose precision.
MAR, conditional on observed dataFollow-up is less likely among people with a recorded prior poor score, and that prior score adequately explains the dependence.A correctly specified method can use observed predictors to account for missingness.
MNAR relative to available dataPeople with an unobserved worsening outcome are less likely to return even after accounting for recorded history.The missing values cannot be identified from observed data alone without extra assumptions or evidence.

These labels are relative to which variables are actually observed and included in an analysis. They cannot generally be established by a simple test of the recorded values, because the missing outcomes themselves are unseen. MCAR is a strong assumption; MAR is not the same as random dropout, and MNAR sensitivity analysis asks what follows if MAR is wrong.

See how the observed denominator can mislead

Consider a fictional trial with 100 participants assigned to each of two groups. In the intervention group, outcomes are observed for 80 people, of whom 40 succeed; in control, outcomes are observed for 90 people, of whom 36 succeed. The complete-case success rates are $40/80=0.50$ and $36/90=0.40$, an observed-data difference of 10 percentage points.

There are 20 missing intervention outcomes and 10 missing control outcomes. If every missing intervention outcome were a failure and every missing control outcome a success, the full-group difference would be $40/100-46/100=-0.06$. Under the opposite extreme, it would be $60/100-36/100=0.24$. These extreme assignments are transparent bounds for this binary outcome, not plausible estimates of what actually happened.

CalculationInterventionControlDifference, intervention minus control
Observed cases only40/80 = 50%36/90 = 40%+10 percentage points.
Adverse extreme for intervention40/100 = 40%46/100 = 46%-6 percentage points.
Favourable extreme for intervention60/100 = 60%36/100 = 36%+24 percentage points.

The wide range exposes how much the unobserved outcomes could matter without assigning a mechanism. It does not imply that all missing patients would succeed or fail. If the analysis target is a different estimand, such as success while remaining on treatment, define it carefully rather than silently changing denominators.

Select an analysis tied to defensible assumptions

Complete-case analysis discards incomplete records and can lose precision or introduce bias when missingness relates to outcomes or predictors. Multiple imputation draws several plausible replacements conditional on an imputation model, analyses each completed dataset and combines estimates with uncertainty from missingness. A single filled-in value or mean replacement generally understates uncertainty and may distort associations.

Likelihood-based longitudinal models and inverse-probability weighting may be useful under their respective modelling and positivity assumptions. Include variables predictive of missingness and the missing value, preserve the substantive model's structure, and avoid using information unavailable at the decision point if that would contaminate an applied prediction. The correct method depends on the estimand, data structure and plausible mechanism; no method can recover information that was never collected without assumptions.

Sensitivity analysis should vary assumptions not identified by the observed data. A pattern-mixture approach might lower imputed outcomes for participants who dropped out by a clinically meaningful amount; a tipping-point analysis finds how large a departure would change the conclusion. State the size and direction of the departure, its clinical rationale and the effect on both point estimates and uncertainty.

Consequences for health-economic models

Missing costs and health outcomes may be correlated, and complete cases can exclude people with complex, high-cost care. Imputation or joint modelling should respect skewed costs, bounded utilities, structural zero costs and longitudinal timing where relevant. A model built from observed data should carry uncertainty and plausible missingness-related bias through to incremental costs and effects rather than treating an imputed spreadsheet as fully observed truth.

If people leave an electronic record system, the absence of a later admission may mean no admission or care outside the system. Linkage, follow-up checks and alternative data sources can reduce that ambiguity. Explicitly examine how the economic conclusion changes under plausible missing cost, quality-of-life and event scenarios.

  • Define the estimand: Identify which outcomes and post-treatment events the analysis intends to represent.
  • Describe the pattern: Report missingness by group, variable, time and known reason.
  • Avoid automatic zeros: An unrecorded cost or event is not necessarily zero.
  • Preserve uncertainty: Multiple imputation requires repeated datasets and appropriate combining rules.
  • Test departures: A MAR-based primary analysis should be accompanied by relevant MNAR sensitivity analyses when missingness could change the decision.
  • State limits: No statistical procedure alone proves the missing-data mechanism or restores inaccessible care.

Sources and further reading

The Sterne and colleagues paper on multiple imputation explains assumptions, bias and practical pitfalls, while CONSORT reporting guidance specifies transparent reporting of missing trial data. A methodological paper on missing outcome data in trials connects the analysis to the target treatment effect, and a controlled multiple-imputation tutorial illustrates sensitivity analysis. The trial numbers and extreme assignments above are original teaching examples.

Library

Publications

1
  • Book

    Economic Evaluation in Clinical Trials — Glick, Doshi, Sonnad & Polsky, 2nd Edition ed., 2015 (Oxford University Press)

    Practical guidance on conducting cost-effectiveness analyses alongside controlled trials, covering trial design, measurement of costs and quality-adjusted life years, handling censored and missing data, and reporting stochastic uncertainty. Volume 4 in the Handbooks in Health Economic Evaluation series.

Frequently Asked Questions (6)

  • What is missing data?

    Information intended to be collected but absent for some individuals or variables, arising from dropout, non-response, or data entry errors.

    Source: Rubin 1976

  • What is missing data and where does it come from?

    Missing data is information that was meant to be collected but is absent for some individuals or variables, arising from participants dropping out, declining to answer, or errors in recording. It matters because gaps can bias results and reduce precision, and how much harm they do depends on why the values are missing rather than simply how many. Data missing for reasons tied to the outcome distort conclusions far more than data missing at random. Absent intended information is what it is. Little and Rubin (2002) describe this.

    Source: Little & Rubin 2002

  • Why does missing data matter?

    Missing data matter because they reduce the amount of information, lowering precision, and, more seriously, can bias the results if the reasons for missingness are related to the variables or outcomes, so that the observed data are unrepresentative. Improper handling can distort conclusions. So missing data matter for both the efficiency and the validity of an analysis, since ignoring or mishandling them, for example by naive complete case analysis when missingness is not random, can produce biased and misleading estimates, which is why understanding the missingness and using appropriate methods are important to draw sound conclusions from incomplete data.

    Source: Rubin 1976

  • What are the mechanisms of missing data?

    The mechanisms of missing data are commonly classified as missing completely at random, where missingness is unrelated to any data; missing at random, where it depends on observed variables but not the missing values themselves; and missing not at random, where it depends on the missing values. These mechanisms determine which methods give valid results. So missing data are categorised by why they are missing, since the mechanism governs whether and how the missingness can bias the analysis and which handling methods are valid, with missing completely at random and missing at random allowing standard principled methods and missing not at random requiring more specialised approaches and sensitivity analysis.

    Source: Rubin 1976

  • How is missing data handled?

    Missing data are handled by methods ranging from simple but often biased approaches, such as complete case analysis or single imputation, to principled methods such as multiple imputation and likelihood-based approaches like mixed models, which use the observed data validly under the missing at random assumption. So missing data are handled by choosing methods appropriate to the missingness mechanism and the analysis, with multiple imputation and maximum likelihood generally preferred over simpler approaches because they make fuller use of the data and reflect the uncertainty, and with sensitivity analyses used to examine robustness, since the goal is valid inference despite the incompleteness.

    Source: Rubin 1987

  • Why are the reasons for missing data important?

    The reasons for missing data are important because they determine the missingness mechanism, which governs whether the missingness biases the analysis and which methods are valid: if missingness is unrelated to the data, simple methods suffice, but if it depends on the values, bias can arise and specialised methods are needed. So understanding why data are missing is central to handling them correctly, since the mechanism, whether missing completely at random, at random, or not at random, dictates the appropriate approach and the risk of bias, which is why the causes of missingness are considered and, because the mechanism is often untestable, sensitivity analyses examine the impact of different assumptions.

    Source: Rubin 1976

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 24 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-120

Stable URI · Machine-readable · Resolvable · CC BY 4.0