VerifiedEvidence: highv1.0.0

Outcome Measure

A specified instrument, observation or data rule used to quantify a health, care or resource outcome for a defined population and time point.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Outcome Measure

An outcome measure is an operational way to observe or quantify a result that matters to a study or decision. It connects an outcome concept, such as pain, survival or ability to function, to a particular instrument, definition, time point and scoring rule. A well-chosen measure answers a relevant question with sufficient validity and precision; a number can be reproducible yet measure the wrong thing.

From what matters to what is recorded

The outcome domain names the aspect of health or care of interest. The measure specifies how that domain becomes observable, and an endpoint commonly adds a precise analysis variable, timing and comparison for a study. For example, “fatigue” is a domain; a defined patient-reported fatigue questionnaire is an instrument; change in its score from baseline to week 12 can be an endpoint. The instrument, endpoint and broader outcome should be reported distinctly.

Outcome questionPossible measureEssential specification
Does a treatment extend life?All-cause survival.Time origin, death ascertainment, follow-up and censoring rules.
Does it improve symptoms?Patient-reported symptom scale.Exact instrument, recall period, score direction and assessment schedule.
Does it reduce hospital use?Admissions or hospital days.Eligible events, data source, observation window and recurrent-event handling.
Does it improve functioning?Performance task or functioning questionnaire.Task, scoring, assistive devices, missing assessments and meaningful change.
Does it change resource consequences?Recorded services and cost.Perspective, included resources, unit costs, currency and price year.

An outcome can be beneficial or harmful. A study may need measures of both, since a gain in symptom control can coexist with toxicity or caregiver burden. Patients and other decision makers should help identify which domains matter before an instrument is selected; convenience of collection alone is not a sufficient reason to exclude an important outcome.

Who reports it, and how?

Some outcomes are reported directly by patients; others are judged by clinicians, observed by another person or assessed through a standardized performance task. These modes provide different information and have different sources of error. The FDA distinguishes patient-reported, clinician-reported, observer-reported and performance outcome assessments; laboratory results and device readings may provide additional types of endpoint data but should not be mislabeled as a patient's own report.

  • Patient-reported outcome (PRO): The patient reports their symptoms or function without interpretation of the response by a clinician or anyone else. The wording, recall period, language and accessibility affect what is captured.
  • Clinician-reported outcome: A trained clinician rates a finding that requires clinical judgment. Training and consistency across raters matter.
  • Observer-reported outcome: A caregiver or other observer reports observable behavior, particularly when a patient cannot report reliably. An observer's assumption about an internal symptom is not the same as the patient's experience.
  • Performance outcome: A person completes a standardized task, such as a timed walk, under defined conditions. The task may capture one aspect of function without describing the whole lived experience.

Administrative records can measure events such as admissions or prescriptions, but the reason for coding is often billing or service management rather than research. Confirm the coding definition, whether events outside the system are visible and whether the recorded item represents the intended outcome. A prescription is not necessarily medication taken, and an encounter count is not a direct measure of quality of life.

Choosing a measure fit for purpose

Content validity asks whether the measure covers the important parts of the intended construct for the target population. Reliability concerns consistency when the underlying state has not changed; measurement error describes the imprecision around a score. Responsiveness concerns the ability to detect change in the construct when change occurs. Interpretability asks what a numerical difference means in context. These properties depend on the population, language, setting and intended use rather than belonging permanently to a questionnaire by name.

Check whether the score has a defined direction and range, floor or ceiling effects, feasible administration, accessible versions and an appropriate recall window. A threshold for “response” or a minimally important difference needs evidence and a specified context; statistical significance is not automatically clinical importance. A measure can be reliable yet insensitive to a treatment's relevant effect, and an apparently responsive score can be misleading if the instrument itself has changed or assessment is unblinded.

The measurement schedule is part of the definition. A score at week 12, area under a symptom curve, time to the first event and lifetime QALYs summarize different experiences. A fixed-time score may miss early toxicity or later relapse; a survival endpoint requires event and censoring definitions. Choose timing to reflect the expected course of benefit and harm, not only the visit schedule easiest to extract.

Worked example: direction and timing change interpretation

Suppose a symptom score ranges from 0 (none) to 10 (worst). In an illustrative comparison, the treatment group's mean score is 7 at baseline and 4 at week 12; the comparator group's mean is 7 at baseline and 6 at week 12. Define change as week-12 minus baseline, so negative numbers indicate improvement. The treatment change is $4-7=-3$, comparator change is $6-7=-1$, and the difference in changes is $-3-(-1)=-2$ points.

GroupBaseline meanWeek-12 meanChange: week 12 minus baselineInterpretation
Treatment74−3Mean score decreased by 3 points.
Comparator76−1Mean score decreased by 1 point.
Treatment minus comparator0−2−2Treatment has a 2-point larger mean reduction.

The same observed values could instead be reported as a two-point greater improvement by defining improvement as baseline minus week 12; state the convention before interpreting a sign. These group means alone do not provide a standard error, prove a causal effect or establish that two points is important to patients. Study design, variation, missing scores, the instrument's validity and a justified importance threshold are needed for those judgments. Patients who die or are too ill to complete a week-12 measure pose a substantive estimand question, not merely a blank cell to fill by default.

Outcome measure, estimand and analysis are different

The outcome measure supplies observed data, whereas an estimand defines the treatment-effect quantity a study intends to learn. In a trial, that includes the population, treatment conditions, variable, handling of events after treatment starts that affect interpretation, and a population-level summary. Treatment discontinuation, rescue therapy and death can change what a week-12 score comparison means. An estimator and statistical model then attempt to estimate that target from the available data.

For example, “mean symptom score at week 12 among those still attending” is different from “mean week-12 outcome under assignment to treatment in all randomized participants.” Conditioning on attendance can select healthier people. A treatment-policy strategy, hypothetical strategy or composite outcome can answer different questions, but the choice must suit the decision and be planned. Do not describe a measure as if it settled these intercurrent-event and missing-data decisions by itself.

Multiple outcomes also require a clear role. A primary endpoint normally anchors a trial's main objective and error-control plan; secondary outcomes broaden interpretation. A composite endpoint combines specified event types and can conceal opposing component effects if only its total is reported. A surrogate such as a biomarker may predict a clinical outcome in some settings but needs evidence before it stands in for the benefit patients experience. A core outcome set identifies a minimum set of what should be measured in a field; appropriate instruments still have to be chosen for how to measure each domain.

Use in health economic evaluation

Outcome measures supply evidence for effects, harms, resource use and health-related quality of life in economic models. A QALY combines time spent in health states with utility weights; it is a derived outcome, not a direct substitute for every patient-reported symptom or social consequence. A condition-specific scale may detect change relevant to patients but cannot simply be multiplied by survival time to produce QALYs. Mapping to preference-based utilities requires an appropriate validated method and carries uncertainty.

The model's endpoint definitions and time horizon should match the underlying evidence. Event-free survival and overall survival measure different things; an observed hospital admission and a modeled cost are related but distinct; a within-trial measure may require extrapolation to represent the decision horizon. When evidence sources use different instruments or event definitions, make the harmonization and uncertainty explicit rather than treating the numbers as directly interchangeable.

Equity and relevance also matter. An average outcome can hide different effects for subgroups, and an instrument not validated in a particular language or disability group can create biased comparisons. Report who could complete the measure, missingness by group and why the selected domains reflect the decision population.

Reporting and quality checks

Identify the domain, exact instrument or data algorithm, respondent, version, language, range, score direction, event rules, assessment times and clinically meaningful interpretation where justified. Specify the target population and estimand, missing-data strategy, primary or secondary status, and whether a measure was selected before looking at results. Provide sufficient detail for another team to reproduce the endpoint.

  • Check construct alignment. Show why the measure captures the outcome important to patients or the decision maker.
  • Check measurement properties. Evaluate validity, reliability, error and responsiveness for the intended population and context.
  • Check timing and denominators. Confirm that observation windows, event definitions and the people represented by each summary are clear.
  • Check score direction. State whether a lower or higher value is better and how change is calculated.
  • Check incomplete data. Explain deaths, withdrawals, missed assessments and informative loss to follow-up in relation to the estimand.
  • Check economic translation. Distinguish direct outcomes, resource-use measures, costs and derived QALYs; justify any mapping between them.

Sources and further reading

The FDA clinical outcome assessment guidance page distinguishes patient, clinician, observer and performance measures. The ICH E9(R1) addendum distinguishes the endpoint variable from the full treatment-effect estimand. COSMIN provides methods for selecting and assessing outcome measurement instruments, and the COMET Initiative addresses core outcome sets. The symptom-score example is an original teaching calculation, not patient data.

Institutional Perspectives (5)

  • NICE

    QALY Measured With the EQ-5D

    Health effects are expressed in QALYs, with health-related quality of life measured using the EQ-5D as the preferred instrument and valued on public preferences via a choice-based method. In the reference case an additional QALY receives equal weight regardless of who receives the benefit.

    NICE Health Technology Evaluations: The Manual (PMG36), Section 4 (Economic Evaluation)View source →
  • CADTH (CDA-AMC)

    QALY-Based Cost-Utility Analysis

    The reference case is a cost-utility analysis with outcomes expressed as QALYs, combining length and health-related quality of life into a single measure.

    CADTH (now CDA-AMC), Guidelines for the Economic Evaluation of Health Technologies: Canada, 4th Edition (2017)View source →
  • PBAC

    QALY via Cost-Utility Analysis (Productivity Excluded)

    Where a cost-utility analysis is used, health outcomes are expressed as QALYs; the reference case prioritises health effects and productivity changes must not be included in the primary analysis.

    Pharmaceutical Benefits Advisory Committee, Guidelines for Preparing a Submission to the PBAC, Section 3AView source →
  • ZIN

    QALY via Cost-Utility Analysis (EQ-5D-5L, Dutch Value Set)

    New interventions should be assessed by cost-utility analysis with QALYs as the reference outcome; the EQ-5D-5L with the Dutch value set is the preferred instrument for computing QALYs.

    Zorginstituut Nederland, Guideline for Economic Evaluations in Healthcare (2024)View source →
  • IQWiG

    QALY Not Adopted as Standard; Efficiency-Frontier Approach

    IQWiG does not accept the QALY as its standard outcome for cost-benefit assessment, citing unresolved methodological and distributive-justice concerns. It assesses quality of life separately (e.g. standard gamble, time trade-off, or instruments such as the EQ-5D) and judges cost relative to benefit using an efficiency frontier within each therapeutic area.

    IQWiG, General Methods, Version 7.0 (2023), Chapter on Health Economic EvaluationView source →

Library

Publications

1
  • Report

    Health at a Glance 2023: OECD Indicators — Organisation for Economic Co-operation and Development, 2023 Edition ed., 2023 (OECD Publishing)

    OECD’s comprehensive biennial compendium of comparative indicators on population health and health-system performance across member and partner countries — health status, risk factors, access, quality, resources and spending — the standard cross-country benchmarking reference.

Frequently Asked Questions (6)

  • What is an outcome measure?

    A metric assessing the actual results of healthcare, such as survival or functional status, unlike measures of structural resources or processes.

    Source: Donabedian A. Evaluating the quality of medical care. Milbank Memorial Fund Quarterly. 1966;44(3):166-206. doi:10.2307/3348969.

  • What results of care does an outcome measure assess?

    An outcome measure is a metric assessing the actual results of healthcare. It assesses what care achieved for the patient, the end result rather than the effort. Examples are survival and functional status, the real effects that matter to patients. It differs from a process measure, which tracks whether the right steps were taken, whereas an outcome measure tracks what those steps produced. It is the basis of outcome quality, the dimension of quality concerned with results. What care actually achieved is what it assesses. Donabedian (1966) set out this approach to quality.

    Source: Donabedian 1966

  • What does an outcome measure assess?

    An outcome measure assesses the actual results of healthcare, so it gauges what care actually achieves, such as survival or functional status, unlike measures of structural resources or processes. This assessment of results defines it. So an outcome measure is a metric assessing the actual results of healthcare, such as survival or functional status, unlike measures of structural resources or processes These examples of survival and functional status are actual results an outcome measure assesses.

    Source: Donabedian 1966

  • What are examples of an outcome measure?

    Examples of an outcome measure are survival and functional status, so metrics of these actual results of healthcare illustrate an outcome measure, unlike measures of structural resources or processes. These examples illustrate the concept. So an outcome measure is a metric assessing the actual results of healthcare, such as survival or functional status, unlike measures of structural resources or processes This contrast is what places an outcome measure against a process measure, one on results and one on care given.

    Source: Donabedian 1966

  • How does an outcome measure differ from a process measure?

    An outcome measure differs from a process measure in what it assesses: an outcome measure assesses the actual results of healthcare, such as survival or functional status, while a process measure assesses whether appropriate, evidence-based care processes were followed. So an outcome measure gauges results and a process measure gauges the care given, connected as complementary kinds of measure. So an outcome measure assesses results, not processes.

    Source: Donabedian 1966

  • How does an outcome measure relate to outcome quality?

    An outcome measure relates to outcome quality as the metric to the dimension it captures: outcome quality is a healthcare quality dimension focused on actual care results such as survival, and an outcome measure is a metric assessing those actual results. So an outcome measure gauges outcome quality, connected as the dimension concerned with results and the metric that measures them This relationship is what makes an outcome measure the metric that gauges the outcome quality dimension.

    Source: Donabedian 1966

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 24 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HS-NHS-HQ-048

Stable URI · Machine-readable · Resolvable · CC BY 4.0