VerifiedEvidence: highv1.0.0

Surrogate Endpoint

A trial outcome used as a substitute for a direct clinical benefit measure, such as survival, based on its expected predictive relationship to that outcome.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

How a surrogate endpoint stands in for patient benefit

A surrogate endpoint is measured instead of a direct outcome that describes how a patient feels, functions, or survives. This page explains the pathway from a treatment's biological effect to the surrogate and then to the patient-relevant outcome. It also shows why measuring the surrogate accurately is not enough: the central question is whether changes in it reliably predict changes in the outcome that matters to patients.

The causal pathway behind a surrogate

A surrogate is credible only when its place in the disease and treatment pathway is understood. The treatment may affect the surrogate, the clinical outcome, or both through several mechanisms. A surrogate can therefore be strongly associated with prognosis without capturing the full causal effect of treatment.

  • The treatment effect on the surrogate shows whether the intervention changes the intermediate measure.
  • The relationship between the surrogate and the clinical outcome shows whether patients with different surrogate values tend to have different outcomes.
  • The treatment effect on the clinical outcome may also pass through pathways that the surrogate does not measure.
  • Adverse effects can worsen patient outcomes even when the surrogate changes favourably.

What makes an endpoint a surrogate

An intermediate measure does not become a validated surrogate simply because it occurs earlier or is biologically plausible. Its intended use must identify the clinical outcome, intervention class, disease, population, and decision context for which it is expected to predict benefit. Validation is therefore specific to a defined context rather than a permanent property of the marker itself.

Examples of potential surrogates include:

  • Blood pressure as an intermediate measure used to predict cardiovascular outcomes in a defined treatment context.
  • Viral load as an intermediate measure used to predict disease progression or transmission-related outcomes in a defined infection.
  • Progression-free survival as a time-to-event endpoint used to anticipate overall survival or other patient-relevant outcomes in some oncology settings.
  • Biomarker response as an intermediate measure used to predict morbidity, function, or survival when the required relationship has been established.

Direct, intermediate, and surrogate endpoints are different

A direct clinical endpoint measures a benefit or harm that is meaningful in its own right, such as survival, symptoms, function, or health-related quality of life. An intermediate endpoint occurs before a direct outcome but may or may not predict it. A surrogate endpoint is an intermediate endpoint being used specifically as a substitute for a direct clinical endpoint.

Endpoint roleWhat it measuresWhat must be established
Direct clinical endpointHow a patient feels, functions, or survivesThe endpoint must be valid, reliable, and relevant to patients.
Intermediate endpointA biological or clinical event earlier in a pathwayThe measure must be valid for its stated purpose, but prediction of later benefit is not assumed.
Surrogate endpointAn intermediate measure used in place of a direct clinical outcomeEvidence must support prediction of the treatment effect on the specified clinical outcome.
Biomarker endpointA biological characteristic used as a study outcomeA biomarker may be exploratory, prognostic, predictive, pharmacodynamic, or surrogate; these roles are not interchangeable.

Prognostic association is not surrogate validation

A prognostic factor separates patients with different expected outcomes regardless of treatment. Surrogate validation asks a harder question: whether the effect of treatment on the surrogate predicts the effect of treatment on the clinical outcome. A strong patient-level correlation can coexist with poor prediction of treatment effects across trials.

  • Patient-level association asks whether patients with better surrogate results tend to experience better clinical outcomes.
  • Trial-level association asks whether treatments producing larger surrogate effects also produce larger clinical-outcome effects across studies.
  • A robust validation normally needs evidence at both levels, with particular attention to trial-level prediction.

Levels of evidence for surrogacy

Evidence for a surrogate ranges from biological plausibility to repeated empirical validation. The appropriate standard depends on the consequence of the decision and the uncertainty that remains. A measure accepted for one regulatory or reimbursement purpose should not automatically be treated as validated for every other purpose.

  1. Establish biological plausibility. Explain why the surrogate lies on a pathway relevant to the clinical outcome and identify plausible pathways that bypass it.
  2. Demonstrate measurement validity. Show that the surrogate can be measured accurately, reliably, and consistently in the intended population.
  3. Demonstrate patient-level association. Assess whether individual surrogate values or changes are associated with individual clinical outcomes.
  4. Demonstrate trial-level association. Assess whether treatment effects on the surrogate predict treatment effects on the clinical outcome across multiple trials.
  5. Test transportability. Examine whether the relationship holds across relevant interventions, disease stages, populations, settings, and follow-up periods.
  6. Quantify prediction uncertainty. Report the uncertainty around the predicted clinical effect rather than relying only on a correlation coefficient.

Estimating treatment effects on both endpoints

Surrogate evaluation begins with a treatment-effect estimate for each endpoint in each trial. The effect measure must suit the outcome, such as a mean difference for a continuous biomarker, a risk ratio for a binary response, or a hazard ratio for a time-to-event outcome. Direction and scale must be aligned before comparing effects across trials.

For trial (j), let (\hat{\theta}{Sj}) be the estimated treatment effect on the surrogate and (\hat{\theta}{Tj}) the estimated treatment effect on the true clinical endpoint. A trial-level model can then relate the two effects:

$$ \hat{\theta}{Tj} = \alpha + \beta\hat{\theta}{Sj} + \varepsilon_j $$

Here, (\alpha) is the predicted clinical effect when the surrogate effect is zero, (\beta) describes how clinical effects change with surrogate effects, and (\varepsilon_j) represents unexplained between-trial variation. Estimation should account for uncertainty in both treatment-effect estimates and, when relevant, their within-trial correlation.

Interpreting trial-level predictive performance

A high trial-level coefficient of determination can support surrogacy, but it does not by itself prove that predictions are sufficiently accurate for decisions. The number, size, diversity, and independence of trials affect how much confidence the analysis deserves. Prediction intervals are especially important because they show the range of clinical effects compatible with a new observed surrogate effect.

One summary is the trial-level proportion of variation explained:

$$ R^2_{trial} = 1 - \frac{\sum_j(\hat{\theta}{Tj}-\widehat{\theta}{Tj})^2}{\sum_j(\hat{\theta}_{Tj}-\overline{\theta}_T)^2} $$

Values closer to 1 indicate that the fitted surrogate relationship explains more of the between-trial variation in clinical effects. Interpretation must also consider model fit, calibration, influential studies, prediction error, and whether validation was internal or external.

The surrogate threshold effect

The surrogate threshold effect is the minimum treatment effect on the surrogate required to predict a non-zero benefit on the clinical endpoint at a stated confidence level. It translates a fitted trial-level relationship into a more decision-oriented benchmark. The threshold depends on the model, effect scales, evidence set, and chosen uncertainty level.

The threshold should not be treated as a universal clinical cut-off. It can change when new trials are added or when the analysis is applied to another intervention class or population.

Common validation designs

Surrogate validation can use individual participant data, aggregate trial data, or both. Individual data allow joint modelling of patient-level and trial-level relationships, whereas aggregate data may be more available but can support fewer checks. Meta-analytic approaches are often needed because a single trial cannot establish whether variation in surrogate effects predicts variation in clinical effects across treatments.

  • Joint models can estimate associations between surrogate and clinical outcomes while accounting for their distributions and censoring.
  • Two-stage meta-analytic models first estimate treatment effects within trials and then relate those effects across trials.
  • Landmark and time-dependent methods can reduce some biases when the surrogate is measured after treatment begins.
  • External validation tests predictions in trials not used to fit the surrogate relationship.

Why apparently promising surrogates can fail

A surrogate can fail when it captures only one part of the treatment's effect or when the intervention acts through a different mechanism from those used in validation. Bias can also arise when the surrogate is selected after results are known or when trials with both endpoints are unrepresentative. These failures can lead to an exaggerated, attenuated, or reversed prediction of patient benefit.

  • Off-pathway treatment effects can change the clinical outcome without changing the surrogate.
  • Surrogate-specific effects can change the surrogate without producing clinical benefit.
  • Treatment harms can offset benefits predicted from the surrogate.
  • Short follow-up can miss delayed benefits, harms, or loss of effect.
  • Informative censoring, treatment switching, and post-randomisation selection can distort associations.
  • Publication and availability bias can exclude trials with weak or contradictory surrogate relationships.

Using surrogate endpoints in trials

Surrogates can shorten follow-up, reduce sample-size requirements, or provide an earlier signal when direct outcomes take years to observe. Those advantages can accelerate evidence generation, especially in serious conditions with unmet need. They also transfer uncertainty from observation of the clinical outcome to prediction of that outcome.

Trial protocols should pre-specify the endpoint definition, measurement timing, estimand, missing-data handling, and analysis method. Reports should identify whether the surrogate is validated, reasonably likely, or exploratory for the exact intended context and should not describe surrogate improvement as proven patient benefit.

Regulatory acceptance and confirmatory evidence

Regulators may accept surrogate endpoints under standard or expedited pathways, but the evidentiary basis and post-authorisation obligations differ by jurisdiction and programme. Acceptance means the endpoint is considered adequate for a specific regulatory decision; it does not remove uncertainty about the eventual clinical benefit. Confirmatory studies may be required to verify benefit and reassess the benefit-risk balance.

  • Regulatory status should be reported with the jurisdiction, indication, intervention class, and date.
  • A surrogate described as reasonably likely to predict benefit carries more residual uncertainty than one supported by extensive validation.
  • Failure, delay, or ambiguity in confirmatory evidence is relevant to clinicians, payers, patients, and economic evaluators.

Implications for health technology assessment

Health technology assessment must translate surrogate evidence into expected effects on patient-relevant outcomes when those outcomes drive comparative effectiveness, quality-adjusted life-years, costs, or survival. The translation should be explicit and should preserve parameter, structural, and external-validity uncertainty. An accepted regulatory surrogate is therefore an input to assessment rather than automatic proof of cost-effectiveness.

Assessors should examine:

  • Whether the validation evidence matches the technology's mechanism, indication, population, comparator, and line of therapy.
  • Whether the surrogate-to-outcome relationship is estimated from randomized comparisons rather than prognostic association alone.
  • Whether predictions include uncertainty and allow for treatment effects not mediated by the surrogate.
  • Whether mature direct-outcome data are available and consistent with earlier surrogate-based predictions.
  • Whether additional evidence collection, conditional reimbursement, or managed entry can reduce decision uncertainty.

Representing surrogate uncertainty in an economic model

An economic model may use a surrogate-to-outcome equation to predict transition risks, survival, symptoms, or another patient-relevant input. The model should not insert the predicted mean effect as though it were directly observed. Probabilistic analysis should reflect uncertainty in the surrogate treatment effect, translation parameters, residual prediction error, and correlations among them.

Useful sensitivity analyses include:

  • A base case using the best-supported surrogate relationship for the intended context.
  • A direct-evidence scenario using mature clinical-outcome data when available.
  • Alternative validation models or evidence sets that test structural uncertainty.
  • A conservative scenario that attenuates the predicted clinical benefit or increases prediction error.
  • Scenario analyses for class effects, follow-up duration, and treatment-specific departures from the validation relationship.

Worked example: predicting a clinical effect

Suppose a meta-analytic validation model relates a log treatment effect on a surrogate to a log hazard ratio for a clinical outcome. For illustration, the fitted intercept is (0.02), the slope is (0.70), and a new trial reports a surrogate hazard ratio of (0.75). This example shows the calculation but does not establish that the assumed relationship is valid for any real technology.

First convert the surrogate hazard ratio to the log scale:

$$ \hat{\theta}_S = \ln(0.75) = -0.288 $$

Then apply the fitted relationship:

$$ \widehat{\theta}_T = 0.02 + 0.70(-0.288) = -0.182 $$

Convert the predicted log effect back to a hazard ratio:

$$ \widehat{HR}_T = \exp(-0.182) \approx 0.83 $$

The model predicts a clinical-outcome hazard ratio of approximately 0.83. A decision analysis must also use the prediction interval; if that interval includes no benefit or harm, the mean prediction alone gives a misleading impression of certainty.

What should be reported

Transparent reporting allows readers to judge whether the surrogate evidence supports the claim being made. The report should separate observed effects from predicted effects and clearly identify any extrapolation beyond the validation data. It should also explain how surrogate uncertainty affects the final clinical or economic conclusion.

  • The report should name the surrogate and the patient-relevant outcome it is intended to replace.
  • The report should define the population, disease stage, intervention class, comparator, setting, and follow-up period.
  • The report should describe patient-level and trial-level evidence separately.
  • The report should provide effect scales, model assumptions, calibration results, and prediction intervals.
  • The report should disclose missing trials, data limitations, and conflicts between surrogate and clinical evidence.
  • The report should state how confirmatory evidence will be incorporated when it becomes available.

Common misunderstandings

Surrogate endpoints are often discussed as if biological credibility, correlation, validation, and regulatory acceptance were equivalent. They are not. Keeping these claims separate prevents an intermediate result from being overstated as demonstrated patient benefit.

  • A statistically significant treatment effect on a surrogate does not prove a treatment effect on the clinical outcome.
  • A strong association between the surrogate and prognosis does not establish trial-level surrogacy.
  • A validated surrogate in one setting is not automatically transferable to another disease, treatment class, or population.
  • Regulatory acceptance does not guarantee reimbursement or a favourable health technology assessment.
  • Earlier measurement reduces trial duration but does not eliminate long-term uncertainty.
  • A surrogate endpoint can support decision making while still requiring confirmatory direct-outcome evidence.

Interpreting the evidence for a decision

The practical question is not whether a surrogate is simply valid or invalid, but how much confidence it provides for the decision at hand. Confidence is stronger when biological reasoning, measurement quality, patient-level association, trial-level prediction, external validation, and contextual fit point in the same direction. Decision makers should lower confidence when any of these elements is weak or when the consequence of a wrong prediction is substantial.

A well-supported surrogate can make earlier decisions possible. Its use remains an explicit exchange: earlier information is gained at the cost of added uncertainty about the patient benefit that has not yet been directly observed.

Library

Publications

1
  • Book

    Economic Evaluation in Clinical Trials — Glick, Doshi, Sonnad & Polsky, 2nd Edition ed., 2015 (Oxford University Press)

    Practical guidance on conducting cost-effectiveness analyses alongside controlled trials, covering trial design, measurement of costs and quality-adjusted life years, handling censored and missing data, and reporting stochastic uncertainty. Volume 4 in the Handbooks in Health Economic Evaluation series.

Frequently Asked Questions (6)

  • What is a surrogate endpoint?

    A trial outcome used as a substitute for a direct clinical benefit measure, such as survival, based on its expected predictive relationship to that outcome.

    Source: Prentice 1989

  • What is the danger in relying on a surrogate endpoint?

    A surrogate endpoint stands in for a clinical outcome that would take too long or too many patients to measure directly, such as using a blood marker instead of survival. The danger is that a treatment can move the surrogate without delivering the real benefit, or even while causing harm, if the surrogate does not truly lie on the causal path to the outcome. History records treatments that improved a marker yet worsened survival. A surrogate is trustworthy only when firmly validated against the outcome it represents. Its use is a calculated risk. Fleming and DeMets (1996) warn of this.

    Source: Fleming & DeMets 1996

  • Why are surrogate endpoints used?

    Surrogate endpoints are used because they can be measured sooner, more easily, or with smaller samples than the ultimate clinical outcomes, such as survival, allowing trials to be shorter, smaller, and quicker, which speeds the evaluation of treatments and can accelerate access to beneficial ones. A surrogate that reliably predicts clinical benefit lets effect be assessed before the clinical outcome occurs. This efficiency is the main reason for using surrogate endpoints, provided their relationship to the clinical outcome is well established, since their value depends on genuinely reflecting the benefit that matters.

    Source: FDA-NIH Biomarker Working Group 2016

  • What makes a surrogate endpoint valid?

    A surrogate endpoint is valid when strong evidence shows that a treatment's effect on the surrogate reliably predicts its effect on the clinical outcome, so that changes in the surrogate capture the treatment's clinical benefit. This requires more than correlation between the surrogate and the outcome; the surrogate must lie on the causal pathway and capture the full effect of treatment on the outcome, as reflected in Prentice's criteria. Establishing validity is demanding, since a surrogate can be associated with an outcome yet fail to capture a treatment's effect on it, so validation is central to a surrogate's usefulness.

    Source: Prentice 1989

  • What are the risks of using surrogate endpoints?

    The risks of using surrogate endpoints arise because a treatment can affect a surrogate without producing the expected clinical benefit, or can have off-target harms the surrogate does not capture, so relying on an inadequately validated surrogate can lead to adopting treatments that do not truly benefit patients or that cause harm. History includes cases where surrogate improvements did not translate into clinical benefit or accompanied worse outcomes. These risks mean surrogate endpoints must be well validated, and treatments approved on surrogates are often confirmed with clinical outcome data, so the surrogate's limitations are recognised.

    Source: FDA-NIH Biomarker Working Group 2016

  • How does a surrogate endpoint relate to clinical outcomes?

    A surrogate endpoint relates to clinical outcomes as a substitute expected to predict them: it stands in for a direct measure of clinical benefit, such as survival, on the basis that a treatment's effect on the surrogate reflects its effect on the clinical outcome. When this relationship is validated, the surrogate can indicate clinical benefit sooner or more easily. However, if the link is weak or the surrogate does not capture the full treatment effect, changes in it may not correspond to clinical benefit. So the usefulness of a surrogate endpoint depends entirely on its validated relationship to the clinical outcomes that matter.

    Source: Prentice 1989

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 22 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-CTM-093

Stable URI · Machine-readable · Resolvable · CC BY 4.0