VerifiedEvidence: highv1.0.0

Quality of Evidence

A judgement about how much confidence can be placed in an effect estimate, based on study design, risk of bias, and precision.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

How evidence quality determines confidence in an effect

Quality of evidence is a judgment about how confidently an estimated effect can support a particular conclusion or decision. In contemporary GRADE terminology, this is usually called certainty of evidence and is rated for each outcome across the relevant body of evidence. This page explains how study limitations, inconsistency, indirectness, imprecision, and publication bias change confidence and how those judgments should inform economic evaluation and policy.

Certainty applies to an outcome and question

Evidence certainty is not a permanent label attached to a publication. The same group of studies can provide high certainty for mortality, moderate certainty for symptoms, and very low certainty for rare harms. The judgment also depends on the population, intervention, comparator, outcome, and decision context being assessed.

LevelMeaning for the effect estimateDecision implication
HighThe true effect is very likely to be close to the estimate.Further evidence is unlikely to change confidence materially.
ModerateThe true effect is likely close to the estimate but may be substantially different.Important residual uncertainty should be considered.
LowThe true effect may be substantially different from the estimate.Decisions should be cautious and uncertainty explored.
Very lowThe true effect is likely substantially different or remains highly uncertain.The estimate provides limited support without strong qualification.

These categories communicate confidence, not whether the effect is large, beneficial, statistically significant, or clinically important. High-certainty evidence can show little or no benefit, while low-certainty evidence can suggest a large benefit that remains uncertain.

Study quality and certainty of a body of evidence are different

Risk-of-bias assessment examines potential flaws within individual studies or results. Certainty assessment integrates risk of bias with consistency across studies, directness to the question, precision, publication bias, and other considerations. A well-conducted single study can still provide imprecise or indirect evidence, and several imperfect studies can sometimes provide a coherent body of evidence.

Related terms should be kept separate:

  • Reporting quality concerns whether methods and results are described adequately.
  • Methodological quality concerns whether the design and conduct are appropriate.
  • Risk of bias concerns systematic error in a study result.
  • Evidence certainty concerns confidence in the effect estimate for a specified outcome.
  • Recommendation strength also considers benefits, harms, values, resources, equity, acceptability, and feasibility.

The question defines what counts as direct evidence

Certainty assessment begins with a clear question, commonly structured by population, intervention, comparator, and outcome. Evidence is direct when it closely matches that question. A precise estimate from another population, dose, comparator, setting, or surrogate outcome may still provide lower certainty for the target decision.

The assessment should specify:

  • Target population and relevant subgroups.
  • Intervention, dose, delivery, and duration.
  • Comparator representing the actual decision alternative.
  • Outcome definition, measurement, and follow-up.
  • Effect measure and threshold for an important benefit or harm.
  • Setting, care pathway, and jurisdiction when they affect applicability.

Starting certainty depends on the evidence design

Within GRADE for intervention effects, randomised trials generally begin at high certainty and non-randomised evidence generally begins at low certainty before domain judgments. This starting point reflects expected protection from confounding, not an automatic final rating. Serious limitations can lower randomised evidence, while strong features can raise certainty from observational evidence.

Other question types, such as diagnosis, prognosis, prevalence, or qualitative evidence, require an appropriate framework and starting logic. Analysts should not force every evidence question into intervention-effect rules.

Risk of bias asks whether study methods distort the result

Risk of bias evaluates whether the design, conduct, analysis, or reporting could systematically move the estimate away from the truth. The relevant domains depend on the study design and outcome. Judgments should focus on the result contributing to the synthesis rather than assigning a vague quality score to the whole paper.

Potential concerns include:

  • Problems in randomisation or allocation concealment.
  • Deviations from intended interventions.
  • Missing outcome data related to prognosis or treatment.
  • Measurement influenced by knowledge of treatment or unsuitable instruments.
  • Selective reporting among outcomes, time points, analyses, or subgroups.
  • Confounding, selection, exposure classification, or immortal-time bias in observational studies.

Risk of bias should not be reduced to a checklist total

Different flaws can have different directions and consequences. Adding item scores assumes equal importance and can allow a critical defect to be cancelled by several minor strengths. Domain-level judgments with reasons are more informative than a numeric study-quality total.

Reviewers should record the signalling evidence, judgment, likely direction, likely magnitude, and impact on the body of evidence. Disagreement should be resolved or reported rather than hidden inside an average score.

Inconsistency concerns unexplained differences across studies

Inconsistency is present when effect estimates differ more than expected from sampling variation and the differences cannot be explained credibly. Variation may arise from populations, interventions, comparators, outcomes, methods, follow-up, or bias. Heterogeneity should be examined clinically and methodologically before relying on a statistical statistic.

A commonly reported statistic is:

$$ I^2=\max\left(0,\frac{Q-df}{Q}\right)\times100% $$

where (Q) is Cochran's heterogeneity statistic and (df) is its degrees of freedom. (I^2) does not measure clinical importance, the direction of effects, or whether heterogeneity is explained, and it can be imprecise when few studies are available.

Prediction intervals help interpret heterogeneity

A confidence interval around the pooled effect describes uncertainty in the mean effect. A prediction interval estimates a range within which the effect in a future comparable setting may lie, conditional on model assumptions. It can reveal decision-relevant heterogeneity that a statistically significant pooled estimate conceals.

When effects vary across studies, reviewers should examine whether all effects favour the same option, whether some cross an important threshold, and whether credible subgroups explain the variation. Post hoc subgroup explanations require caution.

Indirectness is a mismatch with the target decision

Indirectness arises when evidence differs from the target question or relies on an indirect comparison or surrogate outcome. The judgment should explain which element is mismatched and why that difference could change the effect. Geographic difference alone is not automatically indirect if care, populations, and effect modifiers remain applicable.

Sources of indirectness include:

  • A study population with different severity, comorbidity, age, or prior treatment.
  • An intervention delivered at another dose, intensity, or setting.
  • A comparator that is no longer relevant.
  • A surrogate endpoint used in place of patient-relevant benefit.
  • Follow-up too short for durable benefits or delayed harms.
  • Indirect comparison through a network without a direct head-to-head trial.

Surrogate endpoints create specific uncertainty

A surrogate can be measured precisely yet provide indirect evidence for the patient-relevant outcome. Certainty depends on whether treatment effects on the surrogate reliably predict treatment effects on how patients feel, function, or survive in the intended context. Biological plausibility and individual-level association alone are insufficient.

The assessment should identify the target clinical outcome, intervention class, population, validation evidence, and prediction uncertainty. Regulatory acceptance of a surrogate does not automatically establish high certainty for an economic model's long-term health effect.

Imprecision depends on the decision threshold

Imprecision concerns whether the confidence interval allows meaningfully different conclusions. Sample size and event count matter, but the central question is whether the interval crosses thresholds separating important benefit, trivial effect, no effect, and important harm. Statistical significance alone is an inadequate rule.

For an estimated effect (\hat{\theta}) with standard error (SE(\hat{\theta})), a conventional 95% confidence interval is:

$$ \hat{\theta}\pm1.96\times SE(\hat{\theta}) $$

Interpretation requires the effect scale and decision threshold. A narrow relative-effect interval can still imply a wide range of absolute effects when baseline risk is uncertain.

Absolute effects make certainty decision relevant

Relative effects can appear stable across risk groups while absolute benefit varies with baseline risk. Evidence profiles should present absolute and relative effects when possible. Certainty in the absolute effect can be lower when baseline risk is uncertain or poorly matched to the target population.

If baseline risk is (p_0) and the risk ratio is (RR), treated risk is:

$$ p_1=RR\times p_0 $$

The absolute risk difference is:

$$ RD=p_1-p_0=p_0(RR-1) $$

Uncertainty in both (RR) and (p_0) should be represented. Applying a trial-relative effect to an inappropriate baseline risk can misstate population benefit.

Optimal information size is one imprecision check

The optimal information size asks whether the accumulated evidence includes roughly the sample size that a sufficiently powered single trial would require for an important effect. Failure to meet it can support concern about imprecision, especially when events are few. It should be interpreted alongside confidence intervals and decision thresholds rather than as an automatic downgrade.

Rare harms can remain highly uncertain despite a large total sample if follow-up is short, ascertainment is weak, or the event population is restricted. Absence of observed events is not proof of no risk.

Publication bias concerns missing evidence

Publication bias occurs when available evidence differs systematically from all evidence generated, often because study results influence publication, availability, or reporting. Small positive studies, missing registered outcomes, delayed negative results, and sponsor control can exaggerate effects. Statistical asymmetry tests have limited power and can reflect heterogeneity rather than publication bias.

Assessment should examine:

  • Trial registries, protocols, regulatory submissions, and unpublished reports.
  • Discrepancies between planned and published outcomes or analyses.
  • Small-study effects and unexplained funnel-plot asymmetry.
  • Industry sponsorship and control of data or publication.
  • Known completed studies with unavailable results.
  • Selective availability of harms and long-term follow-up.

Large effects can increase certainty under strict conditions

Observational evidence may be upgraded when a large or very large effect is credible and unlikely to be explained by bias or confounding. The effect should be consistent, directly relevant, and based on a comparison that does not exaggerate the estimate. Implausibly large effects can also signal sparse data, selection, measurement error, or publication bias.

Upgrading should not occur when the same limitation is already being used to explain the apparent magnitude. The rationale and effect threshold should be explicit.

A dose-response gradient can support a causal effect

Evidence that larger exposure produces larger benefit or harm can increase confidence when the gradient is credible and alternative explanations are unlikely. Dose can refer to biological dose, treatment intensity, duration, adherence, or exposure level. Confounding by severity or treatment selection can also generate an apparent gradient.

The analysis should examine exposure measurement, functional form, temporality, thresholds, and competing causal explanations. A statistically significant trend alone is insufficient.

Residual confounding can sometimes strengthen confidence

Observational evidence can be upgraded when all plausible residual confounding would reduce an observed effect or create an effect where none is observed. This is a demanding condition and should not be used simply because measured adjustment was extensive. Reviewers must state the confounder, its expected direction, and why the observed result remains conservative.

Unknown confounding cannot be assumed to operate helpfully. Sensitivity analysis can quantify how strong unmeasured confounding would need to be to change the conclusion.

Certainty ratings are outcome-specific

Benefits, harms, quality of life, resource use, and treatment burden often rely on different studies and methods. A single overall label for the intervention can hide high certainty for one outcome and very low certainty for another. Evidence profiles should rate every critical outcome separately.

Decision makers then integrate outcomes rather than averaging certainty ratings. A severe harm with low-certainty evidence can still matter when its plausible consequences are large.

Critical outcomes should be chosen before results

Outcomes should reflect what matters to patients and decisions, not what happens to be reported. Rating importance after seeing favourable results invites selective emphasis. Patients, clinicians, methodologists, and decision makers can contribute to prioritisation.

Outcome sets should distinguish:

  • Critical outcomes that determine the decision.
  • Important but non-critical outcomes that add context.
  • Surrogate, process, or exploratory outcomes.
  • Benefits, harms, burden, quality of life, and resource consequences.

Evidence synthesis and certainty assessment are separate steps

Meta-analysis estimates a combined effect under specified assumptions. Certainty assessment judges how much confidence to place in that estimate for the target question. A precise pooled result can be low certainty because of bias or indirectness, while a narrative synthesis can support moderate certainty when estimates are consistent and limitations are understood.

The statistical model, effect measure, heterogeneity, and certainty rationale should be reported together but not conflated. Failure to meta-analyse does not automatically mean the evidence is very low certainty.

Network meta-analysis adds coherence questions

Network meta-analysis combines direct and indirect comparisons across several interventions. Certainty depends on risk of bias, indirectness, imprecision, heterogeneity, and incoherence between direct and indirect evidence. Transitivity requires that distributions of important effect modifiers are sufficiently comparable across comparisons.

A highly ranked treatment can have low-certainty evidence when estimates are sparse or network assumptions are weak. Ranking probabilities should not substitute for pairwise effect estimates and certainty judgments.

Diagnostic and prognostic evidence need adapted frameworks

Diagnostic accuracy, prognosis, prevalence, and qualitative findings do not fit intervention-effect judgments without modification. Diagnostic certainty must consider patient selection, reference standards, test conduct, thresholds, directness, consistency, and precision for sensitivity and specificity. Prognostic evidence must consider population, outcome measurement, adjustment, attrition, and model performance.

The framework should be named and matched to the question. Using GRADE language without applying the relevant method can create an appearance of rigour without a valid judgment.

Evidence certainty is not recommendation strength

A strong recommendation can sometimes be made with low-certainty evidence when consequences are clear, alternatives are unacceptable, or the intervention prevents catastrophic harm at low burden. A weak or conditional recommendation can accompany high-certainty evidence when benefits and harms are closely balanced or preferences vary. Certainty is one input to recommendation, not the recommendation itself.

Other considerations include:

  • Magnitude and balance of desirable and undesirable effects.
  • Values and preferences and their variability.
  • Resource requirements and cost-effectiveness.
  • Equity, acceptability, and feasibility.
  • Urgency, reversibility, and consequences of delay.

Economic evaluation needs outcome-specific certainty

Health-economic models combine treatment effects, baseline risks, costs, utilities, extrapolation, and structural assumptions. A single evidence grade cannot describe the whole model. Each decision-critical input should have traceable evidence and uncertainty, while model validation and structural credibility are assessed separately.

Low-certainty treatment-effect evidence should influence probability distributions, scenario analysis, interpretation, and research recommendations. It should not be converted mechanically into an arbitrary wider range without considering the domain causing uncertainty.

Certainty affects value-of-information analysis

Evidence uncertainty can create a risk of adopting the wrong option. Value-of-information analysis estimates the expected benefit of reducing uncertainty, conditional on model structure and current evidence. It complements rather than replaces qualitative certainty assessment.

Expected value of perfect information per patient is:

$$ EVPI=E_{\boldsymbol{\theta}}\left[\max_d NB(d,\boldsymbol{\theta})\right]-\max_d E_{\boldsymbol{\theta}}\left[NB(d,\boldsymbol{\theta})\right] $$

where (d) is a decision option, (NB) is net benefit, and (\boldsymbol{\theta}) represents uncertainty. A high EVPI can justify research, but only if the model includes the important sources of uncertainty credibly.

Worked certainty judgment

Suppose five randomised trials estimate a reduction in hospitalisation. Allocation and outcome measurement are sound, but two trials have substantial missing outcomes; effect estimates vary in direction; the population is less severe than the target population; and the confidence interval includes both a trivial and an important benefit. These are distinct concerns rather than four reasons to lower certainty automatically by four levels.

Reviewers should judge the seriousness of risk of bias, inconsistency, indirectness, and imprecision, identify whether the same underlying problem drives more than one domain, and avoid double counting. The final rating should state the effect estimate, absolute range, reasons for each judgment, and what evidence could increase confidence.

Certainty can change when evidence changes

New studies, longer follow-up, unpublished results, improved measurement, or a new target population can change the rating. The assessment should have a search date, version, and update trigger. A rating copied into another review without rechecking the question and evidence may no longer be valid.

Living reviews and guidance require governance for surveillance, reanalysis, and communication. A change in rating should preserve the previous rationale and identify what new evidence altered the judgment.

Common mistakes

Quality-of-evidence language can create false precision when ratings are treated as mechanical scores. The purpose is a transparent, structured judgment about confidence, not a substitute for understanding the evidence. The following errors should be checked before a rating informs a decision.

  • Assigning one certainty rating to an entire review or intervention.
  • Treating randomised design as automatically high certainty at the end of assessment.
  • Equating statistical significance with high certainty.
  • Downgrading for a small sample without examining confidence intervals and decision thresholds.
  • Using (I^2) alone to decide inconsistency.
  • Treating a surrogate as direct evidence because it is objectively measured.
  • Downgrading twice for the same underlying limitation without justification.
  • Using a numeric checklist total to determine risk of bias or certainty.
  • Converting low certainty into a claim of no effect.
  • Treating high certainty as proof that an intervention should be adopted.

Reporting a certainty assessment

Transparent reporting should allow another reviewer to understand and challenge every judgment. Evidence tables should connect the question, effect estimate, absolute effect, participants, studies, domains, and final certainty. Explanatory footnotes should identify the evidence rather than repeat a category label.

  • State the population, intervention, comparator, outcome, setting, and follow-up.
  • Report relative and absolute effects with confidence intervals.
  • Identify study designs and participants contributing to each outcome.
  • Record risk of bias, inconsistency, indirectness, imprecision, and publication-bias judgments.
  • Record any upgrading for large effect, dose response, or opposing residual confounding.
  • Explain each rating change with outcome-specific evidence and avoid double counting.
  • Distinguish evidence certainty from effect size, importance, and recommendation strength.
  • Record reviewers, conflicts, search date, framework version, and update plan.

The decision standard

Quality of evidence is high when the complete body of directly relevant evidence makes a materially different true effect unlikely. Confidence decreases when study limitations, unexplained inconsistency, indirectness, imprecision, or missing evidence allow different decision-relevant conclusions. A good certainty assessment makes those reasons visible, outcome-specific, and proportionate so decision makers can act without either overstating weak evidence or dismissing uncertainty as an absence of knowledge.

Library

Publications

1
  • Journal article

    GRADE Guidelines: 1. Introduction — GRADE Evidence Profiles and Summary of Findings Tables — Guyatt, Oxman, Akl, Kunz, Vist, Brozek, et al., GRADE Series ed., 2011 (Journal of Clinical Epidemiology)

    The introductory paper of the GRADE (Grading of Recommendations Assessment, Development and Evaluation) series, setting out how to rate certainty of evidence (high/moderate/low/very low) and build evidence profiles and summary-of-findings tables.

Frequently Asked Questions (6)

  • What is quality of evidence?

    A judgement about how much confidence can be placed in an effect estimate, based on study design, risk of bias, and precision.

    Source: Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924-926. doi:10.1136/bmj.39489.470347.AD.

  • What does the quality of evidence express about an effect estimate?

    Quality of evidence expresses how much confidence can be placed in an effect estimate as a reflection of the truth, rather than how large or favourable the effect is. It rests on features such as the study design, the risk of bias, the consistency of results, and the precision of the estimate, which together determine how firmly the number can be believed. A large effect from weak studies may warrant less confidence than a small one from strong ones. It rates belief in the estimate, not its size. Guyatt and colleagues (2008) describe this.

    Source: Guyatt et al. 2008

  • What factors determine quality of evidence?

    The quality of evidence is determined by the study design, with randomised trials starting higher and observational studies lower; the risk of bias in the studies; the consistency of findings across them; the directness or applicability of the evidence to the question; the precision of the estimates; and the likelihood of publication bias. In GRADE, these factors adjust the rating up or down. Stronger designs, unbiased, consistent, precise, and directly relevant evidence raise quality, while limitations lower it. So multiple factors together determine the quality of evidence, reflecting how much the estimate can be trusted.

    Source: Guyatt et al. 2008

  • Why does quality of evidence matter?

    Quality of evidence matters because decisions should reflect not only the estimated effect but how confident one can be in it: high-quality evidence supports firm conclusions and strong recommendations, while low-quality evidence warrants caution and weaker recommendations. Ignoring quality could lead to overconfident decisions based on unreliable estimates. By conveying the trustworthiness of the evidence, quality ratings help decision makers weigh it appropriately and communicate its reliability. So assessing the quality of evidence is central to evidence-based decisions, ensuring that conclusions and their strength reflect how much the underlying evidence can be trusted.

    Source: Guyatt et al. 2008

  • How is quality of evidence graded?

    Quality of evidence is graded using frameworks such as GRADE, which start from the study design and adjust the rating according to risk of bias, inconsistency, indirectness, imprecision, and publication bias, with large effects or a dose-response relationship able to raise it. Combining these judgements yields a grade for each outcome, commonly high, moderate, low, or very low, reflecting the confidence that the estimated effect reflects the truth. The grading applies to the body of evidence, not individual studies. So quality of evidence is graded by systematically appraising these factors and assigning a certainty level.

    Source: Guyatt et al. 2008

  • How does quality of evidence relate to certainty of evidence?

    Quality of evidence and certainty of evidence are closely related terms, both referring to how much confidence can be placed in an effect estimate; in GRADE, the term certainty of evidence has increasingly been preferred to quality of evidence, but they denote the same judgement about the trustworthiness of the estimate. Both are assessed from the same factors and graded into the same levels. So the two are largely synonymous, with certainty of evidence the more current term, emphasising confidence in the estimate rather than the quality of the studies alone, though both convey how reliable the evidence is.

    Source: Guyatt et al. 2008

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 22 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-EA-033

Stable URI · Machine-readable · Resolvable · CC BY 4.0