VerifiedEvidence: highv1.0.0

Benchmarking

The process of comparing an organisation's performance metrics against peer organisations or established standards, identifying strengths and improvement areas.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

How benchmarking turns comparison into improvement

Benchmarking compares an organisation, service, team, or system with relevant peers, standards, or prior performance to identify meaningful variation and opportunities to improve. It is useful only when the measures, populations, and contexts are sufficiently comparable. This page explains how benchmarks are selected, adjusted, interpreted, and translated into action without mistaking a ranking for a diagnosis.

A benchmark needs a clear reference point

A benchmark is the value against which performance is compared. The reference can represent current peers, an external standard, a high-performing group, an evidence-based target, or the organisation's own past performance. Each choice answers a different question and can produce a different conclusion.

Benchmark typeMain questionImportant limitation
Peer benchmarkHow does performance compare with similar organisations?Similarity depends on defensible peer selection.
Best-in-class benchmarkWhat performance has been achieved by strong performers?The result may not be transferable to every setting.
Standard or targetDoes performance meet an agreed expectation?The target may be aspirational, minimum, or evidence-based.
Historical benchmarkIs the organisation improving over time?System-wide change and data revisions can distort the trend.
Process benchmarkHow do work practices differ from effective organisations?A copied process may fail if context and mechanisms differ.

The purpose should be stated before selecting metrics

Benchmarking can support quality improvement, productivity review, accountability, commissioning, regulation, learning, or strategic planning. A metric chosen for public accountability may not be detailed enough to guide operational change. The purpose determines the unit of analysis, comparator, frequency, level of adjustment, and acceptable uncertainty.

A useful specification identifies:

  • The decision or improvement question.
  • The organisation, service, pathway, or population being assessed.
  • The intended users and consequences of the comparison.
  • The reference group or standard.
  • The measurement period and update frequency.
  • The actions that a result could reasonably trigger.

Comparable organisations are not always nearby organisations

Peers should be selected because they face sufficiently similar demand, responsibilities, resources, and operating conditions. Geographic proximity or organisational type alone may not create a fair comparison. Peer selection should occur before results are examined to reduce the risk of choosing a favourable reference group.

Relevant matching dimensions can include:

  • Population size, age structure, deprivation, and disease burden.
  • Service scope, case complexity, and referral role.
  • Rurality, travel patterns, and local market conditions.
  • Workforce supply, teaching status, and specialist functions.
  • Funding arrangements, prices, and accounting rules.
  • Data coverage, coding practice, and reporting maturity.

Every indicator needs a technical definition

A shared label does not guarantee a shared measure. Organisations can calculate waiting time, readmission, cost, productivity, or mortality using different populations, exclusions, dates, or denominators. An indicator dictionary should make every component reproducible before results are compared.

The specification should state:

  • The numerator and denominator.
  • Inclusion and exclusion criteria.
  • Unit of measurement and direction of better performance.
  • Data source, refresh date, and responsible owner.
  • Observation window and follow-up period.
  • Risk adjustment, standardisation, and suppression rules.
  • Treatment of missing, duplicate, and extreme values.
  • Known limitations and permitted uses.

Rates need appropriate denominators

Counts can reflect organisation size rather than performance. Rates relate events to the relevant population, activity, exposure, or time at risk. The denominator must correspond to the process that could generate the event.

A basic rate is:

$$ Rate = \frac{Number\ of\ events}{Population\ or\ activity\ at\ risk}\times k $$

where (k) is a scaling constant such as 100, 1,000, or 100,000. A low-volume service can show a large rate change after only one event, so uncertainty and counts should be reported alongside the rate.

Case mix can make raw comparisons unfair

Outcome and resource measures often depend on patient need, severity, comorbidity, age, deprivation, and other factors outside the immediate control of the provider. Risk adjustment estimates what would be expected given the case mix and enables a more like-for-like comparison. It should not adjust away inequities, poor access, or factors that the service is expected to change.

An observed-to-expected ratio is:

$$ O/E = \frac{Observed\ events}{Expected\ events\ given\ case\ mix} $$

A ratio above 1 indicates more events than expected and a ratio below 1 indicates fewer, assuming a higher event count represents worse performance. Interpretation depends on model calibration, coding quality, omitted variables, and whether the expected value reflects an appropriate reference population.

Standardisation makes adjusted rates interpretable

Direct standardisation applies organisation-specific stratum rates to a common standard population. Indirect standardisation compares observed events with those expected from reference rates applied to the organisation's population. The methods answer related but different questions and should be labelled clearly.

For direct standardisation across strata (g):

$$ R_{std}=\frac{\sum_g w_g r_g}{\sum_g w_g} $$

where (r_g) is the organisation's rate in stratum (g) and (w_g) is the standard population weight. Sparse strata can make direct estimates unstable, while indirect results are not always directly comparable across organisations with different case mixes.

Uncertainty matters more than rank order

Observed differences contain both systematic variation and random variation. A league table can imply precision that the data do not support, especially for small organisations or rare outcomes. Confidence intervals and statistical process methods help distinguish unusual results from expected fluctuation.

A standardised score can be expressed as:

$$ z_i=\frac{x_i-\mu}{\sigma} $$

where (x_i) is organisation (i)'s result, (\mu) is the reference mean, and (\sigma) is an appropriate measure of variation. A z-score should not be used mechanically when distributions are skewed, denominators differ, observations are correlated, or hierarchical structure is important.

Funnel plots show precision and outliers

A funnel plot displays organisational performance against a measure of precision, commonly activity volume. Control limits widen for smaller denominators and narrow for larger ones, reducing the tendency to label small providers as extreme because of random variation. Points outside limits are signals for investigation rather than proof of good or poor practice.

The plot should identify the target or average, control limits, denominator, adjustment method, and data period. Multiple testing, overdispersion, repeated measures, and data-quality differences can create more apparent outliers than the basic model predicts.

Statistical significance is not practical importance

A large dataset can make a small, operationally unimportant difference statistically significant. A smaller service can have a clinically important gap with wide uncertainty. Benchmarking should therefore combine statistical evidence with absolute magnitude, clinical importance, cost, feasibility, and patient impact.

Useful interpretation reports:

  • The absolute difference from the benchmark.
  • The relative difference or ratio.
  • The uncertainty interval.
  • The number of people, events, days, or currency units represented.
  • A meaningful target or minimally important difference when one exists.
  • The likely consequences of acting or not acting.

Productivity benchmarking needs both inputs and outputs

Productivity compares the quantity and quality of outputs with the inputs used to produce them. A simple activity-per-staff measure may be useful operationally but can reward volume without quality or case complexity. A broader assessment should include relevant labour, capital, supplies, outcomes, and patient experience.

A basic productivity ratio is:

$$ Productivity = \frac{Quality\ adjusted\ output}{Resource\ input} $$

Higher output per input is not automatically better if waiting, safety, equity, staff wellbeing, or outcomes worsen. Price differences should also be separated from real resource-use differences when comparing costs.

Efficiency analysis and benchmarking are related but distinct

Benchmarking describes performance relative to a reference. Frontier methods such as data envelopment analysis or stochastic frontier analysis estimate efficiency relative to the best observed production frontier under model assumptions. These methods can strengthen productivity analysis but should not be presented as neutral rankings.

Results depend on input and output selection, scale assumptions, statistical noise, environmental variables, and the quality of observed peers. An organisation can appear efficient because all peers perform poorly or because important outcomes were omitted.

Process benchmarking looks beneath the outcome

An outcome gap does not reveal which practice caused it. Process benchmarking compares workflows, staffing models, scheduling, clinical pathways, information systems, and other operational features to identify plausible mechanisms. The aim is to learn what produces the result, not simply to imitate the visible feature of a high performer.

Process comparison should examine:

  • The sequence and timing of work.
  • Roles, skill mix, handoffs, and decision rights.
  • Capacity, demand management, and scheduling.
  • Standardisation and justified variation.
  • Information flow, digital support, and feedback.
  • Local constraints that affect transferability.

Variation requires diagnosis before action

Performance variation can arise from real differences in care, population need, data quality, coding, random fluctuation, or measurement design. Immediate intervention based on the headline measure can waste resources or create harm. A structured diagnostic phase should confirm the signal and identify its cause.

  1. Verify the data. Check definitions, completeness, coding, denominators, and reporting changes.
  2. Confirm comparability. Review peers, case mix, service scope, and contextual differences.
  3. Quantify the gap. Examine absolute, relative, and uncertainty measures over time.
  4. Disaggregate the result. Locate variation by pathway, population, team, site, or time.
  5. Investigate mechanisms. Compare processes and gather staff and patient evidence.
  6. Select and test action. Implement a plausible change with measures for outcome, process, and unintended effects.

Trends complement cross-sectional comparisons

A provider may perform below peers but improve rapidly, while a high-ranked provider may be deteriorating. Time series show direction, stability, and the effect of interventions. They also reveal whether an apparent gap results from a temporary disruption or a sustained pattern.

Trend analysis should annotate definition changes, shocks, mergers, coding initiatives, and policy changes. Repeated comparisons should use consistent data vintages or clearly reconcile revisions.

Composite scores can conceal trade-offs

Combining several indicators into one score can simplify reporting but requires decisions about normalisation, direction, weighting, missing data, and compensation between dimensions. A high value in one area may offset a serious weakness in another. The component results should therefore remain visible.

If normalised indicators (z_m) have weights (w_m), an additive score is:

$$ Score_i=\sum_{m=1}^{M}w_m z_{im}, \quad \sum_{m=1}^{M}w_m=1 $$

Weights express priorities and should be justified rather than disguised as technical facts. Sensitivity analysis should test whether plausible alternative weights change the conclusion.

Equity should be benchmarked explicitly

An organisation can show strong average performance while some populations experience poor access, quality, or outcomes. Equity benchmarking disaggregates results by relevant characteristics and compares gaps as well as averages. Case-mix adjustment should not erase differences that reflect unequal treatment or avoidable barriers.

Useful equity comparisons include:

  • Outcomes and access by deprivation, ethnicity, sex, age, disability, geography, or other relevant dimensions.
  • Absolute and relative gaps within each organisation.
  • Whether high-performing organisations achieve both strong averages and small inequities.
  • Differential waiting, cancellation, non-attendance, experience, and digital exclusion.
  • Data completeness and representation for marginalised groups.

Gaming and measurement burden can distort behaviour

When benchmark results affect reputation, payment, or regulation, organisations may focus narrowly on the measured target, change coding, select easier patients, or shift activity outside the indicator. Even well-intended measurement can consume staff time and crowd out care. Balanced measures and audit reduce these risks.

A benchmark system should monitor unexpected denominator changes, exclusions, coding shifts, threshold clustering, substitution, and deterioration in unmeasured outcomes. Metrics should be retired when they no longer add enough value to justify their burden.

Learning requires more than naming top performers

High performance may result from transferable practice, favourable context, measurement artefact, or chance. Learning collaboratives, site visits, qualitative inquiry, and shared improvement methods can identify the mechanisms behind the result. Organisations providing data should receive timely, usable feedback rather than a static rank.

Transfer should be treated as a testable adaptation. The adopting organisation should specify which mechanism is expected to work locally and monitor whether outcomes improve without harmful side effects.

Worked benchmarking example

Suppose a hospital records 84 readmissions among 1,200 eligible discharges, giving a raw rate of 7.0%. Its case-mix model predicts 72 readmissions. The observed-to-expected ratio is therefore:

$$ O/E=\frac{84}{72}=1.17 $$

The hospital recorded 17% more readmissions than expected under the model. This is a signal to inspect uncertainty, model fit, coding, discharge pathways, and subgroup patterns; it does not prove that care quality caused the difference.

Common mistakes

Benchmarking can become misleading when numerical comparison is separated from measurement design and context. A rank is easy to communicate but often weaker than a carefully adjusted estimate with uncertainty. The following errors should be checked before results drive action.

  • Comparing organisations with different definitions, populations, or service responsibilities.
  • Selecting peers after seeing which comparison looks favourable.
  • Ranking small differences without uncertainty intervals.
  • Adjusting for variables that are consequences of inequitable care.
  • Treating an outlier as proof of causation or poor performance.
  • Assuming best observed performance is feasible, desirable, or evidence-based.
  • Using activity as productivity without quality or outcome measures.
  • Combining indicators into a score without transparent weights and sensitivity analysis.
  • Copying a high performer's process without testing its mechanism and context.
  • Rewarding target attainment while ignoring gaming, burden, or unmeasured harm.

Reporting a benchmark transparently

A reproducible benchmark should allow users to understand exactly what was compared and how. Reports should present enough component detail to support diagnosis rather than only publishing a rank. They should also state the permitted use and avoid stronger causal claims than the design supports.

  • Define the purpose, unit of analysis, indicator, numerator, denominator, and period.
  • Describe peer selection, exclusions, data sources, and completeness.
  • Report raw and adjusted results with absolute differences, ratios, and uncertainty.
  • Document standardisation, risk adjustment, suppression, and missing-data methods.
  • Show trends and component or subgroup results where they affect interpretation.
  • Identify data revisions, limitations, and potential unintended incentives.
  • Record the agreed action, owner, timeline, balancing measures, and review date.

The decision standard

Good benchmarking identifies credible variation, explains how much confidence to place in it, and supports learning about what could improve. It does not treat the benchmark as an unquestionable target or the rank as the result. The process succeeds when comparable evidence leads to a tested change that improves outcomes, productivity, access, or equity without creating greater harm elsewhere.

Frequently Asked Questions (6)

  • What is benchmarking?

    The process of comparing an organisation's performance metrics against peer organisations or established standards, identifying strengths and improvement areas.

    Source: Donabedian A. Evaluating the quality of medical care. Milbank Memorial Fund Quarterly. 1966;44(3):166-206. doi:10.2307/3348969.

  • What does benchmarking compare performance against?

    Benchmarking is the process of comparing an organisation's performance against a reference. It compares the organisation's performance metrics against peer organisations or established standards, setting its results beside others. It identifies strengths and areas for improvement, showing where the organisation leads and where it lags. It sets performance against peers or standards, the yardsticks that give the comparison meaning. It is closely related to peer comparison, which likewise judges an organisation against others like it. Measuring performance against a reference is what it does. Donabedian (1966) set out this approach to quality.

    Source: Donabedian 1966

  • What does benchmarking compare?

    Benchmarking compares an organisation's performance metrics against peer organisations or established standards, so it sets the organisation's metrics against those of peers or against standards, identifying strengths and improvement areas. This comparison against peers or standards defines it. So benchmarking is the process of comparing an organisation's performance metrics against peer organisations or established standards, identifying strengths and improvement areas This identification of strengths and improvement areas is what benchmarking produces by comparing against peers or standards.

    Source: Donabedian 1966

  • What does benchmarking identify?

    Benchmarking identifies strengths and improvement areas, so by comparing an organisation's performance metrics against peers or established standards it finds where it does well and where it can improve. This identification of strengths and improvement areas defines its purpose. So benchmarking is the process of comparing an organisation's performance metrics against peer organisations or established standards, identifying strengths and improvement areas This setting against peers or standards is what benchmarking uses to evaluate an organisation's performance metrics.

    Source: Donabedian 1966

  • Against what does benchmarking set performance?

    Benchmarking sets performance against peer organisations or established standards, so an organisation's performance metrics are compared with those of peers or with standards, identifying strengths and improvement areas. This setting against peers or standards defines it. So benchmarking is the process of comparing an organisation's performance metrics against peer organisations or established standards, identifying strengths and improvement areas This setting against peers or standards is what benchmarking relies on to gauge where an organisation stands.

    Source: Donabedian 1966

  • How does benchmarking relate to peer comparison?

    Benchmarking relates to peer comparison as a broad process to a specific form: benchmarking is comparing an organisation's performance metrics against peer organisations or established standards, and peer comparison is an evaluation of an organisation's performance metrics relative to similar peer organisations. So peer comparison is benchmarking against peers, connected in that both compare an organisation's performance to identify strengths and improvement areas This relationship is what makes peer comparison benchmarking directed specifically at similar peer organisations.

    Source: Donabedian 1966

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 22 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HS-NHS-PM-003

Stable URI · Machine-readable · Resolvable · CC BY 4.0