VerifiedEvidence: highv1.0.0

P-Score

A frequentist network meta-analysis measure ranking treatments by the extent one is expected to outperform every other treatment in the network.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, P-Score is a frequentist treatment ranking measure used in network meta-analysis to quantify the certainty that one intervention is better than competing interventions based on estimated treatment effects and their precision. It provides a numerical ranking of treatments without requiring Bayesian posterior probabilities and is conceptually related to the Surface Under the Cumulative Ranking Curve (SUCRA). The method exists to summarise comparative treatment performance across an entire treatment network.

Mathematically, the P-score is derived from the estimated pairwise treatment effects and their variances under the frequentist network meta-analysis framework. It is calculated from the mean probabilities that one treatment is superior to all competing treatments, assuming normally distributed treatment-effect estimates. P-scores range from 0 to 1, with higher values indicating greater certainty that a treatment performs better than competing interventions.

In practice, P-scores are calculated following a frequentist network meta-analysis and are presented alongside estimated treatment effects and confidence intervals. They are widely used in comparative effectiveness research and health technology assessment to summarise treatment rankings, although reimbursement decisions are based primarily on comparative treatment effects, uncertainty and cost-effectiveness rather than rankings alone.


Purpose

Used to rank competing interventions within a frequentist network meta-analysis by quantifying the certainty that each treatment is superior to alternative interventions.


Mathematical Formulae

Primary Formula

P-score:

P? = (1 / (m ? 1)) ? ???? �((d??? / SE(d???)))

where:

  • P? = P-score for treatment k
  • m = number of treatments
  • � = cumulative standard normal distribution
  • d??? = estimated treatment effect comparing treatment k with treatment j
  • SE(d???) = standard error of the treatment effect

Supporting Formulae

Standardised treatment comparison:

Z?? = d??? / SE(d???)

Probability of superiority:

P = �(Z??)

Related Mathematical Methods

  • Frequentist Network Meta-Analysis
  • Network Meta-Analysis
  • SUCRA
  • Indirect Comparison
  • Mixed Treatment Comparison
  • Random-Effects Meta-Analysis
  • Fixed Effect Meta-Analysis

Example

A frequentist network meta-analysis compares seven treatments for chronic migraine prevention. Treatment A achieves a P-score of 0.94, Treatment B a P-score of 0.82 and Treatment C a P-score of 0.37. The results indicate that Treatment A has the greatest certainty of being among the most effective interventions within the treatment network, although comparative effect estimates and confidence intervals remain the primary basis for decision-making.


Excel Implementation

FunctionExample FormulaHealth Economics Application
NORM.S.DIST=NORM.S.DIST(A2/B2,TRUE)Calculate the probability that one treatment is superior to another
AVERAGE=AVERAGE(C2:C8)Calculate the average probability across all treatment comparisons
RANK.AVG=RANK.AVG(D2,$D$2:$D$8,0)Rank treatments according to their P-scores
SORT=SORT(A2:B8,2,-1)Order treatments from highest to lowest P-score

VBA (Optional)

Automate calculation of treatment rankings and P-scores from frequentist network meta-analysis output and generate ranking tables for health technology assessment reports.


Sources

  • R�cker G, Schwarzer G. Ranking Treatments in Frequentist Network Meta-Analysis Works Without Resampling Methods. BMC Medical Research Methodology. 2015.
  • Chaimani A, Caldwell DM, Li T, et al. Chapter 11: Undertaking Network Meta-Analyses. Cochrane Handbook for Systematic Reviews of Interventions.
  • Schwarzer G, Carpenter JR, R�cker G. Meta-Analysis with R.
  • NICE. Health Technology Evaluation Manual.
  • ISPOR Good Practice Task Force Report on Network Meta-Analysis.

Library

Publications

1
  • BookFeatured

    Network Meta-Analysis for Decision Making — Dias, Ades, Welton, Jansen & Sutton, 1st Edition ed., 2018 (John Wiley & Sons)

    The definitive text on network meta-analysis (mixed treatment comparisons) for decision making, presenting a coherent Bayesian framework (implemented in WinBUGS) for synthesising evidence across multiple treatments, including inconsistency, bias adjustment, and use in cost-effectiveness models.

Frequently Asked Questions (6)

  • What is the P-score?

    A frequentist network meta-analysis measure ranking treatments by the extent one is expected to outperform every other treatment in the network.

    Source: Rücker & Schwarzer 2015

  • What does the P-score tell us about a treatment's ranking?

    The P-score is a frequentist measure that ranks the treatments in a network meta-analysis by the degree to which each is expected to outperform the others, summarising a treatment's relative standing in a single number between zero and one. A score near one means a treatment is likely to be among the best in the network, and one near zero among the worst, with the uncertainty of the estimates built into the calculation. It offers a compact way to order many treatments. Placing each treatment in the ranking is its purpose. Rücker and Schwarzer (2015) describe this measure.

    Source: Rücker & Schwarzer 2015

  • How is the P-score calculated?

    The P-score is calculated from the relative effect estimates and their standard errors in a frequentist network meta-analysis, quantifying for each treatment the mean extent of certainty that it is better than the other treatments, averaged over all comparisons. It is derived analytically from the estimates without needing resampling or simulation, distinguishing it from the Bayesian SUCRA, which is computed from posterior samples. The result is a value between zero and one for each treatment. So the P-score is calculated directly from the network estimates and their uncertainty, summarising each treatment's relative ranking across the network.

    Source: Rücker & Schwarzer 2015

  • What does the P-score indicate?

    The P-score indicates each treatment's relative ranking in the network, with a higher value, closer to one, meaning the treatment tends to outperform the others, and a lower value, closer to zero, meaning it tends to be outperformed. It summarises the treatment's performance across all pairwise comparisons into a single ranking measure. So the P-score indicates how well each treatment ranks relative to the others in the network, providing a concise summary that helps identify which treatments appear most and least effective, while reflecting the uncertainty in the estimates through its construction.

    Source: Rücker & Schwarzer 2015

  • How does the P-score relate to SUCRA?

    The P-score relates to SUCRA as its frequentist counterpart: both summarise a treatment's ranking in a network meta-analysis into a single value between zero and one, with higher values indicating better relative performance, and the two give very similar results. SUCRA is computed from the posterior samples of a Bayesian network meta-analysis, while the P-score is derived analytically from frequentist estimates without simulation. So the P-score and SUCRA are analogous ranking measures, one for frequentist and one for Bayesian network meta-analysis, providing equivalent summaries of how treatments rank, with the P-score offering the advantage of not requiring resampling.

    Source: Salanti, Ades & Ioannidis 2011

  • What are the limitations of the P-score?

    The limitations of the P-score, shared with other ranking measures, include that it summarises ranking without conveying the magnitude of the differences between treatments, so treatments with similar effects can differ in rank while being clinically similar; that rankings can be unstable when estimates are imprecise or the network is sparse; and that a high rank does not by itself establish clinical importance or certainty. So the P-score is interpreted alongside the relative effect estimates, their uncertainty, and the quality of the evidence, since rankings alone can mislead if taken without regard to the size of the differences and the reliability of the network.

    Source: Rücker & Schwarzer 2015

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 3 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-ESM-046

Stable URI · Machine-readable · Resolvable · CC BY 4.0