VerifiedEvidence: highv1.0.0

Statistical Power

The probability that a statistical test will correctly detect a true effect of a specified size, calculated as one minus the beta level.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept


Theoretically, Statistical Power is the probability that a statistical test correctly rejects a false null hypothesis. It quantifies a study's ability to detect a true effect when one exists and is a fundamental concept in hypothesis testing, study design, and estimation. Statistical power depends on the true effect size, sample size, variability of the outcome, significance level, and the statistical test employed. High statistical power reduces the likelihood of Type II errors and increases confidence that meaningful effects will be detected.

Mathematically, statistical power is defined as one minus the probability of a Type II error (?). It is calculated from the distribution of the test statistic under the alternative hypothesis and therefore depends upon the degree of separation between the null and alternative distributions. Power analysis is commonly performed prospectively during sample size determination to ensure that studies are capable of detecting clinically or economically meaningful effects.

In practice, statistical power is estimated before a study begins using anticipated effect sizes, estimates of variability, significance levels, and desired precision. In health economics it is used to design clinical trials, observational studies, economic evaluations, cost-effectiveness analyses, and health technology assessments. Power calculations are routinely reported to justify sample sizes and demonstrate that studies are adequately designed to detect important differences.

Purpose


Used to quantify the probability of detecting true effects, determine appropriate sample sizes, minimise Type II errors, and support robust study design and statistical inference in health economic research.


Mathematical Formulae

Primary Formula

Power = 1 ? ?

where:

  • ? = probability of a Type II error

Supporting Formulae

For a two-sided z-test:

Power = P(|Z| > z??�?? | H?)

Sample size relationship for comparing two means:

n = 2��(z??�?? + z???)� � ?�

Standardised effect size:

d = ? � �

where:

  • ? = minimum detectable difference
  • � = standard deviation

Related Mathematical Methods

  • Sample Size Calculation
  • Type I Error
  • Type II Error
  • Significance Level
  • Effect Size
  • Hypothesis Testing
  • Confidence Interval
  • Neyman?Pearson Framework

Example

A health economic trial compares annual healthcare costs between two interventions.

Investigators specify:

  • Significance level: � = 0.05
  • Desired power: 90%
  • Standard deviation: �2,500
  • Minimum detectable cost difference: �750

Using the recognised sample size formula, the required sample size is calculated before recruitment. Selecting 90% power means that, if a true cost difference of �750 exists, the study has a 90% probability of detecting that difference while accepting a 10% probability of a Type II error.


Excel Implementation

FunctionExample FormulaHealth Economics Application
NORM.S.INV=NORM.S.INV(0.90)Calculate the z-value corresponding to 90% statistical power.
NORM.S.INV=NORM.S.INV(1-0.05/2)Obtain the critical value for a two-sided 5% significance level.
POWER=2*POWER(2500,2)*POWER(NORM.S.INV(1-0.05/2)+NORM.S.INV(0.90),2)/POWER(750,2)Calculate the required sample size for comparing two means.
ROUNDUP=ROUNDUP(B2,0)Round calculated sample sizes to the nearest whole participant.
IF=IF(CurrentSample>=RequiredSample,"Adequately Powered","Underpowered")Assess whether the planned study meets the required statistical power.

VBA (Optional)

Automate power analyses across multiple effect sizes, significance levels, and sample size scenarios for clinical and health economic study planning.


Sources

  • Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed.
  • Chow SC, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research.
  • Julious SA. Sample Sizes for Clinical Trials.
  • Machin D, Campbell MJ, Tan SB, Tan SH. Sample Size Tables for Clinical Studies.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.
  • Drummond MF, et al. Methods for the Economic Evaluation of Health Care Programmes.
  • NICE Health Technology Evaluation Manual.

Library

Publications

2
  • Book

    Bayesian Methods in Health Economics — Gianluca Baio, 1st Edition ed., 2012 (Chapman & Hall / CRC Press)

    An overview of Bayesian statistical methods for the analysis of health economic data, covering economic evaluation concepts, statistical cost-effectiveness analysis, Bayesian computation and MCMC, and applied health economic evaluation.

  • Journal articleFeatured

    Noninferiority Testing in Cost-Minimization Studies: Practical Issues Concerning Power Analysis — Mark M. Span, Elisabeth M. TenVergert, Christian S. van der Hilst and Ronald P. Stolk, 22(2):261–266 ed., 2006 (International Journal of Technology Assessment in Health Care)

    Examines non-inferiority testing and statistical power requirements needed to support the clinical-effect assumption underlying a valid CMA.

Frequently Asked Questions (6)

  • What is statistical power?

    The probability that a statistical test will correctly detect a true effect of a specified size, calculated as one minus the beta level.

    Source: Cohen 1988

  • What is the chance that statistical power describes?

    Statistical power is the probability that a test will detect a true effect of a given size when that effect really exists, calculated as one minus the risk of a false negative. It describes a study's ability to find what is there, so a study with low power may well miss a genuine benefit and wrongly conclude that a treatment does nothing. Power rises with the sample size, the size of the effect, and lower variability, which is why adequate numbers are planned in advance. The chance of catching a real effect is what it captures. Kirkwood and Sterne (2003) describe this.

    Source: Kirkwood & Sterne 2003

  • What factors affect statistical power?

    Statistical power is affected by the sample size, with larger samples giving greater power; the size of the effect to be detected, with larger effects easier to detect; the significance level, with a stricter level reducing power; and the variability of the outcome, with greater variability reducing power. So statistical power depends on the sample size, effect size, significance level, and variability, which is why these are the inputs to power and sample size calculations, and why studies are designed with enough participants, given the expected effect and variability, to achieve adequate power, since insufficient power risks missing a true effect.

    Source: Cohen 1988

  • Why is statistical power important?

    Statistical power is important because a study with low power is likely to miss a true effect, producing a false negative and possibly wrongly concluding there is no effect, which wastes resources and can mislead. So statistical power matters for a study's ability to answer its question, since adequate power is needed to detect a meaningful effect reliably, and underpowered studies are both scientifically and ethically problematic, which is why power is calculated in advance to ensure the sample size is sufficient, and why low power is a recognised threat to the reliability and interpretability of research findings.

    Source: Cohen 1988

  • How does statistical power relate to sample size?

    Statistical power relates to sample size in that increasing the sample size increases the power, since a larger sample gives more precise estimates and a better chance of detecting a true effect; conversely, achieving a desired power requires a sufficient sample size. So power and sample size are directly linked, which is why sample size calculations use the desired power, along with the effect size, significance level, and variability, to determine the number of participants needed, and why studies are sized to reach adequate power, commonly eighty or ninety per cent, ensuring they can detect the effect of interest.

    Source: Cohen 1988

  • What is considered adequate statistical power?

    Adequate statistical power is conventionally set at eighty per cent or higher, meaning at least an eighty per cent chance of detecting the specified effect if it exists, with ninety per cent used when a lower risk of missing an effect is wanted. So adequate power is commonly taken as eighty per cent, corresponding to a beta of twenty per cent, though higher power is sometimes required, and this convention guides the sample size chosen for a study, since designing for adequate power ensures a reasonable probability of detecting a meaningful effect, whereas power much below this leaves the study prone to false negative conclusions.

    Source: Cohen 1988

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 25 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-203

Stable URI · Machine-readable · Resolvable · CC BY 4.0