VerifiedEvidence: highv1.0.0

Statistical Significance

A judgement that an observed result is unlikely to have occurred by chance under the null hypothesis, based on a p-value below a threshold.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept


Theoretically, Statistical Significance is the conclusion that an observed result is sufficiently incompatible with the null hypothesis that it is unlikely to have arisen solely through random sampling variation, according to a pre-specified significance level. It is a decision concept within the Neyman?Pearson framework of hypothesis testing and provides evidence regarding whether an observed effect is distinguishable from chance. Statistical significance does not measure the magnitude, clinical importance, or economic relevance of an effect.

Mathematically, statistical significance is determined by comparing a calculated p-value with a pre-specified significance level (�), or equivalently by comparing an observed test statistic with its corresponding critical value. The decision depends on the sampling distribution of the test statistic under the null hypothesis and is therefore influenced by the effect size, sample size, outcome variability, and chosen significance level.

In practice, statistical significance is assessed following hypothesis testing in clinical trials, observational studies, economic evaluations, and regression analyses. In health economics it is commonly used when evaluating differences in healthcare costs, quality-adjusted life years, treatment effects, resource utilisation, and model parameters. Interpretation is typically accompanied by confidence intervals and effect estimates to distinguish statistical evidence from practical importance.

Purpose


Used to determine whether observed differences or associations are unlikely to have occurred by chance alone, supporting statistical inference, hypothesis testing, and interpretation of clinical and economic evidence.


Mathematical Formulae

Primary Formula

Statistically significant if:

p � �

where:

  • p = p-value
  • � = pre-specified significance level

Supporting Formulae

For a two-sided z-test:

Z = (Estimate ? Null Value) � SE

Reject H? if:

|Z| � z??�??

For a Wald test:

W = ((?? ? ??) � SE(??))�

Reject H? if:

W � ?�?,??�

Related Mathematical Methods

  • Hypothesis Testing
  • P-value
  • Significance Level
  • Type I Error
  • Type II Error
  • Confidence Interval
  • Wald Test
  • Likelihood Ratio Test
  • Statistical Power

Example

A health economic evaluation compares annual healthcare costs between two treatment groups.

Estimated mean cost difference:

�1,250

Standard error:

�500

The calculated test statistic is:

Z = 1,250 � 500

= 2.50

The corresponding two-sided p-value is:

p = 0.012

Using a significance level of � = 0.05:

0.012 � 0.05

The difference in annual healthcare costs is therefore statistically significant at the 5% level.


Excel Implementation

FunctionExample FormulaHealth Economics Application
IF=IF(B2<=0.05,"Statistically Significant","Not Statistically Significant")Determine statistical significance using the chosen significance level.
T.DIST.2T=T.DIST.2T(A2,98)Calculate the p-value from a t-statistic.
NORM.S.DIST=2*(1-NORM.S.DIST(ABS(A2),TRUE))Calculate the p-value from a z-statistic.
CHISQ.DIST.RT=CHISQ.DIST.RT(A2,1)Calculate the p-value for a chi-square statistic.
T.TEST=T.TEST(B2:B51,C2:C51,2,2)Test for statistically significant differences between treatment groups.

VBA (Optional)

Automate hypothesis-testing workflows and generate reports identifying statistically significant results across multiple health economic outcomes.


Sources

  • Neyman J, Pearson ES. On the Problem of the Most Efficient Tests of Statistical Hypotheses. Philosophical Transactions of the Royal Society A. 1933.
  • Fisher RA. Statistical Methods for Research Workers.
  • Altman DG. Practical Statistics for Medical Research.
  • Casella G, Berger RL. Statistical Inference.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.
  • Drummond MF, et al. Methods for the Economic Evaluation of Health Care Programmes.
  • NICE Health Technology Evaluation Manual.

Library

Publications

1
  • Book

    Statistical Analysis of Cost-Effectiveness Data — Willan & Briggs, 1st Edition ed., 2006 (John Wiley & Sons)

    A synthesis of statistical methods for analysing cost-effectiveness data, including net-benefit regression, confidence intervals for the ICER, cost-effectiveness acceptability curves, and covariate adjustment. Part of the Wiley Statistics in Practice series.

Frequently Asked Questions (6)

  • What is statistical significance?

    A judgement that an observed result is unlikely to have occurred by chance under the null hypothesis, based on a p-value below a threshold.

    Source: Fisher 1925

  • What does statistical significance actually establish about a result?

    Statistical significance establishes only that an observed result would be unlikely to arise by chance if the null hypothesis of no effect were true, judged by a p-value falling below a chosen threshold. It says the result is hard to dismiss as a fluke, but it does not say the effect is large or clinically important, since a huge study can render a trivial difference significant. Nor does it prove the effect is real. Unlikeliness under chance, rather than importance, is what it captures. Kirkwood and Sterne (2003) describe this.

    Source: Kirkwood & Sterne 2003

  • How is statistical significance determined?

    Statistical significance is determined by computing the p-value, the probability of a result as extreme as the one observed if the null hypothesis were true, and comparing it with the significance level: if the p-value is below the threshold, the result is declared statistically significant. So statistical significance is determined by whether the p-value falls below the chosen significance level, which provides the decision rule, and this is equivalent to the test statistic falling in the rejection region, meaning a significant result is one where the data are sufficiently improbable under the null hypothesis to reject it at the set level.

    Source: Fisher 1925

  • Does statistical significance mean an effect is important?

    Statistical significance does not mean an effect is important or large, since it depends on the sample size as well as the effect size, so a large study can find a trivial effect statistically significant, and a small study can miss a meaningful one. So statistical significance is distinct from practical importance, which is why a significant result should not be assumed to be clinically or practically meaningful, and effect sizes and confidence intervals are considered alongside significance, since statistical significance indicates only that an effect is unlikely to be due to chance, not that it is large enough to matter.

    Source: Fisher 1925

  • What are the limitations of statistical significance?

    The limitations of statistical significance include its dependence on sample size, so significance does not indicate importance; the arbitrariness of the threshold, creating a false dichotomy between significant and non-significant; frequent misinterpretation, such as treating a non-significant result as proof of no effect; and the neglect of effect size. So statistical significance is used with awareness of these limitations, which is why it is interpreted alongside effect sizes and confidence intervals rather than as a standalone verdict, since over-reliance on significance can mislead, and a fuller account of results considers the magnitude and precision of effects, not merely whether they cross the significance threshold.

    Source: Fisher 1925

  • How should statistical significance be interpreted?

    Statistical significance should be interpreted as indicating that an observed result is unlikely under the null hypothesis, providing evidence against it, but not as a measure of the size or importance of an effect or the probability that the null is true. So statistical significance should be interpreted together with effect sizes and confidence intervals, which convey the magnitude and precision, since significance alone tells only whether an effect is detectable, not whether it matters, and interpreting it in isolation, or treating the threshold as a sharp divide, can mislead, which is why a careful interpretation considers the evidence, the effect size, and the context together.

    Source: Fisher 1925

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 25 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-204

Stable URI · Machine-readable · Resolvable · CC BY 4.0