VerifiedEvidence: highv1.0.0

Correlation Coefficient

A measure quantifying the strength and direction of the linear relationship between two variables, ranging from negative one to positive one.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, the Correlation Coefficient is a standardised statistical measure that quantifies the strength and direction of association between two variables. The most widely used form is the Pearson product-moment correlation coefficient, which measures the degree of linear relationship between continuous variables. In health economics, correlation coefficients are used to assess relationships among healthcare costs, health outcomes, utilities, clinical measures and model parameters, and to evaluate assumptions underlying statistical and economic models.

Mathematically, the Pearson correlation coefficient is defined as the covariance between two variables divided by the product of their standard deviations. The coefficient ranges from ?1 to +1, where +1 indicates a perfect positive linear relationship, ?1 indicates a perfect negative linear relationship and 0 indicates no linear association. Because the statistic is standardised, it is independent of the measurement units of the variables being compared.

In practice, the correlation coefficient is estimated from paired observations and is routinely reported in descriptive analyses, regression diagnostics, psychometric validation, meta-analysis and probabilistic sensitivity analysis. It is also used to construct covariance matrices for multivariate modelling and to preserve parameter dependence in probabilistic health economic models.


Purpose

Used to quantify the strength and direction of association between variables, assess linear relationships, support regression modelling, construct covariance matrices and preserve parameter dependence in health economic analyses.


Mathematical Formulae

Primary Formula

? = Cov(X,Y) / (�?�?)

Sample estimate:

r = ?((x? ? x?)(y? ? ?)) / �(?(x? ? x?)� ? ?(y? ? ?)�)

Supporting Formulae

Cov(X,Y) = (1 / (n ? 1)) ? ?((x? ? x?)(y? ? ?))

?1 � r � 1

r� = Coefficient of Determination (simple linear regression)

Related Mathematical Methods

Covariance

Linear Regression

Pearson Correlation

Spearman Rank Correlation

Variance-Covariance Matrix

Principal Component Analysis

Factor Analysis


Example

A health economist investigates the relationship between annual healthcare expenditure and EQ-5D utility scores for 500 patients.

The estimated Pearson correlation coefficient is:

r = ?0.62

This indicates a moderately strong negative linear relationship, suggesting that higher healthcare expenditure is associated with lower health-related quality of life in the study population.


Excel Implementation

FunctionExample FormulaHealth Economics Application
CORREL=CORREL(B2:B501,C2:C501)Calculate Pearson correlation coefficient
COVARIANCE.S=COVARIANCE.S(B2:B501,C2:C501)Calculate sample covariance
STDEV.S=STDEV.S(B2:B501)Estimate standard deviation of first variable
STDEV.S=STDEV.S(C2:C501)Estimate standard deviation of second variable
RSQ=RSQ(B2:B501,C2:C501)Calculate coefficient of determination (R�)

VBA (Optional)

Automate calculation of correlation matrices for large health economic datasets and generate summary reports for multivariate analyses and probabilistic sensitivity models.


Sources

Pearson K. Notes on Regression and Inheritance in the Case of Two Parents. Proceedings of the Royal Society of London. 1895.

Kutner MH, Nachtsheim CJ, Neter J, Li W. Applied Linear Statistical Models.

Casella G, Berger RL. Statistical Inference.

Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.

Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes.

Library

Publications

1
  • Book

    Statistical Analysis of Cost-Effectiveness Data — Willan & Briggs, 1st Edition ed., 2006 (John Wiley & Sons)

    A synthesis of statistical methods for analysing cost-effectiveness data, including net-benefit regression, confidence intervals for the ICER, cost-effectiveness acceptability curves, and covariate adjustment. Part of the Wiley Statistics in Practice series.

Frequently Asked Questions (6)

  • What is a correlation coefficient?

    A measure quantifying the strength and direction of the linear relationship between two variables, ranging from negative one to positive one.

    Source: Pearson 1896

  • What does a correlation coefficient quantify between two variables?

    A correlation coefficient quantifies the strength and direction of the linear relationship between two variables in a single number between negative one and positive one. A value near positive one means the two rise together closely, near negative one that one falls as the other rises, and near zero that no linear relationship is apparent. It captures how tightly the points cluster around a straight line, though a strong correlation does not by itself establish that one variable causes the other. Summarising a linear association is what it does. Kirkwood and Sterne (2003) describe this measure.

    Source: Kirkwood & Sterne 2003

  • How is a correlation coefficient interpreted?

    A correlation coefficient is interpreted by its sign and magnitude: a positive value means the variables tend to increase together, a negative value means one tends to decrease as the other increases, and the magnitude, from zero to one in absolute terms, indicates the strength of the linear relationship, with values near one being strong and near zero weak. So a correlation coefficient is interpreted as showing both the direction and the closeness of a linear association, though it does not imply causation and captures only linear relationships, so a value near zero indicates weak linear association but does not rule out a strong non-linear relationship between the variables.

    Source: Pearson 1896

  • What are the limitations of a correlation coefficient?

    The limitations of a correlation coefficient include that it measures only linear relationships, so it can miss strong non-linear associations; that it is sensitive to outliers, which can inflate or deflate it; that it does not imply causation, since correlated variables may both depend on a third factor; and that a correlation can be spurious. So a correlation coefficient is interpreted with awareness that it captures linear association only, can be distorted by outliers, and does not establish cause, which is why it is examined alongside a scatterplot of the data and other evidence, since a single coefficient can mislead about the true relationship between the variables.

    Source: Pearson 1896

  • How does correlation differ from causation?

    Correlation differs from causation in that a correlation coefficient measures the association between two variables without establishing that one causes the other; the variables may be associated because one causes the other, because both are influenced by a common factor, or by chance. So correlation does not imply causation, since an observed association can arise from confounding or coincidence rather than a causal link, which is why causal conclusions require more than correlation, such as evidence from experiments or careful causal analysis, and treating a correlation as proof of cause is a common and serious error in interpreting data.

    Source: Rothman, Greenland & Lash 2008

  • What is the difference between Pearson and Spearman correlation?

    The difference between Pearson and Spearman correlation is that Pearson's coefficient measures the strength of a linear relationship between two quantitative variables, while Spearman's coefficient measures the strength of a monotonic relationship based on the ranks of the values, making it suitable for ordinal data or when the relationship is monotonic but not linear and more robust to outliers. So Pearson and Spearman correlation differ in what they capture and their assumptions, with Pearson suited to linear relationships in quantitative data and Spearman to monotonic relationships and ranked or non-normal data, and the choice depends on the data and the nature of the relationship being assessed.

    Source: Pearson 1896

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 12 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-038

Stable URI · Machine-readable · Resolvable · CC BY 4.0