VerifiedEvidence: highv1.0.0

Mean

A central tendency measure calculated as the sum of a set of values divided by the number of values.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept


Theoretically, the Mean is a measure of central tendency representing the arithmetic average of a set of numerical observations. It is defined as the total of all observed values divided by the number of observations and provides a measure of the central location of a distribution. The mean exists to summarise quantitative data using a single representative value and forms the foundation of numerous statistical methods used throughout health economics.

Mathematically, the mean is calculated by summing all observations and dividing by the sample size for sample data or by the population size for complete populations. It is an unbiased estimator of the population expectation under random sampling and possesses important optimality properties, including minimising the sum of squared deviations. The mean serves as the basis for calculating variance, standard deviation, covariance, regression coefficients and many other statistical quantities.

In practice, the mean is calculated for variables such as healthcare costs, quality-adjusted life years, utility scores, life expectancy and resource utilisation. In health economics it is routinely used to estimate average costs, average health outcomes, incremental differences between interventions and expected values in decision models. Because the mean is sensitive to extreme observations, analysts should assess the underlying distribution and consider complementary summary measures when data are highly skewed.

Purpose


Used to summarise quantitative data with a single measure of central tendency, estimate population expectations, compare groups and provide the basis for statistical inference and health economic evaluation.


Mathematical Formulae

Primary Formula

Sample mean:

x? = (1/n) ? ????� x?

Population mean:

? = (1/N) ? ????? x?

Supporting Formulae

Weighted mean:

x?w = ?w?x? / ?w?

Expected value:

E(X) = ?

Sum of squared deviations:

SS = ?(x? ? x?)�

Related Mathematical Methods

  • Weighted Mean
  • Geometric Mean
  • Harmonic Mean
  • Trimmed Mean
  • Variance
  • Standard Deviation
  • Standard Error of the Mean

Example

A health economist evaluates annual healthcare costs for five patients:

�2,100, �2,450, �2,700, �2,900 and �3,350.

The sample mean is:

x? = (2,100 + 2,450 + 2,700 + 2,900 + 3,350) � 5

x? = �2,700

This value represents the average annual healthcare cost per patient and may be compared with the corresponding mean cost in an alternative treatment group during an economic evaluation.


Excel Implementation

FunctionExample FormulaHealth Economics Application
AVERAGE=AVERAGE(B2:B501)Calculate the average healthcare cost, utility or outcome.
AVERAGEIF=AVERAGEIF(A2:A501,"Treatment",B2:B501)Calculate the mean for a treatment group only.
AVERAGEIFS=AVERAGEIFS(B2:B501,A2:A501,"Treatment",C2:C501,"Female")Calculate subgroup-specific means.
SUM=SUM(B2:B501)Calculate the total value before dividing by the sample size.
COUNT=COUNT(B2:B501)Determine the number of observations contributing to the mean.

VBA (Optional)

A VBA routine can automatically calculate means for multiple variables and treatment groups and generate summary tables for health economic analyses.


Sources

  • Casella G, Berger RL. Statistical Inference. Cengage Learning.
  • Mood AM, Graybill FA, Boes DC. Introduction to the Theory of Statistics. McGraw-Hill.
  • Rice JA. Mathematical Statistics and Data Analysis. Cengage Learning.
  • Agresti A. Statistical Methods for the Social Sciences. Pearson.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation. Oxford University Press.
  • Drummond MF, et al. Methods for the Economic Evaluation of Health Care Programmes. Oxford University Press.

Library

Publications

1
  • Book

    Statistical Analysis of Cost-Effectiveness Data — Willan & Briggs, 1st Edition ed., 2006 (John Wiley & Sons)

    A synthesis of statistical methods for analysing cost-effectiveness data, including net-benefit regression, confidence intervals for the ICER, cost-effectiveness acceptability curves, and covariate adjustment. Part of the Wiley Statistics in Practice series.

Frequently Asked Questions (6)

  • What is the mean?

    A central tendency measure calculated as the sum of a set of values divided by the number of values.

    Source: Casella G, Berger RL. Statistical Inference. 2nd ed. Duxbury; 2002.

  • What does the mean summarise about a set of values?

    The mean summarises a set of values by their arithmetic average, adding them all together and dividing by how many there are. It gives the balancing point of the data, using every value in its calculation, which makes it the natural summary for roughly symmetric distributions. Because it uses every value, it is pulled toward extreme observations, so a few very high or low figures can drag it away from what is typical. The average of all the values is what it captures. Kirkwood and Sterne (2003) describe this measure.

    Source: Kirkwood & Sterne 2003

  • How is the mean calculated?

    The mean is calculated by adding up all the values in a dataset and dividing the total by the number of values. So the mean is calculated as the sum of the observations divided by their count, giving the arithmetic average, which uses every value equally in its computation, and this is why the mean reflects the whole dataset but is also affected by any unusually large or small values, since these enter the sum and can pull the average toward them, in contrast to measures such as the median that depend only on the rank order.

    Source: Casella & Berger 2002

  • When is the mean an appropriate measure?

    The mean is an appropriate measure of central tendency when the data are reasonably symmetric and free of extreme outliers, since it then represents the typical value well and uses all the information in the data. It is less suitable for skewed data or data with outliers, which distort it. So the mean is appropriate for symmetric, outlier-free distributions, where it is an efficient and informative summary, but for skewed data or data with extreme values the median is often preferred, since the mean can be pulled away from the bulk of the data by the extremes, giving a less representative picture of the typical value.

    Source: Casella & Berger 2002

  • How does the mean differ from the median?

    The mean is the arithmetic average, using all the values and thus sensitive to extreme values, while the median is the middle value when the data are ordered, resistant to outliers and skew. In a symmetric distribution the two coincide, but in a skewed one they differ, with the mean pulled toward the longer tail. So the mean and median differ in how they summarise the centre and in their sensitivity to the data's shape, with the mean using all the values and the median only the rank order, which is why the median is preferred for skewed or outlier-prone data and the mean for symmetric data.

    Source: Casella & Berger 2002

  • What are the limitations of the mean?

    The limitations of the mean include its sensitivity to extreme values and skew, since outliers can pull it away from the bulk of the data, making it unrepresentative of the typical value in such cases; and that for highly skewed data it can mislead. So the mean is used with awareness that it can be distorted by outliers and skew, which is why it is reported alongside a measure of spread and, for skewed data, often accompanied or replaced by the median, since relying on the mean alone for asymmetric distributions can give a misleading impression of the centre, whereas for symmetric data it is an efficient and appropriate summary.

    Source: Casella & Berger 2002

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 18 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-111

Stable URI · Machine-readable · Resolvable · CC BY 4.0