Topic
Statistical methods
Health economic analyses draw on standard statistical methods to estimate effects and describe their precision. Concepts range from hypothesis testing and regression to multiple imputation for missing data and mixed effects models for repeated measurements. Bayesian analysis, with its prior and posterior distributions, appears alongside frequentist approaches.
Concepts in this topic
- Accelerated Failure Time ModelAn accelerated failure time (AFT) model is a survival model in which treatment or another factor stretches or shrinks time to an event by a time ratio.
- Alpha SpendingAlpha spending is a group sequential trial method that spreads the type I error rate (alpha) across interim and final analyses as information accrues.
- Area Under CurveArea under the curve (AUC), or C-statistic, is the probability a test or risk model scores a random case above a random non-case (0.5 chance, 1 perfect).
- Asymptotic NormalityAsymptotic normality means an estimator's sampling distribution is approximately normal in large samples, the basis of Wald confidence intervals.
- Autoregressive CorrelationAutoregressive correlation is a structure for repeated measures in which correlation falls by a constant factor with each time step between measurements.
- Bayesian AnalysisBayesian analysis updates a prior distribution with data via Bayes' theorem to give a posterior, used in HTA for evidence synthesis and decision models.
- Beta LevelThe predetermined acceptable probability of failing to detect a true effect, with one minus beta defining a study's statistical power.
- Between-Subject VariationBetween-subject variation is how much costs or outcomes differ from one patient to another, as opposed to variation in one person's repeated measurements.
- Bias-Variance TradeoffThe bias-variance tradeoff is how cutting systematic error (bias) tends to raise sampling variance, as in choosing HTA survival curves or subgroup inputs.
- Binomial DistributionThe binomial distribution models how many of a fixed number of patients have an event when each has the same probability and outcomes are independent.
- Bootstrap Standard ErrorBootstrap standard error is the standard deviation of an estimate, such as mean trial cost, recalculated across many resamples of the observed data.
- BootstrappingBootstrapping is a resampling method that estimates sampling uncertainty in trial costs, QALYs and net benefit from patient samples drawn with replacement.
- Brier ScoreThe Brier score is the mean squared difference between predicted risks and observed binary outcomes, used to judge prediction models; lower is better.
- Central Limit TheoremThe central limit theorem states that, under suitable independence and finite-variance conditions, a properly standardised sample mean approaches a normal distribution as the number of observations increases.
- Chi-Square DistributionThe chi-square distribution describes a sum of squared independent standard normals and is the reference for likelihood ratio, heterogeneity and fit tests.
- Classification AccuracyClassification accuracy is the proportion of cases a test, coder or prediction rule assigns to the correct category, judged against a reference standard.
- Cluster AnalysisCluster analysis is an unsupervised method that sorts patients, providers or countries into groups whose members are more alike than those in other groups.
- Cohen's dCohen's d is a standardised effect size: the difference between two group means divided by their pooled standard deviation, used to compare studies.
- CommunalityCommunality is the portion of an observed variable's variance attributed to the common factors retained in a factor-analysis model.
- Condition IndexA condition index is the ratio of the largest singular value of a scaled regression design matrix to another singular value, used to assess near-linear dependence among predictors.
- Confidence IntervalA confidence interval gives a range of parameter values that are compatible with the observed data under a specified statistical model and confidence procedure.
- Confidence Interval EstimationThe statistical process of calculating a range of plausible values for an unknown parameter based on sample data and a specified confidence level.
- Confirmatory Factor AnalysisConfirmatory factor analysis tests a prespecified measurement model in which observed variables are linked to hypothesised latent factors and the implied relationships are compared with data.
- Cook's DistanceCook's distance measures how much an ordinary-least-squares regression's fitted values change when one observation is removed, scaled by residual variation and model size.
- Correlation CoefficientA correlation coefficient is a dimensionless statistic that quantifies the direction and strength of a specified form of association between two paired variables.
- Credible IntervalA credible interval is a set of parameter values containing a stated share of the posterior probability under a specified Bayesian model and observed data.
- Credible Interval EstimationThe Bayesian process of calculating a range of plausible values for a parameter based on its estimated posterior probability distribution.
- Critical ValueA critical value is a cutoff on a test statistic's reference distribution that defines a rejection boundary for a specified hypothesis test and error level.
- DecileA descriptive statistic dividing a ranked dataset into ten equal-sized groups, the first decile being the bottom ten percent and the tenth the top.
- Delta AdjustmentA sensitivity analysis technique shifting the assumed outcome for patients with missing trial data by a specified amount to test conclusions' robustness.
- Delta MethodThe delta method approximates the variance of a smooth function of estimated parameters by propagating their covariance through the function’s local derivatives.
- Design EffectThe design effect is the ratio of an estimator's variance under an actual sampling design to its variance under a comparable simple random sample of the same nominal size.
- DFBETASA standardised version of the DFBETA diagnostic, scaled by its standard error, allowing influence of observations to be compared consistently across models.
- Diagnostic AccuracyDiagnostic accuracy is how well a test agrees with a reference standard in detecting a target condition, usually reported as sensitivity and specificity.
- Effect SizeAn effect size is a measure of the magnitude of an association or intervention effect for a specified outcome, comparator, population, and scale.
- Effect Size EstimationThe statistical process of calculating a quantitative measure of a treatment effect's magnitude, often standardised to allow comparison across differing measurement scales.
- Effective Sample SizeThe sample size a simple random sample would need to match the precision of an actual sample from a more complex design, such as clustered.
- EigenvalueAn eigenvalue describes how a matrix scales a particular direction without changing that direction.
- Exchangeable CorrelationA correlation structure for repeated measures assuming any two observations within the same cluster are equally correlated, regardless of their timing or order.
- Exploratory Factor AnalysisExploratory factor analysis is a latent-variable method that examines whether correlations among observed variables can be represented by a smaller, initially unspecified set of common factors.
- F-DistributionA continuous probability distribution arising from the ratio of two independent chi-square variables divided by their degrees of freedom, underlying analysis of variance.
- Factor AnalysisFactor analysis explains correlations among observed variables using a smaller number of unobserved latent factors.
- Factor LoadingA factor loading is a model-estimated coefficient that relates an observed variable to a latent factor under a specified factor-analytic measurement model.
- Factor RotationFactor rotation transforms an extracted factor solution to make its loading pattern easier to interpret while preserving the model-implied common covariance under an equivalent transformation.
- Fieller MethodA technique for calculating a confidence interval around the ratio of two normally distributed variables, useful for measures such as the ICER.
- Finite Mixture ModelA statistical model representing a population as a combination of a finite number of distinct subgroups, each following its own probability distribution.
- Fixed EffectsA panel data modelling approach estimating a separate intercept for each individual or group, controlling for all its time-invariant characteristics.
- Frequentist AnalysisA statistical inference approach based on the long-run frequency properties of estimators, calculated from observed data alone, unlike a Bayesian approach.
- Generalized Estimating EquationsGeneralized estimating equations are a regression approach for estimating population-averaged associations from correlated observations by specifying a marginal mean model and accounting for within-cluster dependence.
- Generalized Linear ModelA modelling framework extending linear regression to non-normal outcomes, such as binary or count data, via a specified link function.
- Group-Based Trajectory ModelA technique identifying distinct subgroups following similar patterns of change over time, without assuming everyone follows the same single trajectory.
- Growth Curve ModelA growth curve model estimates how a repeatedly measured outcome changes over time, allowing the average trajectory and individual trajectories to differ under specified assumptions.
- Growth Mixture ModelAn extension of growth curve modelling allowing unobserved subgroups to follow qualitatively different average trajectories, rather than varying only in degree.
- Harrell C-IndexA survival model discrimination measure generalising the ROC area under curve to time-to-event data, the probability the model ranks patients correctly.
- Hedges' GA standardised effect size measure similar to Cohen's d but corrected for small sample size bias.
- Hierarchical ModelA hierarchical model represents observations or parameters at multiple related levels and estimates variation within and between those levels.
- Hierarchical TestingA multiple testing strategy organising hypotheses into a prespecified sequence, so later tests are only conducted if earlier, higher-priority tests succeed.
- Hosmer-Lemeshow TestA goodness of fit test assessing how well a logistic regression model's predicted probabilities match observed outcome frequencies across risk groups.
- Hypothesis TestingHypothesis testing is a statistical decision framework that evaluates how compatible observed data are with a prespecified null hypothesis by comparing a test statistic with its sampling distribution under stated assumptions.
- ImputationA technique replacing missing data values with estimated substitutes, ranging from simple mean substitution to sophisticated multiple imputation approaches.
- Interquartile RangeA dispersion measure calculated as the difference between the seventy-fifth and twenty-fifth percentiles, the range containing the middle half of observations.
- Intraclass CorrelationA measure quantifying the proportion of total variance in an outcome attributable to differences between groups or clusters, relative to variance within them.
- JackknifeA resampling technique estimating a statistic's variability by systematically recalculating it after removing one observation at a time.
- Jackknife Standard ErrorAn estimate of a parameter's statistical uncertainty calculated using the jackknife technique of omitting one observation at a time and examining resulting variability.
- K-Fold Cross-ValidationAn internal validation procedure that repeatedly fits a model on all but one of K data folds and evaluates it on the held-out fold.
- Kaiser CriterionA factor analysis rule retaining only factors with an eigenvalue greater than one, since these explain more variance than a single original variable.
- Kendall TauA non-parametric measure of association between two ranked variables, based on the number of concordant and discordant pairs of observations.
- KurtosisA measure describing a distribution's tail shape relative to a normal distribution, higher values indicating heavier tails and more extreme outliers.
- Last Observation Carried ForwardLast observation carried forward (LOCF) fills missing trial follow-up values with each participant's last observed value, assuming no change after dropout.
- Latent Class AnalysisA technique identifying unobserved, categorical subgroups within a population from patterns of responses, assuming these fully explain the observed associations.
- Latent Growth ModelA modelling framework representing an individual's trajectory using unobserved latent variables for starting level and rate of change, both varying across people.
- Latent Transition AnalysisA technique modelling how individuals move between unobserved categorical subgroups over time, combining latent class analysis with a longitudinal framework.
- Latent VariableA latent variable is an unobserved quantity, such as health-related quality of life or preference class membership, estimated from observed indicators.
- Law of Large NumbersA theorem stating that as sample size increases, the sample average of a variable converges toward the true underlying population average.
- Leave-One-OutA cross-validation form repeatedly refitting a model using all but one observation, using that observation to test accuracy, until each has served once.
- Leverage PointAn observation with an unusual combination of predictor values, distinct from the bulk of the data, capable of disproportionately influencing a regression model.
- Linear RegressionA modelling technique estimating the linear relationship between a continuous outcome and one or more predictors, giving expected change per unit predictor.
- Listwise DeletionA missing data method excluding any observation with a missing value on any variable in an analysis, using only complete cases.
- Logistic RegressionA modelling technique estimating the relationship between predictor variables and a binary outcome, expressing results as odds ratios.
- Maximum Likelihood EstimationA method estimating a model's parameters by finding the values that make the observed data most probable.
- MeanA measure of central tendency obtained, in its arithmetic form, by dividing the sum of values by their count; a population mean is the corresponding expected value.
- Mean Absolute ErrorThe arithmetic average of the absolute differences between numerical predictions and their corresponding observed values, expressed in the outcome’s units.
- Mean Squared ErrorMean squared error is the average of squared differences between predictions or estimates and their corresponding observed or reference values, evaluated over a specified set or sampling process.
- MedianA central tendency measure representing the middle value of an ordered dataset, dividing it into two equal halves.
- Mediation AnalysisA statistical approach examining whether an exposure's effect on an outcome operates through an intermediate variable, called a mediator, rather than directly.
- Method of MomentsA technique estimating a distribution's parameters by equating sample moments, such as mean and variance, to their theoretical population equivalents.
- Missing at RandomMissing at random is a missing-data mechanism in which, conditional on observed data, the probability of a missingness pattern does not depend on the values that remain unobserved.
- Missing Completely at RandomA missing data classification indicating the probability of missingness is entirely unrelated to any observed or unobserved variables, the strongest assumption.
- Missing DataUnobserved or unavailable values required for a specified analysis, whose consequences depend on the pattern, mechanism and target estimand.
- Missing Not at RandomMissing not at random describes a missing-data mechanism in which the probability that a value is unobserved depends on that unobserved value even after conditioning on the available observed information.
- Mixed Effects ModelA mixed effects model combines population-level fixed coefficients with probability-modelled variation across groups or individuals to analyse outcomes with clustered or repeated observations.
- Mixed Model Repeated MeasuresA longitudinal trial analysis method using a mixed effects framework under the missing at random assumption, an alternative to last observation carried forward.
- ModeA central tendency measure representing the most frequently occurring value within a dataset.
- Moderation AnalysisA statistical approach examining whether an exposure-outcome relationship's strength or direction differs depending on the level of a third, moderating variable.
- MulticollinearityA condition in which predictors in a model are linearly dependent or nearly so, making their separate coefficients difficult or impossible to estimate precisely.
- Multiple ImputationMultiple imputation is a missing-data method that creates several plausible completed datasets, analyses each for the same target, and pools the estimates and their uncertainty to reflect uncertainty about unobserved values.
- Negative BinomialA discrete probability distribution modelling count data with more variability than a Poisson distribution would predict, known as overdispersion.
- Negative Binomial RegressionA modelling technique analysing count outcome data, such as hospitalisations, that shows overdispersion relative to what a Poisson model would assume.
- Null HypothesisIn hypothesis testing, the default assumption of no true effect or difference between the groups or conditions being compared.
- Oblique RotationA factor analysis technique transforming a solution into a more interpretable structure while allowing the resulting factors to correlate with each other.
- One-Tailed TestA hypothesis test assessing evidence for an effect in only one specified direction, requiring strong prior justification for that direction.
- Orthogonal RotationA factor analysis technique transforming a solution into a more interpretable structure while constraining the resulting factors to remain uncorrelated.
- OutlierAn observation differing markedly from other values in a dataset, either from genuine unusual variation or a measurement or entry error.
- OverfittingA situation in which a model is fit too closely to the noise in its training data, performing poorly on new, independent data.
- P-ValueA tail probability, calculated under a specified null hypothesis and statistical model, of obtaining a test statistic at least as extreme as the observed statistic according to the chosen test.
- Pairwise DeletionA missing data method in which each specific calculation uses all observations with complete data for that particular pair of variables.
- Partial CorrelationA measure of the relationship between two variables after statistically controlling for the influence of one or more additional variables.
- Pattern Mixture ModelA missing data modelling approach stratifying analysis by the observed pattern of missingness, rather than assuming one model applies to everyone.
- Pearson CorrelationA measure quantifying the strength and direction of the linear relationship between two continuous variables, ranging from negative one to positive one.
- PercentileA descriptive statistic indicating the value below which a specified percentage of observations in a dataset fall.
- Poisson DistributionA discrete probability distribution describing the number of times an event occurs within a fixed interval, assuming independence and a constant average rate.
- Poisson RegressionA modelling technique analysing count outcome data, such as clinical events per patient, based on the assumption that outcomes follow a Poisson distribution.
- Posterior DistributionIn Bayesian statistics, the updated probability distribution for a parameter after combining a prior distribution with the likelihood of observed data.
- Prediction ErrorThe difference between a value predicted by a statistical model and the value that is actually observed.
- Prediction IntervalA range calculated to contain a future individual observation with a specified confidence, unlike a confidence interval describing a population parameter.
- Principal ComponentA principal component is an uncorrelated linear combination of centered variables whose loading vector identifies a successive direction of maximum variance in the data.
- Prior DistributionIn Bayesian statistics, the probability distribution representing existing beliefs about a parameter before observing new data.
- QuartileA descriptive statistic dividing a ranked dataset into four equal-sized groups, marking the twenty-fifth, fiftieth, and seventy-fifth percentiles.
- RangeA dispersion measure calculated as the difference between the maximum and minimum values in a dataset, sensitive to extreme outliers.
- Regression AnalysisRegression analysis estimates how an outcome varies with one or more explanatory variables through a specified statistical model.
- Repeated Measures AnalysisA statistical approach analysing data in which the same outcome has been measured multiple times on the same individuals, accounting for correlation.
- ResidualThe difference between an observed value and the value predicted for it by a fitted statistical model.
- Robust Standard ErrorAn estimate of a regression coefficient's statistical uncertainty that remains valid even when standard modelling assumptions, such as constant error variance, are violated.
- ROC CurveA graph depicting the trade-off between a test's true positive rate and false positive rate across a full range of decision thresholds.
- Root Mean Square ErrorThe square root of the mean squared difference between numerical predictions and their corresponding observed values, expressed in the outcome’s original units.
- Sample Size CalculationThe process of determining the number of participants a study needs to have an adequate chance of detecting a specified treatment effect.
- Sandwich EstimatorA method calculating an estimator's variance that remains valid even when standard model assumptions are violated, forming the basis of robust standard errors.
- Selection ModelA missing data modelling approach explicitly modelling the process by which observations become missing, alongside the model for the outcome itself.
- Significance LevelThe significance level is a prespecified upper bound on the probability that a valid statistical test rejects its null hypothesis when that hypothesis is true.
- SkewnessA measure describing a distribution's asymmetry around its mean, positive skewness meaning a longer tail toward higher values.
- Spearman CorrelationA non-parametric measure of the monotonic relationship between two ranked variables, calculated from data ranks rather than raw values.
- Standard DeviationStandard deviation is the square root of variance, expressing how dispersed values are around their mean in the variable’s original units under a specified population or sample convention.
- Standard ErrorA measure of an estimated parameter's precision, the standard deviation of its sampling distribution across repeated samples of the same size.
- Standard NormalA specific normal distribution with a mean of zero and standard deviation of one, used as a reference for standardising other variables.
- Standardised ResidualA regression diagnostic dividing an observation's raw residual by its estimated standard deviation, allowing comparison on a common scale.
- Statistical PowerStatistical power is the probability that a study will reject a null hypothesis when a specified alternative hypothesis is true.
- Statistical SignificanceStatistical significance is a threshold-based conclusion that an observed result is sufficiently incompatible with a prespecified null hypothesis under a stated statistical model to meet the chosen significance level.
- Structural Equation ModelA system of linked equations representing hypothesised relationships among observed and possibly latent variables, often combining a measurement model with structural paths.
- Studentized ResidualA standardised residual excluding an observation's own contribution to the model fit, a more robust diagnostic for identifying influential data points.
- Survival RegressionA modelling approach estimating the relationship between predictor variables and a time-to-event outcome, including the Cox model and parametric survival models.
- Synthetic ControlA quasi-experimental method constructing a weighted combination of untreated units to serve as a comparison group resembling a treated unit before intervention.
- T-DistributionA continuous probability distribution similar to the normal but with heavier tails, used when estimating a mean from a small sample with unknown variance.
- Taylor Series ExpansionA mathematical technique approximating a complex function using a polynomial series based on its derivatives at a single point, underlying the delta method.
- Taylor Series MethodA local approximation that uses derivatives around a reference point to estimate a function and propagate input uncertainty into the variance of a derived result.
- Test StatisticA single numerical value calculated from sample data during a hypothesis test, used to judge whether to reject the null hypothesis.
- Tolerance IntervalAn interval calculated to contain a specified proportion of an entire population's values with a given confidence, unlike a confidence or prediction interval.
- Trajectory AnalysisA statistical approach identifying and characterising distinct patterns of change in an outcome over time, often revealing subgroups with different courses.
- Two-Tailed TestA hypothesis test assessing evidence for an effect in either direction, without specifying in advance whether it is expected to be better or worse.
- Type I ErrorIn hypothesis testing, the error of incorrectly rejecting a true null hypothesis, concluding a genuine effect exists when it does not.
- Type II ErrorIn hypothesis testing, the error of failing to reject a false null hypothesis, concluding no effect exists when a genuine effect is present.
- UnbiasednessA property of an estimator whose expected value, across repeated samples, equals the true underlying value of the parameter being estimated.
- Uniform DistributionA probability distribution in which every value within a defined range is equally likely to occur.
- Unstructured CorrelationA correlation structure for repeated measures allowing the correlation between every pair of time points to be estimated freely and independently.
- VarianceA dispersion measure representing the average squared deviation of individual data points from the mean of a dataset.
- Variance Inflation FactorA regression diagnostic quantifying how much an estimated coefficient's variance is inflated by correlation with other predictors, higher values meaning worse multicollinearity.
- Varimax RotationAn orthogonal factor rotation maximising the variance of squared loadings within each factor, tending to make each variable load strongly on only one factor.
- Within Subject VariationThe degree to which repeated measurements on the same individual fluctuate over time, unlike between-subject variation reflecting differences across individuals.
- Working CorrelationIn generalised estimating equations, an assumed correlation structure among repeated observations, improving statistical efficiency even if not perfectly correct.
- Z-ScoreA standardised statistic representing how many standard deviations an observation lies above or below the mean of its distribution.