VerifiedEvidence: highv1.0.0

Generalized Estimating Equations

A method analysing repeated measures or clustered data, estimating population-average effects while accounting for within-cluster correlation, without a full random effects specification.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, Generalised Estimating Equations (GEE) are a semiparametric regression framework used to estimate population-averaged relationships when observations are correlated within clusters or across repeated measurements. They are founded on quasi-likelihood theory and extend generalised linear models by allowing the analyst to specify a working correlation structure for non-independent outcomes. GEE exist to obtain consistent regression estimates without requiring full specification of the joint probability distribution of the repeated or clustered observations.

Mathematically, GEE estimate regression parameters by solving a system of estimating equations that weights residuals by the inverse of the model-based covariance matrix. The mean response is linked to explanatory variables through a specified link function, while the covariance matrix combines a variance function with a working correlation matrix. Under regularity conditions, coefficient estimates remain consistent even when the working correlation structure is misspecified, provided that the mean model is correctly specified, with robust sandwich standard errors used for inference.

In practice, GEE are implemented by selecting an appropriate outcome distribution, link function and working correlation structure, such as independent, exchangeable, autoregressive or unstructured correlation. Regression coefficients are estimated iteratively, and robust standard errors are commonly reported. In health economics, GEE are used to analyse repeated cost, utilisation, quality-of-life and clinical outcome data collected from patients, healthcare providers or geographical clusters.


Purpose

Used to estimate population-averaged associations and treatment effects from longitudinal or clustered health economic data while accounting for within-cluster correlation.


Mathematical Formulae

Primary Formula

?? D??V???(y? ? ??) = 0

where:

  • y? = vector of observed outcomes for cluster i
  • ?? = vector of expected outcomes for cluster i
  • D? = ??? / ???
  • V? = working covariance matrix for cluster i
  • ? = vector of regression coefficients

Supporting Formulae

g(???) = x????

V? = A???�R?(�)A???�

Var(??) = B??MB??

B = ?? D??V???D?

M = ?? D??V???(y? ? ??)(y? ? ??)?V???D?

where:

  • g(�) = link function
  • x?? = vector of explanatory variables
  • A? = diagonal matrix of marginal variances
  • R?(�) = working correlation matrix
  • � = correlation parameter
  • Var(??) = robust sandwich covariance estimator

Related Mathematical Methods

  • Generalised Linear Models
  • Quasi-Likelihood Estimation
  • Sandwich Variance Estimation
  • Exchangeable Correlation
  • Autoregressive Correlation
  • Cluster-Robust Standard Errors
  • Longitudinal Data Analysis

Example

A health economist analyses quarterly healthcare expenditure for 1,000 patients observed over two years. Because repeated expenditure measurements from the same patient are correlated, a GEE model with a log link, gamma variance function and exchangeable working correlation structure is fitted. The estimated treatment coefficient is ?0.12.

exp(?0.12) = 0.887

The treatment group therefore has approximately 11.3% lower mean expenditure than the comparator group, interpreted as a population-averaged effect.


Excel Implementation

FunctionExample FormulaHealth Economics Application
EXP=EXP(B2)Transform a log-link regression coefficient into a mean ratio
MMULT=MMULT(A2:D5,F2:I5)Perform matrix multiplication within estimating-equation calculations
MINVERSE=MINVERSE(A2:D5)Invert the working covariance matrix
TRANSPOSE=TRANSPOSE(A2:D5)Construct transposed design and derivative matrices
SUMPRODUCT=SUMPRODUCT(B2:B101,C2:C101)Calculate weighted residual and score components

VBA (Optional)

Automate GEE estimation by iteratively updating regression coefficients, working correlation matrices and robust covariance estimates until convergence.


Sources

  • Liang KY, Zeger SL. Longitudinal data analysis using generalised linear models. Biometrika. 1986;73(1):13-22.
  • Zeger SL, Liang KY. Longitudinal data analysis for discrete and continuous outcomes. Biometrics. 1986;42(1):121-130.
  • Hardin JW, Hilbe JM. Generalized Estimating Equations. 2nd ed.
  • Diggle PJ, Heagerty P, Liang KY, Zeger SL. Analysis of Longitudinal Data. 2nd ed.
  • Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes. 4th ed.

Library

Publications

1
  • Book

    Bayesian Methods in Health Economics — Gianluca Baio, 1st Edition ed., 2012 (Chapman & Hall / CRC Press)

    An overview of Bayesian statistical methods for the analysis of health economic data, covering economic evaluation concepts, statistical cost-effectiveness analysis, Bayesian computation and MCMC, and applied health economic evaluation.

Frequently Asked Questions (6)

  • What are generalized estimating equations?

    A method analysing repeated measures or clustered data, estimating population-average effects while accounting for within-cluster correlation, without a full random effects specification.

    Source: Liang & Zeger 1986

  • What do generalized estimating equations estimate from clustered data?

    Generalized estimating equations analyse repeated or clustered measurements to estimate the average effect across a population, while accounting for the fact that observations within a cluster are correlated. They focus on the population-average relationship rather than modelling each cluster's own deviation, treating the within-cluster correlation as a nuisance to be adjusted for rather than a quantity of interest. This makes them a robust choice when the goal is a marginal, population-level effect. Estimating population-average effects under correlation is their purpose. Kirkwood and Sterne (2003) describe this method.

    Source: Kirkwood & Sterne 2003

  • How do generalized estimating equations work?

    Generalized estimating equations work by extending generalised linear models to correlated data, specifying a model for the mean response and a working correlation structure for the within-cluster correlation, then estimating the parameters through estimating equations that account for that correlation. Robust standard errors give valid inference even if the working correlation is misspecified. So generalized estimating equations work by fitting the mean model while adjusting for within-cluster correlation via a chosen working structure, and using robust variance estimation, which yields consistent estimates of population-average effects and valid standard errors for clustered or repeated-measures data.

    Source: Liang & Zeger 1986

  • When are generalized estimating equations used?

    Generalized estimating equations are used for analysing clustered or longitudinal data when the interest is in population-average effects, such as the average effect of a treatment across the population, rather than in effects for individual clusters. They are common for repeated measures and grouped data with various outcome types. So generalized estimating equations are used when correlated data must be analysed and average effects are wanted, providing a flexible approach for binary, count, or continuous outcomes that accounts for within-cluster correlation, which is why they are widely applied in longitudinal studies and clustered designs where marginal, population-level effects are of interest.

    Source: Liang & Zeger 1986

  • How do generalized estimating equations differ from mixed models?

    Generalized estimating equations estimate population-average, or marginal, effects and treat the within-cluster correlation as a nuisance handled by a working structure, while mixed models include random effects that model the correlation explicitly and estimate cluster-specific, or conditional, effects. The two answer subtly different questions, and their estimates can differ for non-linear models. So generalized estimating equations and mixed models differ in whether they target average or cluster-specific effects and in how they treat the correlation, with generalized estimating equations focusing on marginal effects with robust inference and mixed models on conditional effects with a full specification, and the choice depends on the question and the interpretation desired.

    Source: Liang & Zeger 1986

  • What are the advantages of generalized estimating equations?

    The advantages of generalized estimating equations include that they give consistent estimates of population-average effects and valid robust standard errors even when the working correlation structure is misspecified; that they handle various outcome types and correlation patterns flexibly; and that they avoid the stronger distributional assumptions of full random effects models. So generalized estimating equations are advantageous for their robustness and flexibility in analysing correlated data for marginal effects, requiring fewer assumptions about the correlation than mixed models, which is why they are a popular choice when population-average effects are of interest and a full model of the correlation structure is not needed or desired.

    Source: Liang & Zeger 1986

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 16 Dec 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-SA-070

Stable URI · Machine-readable · Resolvable · CC BY 4.0