Signature
ATT_gt = (Y_g_t - Y_g_base) - (Y_n_t - Y_n_base)
| Inputs | Definition | Unit |
|---|---|---|
Y_g_t | Mean outcome in period t of units first treated in period g | outcome unit |
Y_g_base | Mean outcome of the same cohort in period g minus 1 | outcome unit |
Y_n_t | Mean outcome in period t of units never treated during the panel | outcome unit |
Y_n_base | Mean outcome of never-treated units in period g minus 1 | outcome unit |
ATT_gt | Average effect in period t for units first treated in period g | outcome unit |
|---|
Function
Difference-in-differences estimate of a policy effect from treated and comparison groups
Maps outcomes observed before and after a policy in a group exposed to it and a group that is not to an estimate of the policy's average effect on the exposed: the change in the exposed group minus the change in the comparison group. Fixed differences between the groups and shocks common to both cancel. The estimate is causal only under parallel trends and no anticipation. The notation follows the Difference-in-Differences article.
Computational function
Computational function: group-time effects and event-time averages from cohort mean outcomes
Takes a table of mean outcomes by adoption cohort and period, with never-treated units in the last row, the adoption period of each cohort and the number of units in each cohort, and returns every post-adoption ATT(g,t) (HE-FM-DID-004) and their averages by time since adoption e, weighting cohorts by size among those observed at e, as in Callaway and Sant'Anna's event-study aggregation. The inputs differ from the formula's: a whole panel of means instead of four.
Inputs and outputs:
Y: Matrix of mean outcomes, one row per adoption cohort and a last row for never-treated units, one column per period; required. Unit: outcome unit.;years: Periods labelling the columns, consecutive; required, vector. Unit: years.;g: Adoption period of each cohort row, each after the first period; required, vector. Unit: years.;n: Number of units in each cohort; required, vector, above zero. Unit: units.;ATT: Group-time effect for each cohort and post-adoption period, withe= t minus g. Unit: outcome unit.;event: Size-weighted average of the effects at each e. Unit: outcome unit.Assumption: Parallel trends between each cohort and the never-treated units, no anticipation, treatment that stays on once adopted, and cohort means measured on the same scale in every period.
Worked example (Two cohorts and a never-treated group, 2021 to 2025): Admissions per 1,000 of 60, 59, 55, 53, 51 for three units adopting in 2023, 55, 54, 53, 50, 49 for two units adopting in 2024 and 50, 49, 48, 47, 46 for never-treated units give effects of minus 3, minus 4 and minus 5 for the 2023 cohort and minus 2 and minus 2 for the 2024 cohort; weighting by 3 and 2 units gives minus 2.6 in the adoption year, minus 3.2 one year later and minus 5 two years later (computed here for illustration).
years = [2021, 2022, 2023, 2024, 2025]; g = [2023, 2024]; n = [3, 2]; ATT = [-3, -4, -5, -2, -2]; event = [-2.6, -3.2, -5]Worked example (Cohorts moving with the never-treated group): If both cohorts fall by 1 a year like the never-treated units, every effect and every event-time average is zero.
ATT = [0, 0, 0, 0, 0]; event = [0, 0, 0]Excel: With cohort means in rows 2 and 3 and never-treated means in row 4, years in B1:F1 and each cohort's base-period column found by MATCH,
=(D2-INDEX($B2:$F2,MATCH($H2-1,$B$1:$F$1,0)))-(D$4-INDEX($B$4:$F$4,MATCH($H2-1,$B$1:$F$1,0)))gives ATT(g,t) for the cohort whose adoption year is in H2 and the period in column D; the size-weighted average at each e uses SUMPRODUCT of effects and cohort sizes divided by SUM of the sizes observed at that e.R:
att_gt <- function(Y, years, g, n) { N <- Y[nrow(Y), ]; out <- NULL; for (k in seq_along(g)) { b <- which(years == g[k]-1); for (j in which(years >= g[k])) out <- rbind(out, data.frame(g = g[k], t = years[j], e = years[j]-g[k], n = n[k], ATT = (Y[k, j]-Y[k, b])-(N[j]-N[b]))) }; list(att = out, event = sapply(split(out, out$e), function(d) sum(d$n*d$ATT)/sum(d$n))) }WithY <- rbind(c(60, 59, 55, 53, 51), c(55, 54, 53, 50, 49), c(50, 49, 48, 47, 46)),att_gt(Y, 2021:2025, c(2023, 2024), c(3, 2))returns the effects and event-time averages of the first example.Python:
def att_gt(Y, years, g, n): N = Y[-1]; att = [(gk, t, t-gk, nk, (Y[k][j]-Y[k][years.index(gk-1)])-(N[j]-N[years.index(gk-1)])) for k, (gk, nk) in enumerate(zip(g, n)) for j, t in enumerate(years) if t >= gk]; ev = {e: sum(r[3]*r[4] for r in att if r[2] == e)/sum(r[3] for r in att if r[2] == e) for e in sorted({r[2] for r in att})}; return {"att": att, "event": ev}Returns the same values as the R function, with Y as a list of rows.Test (Event-time average with one cohort equals that cohort's effect): At e = 2 only the 2023 cohort is observed, so the average equals its effect of minus 5. Expected result: TRUE. Excel check:
=ABS(EventAvg2-EffectCohort2023Year2025)<1E-9Test (No differential change gives zero effects): Shifting every cohort row so that it changes like the never-treated row returns zeros. Expected result: TRUE. Excel check:
=SUMPRODUCT(ABS(ATTRange))=0Common error (Averaging effects without stating weights): An unweighted mean of the five effects is minus 3.2, while weighting each cohort-period effect by cohort size gives minus 44/13, about minus 3.38, and the event-time averages range from minus 2.6 to minus 5; the aggregation scheme changes the summary and must be reported.
Source: Callaway B, Sant'Anna PHC. Difference-in-differences with multiple time periods. Journal of Econometrics. 2021;225(2):200-230 (read as arXiv 1803.09015 version 4). Table 1 (event-study weights); Roth J, Sant'Anna PHC, Bilinski A, Poe J. What's trending in difference-in-differences? A synthesis of the recent econometrics literature. Journal of Econometrics. 2023;235(2):2218-2244 (read as arXiv 2201.01194 version 3). Section 3, equation 8.
ATT_g_t = (Y_g_t - Y_g_(g-1)) - (Y_n_t - Y_n_(g-1)); event_e = sum_(g) [n_g * ATT_g_(g+e)] / sum_(g) [n_g]
Try this function
Implementations
Excel
Group-time effect from named cohort and never-treated means
With the cohort's means in CohortNow and CohortBase and the never-treated means in NeverNow and NeverBase, the formula returns ATT(g,t), held in GroupTimeATT.
=(CohortNow-CohortBase)-(NeverNow-NeverBase)
Assumptions
Parallel trends between each cohort and the never-treated group
Without treatment, each cohort's mean would have changed like the never-treated mean from g minus 1 to t. Callaway and Sant'Anna also allow not-yet-treated units as the comparison, and conditional versions with covariates.
Staggered adoption that stays on and is not anticipated
Units remain treated after adoption, and outcomes in g minus 1 are not affected by the coming policy.
Worked examples
Cohort adopting in 2023, effect in 2025
In an illustrative panel, units adopting in 2023 average 59 admissions per 1,000 in 2022 and 51 in 2025, while never-treated units fall from 49 to 46, so the effect two years after adoption is minus 8 minus minus 3, or minus 5 (computed here for illustration).
Y_g_t = 51; Y_g_base = 59; Y_n_t = 46; Y_n_base = 49; ATT_gt = -5
Cohort adopting in 2024, effect in the adoption year
Units adopting in 2024 move from 53 in 2023 to 50 in 2024 while never-treated units fall from 48 to 47, an effect of minus 2 (computed here for illustration).
Y_g_t = 50; Y_g_base = 53; Y_n_t = 47; Y_n_base = 48; ATT_gt = -2
Common errors
Using already-treated units as controls under staggered adoption
Two-way fixed effects regressions under staggered adoption include comparisons in which already-treated units serve as controls; when effects change over time some of these comparisons receive negative weights, and the coefficient can even take the opposite sign to every underlying effect.
Using the cohort's own first period as the base
The base is the last period before adoption, g minus 1; a base further back adds any pre-policy differential trend to every effect.
Sources
Group-time ATT estimated from cohort and comparison-group changes since g minus 1
Roth J, Sant'Anna PHC, Bilinski A, Poe J. What's trending in difference-in-differences? A synthesis of the recent econometrics literature. Journal of Econometrics. 2023;235(2):2218-2244 (read as arXiv 2201.01194 version 3). Section 3, equation 8: ATT(g,t) is estimated by the mean change in Y from period g minus 1 to t for cohort g minus the mean change for a comparison set; Callaway and Sant'Anna consider never-treated units or all not-yet-treated units as the comparison; TWFE regressions make forbidden comparisons between already-treated units.
Group-time average treatment effects and their aggregation by length of exposure
Callaway B, Sant'Anna PHC. Difference-in-differences with multiple time periods. Journal of Econometrics. 2021;225(2):200-230 (read as arXiv 1803.09015 version 4). Sections 2 and 3: ATT(g,t) under parallel trends with never-treated or not-yet-treated comparison groups; Table 1 weights for aggregation, with the event-study parameter weighting ATT(g, g + e) by P(G = g | G + e <= T).
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0