Difference-in-differences regression with group, period and interaction terms

Writes each group-period mean as a constant, a fixed group gap, a common period change and an interaction that is non-zero only for the treated group after the policy. Fitted by ordinary least squares to unit-level data, the interaction coefficient equals the two-by-two estimate, and the regression also gives its standard error. gamma_D and lambda_P are the article's gamma and lambda, written with subscripts naming the indicator each one multiplies.

Signature

Y = alpha + gamma_D * D + lambda_P * P + tau * D * P
Inputs
InputsDefinitionUnit
alphaConstant: the comparison group's mean before the policyoutcome unit
gamma_DTreated group mean minus comparison group mean before the policyoutcome unit
D1 for the treated group, 0 for the comparison groupnone
lambda_PChange in the comparison group's mean from before to after the policyoutcome unit
P1 for the period after the policy starts, 0 beforenone
tauCoefficient on D x P, equal to the two-by-two difference-in-differences estimateoutcome unit
Output
YFitted mean outcome for the group and period set by D and Poutcome unit

Function

Difference-in-differences estimate of a policy effect from treated and comparison groups

Maps outcomes observed before and after a policy in a group exposed to it and a group that is not to an estimate of the policy's average effect on the exposed: the change in the exposed group minus the change in the comparison group. Fixed differences between the groups and shocks common to both cancel. The estimate is causal only under parallel trends and no anticipation. The notation follows the Difference-in-Differences article.

Try this function

Implementations

  • Excel

    Difference-in-differences regression coefficients from named group means

    With the four means named as in HE-IM-DID-001, the first three formulas give the constant, the group gap and the common change, held in Alpha, GapD and ChangeP, and the fourth the interaction coefficient, held in Interaction. With unit-level data in columns, LINEST on columns D, P and D x P returns the same coefficients.

    =ComparisonBefore; =TreatedBefore-ComparisonBefore; =ComparisonAfter-ComparisonBefore; =TreatedAfter-TreatedBefore-ComparisonAfter+ComparisonBefore

Assumptions

  • Balanced panel or repeated cross-sections for the difference-in-differences regression

    With a balanced panel the regression with a constant, D, P and D x P gives the same coefficient as one with unit and period fixed effects; the form with a constant also applies to repeated cross-sections.

  • Standard errors clustered at the level of policy assignment

    Outcomes are correlated within areas or hospitals over time, so standard errors are clustered at the level at which the policy was assigned, which needs enough treated and comparison clusters.

Worked examples

  • Fitted treated mean after the policy in the admissions example

    With alpha of 48.0, a fixed gap of 4.0, a common change of minus 2.0 and an interaction of minus 3.0, the fitted mean for region A after the policy is 47.0, the observed value.

    alpha = 48; gamma_D = 4; lambda_P = -2; tau = -3; D = 1; P = 1; Y = 47
  • Fitted comparison mean after the policy in the admissions example

    For region B after the policy only the constant and the common change apply: 48.0 minus 2.0 gives 46.0.

    alpha = 48; gamma_D = 4; lambda_P = -2; tau = -3; D = 0; P = 1; Y = 46
  • Fitted New Jersey mean from the Card and Krueger difference-in-differences regression

    From Card and Krueger's rounded means, alpha is 23.33, the gap minus 2.89, the common change minus 2.16 and the interaction 2.75, so the fitted New Jersey mean after the rise is 21.03.

    alpha = 23.33; gamma_D = -2.89; lambda_P = -2.16; tau = 2.75; D = 1; P = 1; Y = 21.03

Common errors

  • Reading gamma_D as the policy effect

    The group coefficient is the gap before the policy (4.0 admissions per 1,000 in the example); only the interaction measures the effect.

  • Conventional standard errors for difference-in-differences with serially correlated outcomes

    Bertrand, Duflo and Mullainathan found a spurious effect significant at the 5 per cent level for up to 45 per cent of placebo laws when correlation within states over time was ignored.

Sources

  • Two-way fixed effects and interaction regression equivalent to the two-by-two estimate

    Roth J, Sant'Anna PHC, Bilinski A, Poe J. What's trending in difference-in-differences? A synthesis of the recent econometrics literature. Journal of Econometrics. 2023;235(2):2218-2244 (read as arXiv 2201.01194 version 3). Section 2.4 and footnote 2: the OLS coefficient on the post-period times treatment interaction in the TWFE regression equals the sample-analogue estimate; with a balanced panel it is numerically identical to the coefficient from a regression with a constant, a treatment indicator, a second-period indicator and their interaction, which generalises to repeated cross-sections; standard errors clustered at the level at which treatment is assigned.

    View source →

  • Spurious difference-in-differences effects from placebo laws with serially correlated outcomes

    Bertrand M, Duflo E, Mullainathan S. Quarterly Journal of Economics. 2004;119(1):249-275. doi:10.1162/003355304772839588 (abstract read). Placebo laws randomly generated in state-level female wage data: conventional DD standard errors severely understate the standard deviation of the estimators, giving an effect significant at the 5 per cent level for up to 45 per cent of the placebo interventions.

    View source →

Canonical Identity