cassionData Analysis

Back to the lessonLesson 3 of 8When nobody randomised

Minus 0.96 points, and the assumption you cannot test

The same deck as the downloads, rendered as a page. Start the slideshow to present it full screen — arrow keys or a click advance one slide, Escape leaves.

Slides · PDFSlides · PowerPoint

  1. Slide 1 / 26

    What this lesson covers

    • Two differences, subtracted
    • Fit it as a regression, because you need the interval
    • The assumption, stated exactly
    • Why two rounds cannot test it
    • Four arguments to make when you cannot test it
    • Where difference-in-differences goes wrong
    • Report it whole
    • What comes next
    Speaker notes
    Subtracting each group's own baseline gives a difference-in-differences of −0.96 points. It rests on the two groups having been about to change by the same amount, which is untestable with two rounds — so the lesson is how to argue for it rather than how to check it.
  2. Slide 2 / 26

    Two differences, subtracted — In Python

    import pandas as pd
    import numpy as np
    
    cells = d.groupby("feeding_programme")[["baseline", "endline"]].mean()
    print(cells.round(4))
    
    did = ((cells.loc[True, "endline"] - cells.loc[True, "baseline"])
           - (cells.loc[False, "endline"] - cells.loc[False, "baseline"]))
    print(f"difference-in-differences: {did:+.4f}")
  3. Slide 3 / 26

    Two differences, subtracted — In R

    library(dplyr)
    
    d |> summarise(baseline = mean(baseline), endline = mean(endline),
                   .by = feeding_programme) |>
      mutate(gain = endline - baseline)
  4. Slide 4 / 26

    Two differences, subtracted

    BaselineEndlineChange
    Feeding programme57.75%64.41%+6.66
    No programme54.67%62.29%+7.62
    Difference+3.09+2.12−0.96
  5. Slide 5 / 26

    Two differences, subtracted

    • The estimate can be read down the last column or across the last row and it is the same number — That symmetry is what…
    • The baseline gap of 3.09 points is what the design removes — A cross-sectional comparison at endline would have…
    Speaker notes
    The estimate can be read down the last column or across the last row and it is the same number. That symmetry is what the name describes: the difference between the two groups' changes, which equals the change in the difference between the two groups. The baseline gap of 3.09 points is what the design removes. A cross-sectional comparison at endline would have reported +2.12 and called it an effect; the difference-in-differences says most of that gap predates the programme.
  6. Slide 6 / 26

    Fit it as a regression, because you need the interval — In Python

    import statsmodels.formula.api as smf
    
    d = d.assign(gain=d["endline"] - d["baseline"])
    
    naive = smf.ols("gain ~ feeding_programme", data=d).fit()
    clustered = naive.get_robustcov_results(cov_type="cluster",
                                            groups=d["school_id"])
    
    by_school = d.groupby(["school_id", "feeding_programme"])["gain"].mean().reset_index()
    aggregated = smf.ols("gain ~ feeding_programme", data=by_school).fit()
  7. Slide 7 / 26

    Fit it as a regression, because you need the interval — In R

    lm(gain ~ feeding_programme, data = d)                       # naive
    lm(gain ~ feeding_programme, data = by_school)               # school level
  8. Slide 8 / 26

    Fit it as a regression, because you need the interval

    ApproachEstimateSE95% CI
    Student level, naive−0.96 pts0.98−2.89 to +0.96
    Student level, clustered on school−0.96 pts1.29−3.48 to +1.56
    School level, 15 vs 9−0.68 pts1.18−2.99 to +1.63
  9. Slide 9 / 26

    Fit it as a regression, because you need the interval

    • The clustered interval is the one to report — for the reason the regression course established: the programme was…
    • Every version crosses zero comfortably — There is no effect to report, and the interval says the study could not have…
    Speaker notes
    The clustered interval is the one to report, for the reason the regression course established: the programme was assigned to 24 schools, not to 585 children. Every version crosses zero comfortably. There is no effect to report, and the interval says the study could not have ruled out anything between a 3.5-point loss and a 1.6-point gain.
  10. Slide 10 / 26

    The assumption, stated exactly

    • Parallel trends: in the absence of the programme, the two groups' outcomes would have changed by the same amount
    • It does not require the groups to be similar — They start 3.09 points apart and that is fine — the design differences…
    • It does not require the trends to be flat — Both groups can be improving; the assumption is that they would have…
    • It does not require the same variance, sample size or composition — Those affect the interval, not the identification
    • What it does require is unobservable — because it is a statement about a world in which the programme did not happen
    Speaker notes
    Parallel trends: in the absence of the programme, the two groups' outcomes would have changed by the same amount. Three things that assumption does not require, and each is a common misreading: It does not require the groups to be similar. They start 3.09 points apart and that is fine — the design differences it away. Balance matters for how plausible parallel trends is, not for whether the estimator works. It does not require the trends to be flat. Both groups can be improving; the assumption is that they would have improved equally. It does not require the same variance, sample size or composition. Those affect the interval, not the identification. What it does require is unobservable, because it is a statement about a world in which the programme did not happen. That is why this lesson is about argument rather than about testing.
  11. Slide 11 / 26

    Why two rounds cannot test it — In Python

    rounds = assessment["assessment_round"].unique()
    print(rounds)
  12. Slide 12 / 26

    Why two rounds cannot test it — In R

    unique(assessment$assessment_round)
  13. Slide 13 / 26

    Why two rounds cannot test it

    • There are two: baseline and endline — With two points you can draw exactly one line through each group, so any…
    • With three or more pre-programme rounds you can look — Plot each group's outcome over the pre-period and see whether…
    • The design implication is a data-collection decision made years earlier — If you expect to evaluate by…
    Speaker notes
    There are two: baseline and endline. With two points you can draw exactly one line through each group, so any pre-programme divergence is unmeasurable by construction. With three or more pre-programme rounds you can look. Plot each group's outcome over the pre-period and see whether the lines move together. That is not a proof — past parallelism does not guarantee future parallelism — but it is evidence, and it is the single most persuasive exhibit an evaluation of this kind can carry. The design implication is a data-collection decision made years earlier. If you expect to evaluate by difference-in-differences, collect more than one pre-round. The cost is one extra survey; the alternative is an assumption nobody can examine.
  14. Slide 14 / 26

    Four arguments to make when you cannot test it

    • Name why the programme schools were chosen — If selection was on something time-invariant — where they are, who runs…
    • Show that other outcomes moved in parallel — Numeracy is measured on the same children and is not what a feeding…
    Speaker notes
    Each is available here, and together they are what the report offers in place of a test. Name why the programme schools were chosen. If selection was on something time-invariant — where they are, who runs them — parallel trends is more plausible than if selection was on something trending, like a recent fall in enrolment. Show that other outcomes moved in parallel. Numeracy is measured on the same children and is not what a feeding programme targets first.
  15. Slide 15 / 26

    Four arguments to make when you cannot test it — In Python

    def panel(domain):
        wide = (assessment[assessment["domain"] == domain]
                .pivot_table(index="student_id", columns="assessment_round",
                             values="pct").dropna())
        out = wide.join(roster.set_index("student_id"), how="inner").reset_index()
        return out.assign(gain=out["endline"] - out["baseline"])
    
    for domain in ("literacy", "numeracy"):
        frame = panel(domain)
        fit = smf.ols("gain ~ feeding_programme", data=frame).fit(
            cov_type="cluster", cov_kwds={"groups": frame["school_id"]})
        print(f"{domain}: {fit.params['feeding_programme[T.True]']:+.4f}")
  16. Slide 16 / 26

    Four arguments to make when you cannot test it — In R

    # A placebo outcome: it should show nothing, and here it does.
  17. Slide 17 / 26

    Four arguments to make when you cannot test it

    • Numeracy gives +0.39 points, 95% CI −1.16 to +1.94 — the same design applied to an outcome the programme was not…
    • Test a placebo period if one exists — Two pre-programme rounds should give a difference-in-differences of zero
    • Show the result is not driven by one unit — Drop each school in turn and refit; if the estimate swings, the design is…
    Speaker notes
    Numeracy gives +0.39 points, 95% CI −1.16 to +1.94 — the same design applied to an outcome the programme was not expected to move first, and it shows nothing either. That is weak evidence and it is the kind available. Test a placebo period if one exists. Two pre-programme rounds should give a difference-in-differences of zero. There are none here, and saying so is better than implying there were. Show the result is not driven by one unit. Drop each school in turn and refit; if the estimate swings, the design is resting on one school's trend.
  18. Slide 18 / 26

    Four arguments to make when you cannot test it — In Python

    for school in sorted(d["school_id"].unique()):
        subset = d[d["school_id"] != school]
        est = smf.ols("gain ~ feeding_programme", data=subset).fit()
        print(f"without {school}: {est.params['feeding_programme[T.True]']:+.4f}")
  19. Slide 19 / 26

    Four arguments to make when you cannot test it — In R

    # 24 refits, one line each. Cheap, and a reviewer will ask.
  20. Slide 20 / 26

    Four arguments to make when you cannot test it

    • Dropping any one school moves the estimate between −1.61 and −0.37 points — No single school carries it, which is worth…
    Speaker notes
    Dropping any one school moves the estimate between −1.61 and −0.37 points. No single school carries it, which is worth one line in the annex and is the check a reviewer runs first on 24 units.
  21. Slide 21 / 26

    Where difference-in-differences goes wrong

    • Different timing — If the programme started in different schools in different months, a single before-after split…
    • Composition change — The 585 children are those with both rounds
    • Anticipation — If schools changed behaviour once they knew the programme was coming, the baseline is already…
    • A control group that is affected — Children moving between schools, teachers transferring, or a nearby school changing…
    Speaker notes
    Different timing. If the programme started in different schools in different months, a single before-after split misclassifies part of the treatment period. The staggered case needs a different estimator, and the naive two-way fixed effects model is now known to be biased for it. Composition change. The 585 children are those with both rounds. If children left differentially between the groups, the panel is not the same population twice — which is the attrition problem the statistics course's exercise found in this very file. Anticipation. If schools changed behaviour once they knew the programme was coming, the baseline is already contaminated and the estimate is too small. A control group that is affected. Children moving between schools, teachers transferring, or a nearby school changing its practice in response all break the assumption that the comparison group shows what would have happened.
  22. Slide 22 / 26

    Report it whole — Example (cont.)

    School feeding and literacy, difference-in-differences
    
      585 students with both assessment rounds, in 24 schools.
    
                            Baseline   Endline   Change
        Feeding (15 sch)      57.75%    64.41%    +6.66
        No programme (9)      54.67%    62.29%    +7.62
        Difference            +3.09     +2.12     -0.96
    
      Estimate -0.96 points, 95% CI -3.48 to +1.56, clustered on 24 schools.
      No effect on literacy is detected.
    
      The estimate assumes the two groups would have changed by the same amount
      without the programme. Only two assessment rounds exist, so this cannot be
      examined and is argued rather than tested: assignment appears to follow
      district and school size rather than a trend in results, and the same
  23. Slide 23 / 26

    Report it whole — Example (cont.)

      design applied to numeracy also shows nothing.
    
      The interval excludes effects larger than 1.6 points in either direction
      but not smaller ones. The design's minimum detectable effect was 4.9
      points, so this is a null result about large effects only.
  24. Slide 24 / 26

    Report it whole

    • The last paragraph converts a null into a statement with a boundary — which is what the statistics course asked for and…
    Speaker notes
    The last paragraph converts a null into a statement with a boundary, which is what the statistics course asked for and what the power lesson makes precise.
  25. Slide 25 / 26

    What comes next

    • Difference-in-differences uses the baseline by subtracting it.
    Speaker notes
    Difference-in-differences uses the baseline by subtracting it. The next lesson uses it a different way — to choose which comparison schools to keep — and finds that improving the balance means discarding six of the fifteen programme schools.
  26. Slide 26 / 26

    Where this goes next

    Read the full lesson, with runnable code Back to the lesson