cassionData Analysis

Lesson 2 of 8

Unit · A model is a comparison

Adjustment moves the estimate; clustering moves the standard error

Adding every child-level covariate in the register moves the feeding coefficient from 4.47 to 4.37 and leaves the standard error at 0.9. Adding a term for the school leaves the coefficient alone and moves the standard error to 1.5. These are two different problems and they are routinely confused.

PythonR150 minSMART surveyUNICEF indicator definitionsOECD DAC evaluation criteria

The one table this course is built around

import pandas as pd
import statsmodels.formula.api as smf

attendance = pd.read_csv("school-attendance-2024.v1.csv")
roster = pd.read_csv("school-roster-2024.v1.csv").drop_duplicates(
    "student_id", keep="last")
MARKS = {"true": True, "Y": True, "false": False, "N": False}
marked = attendance[attendance["present"].isin(MARKS)].copy()
marked["attended"] = marked["present"].map(MARKS)
per_student = (marked.groupby("student_id")["attended"].mean()
               .rename("rate").reset_index().merge(roster, on="student_id"))

enrolment = pd.read_csv("school-enrolment-2024.v1.csv")
current = enrolment[enrolment["school_year"] == 2024]

d = per_student.merge(
    current[["student_id", "sex", "age_years", "disability_reported",
             "displacement_status", "admin2"]], on="student_id").dropna()
print(f"complete cases: {len(d)}")

m1 = smf.ols("rate ~ feeding_programme", data=d).fit()
m2 = smf.ols("rate ~ feeding_programme + sex + age_years"
             " + disability_reported + displacement_status", data=d).fit()
m3 = smf.ols("rate ~ feeding_programme + sex + age_years"
             " + disability_reported + displacement_status + admin2", data=d).fit()
m4 = m1.get_robustcov_results(cov_type="cluster", groups=d["school_id"])

for name, m in [("child covariates", m2), ("+ district", m3)]:
    print(f"{name:18} {m.params['feeding_programme[T.True]']:+.4f}")
library(dplyr)

m1 <- lm(rate ~ feeding_programme, data = d)
m2 <- lm(rate ~ feeding_programme + sex + age_years +
           disability_reported + displacement_status, data = d)
m3 <- update(m2, . ~ . + admin2)
Model Feeding coefficient SE t
M1 one predictor +0.0447 0.0089 5.02
M2 + sex, age, disability, displacement +0.0437 0.0090 4.85
M3 + district +0.0360 0.0100 3.60
M4 = M1 with school-clustered SEs +0.0447 0.0154 2.90

Read the columns rather than the rows, because they say two different things.

Adding covariates changed the coefficient and left the standard error alone. M1 to M3 moves the estimate from 4.5 points to 3.6 while the standard error goes from 0.9 to 1.0.

Clustering changed the standard error and left the coefficient untouched. M4 is the same fitted line as M1 — identical coefficient to every decimal — with a standard error 1.7 times larger.

These are two different problems. One is about whether the comparison is between comparable groups; the other is about how much information the sample actually contains. Fixing one does nothing for the other, and a model that fixes neither is wrong in two independent ways.

What “adjusting for” actually does

The phrase appears in every programme report and means something narrower than readers assume.

A coefficient with covariates in the model is the comparison within levels of those covariates, pooled across them. “Adjusted for district” means: compare fed and unfed students inside each district, then average those comparisons.

within = (d.groupby("admin2")
          .apply(lambda g: g[g["feeding_programme"]]["rate"].mean()
                 - g[~g["feeding_programme"]]["rate"].mean()))
print(within.round(4))
print(f"unadjusted: {m1.params['feeding_programme[T.True]']:.4f}")
print(f"adjusted:   {m3.params['feeding_programme[T.True]']:.4f}")
d |> summarise(gap = mean(rate[feeding_programme]) -
                     mean(rate[!feeding_programme]), .by = admin2)
District Gap Fed Unfed
Artibonite +3.8 pts 146 141
Centre −4.1 pts 244 42
Nord-Ouest +4.9 pts 251 45
Sud +7.0 pts 104 178

The adjusted coefficient is a weighted average of those four gaps, and seeing them separately is worth the two lines it costs — because one of them runs the other way. Centre’s fed schools attend 4.1 points worse than its unfed schools, against a pooled estimate of +3.6.

A single coefficient averaged over a disagreement is still a number, and it is no longer a summary. Before reporting the adjusted 3.6, say that Centre disagrees and that its comparison rests on 42 unfed students — which is thin enough that the reversal may be noise, and thin enough that it should be checked rather than averaged away.

Why the covariates barely moved anything

M2 adds sex, age, disability and displacement status and moves the coefficient by one tenth of a point. That is not a failure of the covariates — it is a fact about the design.

The feeding programme was assigned by school. Nothing about a child decided whether they got school meals, so child characteristics cannot be confounders of the feeding comparison, and adjusting for them cannot move the estimate much.

District did move it, from 4.5 to 3.6, and for the same reason in reverse: schools with a feeding programme are not evenly spread across districts, so district is a genuine confounder of a school-level exposure.

The rule the design implies: confounders live at the level of assignment. For a school-level programme, look for school-level and district-level differences; adding child-level covariates is neither harmful nor useful here, and mostly signals that the analyst reached for whatever the file happened to carry.

Reading a coefficient table without being misled

Four habits, each of which prevents a specific misreading.

Report the sample size of the model, not of the file. Complete-case fitting dropped 49 students here — 1,200 rows became 1,151 — and the drop is silent. Every model in the table above must be fitted on the same rows or the comparison across models is contaminated by which rows each one kept.

print(f"file {len(per_student)}, model {int(m3.nobs)}")
c(file = nrow(per_student), model = nobs(m3))

Do not read the covariate coefficients as findings. They are adjusted for a set of variables chosen to answer a question about feeding, not about sex or displacement. A coefficient is interpretable only for the exposure the model was built around — this is the “Table 2 fallacy”, and it is how a displacement coefficient fitted as a control ends up quoted as a displacement finding.

Check whether an added covariate changed the sample. A covariate with missing values silently drops rows, so a coefficient that moves when you add it may have moved because the sample changed rather than because the adjustment did anything.

Report the unadjusted estimate too. The pair is what shows the reader how much work the adjustment did, and a report giving only the adjusted number has hidden the one comparison that needs no assumptions.

When adjustment cannot help

Three cases where adding a covariate is not the answer, all of which appear later in this course.

The confounder is not in the file. Adjusting for what you measured says nothing about what you did not, and “adjusted” is not a synonym for “unconfounded”. The epidemiology course made this point with age standardisation; regression does not improve on it.

The covariate is on the causal path. Lesson 5 shows a covariate that removes 42% of the effect by sitting between exposure and outcome, and adding it makes the answer worse in a way no fit statistic reveals.

The problem is the standard error, not the estimate. No covariate fixes clustering. That is the M1-to-M4 row, and lesson 6 is the whole answer to it.

Report it whole

Attendance and school feeding, adjusted analysis

  Linear model, 1,151 students with complete covariates (49 dropped).

  Unadjusted                     +4.47 points   95% CI +2.73 to +6.22
  Adjusted for district          +3.60 points   95% CI +1.64 to +5.56
  Child-level covariates (sex, age, disability, displacement) change the
  estimate by less than 0.1 points and are retained only to show that.

  The district gaps are +3.8, -4.1, +4.9 and +7.0 points. Centre runs the
  other way on 42 unfed students and is not averaged away silently.

  Standard errors above assume independent students and are too small; the
  school-clustered version of the unadjusted model gives +4.47 (95% CI
  +1.45 to +7.49).

  Observational. Schools were not randomised into the programme, and
  district adjustment does not make the remaining comparison causal.

Both estimates, and the clustered interval, in six lines. A reader who sees only the adjusted number cannot tell whether adjustment mattered, and a reader who sees only the naive interval is being told the result is twice as precise as it is.

What comes next

Everything so far has had a continuous outcome. The next lesson takes an outcome that is yes or no — did this referral reach a service — and meets the number this audience misreads more reliably than any other.

Teach this lesson

The lesson as a slide deck, with the prose kept in the speaker notes rather than on the slide. Generated from this page, so it cannot fall out of step with it.

Start the slideshowRead the slides

The PDF needs no software and projects from any machine. The PowerPoint file is there to be edited — add your organisation's branding, cut a section for a shorter session, or merge two lessons into a workshop.