Lesson 8 of 8
Unit · What the evaluation claims
The evaluation that declines the question
The commission asked whether school feeding raises learning. The honest answer is that this design could not have detected an effect of the size the programme cares about, that what it did detect is nothing, and that both statements have to appear in the summary rather than the annex.
Assemble what the course produced
Seven lessons, one programme, one file. Here is everything the analysis established.
| Question | Answer | Where it came from |
|---|---|---|
| Did scores rise? | +6.97 pts, 95% CI +6.1 to +7.9 | Before-after, lesson 1 |
| Is that the programme? | No — comparison schools rose 7.62 | Lesson 1 |
| Are the groups comparable? | No — six of six characteristics imbalanced | Balance table, lesson 2 |
| What is the effect? | −0.96 pts, 95% CI −3.5 to +1.6 | Difference-in-differences, lesson 3 |
| Does matching change it? | −0.71 pts, on 9 of 15 treated schools | Lesson 4 |
| Could a better design run? | Not on this data — no outcome past the cut-off | Lesson 5 |
| What could it have detected? | Nothing below 4.9 points | Power, lesson 6 |
Two of those seven rows are the report and the other five are the annex. The finding is the fourth row and the seventh, and the seventh is the one that has to lead.
The summary that is honest
School feeding and learning outcomes: evaluation summary
FINDING
This evaluation cannot answer whether the school feeding programme
improves literacy, and the reason is in its design rather than its
results.
The programme is delivered in 15 schools and compared against 9. With
fifty children per school and the observed clustering, the smallest
effect this comparison could have detected is 4.9 percentage points. The
programme team's threshold for expansion is 3 points. The evaluation was
therefore incapable of answering the question it was commissioned to
answer, and this was knowable before it began.
What it does establish: literacy rose 7.0 points over the school year in
both programme and comparison schools, with no difference between them
(-0.96 points, 95% CI -3.5 to +1.6). Effects larger than 3.5 points in
either direction are ruled out. Effects between 0 and 3 points -- the
range the programme cares about -- are not.
RECOMMENDATION
Do not discontinue the programme on the basis of this evaluation. Do not
expand it on the basis of the 7-point gain, which is the school year and
is present in schools without the programme.
To answer the question, an evaluation needs about 60 schools rather than
24. If the programme is expanding to new schools, randomising the order of
the roll-out costs nothing and produces a comparison group that does not
have to be argued for.
The first paragraph refuses the question and the last paragraph says what would answer it. Between them there is a finding, an interval and a boundary, and no number is presented as an effect that is not one.
Four sentences an evaluation must be able to write
Each one is a check on the report, and a report that cannot produce all four has a gap.
“The counterfactual is X.” Here: schools without the programme, assumed to have been on the same trajectory.
“The estimate is Y, with interval Z, in units the programme uses.” Here: −0.96 percentage points of literacy, −3.5 to +1.6, where 3 points is the threshold of interest.
“The smallest effect this design could detect is W.” Here: 4.9 points. This is the sentence most evaluations lack, and its absence is what lets a null be read as a refutation.
“This analysis cannot rule out V.” Here: any effect between zero and about three points, which is the whole range the programme is arguing about.
What “no effect detected” is worth
It is worth something, and being precise about how much is the difference between a useful null and a damaging one.
| Reading | Supported? |
|---|---|
| “The programme does not work” | No — the design could not detect the relevant effect |
| “The programme has no large effect on literacy” | Yes — above 3.5 points is ruled out |
| “The 7-point gain is not attributable to the programme” | Yes — comparison schools gained more |
| “Funding should be redirected” | No — that requires an effect estimate this study does not have |
| “The next evaluation needs 60 schools” | Yes — computed, and actionable |
The two “yes” rows in the middle are real findings and they are unwelcome ones: they take away a headline number a programme had been using. Delivering that is the job, and doing it alongside the recommendation in the last row is what makes it receivable.
Writing for the reader who will quote one sentence
Someone will extract one line from this report and put it in a slide. Decide now which line that is.
Not: “No statistically significant effect of school feeding on literacy was found.” True, and it will be read as “the programme does not work”.
Not: “Literacy rose 7 percentage points in programme schools.” True, and it will be read as the effect.
This: “Literacy rose 7 points in both programme and comparison schools; this evaluation was not large enough to detect the 3-point difference the programme cares about.”
Put that sentence in the summary, the conclusion and the covering email. The quotable line is chosen by you or it is chosen for you.
What the course has been about
Eight lessons and one idea: an impact claim is a comparison with an absent half, and the design is the argument about what fills it.
Every method here — randomisation, difference-in-differences, matching, discontinuity — is a different argument for the same missing quantity, and each is stronger or weaker depending on facts about how the programme was rolled out rather than on anything in the analysis.
Which means the most consequential decisions were made before any data existed: who got the programme and in what order, what was measured and on whom, and how many units were assigned. An analyst who arrives afterwards is choosing among the arguments the roll-out left available.
So the useful thing this course asks for is not a technique. It is being in the room in January, with one page, asking how the schools will be chosen and whether the order can be randomised — which is a question that costs nothing to ask and is unanswerable a year later.
Report it whole
Evaluation of the school feeding programme: methods and limitations
Design Difference-in-differences, 15 programme and 9
comparison schools, baseline and endline literacy on
585 students with both rounds.
Counterfactual Comparison schools, assumed to have been on the same
trajectory. Two rounds exist, so parallel trends is
argued (numeracy shows the same null; no single school
drives the result) and not tested.
Estimate -0.96 percentage points, 95% CI -3.5 to +1.6,
clustered on 24 schools.
Power Minimum detectable effect 4.9 points at 80% power,
against a 3-point threshold of interest. Underpowered
for the relevant effect size.
Balance All six baseline characteristics imbalanced beyond 0.25
standardised. Adjusted for baseline and district;
unmeasured selection is not addressed.
Attrition 585 of 741 baseline students have an endline. Those who
left scored 39.3% at baseline against 56.8% for those
who stayed.
Not claimed Any effect on attendance, enrolment or nutrition. Any
causal effect beyond the difference-in-differences
assumption. Any conclusion about effects below 3 points.
Plan Written 2024-01-15, one deviation listed in Annex A,
three exploratory analyses in Annex B.
Eight rows, and a reviewer can reconstruct every decision from them. That is the deliverable this course has been building toward: not a number, but a number surrounded by everything a reader needs in order to know what it is.
What comes next
Module 5 is complete: uncertainty, then models, then designs. Module 6 is about delivery — turning what you have established into something a programme can see, use and rebuild, which is the last thing standing between an analysis and a decision.