# Indicator definitions — WASH coverage and water quality by district

Technical documentation for the project *WASH coverage and water quality by
district*. It is the document to open when someone challenges a number in the
report: every indicator here states its numerator, its denominator, the
disaggregation it is valid at, and the decision it informs.

- **Source dataset:** `wash-household-survey-2024.v1.csv` (2 403 households, 18
  communities, 3 districts)
- **Analysis:** `notebooks/wash-ladders.python.en.ipynb`
- **Decision informed:** which six of eighteen communities receive water point
  rehabilitation in the next funding cycle
- **Standards applied:** WHO/UNICEF JMP service ladders, Sphere water quantity
  and access standards, SDG 6.1.1 and 6.2.1

Every dataset on this platform is synthetic. No real household is described.

---

## 1. Why this document exists

The programme previously reported coverage as *the share of households using an
improved water source*. That number is 82.3% in this survey. The share with at
least **basic** drinking water service — the JMP definition, which adds the
collection-time condition — is 56.3%.

The 26.0 points between them are households with a functioning improved source
more than thirty minutes' round trip away. Under the old measure every one of
them was counted as served, and rehabilitation budget went to the districts with
the fewest boreholes rather than to the households with the worst service. The
definitions below are the correction.

---

## 2. Access indicators

### 2.1 Improved source (reported for comparison only)

| | |
| --- | --- |
| **Numerator** | Households whose main drinking water source is piped, borehole, protected well, protected spring, rainwater or packaged water |
| **Denominator** | All surveyed households |
| **Value** | 0.823 |
| **Valid at** | Household, community, district |
| **Decision informed** | None. It is reported so the report can show what changed. |

This is a property of the **source**, not of the household's service. It is
retained in the analysis solely to quantify the gap against the ladder, and the
report must never present it as coverage.

### 2.2 At least basic drinking water service — **the coverage indicator**

| | |
| --- | --- |
| **Numerator** | Households using an improved source **and** a round-trip collection time of 30 minutes or less, including queuing |
| **Denominator** | All surveyed households |
| **Value** | 0.563 |
| **Valid at** | Household, community, district |
| **Decision informed** | The rehabilitation ranking (§4) |

Both conditions must hold. An improved source at more than 30 minutes is
**limited** service, not basic — this is the JMP ladder, not a local convention,
and the distinction is the reason the project was commissioned.

### 2.3 Over the walk limit

| | |
| --- | --- |
| **Numerator** | Households whose round-trip collection time exceeds 30 minutes |
| **Denominator** | All surveyed households |
| **Valid at** | Household, community, district |
| **Decision informed** | The rehabilitation ranking — this is the component a rehabilitation actually moves |

### 2.4 Below the Sphere quantity standard

| | |
| --- | --- |
| **Numerator** | Households with fewer than 15 litres per person per day |
| **Denominator** | All surveyed households |
| **Valid at** | Household, community, district |
| **Decision informed** | The rehabilitation ranking |

15 l/p/d is the Sphere minimum for survival needs including water, sanitation
and hygiene. It is a floor, not a target; a household just above it is not
adequately served.

### 2.5 Unimproved source

| | |
| --- | --- |
| **Numerator** | Households whose main source is unprotected well, unprotected spring or surface water |
| **Denominator** | All surveyed households |
| **Valid at** | Household, community, district |
| **Decision informed** | The rehabilitation ranking |

---

## 3. Water quality indicators — a different denominator

Quality was tested on a subsample. **These indicators do not share a denominator
with the access indicators and must never be placed in the same table without
the sample size beside them.**

| Indicator | Numerator | Denominator |
| --- | --- | --- |
| Chlorine tested | Households with a free residual chlorine reading | All surveyed households (≈33%) |
| *E. coli* tested | Households with an *E. coli* result | All surveyed households (≈32%) |
| *E. coli* detected | Households with any *E. coli* detected | Households **with an *E. coli* result** |

The detection rate is conditional on being tested. Reporting "57% *E. coli*
positive in SU-02" without "of the 41% of households tested" invites the reader
to apply it to the whole community.

**Quality is reported beside the selection, never inside it.** Contaminated
water needs treatment or a new source; that is a different intervention from
shortening a walk, funded from a different line.

---

## 4. The rehabilitation need score

### 4.1 Definition

```
need = z(over_walk_limit) + z(below_sphere) + z(unimproved) - z(basic_water)
```

where `z(·)` standardises the community-level rate across the 18 communities
(mean 0, standard deviation 1). Higher is worse.

| | |
| --- | --- |
| **Unit of analysis** | Community (18) |
| **Denominator per community** | Households surveyed in it (133–136; median 133) |
| **Decision informed** | Which six communities are funded |

### 4.2 The two design choices a reviewer will ask about

**Why sanitation and hygiene are excluded.** A water point rehabilitation
shortens the walk and restores function. It does not build latrines or supply
soap. Including those components would let a community with poor sanitation
displace one with a broken water point — funding a need this budget cannot meet.
Sanitation and hygiene rates are in the report; they are not in the score.

**Why the components are not weighted.** Equal weights say the four components
matter equally. That is an assumption, not a finding. It is stated here rather
than buried, and §4.3 is the check that the conclusion does not depend on it.

### 4.3 Stability check

Each component is dropped in turn and the top six recomputed:

| Component dropped | Of 6 still selected |
| --- | --- |
| basic water | 5 |
| over walk limit | 5 |
| below Sphere | 5 |
| unimproved | 6 |

Five or six of six survive every single-component drop. The list is therefore
defensible in a meeting: no community is on it because of one input. Had the
overlap fallen to three, the composite would not be reportable and the report
would have to say so.

### 4.4 The selection

| Rank | District | Community | Basic water | Over 30 min | Below Sphere | Need |
| ---: | --- | --- | ---: | ---: | ---: | ---: |
| 1 | Sud-Est | SU-02 | 15.0% | 85.0% | 25.6% | 2.020 |
| 2 | Sud-Est | SU-05 | 25.6% | 73.7% | 18.8% | 1.195 |
| 3 | Nord-Ouest | NO-05 | 34.6% | 65.4% | 18.8% | 0.966 |
| 4 | Nord-Ouest | NO-01 | 42.9% | 54.1% | 15.8% | 0.795 |
| 5 | Nord-Ouest | NO-06 | 49.6% | 42.9% | 18.8% | 0.564 |
| 6 | Centre | CE-05 | 37.6% | 60.9% | 13.5% | 0.539 |

Distribution across districts: Nord-Ouest 3, Sud-Est 2, Centre 1. **This is an
output, not a quota.** It is stated in the report because a selection
concentrated in one district reads as a political choice unless the report shows
it came from the ranking.

Note CE-05 ranks sixth on the composite despite a worse walk rate than NO-06 —
its Sphere-quantity and unimproved-source rates are lower. A reviewer who
expects the list to be ordered by any single column should be pointed here.

---

## 5. Cleaning rules applied before any indicator is computed

| Rule | Rationale | Effect |
| --- | --- | --- |
| Normalise `district` case, spacing and hyphenation | One enumerator team wrote Nord-Ouest four ways: `Nord-Ouest` (727), `NORD-OUEST` (33), `Nord Ouest` (22), `nord-ouest` (20) | 6 district values → 3 |

Ungrouped, that split the worst-performing district into four fragments, none of
which looked alarming — the district-level ranking was wrong before any
indicator was computed. The normalisation runs before the first `groupby` in the
notebook, and every district figure in the report is post-normalisation.

---

## 6. Limitations

State these in any report drawing on these indicators.

1. **Collection time is reported, not measured.** Where a community
   systematically under- or over-states the walk, the ranking moves with it, and
   nothing in this dataset can detect it. This is the largest single threat to
   the selection.
2. **The survey does not say why a source is distant or unimproved.** A broken
   pump, a dry borehole and a community that never had one are indistinguishable
   here, and only the first is what a rehabilitation budget buys. Confirm by
   site visit before committing funds.
3. **Communities are compared without weighting**, which is valid only because
   the survey allocated 133–136 households to each. Do not reuse the score on a
   survey with uneven allocation without re-deriving it on weighted rates.
4. **Quality rates rest on roughly a third of households** and are not
   representative of a community by design.
5. **Seasonality is not observed.** A single survey round cannot distinguish a
   dry-season walk from a year-round one.

---

## 7. Reproducing this analysis

```bash
pnpm examples:build                     # regenerate the notebook from its Quarto source
```

Then run `notebooks/wash-ladders.python.en.ipynb` — in Colab via the badge, or
locally with pandas and numpy. It reads the CSV over HTTPS, so no local dataset
copy is needed.

The dataset is versioned by filename. `wash-household-survey-2024.v1.csv` is
immutable; a correction ships as `.v2.csv` and this document is revised beside
it.

---

## 8. Change log

| Date | Change |
| --- | --- |
| 2026-07-27 | First issue, against dataset v1. |
