cassionData Analysis

Back to the lessonLesson 4 of 8What Sphere asks that a ladder does not

Tested on a third, censored on a fifth

The same deck as the downloads, rendered as a page. Start the slideshow to present it full screen — arrow keys or a click advance one slide, Escape leaves.

Slides · PDFSlides · PowerPoint

  1. Slide 1 / 26

    What this lesson covers

    • The denominator changes, and the report usually does not
    • The JMP risk classes
    • Improved is not safe, and this is where you can prove it
    • The censored value, and why it is not missing
    • What to do with a censored value
    • The cross-field inconsistency worth naming
    • Report it as its own block
    • What comes next
    Speaker notes
    E. coli was tested on 33.4% of households, chlorine on 40.0%, and both on 13.3%. In the water point register, 271 chlorine readings are the string "<0.1" — and dropping them raises apparent compliance by sixteen points.
  2. Slide 2 / 26

    The denominator changes, and the report usually does not — In Python

    import pandas as pd
    
    households = pd.read_csv("wash-household-survey-2024.v1.csv")
    
    ecoli = households["ecoli_cfu_100ml"].notna()
    chlorine = households["free_residual_chlorine_mgl"].notna()
    
    print(f"E. coli tested:   {ecoli.sum():>5} = {ecoli.mean():.1%}")
    print(f"chlorine tested:  {chlorine.sum():>5} = {chlorine.mean():.1%}")
    print(f"both:             {(ecoli & chlorine).sum():>5} = {(ecoli & chlorine).mean():.1%}")
    Speaker notes
    Water quality is expensive to measure. A survey that interviews 2,403 households tests a fraction of them, and every quality indicator therefore runs on a different denominator from every access indicator.
  3. Slide 3 / 26

    The denominator changes, and the report usually does not — In R

    library(dplyr)
    
    households |> summarise(
      ecoli = mean(!is.na(ecoli_cfu_100ml)),
      chlorine = mean(!is.na(free_residual_chlorine_mgl)),
      both = mean(!is.na(ecoli_cfu_100ml) & !is.na(free_residual_chlorine_mgl))
    )
  4. Slide 4 / 26

    The denominator changes, and the report usually does not

    • 802 households for E. coli, 961 for chlorine, and 320 for both — The last number is the one that matters: any question…
    • Print the analysable n for every quality figure — in the table rather than in a footnote
    Speaker notes
    802 households for E. coli, 961 for chlorine, and 320 for both. The last number is the one that matters: any question relating the two runs on 13.3% of the survey, and an interval on it will be wide. Print the analysable n for every quality figure, in the table rather than in a footnote. A quality percentage next to an access percentage on the same row reads as the same denominator, and here it never is.
  5. Slide 5 / 26

    The JMP risk classes — In Python

    tested = households[ecoli]
    classes = pd.cut(
        tested["ecoli_cfu_100ml"],
        [-1, 0, 10, 100, 10**9],
        labels=["safe <1", "low 1-10", "moderate 11-100", "high >100"],
    )
    print(classes.value_counts(normalize=True).round(3))
  6. Slide 6 / 26

    The JMP risk classes — In R

    households |>
      filter(!is.na(ecoli_cfu_100ml)) |>
      mutate(risk = cut(ecoli_cfu_100ml, c(-1, 0, 10, 100, Inf),
                        labels = c("safe", "low", "moderate", "high"))) |>
      count(risk) |> mutate(share = n / sum(n))
  7. Slide 7 / 26

    The JMP risk classes

    Risk classHouseholdsShare of tested
    Safe (<1 CFU/100 mL)43053.6%
    Low (1–10)17321.6%
    Moderate (11–100)13917.3%
    High (>100)607.5%
  8. Slide 8 / 26

    The JMP risk classes

    • Report the classes, not a mean — E
    Speaker notes
    Report the classes, not a mean. E. coli counts are heavily skewed and a mean of 14 CFU/100 mL describes nothing — the classes are what the guideline is written in and what a decision is taken on.
  9. Slide 9 / 26

    Improved is not safe, and this is where you can prove it — In Python

    IMPROVED = {"piped-into-dwelling", "piped-into-yard", "public-tap",
                "borehole", "protected-well", "protected-spring"}
    
    tested = tested.assign(improved=tested["water_source"].isin(IMPROVED))
    print(tested.groupby("improved")["ecoli_cfu_100ml"].agg(
        n="size", safe=lambda s: (s < 1).mean(), high=lambda s: (s > 100).mean()
    ).round(3))
  10. Slide 10 / 26

    Improved is not safe, and this is where you can prove it — In R

    households |>
      filter(!is.na(ecoli_cfu_100ml)) |>
      mutate(improved = water_source %in% improved) |>
      summarise(n = n(), safe = mean(ecoli_cfu_100ml < 1),
                high = mean(ecoli_cfu_100ml > 100), .by = improved)
  11. Slide 11 / 26

    Improved is not safe, and this is where you can prove it

    SourceTestedSafeHigh risk
    Improved62561.1%3.8%
    Unimproved or surface17727.1%20.3%
    Speaker notes
    Source type predicts quality strongly, and 38.9% of households on an improved source still have detectable E. coli. Lesson 1 said the top rung needs freedom from contamination as well as an improved source; this is the number that shows why it is a separate condition rather than a formality.
  12. Slide 12 / 26

    The censored value, and why it is not missing — In Python

    points = pd.read_csv("water-point-monitoring-2024.v1.csv")
    readings = points["free_residual_chlorine_mgl"].dropna()
    
    censored = readings.astype(str).str.startswith("<")
    print(f"tested {len(readings)}, censored {censored.sum()}")
    print(readings[censored].unique())
    Speaker notes
    The water point register measures chlorine in the field, and the kit has a detection limit.
  13. Slide 13 / 26

    The censored value, and why it is not missing — In R

    points |>
      filter(!is.na(free_residual_chlorine_mgl)) |>
      count(censored = startsWith(free_residual_chlorine_mgl, "<"))
  14. Slide 14 / 26

    The censored value, and why it is not missing — In Python

    numeric = pd.to_numeric(readings, errors="coerce")   # <-- the silent one
    print(f"after coercion: {numeric.isna().sum()} became NaN")
    
    in_range = readings[~censored].astype(float).between(0.2, 0.5)
    print(f"in target range, over all tested:   {in_range.sum() / len(readings):.1%}")
    print(f"in target range, dropping censored: {in_range.mean():.1%}")
    Speaker notes
    811 visits carry a reading and 271 of them are the string <0.1. That is a measurement: the true value lies between zero and the detection limit. It is not a missing value, and the two things it is usually turned into are both wrong.
  15. Slide 15 / 26

    The censored value, and why it is not missing — In R

    points |>
      filter(!is.na(free_residual_chlorine_mgl),
             !startsWith(free_residual_chlorine_mgl, "<")) |>
      summarise(in_range = mean(between(as.numeric(free_residual_chlorine_mgl), 0.2, 0.5)))
  16. Slide 16 / 26

    The censored value, and why it is not missing

    • 30.9% against 46.5% — Dropping the censored readings raises apparent compliance with the chlorination target by nearly…
    Speaker notes
    30.9% against 46.5%. Dropping the censored readings raises apparent compliance with the chlorination target by nearly sixteen points, and it does it in exactly one direction — every censored value is a failure, so removing them removes only failures. pd.to_numeric(..., errors="coerce") does this silently and looks like a type fix.
  17. Slide 17 / 26

    What to do with a censored value

    • Classify rather than average — The target is a range, so a censored reading answers the question perfectly: <0.1 is…
    • Substitute the limit, or half of it — if you genuinely need a mean
    • Report the share below the limit as its own number — 33.4% of these readings are below detection, which is a finding…
    Speaker notes
    Three options, in order of preference for this indicator. Classify rather than average. The target is a range, so a censored reading answers the question perfectly: <0.1 is below 0.2 and therefore non-compliant. Compliance is computable on all 811 readings with no substitution at all. Substitute the limit, or half of it, if you genuinely need a mean. State which you used — LOD/2 is the common convention — and report how many values it applied to. Report the share below the limit as its own number. 33.4% of these readings are below detection, which is a finding about chlorination practice rather than a nuisance in the data.
  18. Slide 18 / 26

    What to do with a censored value — In Python

    compliant = pd.to_numeric(readings.where(~censored), errors="coerce").between(0.2, 0.5)
    result = pd.DataFrame({
        "readings": [len(readings)],
        "below detection": [censored.sum()],
        "in 0.2-0.5 range": [compliant.sum()],
        "compliance": [f"{compliant.sum() / len(readings):.1%}"],
    })
    print(result)
  19. Slide 19 / 26

    What to do with a censored value — In R

    # The denominator is every reading, including the ones below the limit.
  20. Slide 20 / 26

    What to do with a censored value

    • Never let a detection limit shrink a denominator — It is the single most common way a water quality report becomes…
    Speaker notes
    Never let a detection limit shrink a denominator. It is the single most common way a water quality report becomes optimistic, and it survives review because the arithmetic on the remaining values is correct.
  21. Slide 21 / 26

    The cross-field inconsistency worth naming — In Python

    treats = households["water_treated_at_home"]
    print(f"treat at home: {treats.sum()}")
    print(f"  of which no chlorine reading: {(treats & ~chlorine).sum()}")
  22. Slide 22 / 26

    The cross-field inconsistency worth naming — In R

    households |> filter(water_treated_at_home) |>
      count(no_reading = is.na(free_residual_chlorine_mgl))
    Speaker notes
    533 households report treating their water and have no chlorine measurement. The cleaning course established what this is — missingness that correlates with the thing being measured — and the consequence here is specific: dropping those rows removes households that treat their water more often than average, so the tested subsample is not representative of the survey it came from.
  23. Slide 23 / 26

    Report it as its own block — Example (cont.)

    Water quality
    
      E. coli, point of collection          802 households tested (33.4% of survey)
        Safe <1 CFU/100 mL                  53.6%
        Low 1-10                            21.6%
        Moderate 11-100                     17.3%
        High >100                            7.5%
        Improved sources: 61.1% safe (n=625)
        Unimproved and surface: 27.1% safe (n=177)
    
      Free residual chlorine, water points  811 visits tested (30.8% of visits)
        Below detection limit (<0.1 mg/L)   33.4%   counted as non-compliant
        Within 0.2-0.5 mg/L target          30.9%
    
      Quality was tested on a subsample and does not share the denominator of the
      access indicators above. 533 households reporting home treatment have no
  24. Slide 24 / 26

    Report it as its own block — Example (cont.)

      chlorine reading, so the tested subsample over-represents untreated water.
    Speaker notes
    Its own block, its own denominators, its own limitation. A quality figure placed in a table of access figures will be read against their denominator, and nothing in the layout will stop it.
  25. Slide 25 / 26

    What comes next

    • Everything so far is a cross-section: what was true on the day someone called.
    Speaker notes
    Everything so far is a cross-section: what was true on the day someone called. The next unit is the register that visits the same water points twelve times, and the first thing it does is turn one functionality rate into three.
  26. Slide 26 / 26

    Where this goes next

    Read the full lesson, with runnable code Back to the lesson