cassionData Analysis

Back to the lessonLesson 7 of 8What the comparison can claim

What else could explain it

The same deck as the downloads, rendered as a page. Start the slideshow to present it full screen — arrow keys or a click advance one slide, Escape leaves.

Slides · PDFSlides · PowerPoint

  1. Slide 1 / 18

    What this lesson covers

    • The definition, and why it is worth being strict about
    • The four candidates in this outbreak
    • Stratify to see it
    • When stratification runs out
    • The three questions to ask of any comparison
    • What comes next
    Speaker notes
    A confounder causes the outcome and differs between the groups. Four of them sit between the districts in this outbreak, one has been removed, one is unmeasured, and one is not a confounder at all — it is the mechanism.
  2. Slide 2 / 18

    The definition, and why it is worth being strict about

    • it is associated with the exposure — it differs between the groups being compared;
    • it is a cause of the outcome, independently of the exposure;
    • it is not on the causal path between the exposure and the outcome.
    Speaker notes
    A confounder is a variable that satisfies three conditions at once: All three matter. A variable that differs between groups but does not cause the outcome is irrelevant. A variable that causes the outcome but is the same in both groups cannot explain a difference. And a variable on the causal path is not a confounder — it is the mechanism, and adjusting for it destroys the finding.
  3. Slide 3 / 18

    The four candidates in this outbreak

    CandidateDiffers between districts?Causes the outcome?On the path?
    Age structureYes, sharplyYesNo — a confounder, already removed
    Water source and densityAlmost certainlyYesNo — a confounder, unmeasured
    Distance to treatmentYesYes, for fatalityYes — it is the mechanism
    Register qualityYesNoNo — not a confounder, a bias
    Speaker notes
    Nord has a standardised attack rate of 7.25 per 1,000 against Sud's 5.95, and case fatality of 6.30% against 3.11%. What else could explain those gaps?
  4. Slide 4 / 18

    The four candidates in this outbreak

    • Age structure — has been handled
    • Water source, sanitation and crowding — are the actual causes of cholera transmission, they certainly differ between a…
    Speaker notes
    Work down that table and each row needs a different response. Age structure has been handled. Lesson 6 removed it and quantified what it was worth — about half the crude gap. Water source, sanitation and crowding are the actual causes of cholera transmission, they certainly differ between a camp-like district and a rural one, and they are not in this dataset. That is the most important sentence in this lesson. An unmeasured confounder cannot be adjusted for, and its existence has to be stated rather than ignored.
  5. Slide 5 / 18

    The four candidates in this outbreak — In Python

    print("Available for adjustment: age, sex, district")
    print("Known to matter and unavailable: water source, sanitation, crowding, "
          "displacement status")
  6. Slide 6 / 18

    The four candidates in this outbreak — In R

    # The list of what you could not adjust for belongs in the limitations section.
  7. Slide 7 / 18

    The four candidates in this outbreak

    • Distance to treatment — is the interesting one and it is not a confounder at all
    • Register quality — is not a confounder either
    Speaker notes
    Distance to treatment is the interesting one and it is not a confounder at all. The chain is: Nord is further from treatment centres → its cases arrive later → more of them die. Distance causes fatality through delay, so delay is on the causal path. Adjusting for delay would remove the very effect you are trying to demonstrate, which is the classic over-adjustment error. Register quality is not a confounder either. Nord's back-filled onset dates do not cause deaths; they distort the measurement of the explanation. That is information bias, and it is fixed by fixing the register, not by adjusting.
  8. Slide 8 / 18

    Stratify to see it — In Python

    import pandas as pd
    
    cases = pd.read_csv("cholera-line-list-2024.v1.csv")
    with_outcome = cases[cases["outcome"].notna()]
    
    stratified = (
        with_outcome.assign(died=with_outcome["outcome"] == "died")
        .groupby(["district", "age_band"])
        .agg(cases=("died", "size"), deaths=("died", "sum"))
    )
    stratified["cfr"] = stratified["deaths"] / stratified["cases"]
    print((stratified["cfr"].unstack() * 100).round(1))
    Speaker notes
    The simplest adjustment, and the one that shows its working.
  9. Slide 9 / 18

    Stratify to see it — In R

    library(dplyr)
    
    cases |>
      filter(!is.na(outcome)) |>
      summarise(cases = n(), cfr = mean(outcome == "died"), .by = c(district, age_band)) |>
      tidyr::pivot_wider(id_cols = age_band, names_from = district, values_from = cfr)
  10. Slide 10 / 18

    Stratify to see it

    • Does the difference persist within strata? — If Nord's case fatality is higher in every age band, age is not the…
    • Are the strata consistent? — A difference that is large in one band and absent in others is effect modification — the…
    • Effect modification is a finding; confounding is a nuisance — Collapsing the first into a single number destroys the…
    Speaker notes
    Two things to read off a stratified table, and only one of them is the adjusted estimate. Does the difference persist within strata? If Nord's case fatality is higher in every age band, age is not the explanation. If it reverses in some bands, something more interesting is happening. Are the strata consistent? A difference that is large in one band and absent in others is effect modification — the exposure genuinely acts differently in different groups — and it must be reported by stratum rather than averaged into a single adjusted figure. Effect modification is a finding; confounding is a nuisance. Collapsing the first into a single number destroys the result; failing to remove the second manufactures one.
  11. Slide 11 / 18

    When stratification runs out — In Python

    cells = with_outcome.groupby(["district", "age_band", "sex"]).size()
    print(f"{len(cells)} cells, smallest {cells.min()}, {(cells < 30).sum()} below 30")
    Speaker notes
    Two variables and the cells empty fast — the disaggregation lesson from module 3, arriving with a different consequence. Three districts by four age bands by two sexes is twenty-four cells on 974 outcomes, and cholera deaths are rare.
  12. Slide 12 / 18

    When stratification runs out — In R

    cases |> filter(!is.na(outcome)) |> count(district, age_band, sex) |>
      summarise(cells = n(), smallest = min(n), under_30 = sum(n < 30))
    Speaker notes
    That is where regression takes over, and Regression for Programme Data later in the programme is about exactly this — adjusting for several variables at once without running out of cells. Its purpose is the same as stratification's, and a model that adjusts for a mechanism makes the same error as a stratification that does.
  13. Slide 13 / 18

    The three questions to ask of any comparison

    • 1. What differs between these groups other than the thing I am studying? — List them
    • 2. Which of them cause the outcome? — Only those are confounders
    • 3. Which of them are on the causal path? — Do not adjust for those
    Speaker notes
    Before attributing a difference to anything: 1. What differs between these groups other than the thing I am studying? List them. The list is the limitations section, and the ones you cannot measure belong in it too. 2. Which of them cause the outcome? Only those are confounders. A district having more schools does not confound a cholera comparison unless schools cause cholera. 3. Which of them are on the causal path? Do not adjust for those. Delay to treatment, in this outbreak, is how distance kills.
  14. Slide 14 / 18

    The three questions to ask of any comparison — In Python

    adjustment_plan = pd.DataFrame({
        "variable": ["age", "sex", "water source", "delay to treatment", "register quality"],
        "differs": [True, False, True, True, True],
        "causes_outcome": [True, True, True, True, False],
        "on_causal_path": [False, False, False, True, False],
        "action": ["adjust", "no need", "unmeasured; state it", "do not adjust",
                   "fix the register"],
    })
    print(adjustment_plan)
  15. Slide 15 / 18

    The three questions to ask of any comparison — In R

    tibble::tribble(
      ~variable,        ~action,
      "age",            "adjust",
      "water source",   "unmeasured; state it",
      "delay",          "do not adjust; it is the mechanism",
      "register",       "fix it; this is bias, not confounding"
    )
  16. Slide 16 / 18

    The three questions to ask of any comparison

    • Write that table before the analysis, not after — Deciding what to adjust for once you have seen which adjustment gives…
    Speaker notes
    Write that table before the analysis, not after. Deciding what to adjust for once you have seen which adjustment gives a nicer answer is the thing that makes observational epidemiology untrustworthy, and it is indistinguishable from honest work after the fact.
  17. Slide 17 / 18

    What comes next

    • You have named what else could explain a difference and removed what you could.
    Speaker notes
    You have named what else could explain a difference and removed what you could. The last lesson is what remains sayable — the sentence an observational comparison is entitled to, and the several it is not.
  18. Slide 18 / 18

    Where this goes next

    Read the full lesson, with runnable code Back to the lesson