cassionData Analysis

Lesson 8 of 8

Making it run again next quarter

Turn a working analysis into one that reruns on a new export without editing, fails loudly when the input changes, and can be handed to someone who has none of your context.

PythonR60 minResults-Based Management (RBM)

The test

Next quarter’s export arrives. How many things do you have to edit before the report regenerates?

If the answer is more than one — the input path — the analysis is not reproducible, it is a record of one afternoon. This lesson gets the answer down to one.

Split the analysis into stages that hand off files

Three scripts, each with one job, each writing a file the next one reads.

scripts/
  01-read.py        raw export      -> data/interim/muac-typed.parquet
  02-clean.py       typed           -> data/processed/muac-clean.parquet
                                       output/tables/cleaning-log.csv
  03-indicators.py  clean           -> output/tables/gam-by-commune.csv

The value is not tidiness. It is that when the indicator table looks wrong you can rerun stage three in two seconds instead of rerunning a fifteen-minute notebook, and that a colleague debugging stage two does not have to understand stage three.

Use Parquet rather than CSV between stages. It preserves types — which means the work you did in lesson 3 to get child_id read as text is not undone the moment you write it out and read it back.

muac.to_parquet("data/interim/muac-typed.parquet", index=False)
arrow::write_parquet(muac, "data/interim/muac-typed.parquet")

Put the settings in one place

Every threshold, path and cut-off goes at the top of the file or in a config, never buried in the middle of an expression.

# config.py
from pathlib import Path

RAW = Path("data/raw/muac-screening-artibonite-2024.v1.csv")
INTERIM = Path("data/interim/muac-typed.parquet")
PROCESSED = Path("data/processed/muac-clean.parquet")
TABLES = Path("output/tables")

SAM_MM = 115
GAM_MM = 125
PLAUSIBLE_MUAC = (80, 220)
AGE_RANGE_MONTHS = (6, 59)
MISSING_CODE = "-99"
# config.R
RAW       <- "data/raw/muac-screening-artibonite-2024.v1.csv"
INTERIM   <- "data/interim/muac-typed.parquet"
PROCESSED <- "data/processed/muac-clean.parquet"
TABLES    <- "output/tables"

SAM_MM           <- 115
GAM_MM           <- 125
PLAUSIBLE_MUAC   <- c(80, 220)
AGE_RANGE_MONTHS <- c(6, 59)
MISSING_CODE     <- "-99"

A reviewer who wants to know what thresholds you used reads one file. A successor adapting this to a different country changes one file. And when the WHO revises a cut-off, you change one number rather than searching for 125 across four scripts and missing the one written as 12.5 * 10.

Make the pipeline fail loudly

An analysis that silently produces a wrong number is worse than one that crashes. Assert what must be true, at every stage boundary.

def check_input(df):
    expected = {
        "child_id", "commune", "screening_date", "age_months",
        "sex", "muac_mm", "oedema", "outcome",
    }
    missing = expected - set(df.columns)
    if missing:
        raise ValueError(f"export is missing columns: {sorted(missing)}")

    if df["child_id"].isna().any():
        raise ValueError("rows without an identifier")

    measured = df["muac_mm"].dropna()
    if len(measured) == 0:
        raise ValueError("no MUAC measurements at all — wrong file?")

    share_missing = df["muac_mm"].isna().mean()
    if share_missing > 0.20:
        raise ValueError(
            f"{share_missing:.1%} of MUAC missing — above the 20% tolerance. "
            "Investigate before reporting."
        )
check_input <- function(df) {
  expected <- c("child_id", "commune", "screening_date", "age_months",
                "sex", "muac_mm", "oedema", "outcome")
  missing <- setdiff(expected, names(df))
  if (length(missing) > 0) {
    stop("export is missing columns: ", paste(missing, collapse = ", "))
  }

  if (any(is.na(df$child_id))) stop("rows without an identifier")

  share_missing <- mean(is.na(df$muac_mm))
  if (share_missing > 0.20) {
    stop(sprintf(
      "%.1f%% of MUAC missing - above the 20%% tolerance. Investigate before reporting.",
      share_missing * 100
    ))
  }
}

The threshold check is the interesting one. It encodes an editorial judgement — that above 20% missing, the rate is not worth reporting — and it enforces it automatically, on every future export, including the ones processed by someone who never read this lesson.

Write the assertion for the failure you would be embarrassed to miss, not for the failure you think is likely. The likely ones you will notice.

Parameterise instead of copying

Twelve communes, twelve commune reports. Do not copy the script twelve times.

import sys
from pathlib import Path

def build_report(commune: str | None = None):
    df = pd.read_parquet(PROCESSED)
    if commune:
        df = df[df["commune"] == commune]
        if df.empty:
            raise ValueError(f"no rows for commune {commune!r}")

    table = indicator_table(df, "age_band")
    name = commune.lower().replace(" ", "-") if commune else "all"
    table.to_csv(TABLES / f"gam-{name}.csv", index=False)
    return table


if __name__ == "__main__":
    target = sys.argv[1] if len(sys.argv) > 1 else None
    build_report(target)
build_report <- function(commune = NULL) {
  df <- arrow::read_parquet(PROCESSED)
  if (!is.null(commune)) {
    df <- filter(df, commune == !!commune)
    if (nrow(df) == 0) stop("no rows for commune ", commune)
  }

  table <- indicator_table(df, "age_band")
  name <- if (is.null(commune)) "all" else tolower(gsub(" ", "-", commune))
  readr::write_csv(table, file.path(TABLES, paste0("gam-", name, ".csv")))
  table
}

args <- commandArgs(trailingOnly = TRUE)
build_report(if (length(args) > 0) args[1] else NULL)

The same pattern extends to the full narrative report. Quarto renders one template with a parameter, twelve times, producing twelve documents that cannot drift from one another because there is only one document.

What never goes in the repository

This matters more in this sector than in most. A repository holding programme data can hold beneficiary names, GPS points that identify a household, or a protection case list.

  • Never commit raw beneficiary data. Not identified, not pseudonymised. Git keeps every version forever, and deleting a file does not remove it from the history.
  • Never commit credentials. DHIS2 tokens, KoboToolbox API keys, database passwords. Read them from the environment.
  • Add data/raw/ to .gitignore and document where the real file lives — a shared drive, an encrypted volume — in the README.
data/raw/
data/interim/
data/processed/
.env
*.Rhistory
__pycache__/
.Rproj.user/

The analysis code is what belongs in version control. The data belongs wherever your organisation’s data protection policy says it belongs, and the README should say which.

The handover README

Written for someone with your job title and none of your context. Five headings.

# MUAC screening analysis, Artibonite

## What this produces
Quarterly GAM and SAM rates by commune, sex and age band, with 95% intervals.
Feeds the CMAM site allocation decision and the quarterly donor report.

## Where the data comes from
CommCare export, project `artibonite-nutrition`, form `Community screening`.
Exported monthly by the M&E assistant to the shared drive at <path>.
Not in this repository.

## How to run it
    uv sync
    uv run python scripts/01-read.py
    uv run python scripts/02-clean.py
    uv run python scripts/03-indicators.py

## Decisions you should know about
See output/tables/cleaning-log.csv. The one that matters: rows with missing
age are kept, because 60% of the Gros-Morne week of 10-14 June is missing age
and dropping them shifts the commune ranking.

## Who to ask
Indicator definitions: <name>, M&E Manager.
Data collection issues: <name>, Field Coordinator.

That README is the difference between a successor continuing your work and a successor rebuilding it in Excel because they could not tell what your scripts assumed.

Where to go from here

You can now take a routine programme export from arrival to a defended indicator table, in either language, with the reasoning recorded. That is the foundation the rest of the programme builds on:

  • Data Cleaning and Validation goes deeper into the profiling and correction work of lessons 5 and 6, including fuzzy matching and plausibility rule sets.
  • Indicator Design and the LogFrame starts where lesson 7’s reference sheet ends, and works back to the theory of change the indicator is supposed to evidence.
  • Survey Analysis, Sampling and Weighting covers what this course did not: what changes when the data is a probability sample rather than a census of those screened.

Practise on a different dataset before moving on. The WASH household survey and the routine vaccination coverage register on this platform both carry a different mix of defects, and working one of them end to end without the lesson in front of you is the real test of whether this stuck.

Teach this lesson

The lesson as a slide deck, with the prose kept in the speaker notes rather than on the slide. Generated from this page, so it cannot fall out of step with it.

The PDF needs no software and projects from any machine. The PowerPoint file is there to be edited — add your organisation's branding, cut a section for a shorter session, or merge two lessons into a workshop.