Course · Intermediate · Sector Analysis
Public Health and Applied Epidemiology
Incidence, prevalence, cascades and coverage for HIV, TB, malaria and immunisation programmes, plus the outbreak curve you may have to draw at short notice.
What you will be able to do
- Distinguish a rate, a ratio and a proportion, and say what person-time adds that a headcount denominator cannot
- Build a cascade as chained denominators and locate the step where the largest share is lost
- Draw an epidemic curve from a line list, and say what the missing onset dates do to its peak
- Compute attack rates and case fatality with their denominators stated, against the thresholds the response is judged on
- Age-standardise two districts and say how much of the apparent difference between them was structure
- Name the confounders in a routine-data comparison, and state plainly what an observational finding can and cannot claim
Standards and methodologies
Epidemiology is the discipline of denominators, and this course is where the platform’s standing instruction — always state how an indicator is calculated — meets the measures that have names and published thresholds.
It runs on two registers. The cholera line list is 975 cases recorded one at a time over sixteen weeks, shipped with the population they came from, which is the only dataset here that supports an epidemic curve, an attack rate and case fatality at once. The vaccination extract has appeared in six earlier courses and this one finishes it: coverage three ways, and the dropout between doses that says whether children who start a schedule complete it.
Three results shape the course, all computed from those files.
Age structure is half the difference between the worst and best district. Crude cholera attack rates are 7.94 and 5.51 per 1,000. Standardised to a common population they are 7.25 and 5.95, so the gap narrows from 2.43 to 1.30. Half the apparent difference was who lives there, not what happened to them.
Case fatality is 3.90% against a 1% target, and 6.30% in one district against 2.00% in another. The predictor in this data is onset-to-admission delay, which is a service question rather than a clinical one.
The epidemic curve is not the outbreak. Forty-seven cases have no onset date, and one district’s register is back-filled from the admission book, so its curve is really an admission curve shifted right.
Both Python and R throughout. You should have done module 3 — this course assumes you can defend a denominator, put an interval on a proportion, and read a routine figure through its reporting rate.
Practice
Reading the lessons is not the same as having done the work. Each of these applies the course to a dataset it did not teach on.
- Lab · 150 minThe rate the headcount hidTurn a treatment register into a cohort, count person-days instead of children, and watch three sites that share a defaulter proportion separate once you account for how long each one held its children.
- Exercise · 75 minForty-seven cases with no onsetThe epidemic curve is due at four o'clock. Forty-seven of 975 cases have no onset date and they are not spread evenly. Draw the curve three ways, say which one you would send, and what the missing dates could still be hiding.
Progress
Take it offline
The whole course as a typeset PDF — every lesson, every code example, the data dictionary and the indicator definitions. Generated from the same source as this page.
The LaTeX source ships alongside each PDF, so an organisation can rebrand the handout or fold a lesson into its own training pack.