Path
Data Analyst
The generalist analyst track — take a messy export from any sector, clean it defensibly, and produce the table or figure that answers the question actually asked.
Competencies
Data wrangling
IntermediateReshape, join and aggregate real programme exports without silently changing the row count or inventing rows that were never recorded.
Defensible cleaning
IntermediateDecide what to correct, what to set to missing and what to leave alone, and write down the rule and its effect for each.
Exploratory analysis
IntermediateFind what is in a dataset before deciding what to report from it, including the defects that will change the answer.
Communicating a result
IntermediateProduce the one table or figure that answers the question, and say what it does not establish.
Before you start
- Comfort with data in a spreadsheet — filters, formulas, a pivot table. No programming background is assumed.
- A dataset of your own that you have failed to get an answer out of. Every stage below works better against a file you already care about.
The route through
Stage 1
Ground the numbers
Know what a figure counts before you compute it, so the cleaning decisions in the next stage have something to be right or wrong about.
Stage 2Choose one
Pick the language your team uses
Go deep in one of Python or R until you can read any export and hand over a script that reruns. The second language is far cheaper once you have the first.
These are alternatives, not a sequence. Take the one your team already uses — the second is far cheaper to add once you have the first.
Stage 3
Make the data trustworthy
Turn a raw export into a table you would defend line by line, and hand over a cleaning log that answers the auditor's question before it is asked.
Stage 4
Put the files together
Assemble several exports into one analysis table you would defend column by column, with every join proved, every grain stated and every denominator sourced.
Stage 5
Check it like an auditor
Assess your own data before a donor does — five dimensions with measures, a recount against the source, and a report where every finding has an owner and a date.
Stage 6
Define what you report
Write indicator definitions two analysts compute the same way, defend the denominator, and set a baseline and target that survive a mid-term review.
Stage 7
Put an interval on it
Analyse a cluster survey the way its design requires — weights from the frame, a measured design effect, and an interval that says what the sample can and cannot settle.
Stage 8
Know where it came from
Read a routine reporting system as the database it is, pull an extract you can point at months later, and answer what was counted for any figure it produces.
Stage 9
Apply it to nutrition
Take the methods into one sector — WHO growth standards, SMART plausibility, the IPC phases and the Sphere performance thresholds, on a survey and a treatment register.
Stage 10
Apply it to public health
Rates with person-time denominators, treatment cascades, coverage three ways, and an outbreak line list turned into a curve, attack rates and a case fatality you can defend.
Stage 11
Apply it to WASH
The JMP service ladders and the Sphere minimums on a household survey, then a monitoring register with repeat visits that turns one functionality rate into three and bounds each against the rounds nobody drove.
Stage 12
Apply it to food security
FCS, HHS, rCSI and the livelihood coping module built from raw components and found to disagree by a factor of seven, then a price series whose seasonality dwarfs the programme effect and the evidence table that reports both.
Stage 13
Apply it to protection
The analyses you must decline to publish, alongside the ones that matter — a consent-gated referral pathway, a nineteen-point equity gap located at a specific gate, and a caseload that explains two other tables.
Stage 14
Apply it to education
Gross against net enrolment on a projected denominator, two attendance numbers twenty-six points apart, a cohort through promotion and repetition, and two assessment rounds whose instruments differ.
Stage 15
Say how sure you are
An interval on every proportion, the right test for a comparison, an effect size beside every p-value, and the count of comparisons that turns two striking schools back into noise.
Stage 16
Model more than one thing at once
A coefficient is a comparison — which one, between which units, adjusted for what. An odds ratio your reader will misread, a covariate that removes 42% of the effect, and a model that explains three per cent and settles a targeting decision.
Stage 17
Say what caused it
A seven-point gain that is entirely the school year, a comparison group imbalanced on every characteristic measured, and the minimum detectable effect that decided the answer before any data existed.
Stage 18
Put it in front of them
The mark the comparison implies, an interval that stops a ranking, a palette that already means something to this audience, and a figure generated from the dataset so the chart and the sentence cannot drift.
Stage 19
Make it rerunnable
Raw data read-only and code the only thing edited, an environment pinned so a colleague's laptop gives your numbers, checks that stop the pipeline rather than producing a plausible wrong one, and a handover a successor can act on.
Stage 20
Put it in front of the decision
A page that answers one question rather than twenty, the definition panel that stops the monthly argument, and the report whose findings, limitations and recommendation survive being read separately.
Recommended projects
How the skill is assessed
Can join two datasets and prove the join did not change the row count or duplicate a record.
EvidenceA notebook using an explicit join validation that fails loudly rather than inflating silently.
Can document a cleaning rule with its rationale and its measured effect on the data.
EvidenceA cleaning table naming each rule, why it exists and how many records it touched.
Can distinguish a missing value from a structurally absent one and treat them differently.
EvidenceAn analysis in which a school closure, a non-response and an impossible value are handled as three different things.
Most of this job is not analysis. It is getting from what arrived to something that can be analysed at all, and being able to say afterwards exactly what you changed.
The three projects here were chosen because each turns on one such decision, and in each case the decision changes the answer rather than tidying it. A school closure filled as absence moves the dropout list from 85 students to 130. A district name spelled four ways splits the worst-performing district into four fragments, none of which looks alarming. A partial hunger scale filled with zero scores a hungry household as food secure and drops it out of the caseload.
None of those are edge cases. They are what real exports look like.
What this path still needs
Three courses are published: Data Analysis Foundations, Python for Programme Data and R and the Tidyverse for Programme Data. The bulk of this role sits in the Data Preparation module — cleaning and validation, joining and reshaping, data quality assessment — which is on the programme roadmap and not yet written. See the programme for its syllabus.