cassionData Analysis

Course · Advanced · Sector Analysis

Protection and GBV Data

Analyse referral, case management and incident data under the safety and confidentiality rules the sector imposes — including the analyses you must decline to publish.

PythonR16 h8 lessons

What you will be able to do

  • State what a protection dataset must not contain, and justify each omission by what the analysis actually needs
  • Assess disclosure risk in a disaggregated table and apply a suppression rule before the table leaves the building
  • Build a consent-gated referral pathway and locate the node where the largest share is lost
  • Compute caseload per caseworker and relate it to the quality measures that fall as it rises
  • Handle a censored case register without biasing time to closure downward
  • Say in writing that a case curve measures reporting rather than prevalence, and show the evidence for it

Standards and methodologies

Core Humanitarian Standard (CHS)Sphere StandardsUNICEF indicator definitionsOECD DAC evaluation criteria

This is the only course on the platform where the correct answer to an analytical question is sometimes decline to produce it. That is not a caveat attached to the material; it is the material.

The audience is real. Protection and GBV information managers are asked weekly for a table that would identify a survivor, usually by someone with entirely good intentions who has not thought it through, and the skill being taught is to recognise the request, refuse the table, and offer the thing that answers the underlying question safely.

The course runs on two files. The referral dataset has appeared in earlier courses as a joining and disaggregation problem; here it is analysed as what it is, and what it does not contain is the first lesson. It gains a case management register — caseworker, month opened, month closed, closure reason — because a pathway analysis cannot answer the two questions a supervisor is judged on.

Three results shape the course, all computed from those files.

38.8% of cases reach a service and 43.8% of consenting cases do, and the difference between those denominators is a person’s decision rather than a programme failure. Consent gates the pathway, and counting a non-consenting case as a failure misstates performance and misrepresents the person.

Cases with a reported disability complete at 26.7% against 46.2%. That gap is nineteen points, it is the reason to disaggregate at all, and it is also the finding most likely to be lost in a table too small to publish.

One area’s caseworkers carry 33.2 open cases each against another’s 14.8, and the overloaded area holds 1.78 case plan reviews per case against 2.45, and loses contact with 38.9% of the cases it closes. Three tables, one cause.

Both Python and R throughout. You should have done module 3 and the earlier module 4 courses — this course assumes you can defend a denominator, read a censored register, and write a limitation as a decision.

Start the course — The columns that are not there

Practice

Reading the lessons is not the same as having done the work. Each of these applies the course to a dataset it did not teach on.

Progress

Enrolling is free and only records your progress — the whole course is readable without it.

Take it offline

The whole course as a typeset PDF — every lesson, every code example, the data dictionary and the indicator definitions. Generated from the same source as this page.

The LaTeX source ships alongside each PDF, so an organisation can rebrand the handout or fold a lesson into its own training pack.