# Data protection note — Protection referral pathway performance

Technical documentation for the project *Protection referral pathway
performance*. It is deliberately the project's **first** deliverable, not an
annex to the analysis: what you do not collect is part of the design, and a
performance dataset assembled without these decisions cannot be made safe
afterwards.

- **Source dataset:** `protection-referrals-2024.v1.csv` (1 850 cases,
  6 admin2 areas)
- **Analysis:** `notebooks/referral-pathway.python.en.ipynb`
- **Decision informed:** which service line and which area receive additional
  capacity, and whether the disability gap needs a separate response
- **Standards applied:** CHS, Sphere, GBV information management principles,
  IASC guidance on data responsibility

This is synthetic protection and GBV data. No real person is described, and this
file must never be used as a template for storing real case data — the safe form
of that is a consent-governed case management system.

---

## 1. Purpose limitation

This dataset exists to answer one question: **where does the referral pathway
lose people?** Every field is justified against that question, and any field
that cannot be is not collected.

The question is about the *pathway*, not about the *people*. That distinction is
what makes a safe dataset possible at all: measuring whether a referral reached a
service needs no detail about the incident that produced it.

**Prohibited secondary uses.** This dataset must not be used for case follow-up,
service verification, beneficiary identification, or any purpose requiring a
person to be located. It cannot support them — by design — and an attempt to use
it that way is a signal that the wrong dataset is being used.

---

## 2. What is collected, and why each field earns its place

| Field | Why it is needed |
| --- | --- |
| `case_id` | Deduplication and record linkage within the analysis. Not derived from any personal attribute. |
| `admin2` | The unit at which capacity is allocated. |
| `service_requested` | The service line, which is what a capacity decision funds. |
| `consent_to_refer` | Gates the denominator (§4). Without it, performance and consent are conflated. |
| `referral_made` | The first pathway node. |
| `referral_accepted` | The pathway outcome. |
| `days_to_first_service` | Timeliness, on the pathway's own clock (§6). |
| `disability_reported` | Equity disaggregation, which is a CHS commitment. |
| `age_group` | Broad band only. Sufficient for equity; insufficient to identify. |
| `sex` | Equity disaggregation. |

---

## 3. What is deliberately **not** collected

| Excluded | Reason |
| --- | --- |
| Names, initials, any identifier | Never required to measure a pathway. |
| Contact details | Same. Their presence would make the file a locating tool. |
| Free text / case narrative | The single highest re-identification risk in protection data, and never analysable at scale anyway. |
| Incident date | Combined with location it identifies. Timeliness is measured from referral instead (§6). |
| Location below admin2 | A village plus a service line plus an age band is identifying in a small population. |
| Exact age | Age band is sufficient for equity analysis. |
| Incident type | Not needed to measure whether a referral completed, and among the most sensitive fields that exists. |
| Perpetrator detail | Never leaves the case management agency, under any circumstances. |

**Under GBV information management principles the incident-level fields never
leave the case management agency.** This dataset is downstream of that boundary,
and the boundary is the reason it can be shared with a cluster at all.

The exclusions are not a redaction applied to a fuller extract. **They are
absences at the point of collection.** A field that is collected and then
stripped still existed, was transmitted, and sits in a backup somewhere; a field
that was never collected does not.

---

## 4. Consent gates the denominator

Completion is measured over cases that **consented** to referral.

| | Value |
| --- | --- |
| Cases | 1 850 |
| Consented | 1 638 (88.5%) |
| Reached a service | 757 |
| **Completion (consent-gated)** | **46.2%** |
| Completion if all cases counted | 40.9% |

The two figures differ by 5.3 points, and the wrong one blames the pathway for
people who chose not to be referred. **Counting a non-consenting case as a
pathway failure both misstates performance and misrepresents a person's
decision** — the pathway did what it should when someone declined.

Consent here is consent *to refer*, recorded at intake. It is not consent to
data sharing, which is a separate conversation, and this note does not treat the
first as covering the second.

---

## 5. Small-cell suppression

**Threshold: 20 cases.** Any cell in a cross-tabulation with fewer than 20 cases
is blanked, not rounded and not merged.

**This threshold is a protection decision taken with the case management agency,
not a formatting preference.** It is recorded here so it cannot be changed by
whoever next edits the notebook. In protection work a small cell is a disclosure
risk as much as a statistical one: "2 cases in Mirebalais, legal service, ages
12–17" can identify a child to anyone who works in that district.

Applies to: the area × service grid, and any disaggregation added later. The
notebook computes counts alongside means and blanks the mean where the count is
below threshold.

**Do not defeat suppression by aggregation.** Publishing a suppressed grid and a
row total from which the suppressed cell can be recovered by subtraction is the
same disclosure. Where a total would permit that, suppress the total too.

---

## 6. Timeliness is measured from referral, not from incident

`days_to_first_service` runs from the referral to the first service contact,
because the dataset holds no incident date and under GBV information management
principles it should not.

**The consequence must be stated in any report.** A survivor who reached a
caseworker late and a service quickly appears here as fast. This measure is a
property of the pathway after entry, not of a person's whole route to help.

---

## 7. Data-quality contradictions: flagged, not dropped

| Contradiction | Cases |
| --- | ---: |
| Service time recorded, referral not accepted | 11 |
| Service time recorded, no referral made | 6 |
| Accepted, no service time recorded | 40 |

**In case management a contradictory record is an entry issue to return to the
caseworker, and deleting it destroys the only trace the case existed.** Silent
deletion is a protection failure, not a cleaning step.

The 40 accepted referrals with no service time mean the **timeliness denominator
is smaller than the completion denominator**. The two are reported separately and
must never be placed in a table implying a shared base.

---

## 8. Equity disaggregation

Recording disability is a CHS commitment, and the analysis it supports is the
project's most consequential finding.

**Completion by disability status**

| Status | Cases | Completed | Completion |
| --- | ---: | ---: | ---: |
| Disability reported | 202 | 62 | 30.7% |
| No disability reported | 1 436 | 695 | 48.4% |

**Gap: −17.7 points**, chi-square p = 3.3 × 10⁻⁶.

**By service line** — the gap appears in four of the five. Safety and security is
the exception, at roughly 35% either way.

| Service line | No disability | Disability reported |
| --- | ---: | ---: |
| Health | 66.1% | 32.1% |
| Legal | 32.4% | 21.1% |
| Livelihood support | 23.6% | 15.8% |
| Psychosocial | 57.7% | 38.5% |
| Safety and security | 35.4% | 35.1% |

That it is not concentrated in one line makes it unlikely to be explained by the
mix of services people with disabilities request. **What it does not tell you is
the mechanism** — physical inaccessibility, pathways assuming mobility, or
something else — and that is answered by asking caseworkers, not by this table.

### 8.1 The coding normalisation that made the gap visible

One area recorded disability as `Yes`/`No` rather than `true`/`false`. Without
normalising, the disaggregation fragments into four categories, two of them from
that single area and both below the suppression threshold — **so the gap would
have been invisible**, and the project's most important finding would have been
lost to a coding accident.

The map is an explicit allow-list; anything outside it stays missing and those
cases are excluded from the disability analysis rather than assumed negative.
**A missing disability field is not a "no".**

---

## 9. Access, retention and sharing

- **Access.** Analysis is performed on the de-identified extract described here.
  Nobody needs the case management system to run it, and nobody should be given
  access to that system in order to.
- **Retention.** The extract is retained for the reporting cycle it informs. It
  carries no case follow-up value, because it deliberately cannot support follow-up.
- **Sharing.** The extract may be shared with the protection cluster and the GBV
  sub-cluster at the aggregation levels in this note, with §5 applied. Row-level
  data is not shared outside the case management agency.
- **Onward publication.** Any figure published externally must have passed §5.
  When in doubt, aggregate further; a coarser number that is safe is worth more
  than a precise one that is not.

---

## 10. Limitations

1. **Referral records show what the pathway did, not what people needed.** A
   service line with few referrals may be one nobody needs or one nobody offers,
   and this dataset cannot separate the two.
2. **Consent is recorded, not explained.** A declining consent rate could be
   informed refusal or a trust problem, and the pathway data cannot distinguish
   them.
3. **The shortfall metric assumes the gap is capacity, not need** (see the
   notebook). It is the right shape for a capacity decision and the wrong shape
   for a needs assessment.
4. **De-identification is not anonymisation.** Even with these exclusions, a
   sufficiently small population and sufficient auxiliary knowledge can
   re-identify. §5 is the operating control, and the exclusions in §3 are what
   keep the residual risk low enough to be controllable.

---

## 11. Reproducing this analysis

```bash
pnpm examples:build
```

Then run `notebooks/referral-pathway.python.en.ipynb`. It needs pandas, numpy and
scipy and reads the CSV over HTTPS.

The dataset is versioned by filename and immutable; a correction ships as
`.v2.csv` with this document revised beside it.

---

## 12. Change log

| Date | Change |
| --- | --- |
| 2026-07-27 | First issue, against dataset v1. |
