Executive Summary

Zero-dose children remain a central concern for immunization planning in Nigeria, where an estimated 2.2 million children were unvaccinated as of 2022 (Gavi Zero-Dose Learning Hub 2023), and Kano State carries one of the country’s heaviest burdens. Multistage cluster probability surveys are the methodological gold standard for measuring vaccination coverage, but their cost and timelines often exceed what programmes can sustain for routine decisions. That gap has driven interest in faster, lower-cost alternatives. In late 2024, the Gates Foundation (GF) commissioned Mindset to conduct a baseline study of zero-dose children in Kano State and, alongside it, a head-to-head comparison of five candidate measurement methods against the gold-standard benchmark. The aim was to map the trade-offs between statistical rigour, speed, and cost across the methods most often proposed as substitutes or complements, rather than to identify a single replacement for the gold standard.

The gold-standard survey covered 15 priority Local Government Areas (LGAs): three sentinel LGAs (Gabasawa, Gaya, Nassarawa) powered for change detection, plus 12 combined non-sentinel LGAs. Penta-1 coverage came in at 66.8% in Gabasawa, 62.2% in Gaya, 74.7% in Nassarawa, and 61.7% in the non-sentinel stratum. Fieldwork ran about three months. The sample size was driven by the dual objective of establishing a baseline and powering future impact evaluation in the sentinel LGAs, which pushed our gold-standard implementation to the larger end of what a coverage benchmark would normally require. That is itself a useful reminder of the cost rapid alternatives are trying to displace. It also has a consequence for the analysis: some comparisons in the report apply analytic adjustments, for instance by matching sample sizes via simulation, so that a method’s measured performance reflects the method itself rather than the scale of the survey it was paired with.

The five alternatives evaluated against this benchmark were the Network Scale-Up Method (NSUM), Adaptive Sampling (AS), Lot Quality Assurance Sampling (LQAS), an enhanced Administrative Records Review (ARR) method, and Rapid Convenience Monitoring (RCM). Each was assessed on six dimensions: bias, precision, efficiency, scale/time, analytic complexity, and operational burden. They differ in purpose, not only in performance. Some estimate prevalence, some classify areas against a threshold, and some support rapid follow-up in a defined catchment or intervention area. Treating them as interchangeable coverage estimators would blur the comparison; method choice should start with the decision at hand, the geographic scope it requires, and the uncertainty the programme can tolerate. Table 1 gives the verdict for each.

Table 1: Head-to-head verdict by method

Method

Observed gap vs. gold standard

Precision

Speed / scale

Bottom line

Gold Standard

None (benchmark)

High in this study; design-tunable

Slow / large

Methodological benchmark; pair with rapid tools

AS

Nil or negligible in simulation study

Our implementation was on-par with an equally-sized PPS gold-standard sample

Slow / large

Conditional; promising family, tested variant roughly on par with an equally sized PPS sample

LQAS

Small to modest; implementation-sensitive

Low for coverage; right-sized for classification

Fast / small

Recommended for ward-level classification; secondary use for coverage estimation possible but less efficient

NSUM

High; unstable, not calibratable

Theoretically improved; bias-dominated

Moderate

Not recommended

ARR

High

Not quantifiable (no sampling)

Fast / desk-based

Not recommended

RCM

High; bounded to walked corridor

Dangerously low (n ≈ 20; serial autocorrelation)

Fast / small

Mop-up and within-corridor description only under tested protocol; generalising to unvisited households requires stronger design

The five methods cluster at the extremes of the scale-accuracy spectrum, with no real middle ground. At one end sits the gold-standard probability survey: large, slow, statistically defensible. At the other are the rapid, small-sample tools: RCM visits 20 children per ward, and LQAS in classification mode draws only 19. The two methods most hoped to bridge that gap turned out to fare worst on statistical validity: NSUM, which uses respondents’ social networks to extract more information per interview, and ARR, which estimates coverage from existing administrative data with little or no fieldwork.

Neither NSUM nor ARR produced credible coverage estimates in this study, and we do not recommend either for that purpose. NSUM was hampered by a restricted effective network (roughly 11 eligible alters per respondent, well below the 100+ where the method tends to pay off) and by pervasive transmission errors about contacts’ children’s vaccination status. ARR, even with numerators calibrated against Primary Health Center (PHC) audits and geospatial denominators, left a large, unstable gap to the gold standard that the enhancements we added could not close.

LQAS earns a positive recommendation for the classification task for which it was designed, albeit with some suggestions to improve validity. It worked through 32 sentinel wards against a 50% Penta-1 threshold (n = 19 per lot, both error rates below 10%) and flagged eight wards as potentially below threshold. It can also be configured to produce coverage estimates at higher levels of aggregation, for example by treating lots as strata with appropriate weighting; our pooled aggregates tracked the gold standard reasonably well in most strata, though one diverged materially. Even used this way, LQAS is less efficient for coverage estimation than a properly sized gold-standard sample, and its small per-lot samples leave less room for the non-response adjustments and field protocols that larger probability surveys can sustain.

AS is a more conditional recommendation. It is best understood as a family of designs, and our conclusions apply to the variant we fielded. In a simulation study of that variant, the implementation came out roughly on par with an equally sized conventional Probability Proportional-to-Size (PPS) sample, not consistently better. We also ran an exploratory simulation of an empirical Bayes variant that looked more promising in retrospective testing, suggesting that the adaptive family has further room to be tuned toward implementations that more reliably beat equally sized gold-standard designs. Even if such improvements materialise, the main constraint is analytic complexity: adaptive selection and variance estimation require a skilled survey statistician in the field, which makes routine use unrealistic for many implementers unless the expected gain is large.

RCM occupies a narrow legitimate niche but is not a general-purpose coverage tool under the protocol tested here. A single walk to a 20-child quota produces a route-defined snapshot whose meaning does not extend to the broader catchment area, and vaccination status clusters along the walk (OR ≈ 2.35), further shrinking effective information. The method is well-suited to mop-up activities, by directly identifying and referring unvaccinated children encountered along the walk, and to descriptive claims about the households actually canvassed. What the tested protocol does not support well is generalising from the walked corridor to unvisited households in the wider catchment, except in the unusual case where the walk plausibly covers the entire population of interest. Stronger design principles, such as multiple random walks, tighter specification of the area being characterised, and the other enhancements detailed in the RCM chapter, could broaden its scope. As implemented here, RCM should not underpin programme decisions about higher administrative levels.

The cost analysis complicates the intuition that rapid methods are necessarily cheap. Labor is the dominant input across all methods: field labor drives person-hours, but specialist analytical and technical labor drives financial cost, because higher-priced specialist time is concentrated in those phases. Marginal field costs per additional interview are more similar across methods than total costs, so large total-cost gaps reflect implementation scope and fixed overhead rather than per-unit differences. A practical consequence is that sample-size reduction is a limited lever for total cost: halving sample size does not come close to halving total cost. The useful frame is right-sizing rather than uniform reduction:

The most promising direction is not a new method but a more flexible use of the existing one. Probability sampling does not have to mean a single one-size multi-stage cluster design. The same framework can be right-sized to fill the missing middle between large benchmark surveys and rapid tools, and made more efficient still through model-assisted estimation that borrows strength from open-source geospatial data and prior survey estimates. The model-based geostatistics literature, which we did not directly assess in this report, sits upstream of that direction. It generates granular coverage surfaces that are independently useful for microplanning, and those same surfaces can feed back into the design and analysis of more efficient probability samples.

The gold standard is not replaceable in this setting, but it can be made smaller, faster, and more efficient. The distinction between gold-standard and rapid alternatives is better understood as a continuum of probability designs than as a categorical divide. Right-sized probability surveys, paired with LQAS for routine supervisory classification and informed by open-source geospatial priors, offer a more defensible path to routine zero-dose measurement than any of the indirect or non-probability alternatives tested here.