6  Recommendations

The preceding chapter assessed each method against statistical and logistical criteria and identified where each method fits. This chapter translates that assessment into two complementary sets of recommendations. The methodological recommendations answer how zero-dose prevalence should be measured well — design choices, sampling frame, estimation strategy, and the analytic capacity those choices imply. The Monitoring and Evaluation (M&E) recommendations answer the downstream question: once a measurement system is in place, what should programmes in Kano (and comparable Northern Nigerian states) routinely track to know whether the zero-dose burden is moving, where it is moving, and why. A concrete pathway for measuring intervention impact at endline is taken up separately in the baseline report’s Endline and Impact-Measurement Design section (Mindset 2025).

6.1 Methodological Recommendations

Five methodological recommendations follow from the comparative findings. They apply whether the next measurement event is a routine round, a hotspot re-check, or the endline that closes out the current programme cycle.

  1. M1. Use Nigeria Demographic and Health Survey (NDHS) modeled surfaces for sample size planning. Open-source data such as the NDHS modeled surfaces can already produce fine-grained coverage estimates before fieldwork, making sample size planning more evidence-based even when the priors are several years old. The 2018 NDHS modeled surface agreed with our 2025 gold-standard Local Government Area (LGA) estimates within roughly 7 percentage points on average — already comparable in magnitude to the ±10-percentage-point single-round precision target of the classical World Health Organization (WHO) 30-cluster, 7-child Expanded Programme on Immunization (EPI) survey under standard design assumptions.1

    Our own planning illustrates the cost of not using such priors: lacking LGA-specific coverage estimates for Gabasawa, Gaya, Nassarawa, and the other sentinel areas, sample-size calculations had to assume the most cautious scenario — a change centred around 50% prevalence (for example, 45% to 55% for a 10-point change), which is the scenario that requires the largest sample. With informative LGA-level priors from the 2018 NDHS surface, the calculations could have been anchored to each LGA’s likely prevalence band, and the resulting targets would have been smaller and better right-sized to each area.

  1. M2. Sample directly via building footprints. Open-source building footprint datasets create a path to sampling designs that approach the operational simplicity of a simple random sample — single-stage, frame-based, list-driven — without the logistical overhead of multi-stage cluster selection and household listing. A decade ago, resources like this were rare; today they offer a practical way to partially fill a perennial gap in survey work: the absence of a usable, up-to-date sampling frame. Care is still needed to handle ineligible units (for example, non-residential structures) and to address the multiplicity problem: multiple footprints can point to the same household. Our experience in this baseline study found this harder than expected in practice — both to adjust for analytically and to elicit reliably from respondents. Even so, these data can be downloaded and sampled with a few lines of R or Python, and that changes what is practical. They also strengthen implementation fidelity. Because sampled points are georeferenced, teams can geofence assignments and use Global Positioning System (GPS) traces to verify that enumerators actually visited the intended locations, rather than substituting more convenient households — an audit trail this study put into practice and reports on in the baseline’s protocol-deviations annex (Mindset 2025). That kind of audit trail reduces several common sources of non-sampling error — enumerator bias toward more accessible or cooperative households, protocol drift, and outright fabrication.

    This approach is especially attractive for small-sample methods such as Lot Quality Assurance Sampling (LQAS) and Rapid Convenience Monitoring (RCM), which already rely on pseudo-random field procedures. Sampling from building footprints fits that tradition while reducing reliance on Probability Proportional-to-Size (PPS) cluster designs and listing operations. Footprints can also complement existing field procedures rather than replace them outright — our own LQAS implementation, for instance, used footprints to pre-specify the visitation order of a pseudo-random walk, substituting a reproducible itinerary for enumerator discretion. There can be some frame undercoverage if the building layer is outdated and new construction is missing, but in many settings this is likely a smaller concern than other dominant error sources (for example, nonresponse, measurement error, or implementation deviations). In many contexts, the net result may be a faster, cheaper, and more standardized workflow with a more defensible link between the design and what was actually implemented.

    A pragmatic middle path uses building footprints as a seed list for a listing-and-enumeration step, rather than as the direct sampling frame. Listers visit the footprints, classify each structure’s residential status and household composition, and the survey then samples from the resulting cleaned list. Multiplicity is largely resolved at the listing stage, and the design retains probability-sample properties without forcing respondents to answer complex multiplicity-adjustment questions in the main interview.

  2. M3. Use model-based and model-assisted estimation where assumptions hold. Going beyond strictly design-unbiased estimation can produce more precise results from the same resources, by drawing on what established survey methodology already provides. Two related but distinct families are relevant here, and they have different epistemic profiles. Model-assisted methods (calibration, generalized regression / Generalized Regression Estimator (GREG), ratio estimation) use a working model to reduce variance but remain design-consistent if the model is wrong — the auxiliary information improves precision without putting the headline estimator at risk. Model-based methods (Fay-Herriot, Battese-Harter-Fuller, multilevel shrinkage and other small-area estimators) trade design-consistency for borrowed strength across related areas, and therefore depend on explicit model diagnostics for credibility. Practical options include more informative stratification, stronger use of calibration and ratio estimation, and other techniques that combine primary survey data with credible auxiliary information. When estimates are required across many domains or small areas, model-based approaches can also borrow strength across related areas and improve the stability of sub-aggregate estimates.

    Practitioners should not shy away from methods that deliberately trade a small amount of bias for substantial variance reduction — Bayesian, shrinkage, and multi-level models are common examples. The relevant objective is minimizing total error (mean squared error), not unbiasedness per se. With transparent diagnostics and validation, such methods deliver more useful estimates without requiring prohibitively large sample sizes in every stratum. This logic underpins our own NDHS modeled layer analysis (Section 3.5.7.1), which uses a model-based prevalence surface to produce credible LGA-level coverage estimates where direct survey data alone are too noisy.

    In Nigeria, this opportunity is likely to expand further as the country moves toward its long-delayed population and housing census. Government and the National Population Commission (NPC) have continued to signal readiness and ongoing preparations, with timing subject to a formal presidential proclamation. Once the census is completed, it should provide refreshed population totals and small-area control figures that make calibration, post-stratification, and small-area estimation more defensible and more effective, especially when users need reliable estimates below the national or state level.

  3. M4. Price analytical complexity explicitly in method selection. Methods often presented as simple in teaching materials can require advanced design-based analysis in real implementation. Budgets should reflect the true analytical demands of weighting, variance estimation, and implementation diagnostics. Several methods in this comparison illustrate the gap concretely: Network Scale-Up Method (NSUM) requires custom estimators and uncertainty propagation that no off-the-shelf survey package delivers; RCM relies on bespoke capture-recapture analysis with non-standard assumptions; and even LQAS — often presented as a simple supervisory rule — needed a weighted, design-consistent reanalysis here to recover defensible classifications from a complex sample. Treating complex designs as simple tabulations creates apparent savings at the expense of validity.

  4. M5. Right-size designs rather than defaulting to either extreme. The right question is right-sizing, not uniform reduction: modest upward adjustments for already-small designs where precision gains are steep, and measured downward adjustments for large gold-standard-style designs where marginal precision is cheap. The design space between very small-sample rapid tools and large benchmark surveys is wide, and programmes can tune design parameters to match decisions and constraints, rather than defaulting to perceived “cheap” options or assuming only full gold-standard designs are viable. These tradeoffs should be made explicitly, with clear statements about expected precision and bias — the alternative is binary thinking about scale and rigor, where the only available choices are “rapid and cheap” or “benchmark and expensive.”

6.2 Monitoring & Evaluation Recommendations

The M&E recommendations draw on evidence from two sources: the baseline report (Mindset 2025), which provides the numerator, denominator, hotspot, and missed-opportunity findings that define what to monitor; and this head-to-head comparison, which provides the cost-precision-timeliness trade-offs that define how to monitor it — without having to rebuild a benchmark survey every year.

The recommendations are organised in three themes:

  • Tracking zero-dose prevalence over time — where are the zero-dose children, and how is that pattern changing?
  • Tracking intervention performance — are programmes converting effort into vaccinated children, and why are coverage indicators moving the way they are?
  • Matching measurement methods to monitoring use cases — which instrument is right-sized for which monitoring question?

These themes sit inside the Monitor/Measure stage of Gavi’s Identify, Reach, Monitor, Measure, and Advocate (IRMMA) framework (Gavi, the Vaccine Alliance 2021). The specific demand-, supply-, and access-side interventions whose performance these indicators help monitor are set out in the baseline report’s Prioritized Intervention Strategies section (Mindset 2025).

6.2.1 Tracking Zero-Dose Prevalence over Time

Ideally, routine monitoring of zero-dose prevalence works best when it is organized around the geographic resolution at which programmes actually make decisions (NPHCDA 2021), rather than a single national or state headline figure. The baseline established Penta-1 coverage benchmarks for the three sentinel LGAs — 66.8% in Gabasawa, 62.1% in Gaya, 74.7% in Nassarawa — plus a combined non-sentinel benchmark of 61.7% covering the remaining twelve LGAs (Mindset 2025). The design was powered to produce defensible benchmarks at this four-stratum level.

Annual or semi-annual full re-measurement of all 15 LGAs at gold-standard precision is neither necessary nor affordable, and shorter intervals framed as universal requirements tend to overstate what programmes can realistically resource. A more practical approach treats the full gold-standard design as a periodic anchor, aligned where possible to NDHS or Multiple Indicator Cluster Survey (MICS) rounds (WHO 2018), and uses lighter, classification-oriented instruments to fill the years in between.

  1. E1. Pace the anchor benchmark to programme cycles, not the calendar. The two headline indicators — zero-dose prevalence and Penta-1 coverage, both reported with 95% confidence intervals — should be refreshed against a full benchmark roughly once per programme cycle. Where resources allow, that means every three to five years, timed to coincide with the next NDHS or MICS round where possible (WHO 2018).

    Treating a fixed three-year interval as a binding default risks working against how programmes actually plan and budget for measurement. It is better presented as a target to argue toward. Within any given cycle, the anchor round is what calibrates the lighter monitoring instruments described below; skipping a planned anchor to fund routine activity is inadvisable, because the rest of the architecture depends on having a recent benchmark to calibrate against.

  2. E2. Report at the right strata, and label uncertainty honestly at lower levels. The reporting levels below reflect the strata this baseline was powered for; studies designed for different purposes will have their own defensible reporting levels. For the Kano programme, three levels carry decision relevance.

    At the state and stratum level, report zero-dose prevalence and Penta-1 coverage with 95% confidence intervals against the four baseline benchmarks — the three sentinel LGAs and the combined non-sentinel stratum. These are the only four LGA-level estimates the baseline established with usable precision, and are therefore the defensible pre-post comparators.

    At the LGA level for the remaining twelve non-sentinel LGAs, point estimates are noisier and should be reported with their uncertainty interval rather than as standalone figures.

    At the ward level, the right approach depends on whether the survey was powered for ward-level point estimates. For the sentinel LGAs in this baseline, where wards were oversampled, ward-level coverage estimates with confidence intervals are defensible (Mindset 2025). For other settings, including the non-sentinel LGAs in this study and routine monitoring rounds not designed for ward-level point estimation, supervisory classification is the more appropriate framing: each ward should be flagged as acceptable, marginal, or potentially unacceptable against a programme-relevant threshold (see E9 on LQAS).

  3. E3. Decompose hotspots into prevalence and absolute-burden categories. A single hotspot definition obscures a planning decision that comes down to programme theory. Prevalence hotspots — wards or 1 km cells with high zero-dose rates — are the right target when the theory of change rests on local immunity: closing a high-prevalence pocket protects the surrounding population. Absolute-burden hotspots — areas with the largest counts of zero-dose children regardless of rate — are the right target when the theory of change is to reach the most children for the least cost, which often points to dense, lower-prevalence settlements.

    The baseline identified clusters of elevated zero-dose prevalence that can span ward boundaries within a sentinel LGA (Mindset 2025). Routine monitoring should preserve a hotspot layer — refreshed when new data are available — and report indicator values both inside and outside the current hotspots, decomposing results into both categories so that planning decisions can be defended on either theory of change (Utazi et al. 2024). A single LGA-level number that averages over a hotspot and its surroundings will mask exactly the variation programmes most need to see (Utazi et al. 2024).

6.2.2 Tracking Intervention Performance

Coverage and zero-dose prevalence are downstream outcomes. Programmes also need process indicators that move on a faster cadence and can tell managers why outcomes are changing (Gavi, the Vaccine Alliance 2021). The baseline organised barriers along three drivers — demand, supply, and access — and that structure carries naturally into the monitoring backbone.

  1. E4. Demand-side indicators, with Kano-specific calibration on missed opportunities. The most actionable demand-side indicators are caregiver-reported reasons for non-vaccination and proxies for community engagement intensity. For the former, the WHO Behavioral and Social Drivers (BeSD) framework organises reasons into four domains — social norms, beliefs, access, and health-facility issues — with social norms dominating in this baseline (cited by around 41% of caregivers), followed by beliefs (23%), access (15%), and health-facility issues (5%) (Mindset 2025). For community engagement, useful proxies include trainings delivered to traditional and religious leaders and community meetings held; trial evidence suggests that structured community engagement can shift vaccination outcomes such as timeliness and dropout even when effects on full coverage are null (Oyo-Ita et al. 2021).

    The missed-opportunity rate — any contact with health services by an unvaccinated or partially vaccinated child that does not result in receipt of all recommended doses (Adamu et al. 2019) — straddles both demand and supply: it reflects a caregiver who reached the facility and a service that failed to convert the contact. In principle it is a useful indicator. In Kano specifically, however, the headroom from fixing missed opportunities alone is modest: the baseline’s counterfactual Diphtheria-Pertussis-Tetanus (DPT)-1 analysis shows that eliminating Missed Opportunities for Simultaneous Vaccination (MOSVs) and early doses would raise overall DPT-1 coverage by only around 1.8 percentage points — roughly 2.7 pp in Nassarawa, where missed-opportunity rates run highest (Mindset 2025). That limited headroom is reason enough to treat missed-opportunity correction as a quality-of-service indicator and a defaulter-tracing trigger, rather than a primary coverage lever in this setting. Where the monitoring instrument already captures full vaccine history per child, programmes should track the rate on that basis.

  2. E5. Supply-side indicators, narrowly scoped and admin-data-aware. The most defensible supply-side indicator is session-level fidelity — sessions actually held against sessions planned in the microplan — because it is reported at the operational unit where corrective action lives: the facility-month. Process targets here should be benchmarked against the state’s own microplanning expectations, not against any external standard.

    We do not recommend building a routine indicator around District Health Information System (DHIS2)-derived administered doses, stockout days or cold-chain functionality as standalone monitoring signals. This study’s findings and the broader literature give us low confidence in administrative numerators as a reliable signal of vaccination coverage (Cutts et al. 2016). These should be tracked operationally at the facility level for management purposes, but they do not belong in the headline M&E dashboard.

  3. E6. Access-side indicators are where the spatial signal lives. The indicators that emerged most clearly from the baseline are the rural–urban gradient in coverage and the caregiver-education gradient (Mindset 2025); the access section also documents marked variation in service-availability awareness — zero-dose children’s caregivers are substantially more likely to not know whether vaccination is offered at their nearest facility — and heterogeneity in facility distance across LGAs that the geospatial travel-time maps capture for the three sentinel LGAs.

    The baseline’s rural–urban and education gradients were strong enough to merit separate reporting tracks; treating them as routine equity disaggregations rather than annex-only analysis keeps equity visible to programme staff (Sato 2023; Akwataghibe et al. 2019). Where the partner has the geospatial analytic capacity, travel time and distance should be tracked off the geospatial layer — building footprints plus facility list. Where caregiver-reported access measures are used instead, programmes should interpret them with the caveat that they conflate distance with willingness to travel.

  4. E7. A small number of cross-cutting equity indicators. Alongside the coverage headline, three equity gaps in zero-dose prevalence are worth tracking routinely: between rural and urban residents, between children of caregivers with and without secondary education, and between the bottom and top wealth quintiles (Arsenault et al. 2017). All three are derived from the anchor probability survey (E8) on a programme-cycle cadence.

    The rural–urban and caregiver-education gradients are cheap to elicit (a single survey question each) and reflect the two largest movers in the baseline’s vaccination-ladder model — settlement type and caregiver education — with rural–urban being the practical binary collapse of settlement type (Mindset 2025). The wealth-quintile gradient has long been a core Gavi equity-monitoring dimension (Arsenault et al. 2017), but it requires a full Demographic and Health Survey (DHS)-style asset roster and the associated respondent and analytic burden; programmes should weigh that cost before including it. Other equity cuts (e.g., caregiver age, household size, religion) are worth exploring but do not need to live in the routine dashboard.

6.2.3 Matching Measurement Methods to Monitoring Use Cases

The head-to-head comparison provides the basis for a tiered measurement architecture (see Table 5.1 and Table 5.2). The underlying principle is that no single method should carry all the monitoring load: each instrument has a use case where it is right-sized, and the cost should be paid where it buys the most information (WHO 2018).

  1. E8. Anchor benchmark layer: a gold-standard probability survey at programme-cycle cadence. The anchor benchmark — a gold-standard multistage cluster survey — belongs on a programme-cycle cadence (see E1), sized to the strata that drive funding decisions (Table 5.1; Table 5.2). This is the only layer of the architecture that produces defensible point estimates for the state and stratum headline indicators (Lohr 2021).

  2. E9. Routine supervisory monitoring: LQAS done as full-coverage ward classification. LQAS is the right tool for the supervisory pass-fail question at the ward level, refreshed annually or semi-annually (Valadez et al. 2002; Rhoda et al. 2010). At \(n = 19\) per ward, it is cheap enough to repeat.

    However, the recommendation is contingent on two implementation requirements. First, wards must be used as strata and the entire LGA must be covered: cherry-picking a few wards breaks the inference and produces figures that cannot be aggregated to the LGA level. Second, the classification threshold should be set by the state task force against an explicit numeric coverage target per stratum, not adopted by default from the 50% cutoff used in this study.

    The Decentralized Immunization Monitoring (DIM) programme implemented by The African Field Epidemiology Network (AFENET) shows that this every-ward-as-a-lot approach is operationally feasible at scale. The Kumbotso pilot fielded 19 interviews in each of the LGA’s 11 wards — 209 interview locations in total — and was designed to scale to eight LGAs across Kano, Bauchi, Borno, and Sokoto States (Attahiru et al. 2025). The subsequent Round 3 multi-state rollout covered 107 wards across all 8 LGAs, each sampled as its own LQAS lot (Zero Dose Learning Hub Nigeria 2025). Programmes that cannot match that resource intensity should report classifications only for the wards actually fielded, rather than presenting partial-coverage LQAS as an LGA-level signal.

    Importantly, LQAS is not the right tool for ward-level coverage point estimation. Reporting aggregated LQAS coverage figures alongside the gold-standard benchmark risks conflating two inferential products that answer different questions.

  3. E10. Use model-assisted approaches for hotspot identification and re-checks. This layer has two steps. First, to identify candidate hotspots between anchor surveys, use NDHS-style modelled coverage surfaces of the kind produced by Utazi et al. (2023) and Sbarra et al. (2021), and refresh them whenever new DHS or MICS rounds are released. The work is desk-based — no fresh field data needed. Second, before committing resources to a deployment in a candidate hotspot, run a small probability sub-sample using the building-footprint frame to confirm the hotspot is still active. This is a quick verification step at the hotspot scale, not a coverage survey of the LGA.

    Both steps feed the prevalence-vs-absolute-burden decomposition in E3, and both are designed for decision moments — specifically, when a programme is choosing where to concentrate next-quarter mop-up effort — rather than for routine reporting.

  4. E11. A hybrid LQASRCM variant as a future-study candidate. The single-walk, 20-child-quota RCM tested here is not a monitoring instrument (see Section 3.6), and NSUM is not a viable substitute for probability-based coverage measurement in this setting (see Section 3.2). Neither belongs in the routine monitoring backbone.

    A more interesting question is whether a systematic RCM variant could earn a place. The vaccination-coverage literature documents one that overlays LQAS decision rules on a door-to-door frame: enumerators canvass blocks until they reach 20 eligible children, and the supervisory call follows from the count of unvaccinated children found (Dietz et al. 2004). A hybrid of this kind would be a practical complement to the ward-level LQAS layer (E9) — it keeps the door-to-door speed while replacing the informal “did the team see enough vaccinated children” judgement with a pre-specified acceptance rule with a known classification error rate.

    We did not field this hybrid in the head-to-head comparison, and the classification thresholds in (Dietz et al. 2004) were derived for a specific operating point — zero unvaccinated children as the classification target — that may not transfer to a different programme-defined threshold without re-derivation. On that basis, we recommend the hybrid LQASRCM variant as a future-study candidate rather than as a current-cycle deployment.

  5. E12. Build a central data repository and feed research findings back into routine priors. A recurring gap in this work is that field data is scattered: routine RCM mop-ups, LQAS supervisory rounds, and ad hoc surveys funded by donors or partners all end up in separate systems held by whoever ran the collection. That fragmentation makes it harder than it should be to update programme priors about where unvaccinated children are, and it forces each new study to rediscover what previous studies have already established.

    A central holder for routinely-collected zero-dose datasets would pay off well above its cost — likely a Gavi-coordinated role, or another entity with the convening authority and data-stewardship capacity to take it on. The scope of what lives there should extend beyond routine programme data to include research outputs: NDHS-aligned modelled surfaces, de-identified study microdata, and the codebooks needed to make that data reusable.


  1. The two quantities are not strictly the same estimand: the 7-percentage-point figure is a mean absolute deviation across LGAs between the 2018 modeled surface and the 2025 gold-standard estimates, while the ±10-percentage-point EPI target is a confidence-interval half-width (margin of error). We compare their magnitudes here as a rough order-of-magnitude check on prior-quality.↩︎