Interpretable Cross-Stage Quality Control for AI Medical Imaging Pipelines
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Multi-stage AI pipelines for medical image analysis can produce structurally implausible outputs (disconnected segmentation masks, lesions placed entirely outside the organ) that propagate to downstream analysis and scoring. We evaluate whether simple, deterministic quality control checks at multiple pipeline stages can catch such failures in a model-agnostic, interpretable, and low-cost manner. We implemented deterministic quality gates, pure functions requiring no GPU or learned parameters, at two stages of a prostate cancer detection pipeline: organ segmentation (8 gates) and lesion detection (3 gates). Gates encode anatomical plausibility constraints ranging from hard structural rules (e.g., the gland must be a single connected component) to empirical bounds (e.g., organ volume within calibrated range). Thresholds were calibrated on 50 PROMISE12 [8] expert segmentations and applied without modification to 1,500 PI-CAI [10] cases across four segmentation models of widely different quality. Gate rejection rates scaled with model weakness: 4.8% (multi-center nnU-Net ensemble), 6.3% (prostate-specific nnU-Net), 11% (MONAI prostate model), 93% (TotalSegmentator multi-organ model), all using identical gate code and thresholds. The central finding was cross-stage complementarity: Stage 2 (organ) and Stage 4 (lesion) gates caught largely non-overlapping failures, with only 2 shared cases out of 143 total rejections for Bosma22b. Gate-rejected cases showed degraded downstream lesion detection metrics (Mann-Whitney, p < 0.05 on three independent tests across models), and gate filtering improved lesion containment more than random case removal for two of three models (bootstrap test, 10,000 iterations). Cross-domain validation on liver tumor segmentation (CT, N=131) and kidney tumor segmentation (CT, N=489) replicated the cross-stage complementarity pattern with 0–2% overlap in all three domains. The gate interface and composition logic are universal, while gate selection and thresholds are domain-specific. These results suggest that deterministic QC gates, approximately 200 lines of Python per domain, can provide interpretable, auditable failure detection across AI imaging pipeline stages. They complement rather than replace model improvement or uncertainty estimation.
Publication details
- DOI
- 10.5281/zenodo.19362420
- OpenAlex
- W7147401215
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.