Beyond Binary Verdicts: Aleatoric Uncertainty Quantification in Agentic CI/CD Quality Pipelines
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Öz
Continuous integration pipelines that rely on LLM-based quality agents produce binary pass/fail verdicts that discard the probabilistic uncertainty inherent in model inference. This paper extends the K11tech Agentic AI QA System with six uncertainty-aware features (F1–F6) that propagate per-agent confidence through the pipeline and expose it to human reviewers in a principled way. The contributions are: (F1) per-agent confidence scoring via verbally elicited LLM self-assessment; (F2) weighted aggregation of agent scores into a pipeline-level uncertainty score; (F3) verdict downgrading — automatically converting PASS verdicts to PASS_UNCERTAIN when confidence falls below a calibrated threshold; (F4) a second HITL escalation path triggered by high uncertainty rather than breaking-change detection alone; (F5) isotonic regression recalibration of raw LLM confidence scores against historical ground truth; and (F6) correlation-aware aggregation that discounts redundant agent signals using a learned inter-agent correlation matrix. Evaluation on a 120-PR synthetic dataset calibrated to empirical agent accuracy distributions demonstrates that the combined features reduce unnecessary HITL escalation by 18% compared to a binary-verdict baseline while maintaining a false negative rate below 0.10, with isotonic recalibration reducing expected calibration error from 0.14 to 0.06. Keywords: Agentic AI, Uncertainty Quantification, Aleatoric Uncertainty, Conformal Prediction, CI/CD, Human-in-the-Loop, LLM Calibration, API Contract Testing, Software Quality Assurance, LangGraph
Publication details
- DOI
- 10.5281/zenodo.20684082
- OpenAlex
- W7164645267
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Oturum Açın to join the discussion.