Model calibration and evaluation via optimal subsampling using electronic health record data
At a glance
- Citations
- 0
- References
- 49
- Comments
- 0
Öz
Abstract A common challenge for validating risk prediction models using electronic health record (EHR) data is that labels for the predicted outcome are not directly available. Towards efficient and unbiased model validation, we study optimal sampling designs for efficiently labelling an informative subset of patients in an EHR cohort. Given a pre-specified number of outcome labels, our design aims to minimize the asymptotic variance of an improved inverse probability weighted (‘I-IPW’) estimator for predictive accuracy metrics. Implementation of the sampling requires accurate risk estimates and the predictive accuracy metric of interest. We therefore propose to implement sampling in two steps. First a portion of the target number of labels is acquired by applying entropy sampling to a random subset of the cohort. These initial labels are used to calibrate risk estimates and obtain an initial estimate of the predictive accuracy metric, which are used to inform optimal sampling of the remaining target number of labels. The final estimate of the predictive accuracy metrics is obtained by applying the I-IPW estimator to the cohort and all acquired labels pooled together. Results from simulation studies and application to a real EHR dataset indicate superior efficiency of the proposed sampling design and I-IPW estimator.
Publication details
- DOI
- 10.1093/jrsssa/qnag036
- OpenAlex
- W7133916474
- Document type
- article
- Language
- EN
- Source
- Journal of the Royal Statistical Society Series A (Statistics in Society)
- Last metadata update
Comments
Oturum Açın to join the discussion.