article

Model calibration and evaluation via optimal subsampling using electronic health record data

  • Journal of the Royal Statistical Society Series A (Statistics in Society)
  • Royal Statistical Society
Research footprint

At a glance

Citations
0
References
49
Comments
0
Paper overview

Abstract

Abstract A common challenge for validating risk prediction models using electronic health record (EHR) data is that labels for the predicted outcome are not directly available. Towards efficient and unbiased model validation, we study optimal sampling designs for efficiently labelling an informative subset of patients in an EHR cohort. Given a pre-specified number of outcome labels, our design aims to minimize the asymptotic variance of an improved inverse probability weighted (‘I-IPW’) estimator for predictive accuracy metrics. Implementation of the sampling requires accurate risk estimates and the predictive accuracy metric of interest. We therefore propose to implement sampling in two steps. First a portion of the target number of labels is acquired by applying entropy sampling to a random subset of the cohort. These initial labels are used to calibrate risk estimates and obtain an initial estimate of the predictive accuracy metric, which are used to inform optimal sampling of the remaining target number of labels. The final estimate of the predictive accuracy metrics is obtained by applying the I-IPW estimator to the cohort and all acquired labels pooled together. Results from simulation studies and application to a real EHR dataset indicate superior efficiency of the proposed sampling design and I-IPW estimator.

Record transparency

Publication details

DOI
10.1093/jrsssa/qnag036
OpenAlex
W7133916474
Document type
article
Language
EN
Source
Journal of the Royal Statistical Society Series A (Statistics in Society)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.