z-logo
open-access-imgOpen Access
Semisupervised Calibration of Risk with Noisy Event Times (SCORNET) using electronic health record data
Author(s) -
Yuri Ahuja,
Liang Liang,
Doudou Zhou,
Sicong Huang,
Tianxi Cai
Publication year - 2022
Publication title -
biostatistics
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 3.493
H-Index - 82
eISSN - 1468-4357
pISSN - 1465-4644
DOI - 10.1093/biostatistics/kxac003
Subject(s) - computer science , estimator , event (particle physics) , nonparametric statistics , set (abstract data type) , machine learning , estimation , data mining , parametric statistics , flexibility (engineering) , artificial intelligence , statistics , mathematics , physics , management , quantum mechanics , economics , programming language
Leveraging large-scale electronic health record (EHR) data to estimate survival curves for clinical events can enable more powerful risk estimation and comparative effectiveness research. However, use of EHR data is hindered by a lack of direct event time observations. Occurrence times of relevant diagnostic codes or target disease mentions in clinical notes are at best a good approximation of the true disease onset time. On the other hand, extracting precise information on the exact event time requires laborious manual chart review and is sometimes altogether infeasible due to a lack of detailed documentation. Current status labels-binary indicators of phenotype status during follow-up-are significantly more efficient and feasible to compile, enabling more precise survival curve estimation given limited resources. Existing survival analysis methods using current status labels focus almost entirely on supervised estimation, and naive incorporation of unlabeled data into these methods may lead to biased estimates. In this article, we propose Semisupervised Calibration of Risk with Noisy Event Times (SCORNET), which yields a consistent and efficient survival function estimator by leveraging a small set of current status labels and a large set of informative features. In addition to providing theoretical justification of SCORNET, we demonstrate in both simulation and real-world EHR settings that SCORNET achieves efficiency akin to the parametric Weibull regression model, while also exhibiting semi-nonparametric flexibility and relatively low empirical bias in a variety of generative settings.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here