Scalable and accurate deep learning for electronic health records

Alvin Rajkomar; Eyal Oren; Kai Chen; Andrew M. Dai; Nissan Hajaj,; Peter J. Liu; Xiaobing Liu; Mimi Sun; Patrik Sundberg; Hector Yee; Kun Zhang,; Gavin E. Duggan; Gerardo Flores; Michaela Hardt; Jamie Irvine; Quoc Le; Kurt; Litsch; Jake Marcus; Alexander Mossin; Justin Tansuwan; De Wang; James; Wexler; Jimbo Wilson; Dana Ludwig; Samuel L. Volchenboum; Katherine Chou,; Michael Pearson; Srinivasan Madabushi; Nigam H. Shah; Atul J. Butte; Michael; Howell; Claire Cui; Greg Corrado; Jeff Dean

arXiv:1801.07860·cs.CY·May 14, 2018

Scalable and accurate deep learning for electronic health records

Alvin Rajkomar, Eyal Oren, Kai Chen, Andrew M. Dai, Nissan Hajaj,, Peter J. Liu, Xiaobing Liu, Mimi Sun, Patrik Sundberg, Hector Yee, Kun Zhang,, Gavin E. Duggan, Gerardo Flores, Michaela Hardt, Jamie Irvine, Quoc Le, Kurt, Litsch, Jake Marcus, Alexander Mossin, Justin Tansuwan

PDF

TL;DR

This paper introduces a scalable deep learning approach using raw electronic health record data in FHIR format, achieving high accuracy in predicting various medical events across multiple centers without data harmonization.

Contribution

It presents a novel method for representing entire raw EHRs with deep learning, enabling accurate multi-center predictions without site-specific data normalization.

Findings

01

Deep learning models achieved AUROC 0.93-0.94 for in-hospital mortality

02

Models outperformed traditional predictive models in all tasks

03

Case-study demonstrates model transparency for clinicians

Abstract

Predictive modeling with electronic health record (EHR) data is anticipated to drive personalized medicine and improve healthcare quality. Constructing predictive statistical models typically requires extraction of curated predictor variables from normalized EHR data, a labor-intensive process that discards the vast majority of information in each patient's record. We propose a representation of patients' entire, raw EHR records based on the Fast Healthcare Interoperability Resources (FHIR) format. We demonstrate that deep learning methods using this representation are capable of accurately predicting multiple medical events from multiple centers without site-specific data harmonization. We validated our approach using de-identified EHR data from two U.S. academic medical centers with 216,221 adult patients hospitalized for at least 24 hours. In the sequential format we propose, this…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.