Model Evaluation in Medical Datasets Over Time

Helen Zhou; Yuwen Chen; Zachary C. Lipton

arXiv:2211.07165·cs.LG·November 15, 2022

Model Evaluation in Medical Datasets Over Time

Helen Zhou, Yuwen Chen, Zachary C. Lipton

PDF

Open Access

TL;DR

This paper introduces the EMDOT framework and Python package to evaluate how machine learning models in healthcare perform over time, highlighting the importance of temporal evaluation strategies.

Contribution

The paper presents a novel framework and tool for assessing model performance over time in medical datasets, addressing the limitations of time-agnostic evaluation methods.

Findings

01

Model performance varies over time in medical datasets.

02

Using recent data for training can improve temporal robustness.

03

Temporal evaluation reveals performance shocks not seen in static assessments.

Abstract

Machine learning models deployed in healthcare systems face data drawn from continually evolving environments. However, researchers proposing such models typically evaluate them in a time-agnostic manner, with train and test splits sampling patients throughout the entire study period. We introduce the Evaluation on Medical Datasets Over Time (EMDOT) framework and Python package, which evaluates the performance of a model class over time. Across five medical datasets and a variety of models, we compare two training strategies: (1) using all historical data, and (2) using a window of the most recent data. We note changes in performance over time, and identify possible explanations for these shocks.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning in Healthcare · Artificial Intelligence in Healthcare and Education · Explainable Artificial Intelligence (XAI)

MethodsTest