HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic Reasoning

Jinning Yang; Wenjie Sun; Wen Shi

arXiv:2508.15338·cs.AI·January 27, 2026

HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic Reasoning

Jinning Yang, Wenjie Sun, Wen Shi

PDF

Open Access 1 Video

TL;DR

HeartLLM introduces a novel method to enable large language models to process ECG signals by discretizing continuous ECG data into tokens, improving diagnostic reasoning and generalization across clinical tasks.

Contribution

The paper presents a new framework that integrates ECG signal processing with LLMs through discretization and tokenization, facilitating open-ended medical reasoning without modifying core models.

Findings

01

Achieves strong performance on ECG question answering and report generation

02

Maintains generalization to out-of-distribution data

03

Demonstrates effectiveness of discretized ECG tokens in medical reasoning

Abstract

Electrocardiography (ECG) plays a central role in cardiovascular diagnostics, yet existing automated approaches often struggle to generalize across clinical tasks and offer limited support for open-ended reasoning. We present HeartLLM, a novel framework that integrates time-series (TS) and language modeling by enabling large language models (LLMs) to process 12-lead ECG signals for clinical text generation tasks. Our approach discretizes continuous ECG embeddings into quantized codes using a lead-wise encoder and quantization module. These quantized codes are then mapped to an extended ECG vocabulary to form ECG tokens, enabling the model to process both ECG and natural language inputs within a unified framework. To bridge the modality gap, we pretrain the model on an autoregressive ECG token forecasting task, allowing the LLM to capture temporal dynamics through its inherent language…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic Reasoning· underline

Taxonomy

TopicsECG Monitoring and Analysis · Machine Learning in Healthcare · Topic Modeling