SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

Sirry Chen; Jieyi Wang; Wei Chen; Zhongyu Wei

arXiv:2601.04638·cs.CL·April 21, 2026

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

Sirry Chen, Jieyi Wang, Wei Chen, Zhongyu Wei

PDF

1 Models 1 Datasets

TL;DR

SpeechMedAssist introduces a novel two-stage training paradigm for SpeechLMs, enabling effective speech-based medical consultations with limited real speech data, outperforming baselines in simulated interactions.

Contribution

The paper proposes a two-stage training approach for SpeechLMs that reduces data requirements and enhances performance in medical speech interactions.

Findings

01

Outperforms baselines in effectiveness and robustness.

02

Requires only 10k synthesized speech samples for training.

03

Excels in both single-turn and multi-turn medical interactions.

Abstract

Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly. Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation. In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients. By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

🤗
SII-Sirry/SpeechMedAssist
model· 7 dl
7 dl

Datasets

SII-Sirry/SpeechMedDataset
dataset· 116 dl
116 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.