Medical Spoken Named Entity Recognition

Khai Le-Duc; David Thulke; Hung-Phong Tran; Long Vo-Dang; Khai-Nguyen; Nguyen; Truong-Son Hy; Ralf Schl\"uter

arXiv:2406.13337·eess.AS·April 3, 2025

Medical Spoken Named Entity Recognition

Khai Le-Duc, David Thulke, Hung-Phong Tran, Long Vo-Dang, Khai-Nguyen, Nguyen, Truong-Son Hy, Ralf Schl\"uter

PDF

Open Access 1 Repo 1 Models 1 Datasets 1 Video

TL;DR

This paper introduces VietMed-NER, the largest Vietnamese spoken medical NER dataset with 18 entity types, and evaluates various pre-trained models, highlighting the superiority of multilingual encoders for speech and text NER tasks.

Contribution

The creation of VietMed-NER, the first large-scale Vietnamese spoken medical NER dataset, and comprehensive baseline evaluations of state-of-the-art models.

Findings

01

Multilingual models outperform monolingual ones on speech and text NER.

02

Encoders outperform sequence-to-sequence models in NER tasks.

03

Translating transcripts enables cross-lingual application of the dataset.

Abstract

Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types. Furthermore, we present baseline results using various state-of-the-art pre-trained models: encoder-only and sequence-to-sequence; and conduct quantitative and qualitative error analysis. We found that pre-trained multilingual models generally outperform monolingual models on reference text and ASR output and encoders outperform sequence-to-sequence models in NER tasks. By translating the transcripts, the dataset can also be utilised for text NER in the medical domain in other…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

leduckhai/multimed
noneOfficial

Models

🤗
leduckhai/VietMed-NER
model

Datasets

leduckhai/VietMed-NER
dataset· 37 dl
37 dl

Videos

Medical Spoken Named Entity Recognition· underline

Taxonomy

TopicsTopic Modeling · Text Readability and Simplification · Natural Language Processing Techniques

MethodsXLM-R