Continual Learning for Monolingual End-to-End Automatic Speech Recognition

Steven Vander Eeckt; Hugo Van hamme

arXiv:2112.09427·eess.AS·January 22, 2026

Continual Learning for Monolingual End-to-End Automatic Speech Recognition

Steven Vander Eeckt, Hugo Van hamme

PDF

Open Access 1 Repo

TL;DR

This paper evaluates various Continual Learning methods for monolingual end-to-end ASR models, demonstrating significant performance improvements in adapting to new tasks while minimizing data retention.

Contribution

It provides a comprehensive comparison of CL methods for monolingual ASR, highlighting the most effective approach in reducing catastrophic forgetting with minimal data.

Findings

01

Best CL method reduces performance gap by over 40%

02

Achieves this with only 0.6% of original data

03

Extends monolingual ASR to new tasks effectively

Abstract

Adapting Automatic Speech Recognition (ASR) models to new domains results in a deterioration of performance on the original domain(s), a phenomenon called Catastrophic Forgetting (CF). Even monolingual ASR models cannot be extended to new accents, dialects, topics, etc. without suffering from CF, making them unable to be continually enhanced without storing all past data. Fortunately, Continual Learning (CL) methods, which aim to enable continual adaptation while overcoming CF, can be used. In this paper, we implement an extensive number of CL methods for End-to-End ASR and test and compare their ability to extend a monolingual Hybrid CTC-Transformer model across four new tasks. We find that the best performing CL method closes the gap between the fine-tuned model (lower bound) and the model trained jointly on all tasks (upper bound) by more than 40%, while requiring access to only 0.6%…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

stevenvdeeckt/cgn_cl_dialect
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Domain Adaptation and Few-Shot Learning · Multimodal Machine Learning Applications