Conformer with dual-mode chunked attention for joint online and offline   ASR

Felix Weninger; Marco Gaudesi; Md Akmal Haidar; Nicola Ferri; Jes\'us; Andr\'es-Ferrer; Puming Zhan

arXiv:2206.11157·eess.AS·June 23, 2022·Interspeech·1 cites

Conformer with dual-mode chunked attention for joint online and offline ASR

Felix Weninger, Marco Gaudesi, Md Akmal Haidar, Nicola Ferri, Jes\'us, Andr\'es-Ferrer, Puming Zhan

PDF

Open Access

TL;DR

This paper introduces a dual-mode Conformer Transducer with chunked attention and knowledge distillation, improving online and offline speech recognition accuracy on diverse datasets.

Contribution

It proposes a novel dual-mode Conformer model with chunked attention and mode-specific components, enhancing online ASR performance with minimal complexity increase.

Findings

01

Chunked attention improves accuracy over autoregressive attention.

02

Knowledge distillation from offline to online mode enhances online accuracy.

03

Achieved 4-5% relative WER reduction on Librispeech and medical datasets.

Abstract

In this paper, we present an in-depth study on online attention mechanisms and distillation techniques for dual-mode (i.e., joint online and offline) ASR using the Conformer Transducer. In the dual-mode Conformer Transducer model, layers can function in online or offline mode while sharing parameters, and in-place knowledge distillation from offline to online mode is applied in training to improve online accuracy. In our study, we first demonstrate accuracy improvements from using chunked attention in the Conformer encoder compared to autoregressive attention with and without lookahead. Furthermore, we explore the efficient KLD and 1-best KLD losses with different shifts between online and offline outputs in the knowledge distillation. Finally, we show that a simplified dual-mode Conformer that only has mode-specific self-attention performs equally well as the one also having…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Blind Source Separation Techniques · Machine Learning and ELM

MethodsKnowledge Distillation