Integrating Pre-Trained Speech and Language Models for End-to-End Speech   Recognition

Yukiya Hono; Koh Mitsuda; Tianyu Zhao; Kentaro Mitsui; Toshiaki; Wakatsuki; Kei Sawada

arXiv:2312.03668·eess.AS·June 7, 2024·1 cites

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

Yukiya Hono, Koh Mitsuda, Tianyu Zhao, Kentaro Mitsui, Toshiaki, Wakatsuki, Kei Sawada

PDF

Open Access 1 Models 1 Video

TL;DR

This paper presents an integrated end-to-end speech recognition model combining pre-trained speech and language models, enabling efficient optimization and achieving performance comparable to state-of-the-art methods.

Contribution

It introduces a novel approach to combine pre-trained speech and language models for end-to-end ASR, facilitating comprehensive optimization and leveraging recent LLM advancements.

Findings

01

Achieves performance comparable to modern E2E ASR models

02

Enables parameter-efficient domain adaptation

03

Facilitates inference optimization

Abstract

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, most of their applications in ASR involve only one of either a pre-trained speech or a language model. This paper proposes integrating a pre-trained speech representation model and a large language model (LLM) for E2E ASR. The proposed model enables the optimization of the entire ASR process, including acoustic feature extraction and acoustic and language modeling, by combining pre-trained models with a bridge network and also enables the application of remarkable developments in LLM utilization, such as parameter-efficient domain adaptation and inference optimization. Experimental results…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

🤗
yky-h/nue-asr
model· 9 dl· ♡ 4
9 dl♡ 4

Videos

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition· underline

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques · Topic Modeling