Mu$^{2}$SLAM: Multitask, Multilingual Speech and Language Models

Yong Cheng; Yu Zhang; Melvin Johnson; Wolfgang Macherey; Ankur Bapna

arXiv:2212.09553·cs.CL·June 28, 2023·5 cites

Mu$^{2}$SLAM: Multitask, Multilingual Speech and Language Models

Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey, Ankur Bapna

PDF

Open Access 1 Video

TL;DR

Mu$^{2}$SLAM is a comprehensive multilingual model trained on speech and text data, achieving state-of-the-art results in speech translation and competitive performance in text understanding across over 100 languages.

Contribution

The paper introduces Mu$^{2}$SLAM, a novel multitask, multilingual sequence-to-sequence model that jointly pre-trains on speech and text data for multiple tasks, improving cross-lingual and cross-modal understanding.

Findings

01

Sets new state-of-the-art on CoVoST AST translation tasks.

02

Matches performance of specialized ASR models on Voxpopuli.

03

Improves text understanding benchmarks by over 6%.

Abstract

We present Mu $^{2}$ SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AST) and Machine Translation (MT), in over 100 languages. By leveraging a quantized representation of speech as a target, Mu $^{2}$ SLAM trains the speech-text models with a sequence-to-sequence masked denoising objective similar to T5 on the decoder and a masked language modeling (MLM) objective on the encoder, for both unlabeled speech and text, while utilizing the supervised tasks to improve cross-lingual and cross-modal representation alignment within the model. On CoVoST AST, Mu $^{2}$ SLAM establishes a new state-of-the-art for models trained on public datasets, improving on xx-en translation over the previous best by 1.9 BLEU points and on en-xx translation by 1.1 BLEU…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Mu$^2$SLAM: Multitask, Multilingual Speech and Language Models· slideslive

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques · Topic Modeling

MethodsGated Linear Unit · Refunds@Expedia|||How do I get a full refund from Expedia? · Multi-Head Attention · Attention Is All You Need · Linear Layer · Inverse Square Root Schedule · Dense Connections · Attention Dropout · Residual Connection · Dropout