MAC: A unified framework boosting low resource automatic speech   recognition

Zeping Min; Qian Ge; Zhong Li; Weinan E

arXiv:2302.03498·cs.CL·February 16, 2023

MAC: A unified framework boosting low resource automatic speech recognition

Zeping Min, Qian Ge, Zhong Li, Weinan E

PDF

Open Access

TL;DR

This paper introduces MAC, a unified Bayesian sampling-based framework that enhances low-resource automatic speech recognition by integrating a novel concatenative synthesis TTS system, significantly reducing error rates across multiple languages.

Contribution

The paper presents MAC, a novel low-resource ASR framework that combines Bayesian sampling with concatenative synthesis TTS, improving recognition accuracy and handling diverse languages and scenes.

Findings

01

MAC reduces CER by over 15% in multiple language ASR tasks.

02

MAC outperforms wav2vec2 on Cantonese, Taiwanese, and Japanese datasets.

03

Achieves 10.9% CER on Cantonese ASR, a 30% relative improvement.

Abstract

We propose a unified framework for low resource automatic speech recognition tasks named meta audio concatenation (MAC). It is easy to implement and can be carried out in extremely low resource environments. Mathematically, we give a clear description of MAC framework from the perspective of bayesian sampling. In this framework, we leverage a novel concatenative synthesis text-to-speech system to boost the low resource ASR task. By the concatenative synthesis text-to-speech system, we can integrate language pronunciation rules and adjust the TTS process. Furthermore, we propose a broad notion of meta audio set to meet the modeling needs of different languages and different scenes when using the system. Extensive experiments have demonstrated the great effectiveness of MAC on low resource ASR tasks. For CTC greedy search, CTC prefix, attention, and attention rescoring decode mode in…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Music and Audio Processing · Speech and Audio Processing