Semantic Data Augmentation for End-to-End Mandarin Speech Recognition

Jianwei Sun; Zhiyuan Tang; Hengxin Yin; Wei Wang; Xi Zhao; Shuaijiang; Zhao; Xiaoning Lei; Wei Zou; Xiangang Li

arXiv:2104.12521·eess.AS·April 27, 2021·Interspeech

Semantic Data Augmentation for End-to-End Mandarin Speech Recognition

Jianwei Sun, Zhiyuan Tang, Hengxin Yin, Wei Wang, Xi Zhao, Shuaijiang, Zhao, Xiaoning Lei, Wei Zou, Xiangang Li

PDF

Open Access

TL;DR

This paper introduces a semantic transposition data augmentation method for Mandarin end-to-end speech recognition, improving model performance by syntactically rearranging transcriptions and reassembling acoustic features.

Contribution

It presents a novel augmentation technique using syntax-based transposition of transcriptions and acoustic reassembly, enhancing Mandarin ASR accuracy.

Findings

01

Consistent performance improvements on Transformer and Conformer models.

02

Effective augmentation strategies and data ratios identified.

03

Semantic transposition enhances robustness of end-to-end ASR.

Abstract

End-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications. During the training of end-to-end ASR, data augmentation is a quite effective technique for regularizing the neural networks. This paper proposes a novel data augmentation technique based on semantic transposition of the transcriptions via syntax rules for end-to-end Mandarin ASR. Specifically, we first segment the transcriptions based on part-of-speech tags. Then transposition strategies, such as placing the object in front of the subject or swapping the subject and the object, are applied on the segmented sentences. Finally, the acoustic features corresponding to the transposed transcription are reassembled based on the audio-to-text forced-alignment produced by a pre-trained ASR system. The combination of original data and augmented one is used for training a new ASR…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Music and Audio Processing · Speech and Audio Processing