Pretraining End-to-End Keyword Search with Automatically Discovered   Acoustic Units

Bolaji Yusuf; Jan "Honza" \v{C}ernock\'y; Murat Sara\c{c}lar

arXiv:2407.04652·eess.AS·July 8, 2024

Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Bolaji Yusuf, Jan "Honza" \v{C}ernock\'y, Murat Sara\c{c}lar

PDF

Open Access 2 Repos

TL;DR

This paper introduces a pretraining method for end-to-end keyword search systems using automatically discovered acoustic units from untranscribed speech data, significantly improving performance over training from scratch.

Contribution

It proposes a novel pretraining approach for E2E KWS leveraging acoustic unit discovery, enhancing performance across languages and AUD systems.

Findings

01

Pretraining with AUD improves E2E KWS performance.

02

Performance correlates with AUD quality.

03

Finetuning outperforms training from scratch.

Abstract

End-to-end (E2E) keyword search (KWS) has emerged as an alternative and complimentary approach to conventional keyword search which depends on the output of automatic speech recognition (ASR) systems. While E2E methods greatly simplify the KWS pipeline, they generally have worse performance than their ASR-based counterparts, which can benefit from pretraining with untranscribed data. In this work, we propose a method for pretraining E2E KWS systems with untranscribed data, which involves using acoustic unit discovery (AUD) to obtain discrete units for untranscribed data and then learning to locate sequences of such units in the speech. We conduct experiments across languages and AUD systems: we show that finetuning such a model significantly outperforms a model trained from scratch, and the performance improvements are generally correlated with the quality of the AUD system used for…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Advanced Text Analysis Techniques · Speech Recognition and Synthesis