Progressive Voice Trigger Detection: Accuracy vs Latency

Siddharth Sigtia; John Bridle; Hywel Richards; Pascal Clark; Erik; Marchi; Vineet Garg

arXiv:2010.15446·eess.AS·March 3, 2021

Progressive Voice Trigger Detection: Accuracy vs Latency

Siddharth Sigtia, John Bridle, Hywel Richards, Pascal Clark, Erik, Marchi, Vineet Garg

PDF

TL;DR

This paper introduces a progressive voice trigger detection system that balances accuracy and latency by using additional audio context after trigger phrases, significantly reducing false rejections with minimal delay.

Contribution

The work proposes a two-stage architecture that adaptively delays decisions to improve trigger detection accuracy without substantially increasing latency.

Findings

01

66% relative reduction in false rejection rate

02

Only 3% of triggers delayed for additional context

03

Negligible increase in overall latency

Abstract

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio context after a detected trigger phrase, we can indeed get a more accurate decision. However, waiting to listen to more audio each time incurs a latency increase. Progressive Voice Trigger Detection allows us to trade-off latency and accuracy by accepting clear trigger candidates quickly, but waiting for more context to decide whether to accept more marginal examples. Using a two-stage architecture, we show that by delaying the decision for just 3% of detected true triggers in the test set, we are able to obtain a relative improvement of 66% in false rejection rate, while incurring only a negligible increase in latency.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.