E2E Spoken Entity Extraction for Virtual Agents

Karan Singla; Yeon-Jun Kim; Srinivas Bangalore

arXiv:2302.10186·eess.AS·November 13, 2023

E2E Spoken Entity Extraction for Virtual Agents

Karan Singla, Yeon-Jun Kim, Srinivas Bangalore

PDF

Open Access

TL;DR

This paper presents a direct speech-based entity extraction method for virtual agents that outperforms traditional transcription followed by text extraction, by fine-tuning pre-trained speech encoders to focus on relevant entity segments.

Contribution

It introduces a novel end-to-end approach for spoken entity extraction that bypasses transcription, improving accuracy and efficiency in virtual agent dialogues.

Findings

01

Direct speech-based extraction outperforms 2-step methods.

02

Fine-tuning speech encoders enhances entity extraction accuracy.

03

Approach reduces processing steps and improves relevance focus.

Abstract

In human-computer conversations, extracting entities such as names, street addresses and email addresses from speech is a challenging task. In this paper, we study the impact of fine-tuning pre-trained speech encoders on extracting spoken entities in human-readable form directly from speech without the need for text transcription. We illustrate that such a direct approach optimizes the encoder to transcribe only the entity relevant portions of speech ignoring the superfluous portions such as carrier phrases, or spell name entities. In the context of dialog from an enterprise virtual agent, we demonstrate that the 1-step approach outperforms the typical 2-step approach which first generates lexical transcriptions followed by text-based entity extraction for identifying spoken entities.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Speech and dialogue systems