Is Lip Region-of-Interest Sufficient for Lipreading?

Jing-Xuan Zhang; Gen-Shun Wan; Jia Pan

arXiv:2205.14295·cs.CV·June 3, 2022

Is Lip Region-of-Interest Sufficient for Lipreading?

Jing-Xuan Zhang, Gen-Shun Wan, Jia Pan

PDF

Open Access

TL;DR

This paper investigates whether using the entire face instead of just the lip region improves lipreading accuracy by leveraging additional facial information through self-supervised learning, showing significant error rate reductions.

Contribution

The study demonstrates that employing the entire face with self-supervised learning enhances lipreading performance compared to traditional lip-only approaches.

Findings

01

16% relative WER reduction with face input

02

Face input outperforms lip input with limited data

03

Face input slightly better with large training data

Abstract

Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redundant and irrelevant to visual speech recognition. However, faces contain much more detailed information than lips, such as speakers' head pose, emotion, identity etc. We argue that such information might benefit visual speech recognition if a powerful feature extractor employing the entire face is trained. In this work, we propose to adopt the entire face for lipreading with self-supervised learning. AV-HuBERT, an audio-visual multi-modal self-supervised learning framework, was adopted in our experiments. Our experimental results showed that adopting the entire face achieved 16% relative word error rate (WER) reduction on the lipreading task, compared with the baseline…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Face recognition and analysis