Audio-Visual Speaker Tracking: Progress, Challenges, and Future   Directions

Jinzheng Zhao; Yong Xu; Xinyuan Qian; Davide Berghi; Peipei Wu; Meng; Cui; Jianyuan Sun; Philip J.B. Jackson; Wenwu Wang

arXiv:2310.14778·cs.MM·April 15, 2025·1 cites

Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions

Jinzheng Zhao, Yong Xu, Xinyuan Qian, Davide Berghi, Peipei Wu, Meng, Cui, Jianyuan Sun, Philip J.B. Jackson, Wenwu Wang

PDF

Open Access

TL;DR

This paper provides a comprehensive survey of audio-visual speaker tracking, covering Bayesian and deep learning methods, datasets, and future challenges, highlighting recent progress and research directions in the field.

Contribution

It is the first extensive survey over five years, summarizing Bayesian filters, deep learning techniques, datasets, and connections to related areas in audio-visual speaker tracking.

Findings

01

Bayesian filters effectively fuse audio-visual data for speaker tracking.

02

Deep learning enhances measurement extraction and state estimation.

03

Performance benchmarks on the AV16.3 dataset demonstrate current methods' effectiveness.

Abstract

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With audio and visual information, the Bayesian-based filter and deep learning-based methods can solve the problem of data association, audio-visual fusion and track management. In this paper, we conduct a comprehensive overview of audio-visual speaker tracking. To our knowledge, this is the first extensive survey over the past five years. We introduce the family of Bayesian filters and summarize the methods for obtaining audio-visual measurements. In addition, the existing trackers and their performance on the AV16.3 dataset are summarized. In the past few years, deep learning techniques have thrived, which also boost the development of audio-visual speaker…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Music and Audio Processing · Advanced Adaptive Filtering Techniques