A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement   in Multimodal Hearing-aids

Abhijeet Bishnu; Ankit Gupta; Mandar Gogate; Kia Dashtipour; Ahsan; Adeel; Amir Hussain; Mathini Sellathurai; and Tharmalingam Ratnarajah

arXiv:2210.13127·eess.AS·October 25, 2022

A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids

Abhijeet Bishnu, Ankit Gupta, Mandar Gogate, Kia Dashtipour, Ahsan, Adeel, Amir Hussain, Mathini Sellathurai, and Tharmalingam Ratnarajah

PDF

Open Access

TL;DR

This paper presents a pioneering transceiver prototype for cloud-based audio-visual speech enhancement tailored for multimodal hearing aids, addressing high data rate, low latency, and real-time processing challenges.

Contribution

It introduces novel uplink and downlink frame structures optimized for AV speech enhancement in hearing aids, with real-time implementation and evaluation.

Findings

01

Supports multiple data rates for uplink channels

02

Achieves low latency suitable for real-time AV processing

03

Demonstrates feasibility with software defined radio hardware

Abstract

In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hearing assistive technology. The innovative design needs to meet multiple challenging constraints including up/down link communications, delay of transmission and signal processing, and real-time AV SE models processing. The transceiver includes device detection, frame detection, frequency offset estimation, and channel estimation capabilities. We develop both uplink (hearing aid to the cloud) and downlink (cloud to hearing aid) frame structures based on the data rate and latency requirements. Due to the varying nature of uplink information (audio and lip-reading), the uplink channel supports multiple data rate frame structure, while the downlink channel has a fixed data…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Advanced Adaptive Filtering Techniques · Hearing Loss and Rehabilitation