Audiopedia: Audio QA with Knowledge

Abhirama Subramanyam Penamakuri; Kiran Chhatre; Akshat Jain

arXiv:2412.20619·cs.LG·December 31, 2024

Audiopedia: Audio QA with Knowledge

Abhirama Subramanyam Penamakuri, Kiran Chhatre, Akshat Jain

PDF

Open Access 1 Repo

TL;DR

Audiopedia introduces a new audio question answering task that combines audio comprehension with external knowledge reasoning, proposing a framework to enhance large audio language models' performance on knowledge-intensive audio tasks.

Contribution

The paper defines a novel knowledge-intensive Audio Question Answering task and proposes a generic framework to improve large audio language models' reasoning capabilities.

Findings

01

Benchmarking shows current models perform suboptimally on Audiopedia tasks.

02

The proposed framework with AEL and KA2LM components improves model performance.

03

First work to address advanced audio understanding with knowledge reasoning.

Abstract

In this paper, we introduce Audiopedia, a novel task called Audio Question Answering with Knowledge, which requires both audio comprehension and external knowledge reasoning. Unlike traditional Audio Question Answering (AQA) benchmarks that focus on simple queries answerable from audio alone, Audiopedia targets knowledge-intensive questions. We define three sub-tasks: (i) Single Audio Question Answering (s-AQA), where questions are answered based on a single audio sample, (ii) Multi-Audio Question Answering (m-AQA), which requires reasoning over multiple audio samples, and (iii) Retrieval-Augmented Audio Question Answering (r-AQA), which involves retrieving relevant audio to answer the question. We benchmark large audio language models (LALMs) on these sub-tasks and observe suboptimal performance. To address this, we propose a generic framework that can be adapted to any LALM, equipping…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Abhiram4572/Audiopedia
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech Recognition and Synthesis · Speech and Audio Processing

MethodsFocus