AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound

Gijs Wijngaard; Elia Formisano; Michele Esposito; Michel Dumontier

arXiv:2505.14142·cs.SD·October 1, 2025

AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound

Gijs Wijngaard, Elia Formisano, Michele Esposito, Michel Dumontier

PDF

Open Access 1 Repo 3 Models 2 Datasets

TL;DR

AudSemThinker advances audio-language models by integrating semantic reasoning inspired by human cognition, supported by a new dataset, AudSem, which improves zero-shot sound understanding and outperforms existing models.

Contribution

The paper introduces AudSemThinker, a novel model for semantic audio reasoning, and AudSem, a curated dataset to facilitate fine-grained sound understanding.

Findings

01

AudSemThinker outperforms state-of-the-art models in multiple settings.

02

AudSem dataset effectively addresses data contamination in zero-shot evaluations.

03

The approach enhances reasoning over sound semantics in audio-language tasks.

Abstract

Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose reasoning is structured around a framework of auditory semantics inspired by human cognition. To support this, we introduce AudSem, a novel dataset specifically curated for semantic descriptor reasoning in audio-language models. AudSem addresses the persistent challenge of data contamination in zero-shot evaluations by providing a carefully filtered collection of audio samples paired with captions generated through a robust multi-stage pipeline. Our experiments demonstrate that AudSemThinker outperforms state-of-the-art models across multiple training settings, highlighting its strength in semantic audio reasoning. Both AudSemThinker and the AudSem…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

gljs/audsemthinker
pytorchOfficial

Models

Datasets

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech Recognition and Synthesis · Natural Language Processing Techniques