GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

Julien Guinot; Elio Quinton; Gy\"orgy Fazekas

arXiv:2506.17886·cs.SD·June 25, 2025

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

Julien Guinot, Elio Quinton, Gy\"orgy Fazekas

PDF

1 Repo

TL;DR

GD-Retriever introduces a diffusion-based framework that enhances controllability and performance in text-music retrieval by generating and manipulating queries within a latent space, enabling interactive and flexible retrieval results.

Contribution

The paper presents a novel diffusion model-based retrieval framework that improves performance and offers controllability and interactivity in text-music retrieval tasks.

Findings

01

Outperforms contrastive teacher models in retrieval accuracy

02

Supports retrieval in audio-only latent spaces with non-joint encoders

03

Enables post-hoc manipulation of retrieval behavior

Abstract

Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems controllable and interactive for users. In text-music retrieval, the ambiguity of freeform language creates a many-to-many mapping, often resulting in inflexible or unsatisfying results. We introduce Generative Diffusion Retriever (GDR), a novel framework that leverages diffusion models to generate queries in a retrieval-optimized latent space. This enables controllability through generative tools such as negative prompting and denoising diffusion implicit models (DDIM) inversion, opening a new direction in retrieval control. GDR improves retrieval performance over contrastive teacher models and supports retrieval in audio-only latent spaces using…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

pliploop/gdretriever
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.