Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

Jiong Yin; Liang Li; Jiehua Zhang; Yuhan Gao; Chenggang Yan; Xichun Sheng

arXiv:2507.21588·cs.AI·July 30, 2025

Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

Jiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao, Chenggang Yan, Xichun Sheng

PDF

TL;DR

This paper introduces a three-stage progressive prompt tuning method for audio-visual multi-task incremental learning, effectively balancing knowledge retention and transfer across tasks.

Contribution

It proposes a novel PHP framework with task-shared, task-specific, and modality-independent prompts for improved continual learning performance.

Findings

01

Achieves state-of-the-art results on four audio-visual tasks.

02

Effectively balances knowledge sharing and task-specific adaptation.

03

Demonstrates robustness across different task orderings.

Abstract

Audio-visual multi-task incremental learning aims to continuously learn from multiple audio-visual tasks without the need for joint training on all tasks. The challenge of the problem is how to preserve the old task knowledge while facilitating the learning of new task with previous experiences. To address these challenges, we introduce a three-stage Progressive Homeostatic and Plastic audio-visual prompt (PHP) method. In the shallow phase, we design the task-shared modality aggregating adapter to foster cross-task and cross-modal audio-visual representation learning to enhance shared understanding between tasks. In the middle phase, we propose the task-specific modality-shared dynamic generating adapter, which constructs prompts that are tailored to individual tasks while remaining general across modalities, which balances the models ability to retain knowledge against forgetting with…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.