CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in   Visual Question Answering

Yuliang Cai; Mohammad Rostami

arXiv:2408.11742·cs.CV·August 22, 2024

CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering

Yuliang Cai, Mohammad Rostami

PDF

Open Access 1 Repo

TL;DR

CluMo introduces a novel prompt-based continual learning method for vision-language models, using cluster-based modality fusion prompts to improve generalization and mitigate catastrophic forgetting in visual question answering tasks.

Contribution

The paper proposes CluMo, a new prompt-based continual learning approach with a cluster-based modality fusion prompt and a two-stage training strategy for VLMs.

Findings

01

Achieves state-of-the-art performance on two benchmarks.

02

Effectively mitigates catastrophic forgetting in continual learning.

03

Enhances generalization in multimodal VQA tasks.

Abstract

Large vision-language models (VLMs) have shown significant performance boost in various application domains. However, adopting them to deal with several sequentially encountered tasks has been challenging because finetuning a VLM on a task normally leads to reducing its generalization power and the capacity of learning new tasks as well as causing catastrophic forgetting on previously learned tasks. Enabling using VLMs in multimodal continual learning (CL) settings can help to address such scenarios. To improve generalization capacity and prevent catastrophic forgetting, we propose a novel prompt-based CL method for VLMs, namely $Clu$ ster-based $Mo$ dality Fusion Prompt (\textbf{CluMo}). We design a novel \textbf{Key-Key-Prompt} pair, where each prompt is associated with a visual prompt key and a textual prompt key. We adopt a two-stage training strategy. During the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yuliangcai2022/clumo
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Advanced Image and Video Retrieval Techniques · Domain Adaptation and Few-Shot Learning