DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

Li Ren; Chen Chen; Liqiang Wang; Kien Hua

arXiv:2505.23694·cs.CV·June 3, 2025

DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

Li Ren, Chen Chen, Liqiang Wang, Kien Hua

PDF

Open Access 1 Repo

TL;DR

DA-VPT introduces a semantic-guided prompt tuning method for Vision Transformers, leveraging metric learning to improve fine-tuning efficiency and performance across recognition and segmentation tasks.

Contribution

The paper proposes a novel framework that guides prompt distributions using semantic information, enhancing parameter-efficient fine-tuning of ViT models.

Findings

01

Improved fine-tuning performance on benchmark datasets.

02

Effective semantic information sharing via prompts.

03

Enhanced efficiency in vision tasks.

Abstract

Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models by partially fine-tuning learnable tokens while keeping most model parameters frozen. Recent research has explored modifying the connection structures of the prompts. However, the fundamental correlation and distribution between the prompts and image tokens remain unexplored. In this paper, we leverage metric learning techniques to investigate how the distribution of prompts affects fine-tuning performance. Specifically, we propose a novel framework, Distribution Aware Visual Prompt Tuning (DA-VPT), to guide the distributions of the prompts by learning the distance metric from their class-related semantic data. Our method demonstrates that the prompts can serve as an effective bridge to share semantic information between image patches and the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

noahsark/da-vpt
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsCCD and CMOS Imaging Sensors · Advanced Memory and Neural Computing · Visual Attention and Saliency Detection

MethodsAttention Is All You Need · Linear Layer · Dense Connections · Vision Transformer · Softmax · Position-Wise Feed-Forward Layer · Absolute Position Encodings · Label Smoothing · Multi-Head Attention · Layer Normalization