Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

Xin Liu; Aoyang Zhou; Aoyang Zhou

arXiv:2511.00831·cs.CV·November 4, 2025

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

Xin Liu, Aoyang Zhou, Aoyang Zhou

PDF

Open Access

TL;DR

This paper introduces LSSA, a novel attack method that improves the transferability of adversarial examples in visual-language models by shuffling image blocks and sampling, leading to more effective cross-modal attacks.

Contribution

The paper proposes LSSA, a new attack strategy that enhances adversarial transferability in VLP models by increasing input diversity through local shuffling and sampling.

Findings

01

LSSA significantly improves transferability across models and tasks.

02

LSSA outperforms existing advanced attack methods.

03

LSSA demonstrates robustness on multiple datasets.

Abstract

Visual-Language Pre-training (VLP) models have achieved significant performance across various downstream tasks. However, they remain vulnerable to adversarial examples. While prior efforts focus on improving the adversarial transferability of multimodal adversarial examples through cross-modal interactions, these approaches suffer from overfitting issues, due to a lack of input diversity by relying excessively on information from adversarial examples in one modality when crafting attacks in another. To address this issue, we draw inspiration from strategies in some adversarial training methods and propose a novel attack called Local Shuffle and Sample-based Attack (LSSA). LSSA randomly shuffles one of the local image blocks, thus expanding the original image-text pairs, generating adversarial images, and sampling around them. Then, it utilizes both the original and sampled images to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Multimodal Machine Learning Applications · Generative Adversarial Networks and Image Synthesis