Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models

Chengyin Hu; Yuxian Dong; Yikun Guo; Xiang Chen; Junqi Wu; Jiahuan Long; Yiwei Wei; Tingsong Jiang; Wen Yao

arXiv:2604.03117·cs.CV·April 6, 2026

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models

Chengyin Hu, Yuxian Dong, Yikun Guo, Xiang Chen, Junqi Wu, Jiahuan Long, Yiwei Wei, Tingsong Jiang, Wen Yao

PDF

TL;DR

This paper introduces a universal physical adversarial patch framework for infrared vision-language models, exposing a significant robustness vulnerability in their semantic understanding capabilities.

Contribution

The authors propose UCGP, a novel physical adversarial patch method tailored for IR-VLMs, enhancing physical deployability and robustness against defenses.

Findings

01

UCGP effectively disrupts IR-VLM semantic understanding across multiple architectures.

02

The method demonstrates high transferability and generalization in real-world scenarios.

03

UCGP reveals a critical robustness vulnerability in infrared multimodal systems.

Abstract

Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception in low-visibility environments, yet their robustness to adversarial attacks remains largely unexplored. Existing adversarial patch methods are mainly designed for RGB-based models in closed-set settings and are not readily applicable to the open-ended semantic understanding and physical deployment requirements of infrared VLMs. To bridge this gap, we propose Universal Curved-Grid Patch (UCGP), a universal physical adversarial patch framework for IR-VLMs. UCGP integrates Curved-Grid Mesh (CGM) parameterization for continuous, low-frequency, and deployable patch generation with a unified representation-driven objective that promotes subspace departure, topology disruption, and stealth. To improve robustness under real-world deployment and domain shift, we further incorporate Meta…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.