Backdoor Directions in Vision Transformers

Sengim Karayalcin; Marina Krcek; Pin-Yu Chen; Stjepan Picek

arXiv:2603.10806·cs.CV·March 12, 2026

Backdoor Directions in Vision Transformers

Sengim Karayalcin, Marina Krcek, Pin-Yu Chen, Stjepan Picek

PDF

Open Access

TL;DR

This paper uncovers a specific trigger direction in Vision Transformers that corresponds to backdoor features, enabling diagnosis and detection of backdoor attacks through mechanistic interpretability.

Contribution

It introduces a linear trigger direction in ViT activations, providing a new diagnostic tool for understanding and detecting backdoor attacks in vision models.

Findings

01

Identified a trigger direction in ViT activations linked to backdoor features

02

Demonstrated causal influence of this direction on backdoor behavior

03

Proposed a weight-based detection scheme for stealthy triggers

Abstract

This paper investigates how Backdoor Attacks are represented within Vision Transformers (ViTs). By assuming knowledge of the trigger, we identify a specific ``trigger direction'' in the model's activations that corresponds to the internal representation of the trigger. We confirm the causal role of this linear direction by showing that interventions in both activation and parameter space consistently modulate the model's backdoor behavior across multiple datasets and attack types. Using this direction as a diagnostic tool, we trace how backdoor features are processed across layers. Our analysis reveals distinct qualitative differences: static-patch triggers follow a different internal logic than stealthy, distributed triggers. We further examine the link between backdoors and adversarial attacks, specifically testing whether PGD-based perturbations (de-)activate the identified trigger…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Advanced Malware Detection Techniques · Security and Verification in Computing