Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

Wenkui Yang; Chao Jin; Haisu Zhu; Weilin Luo; Derek Yuen; Kun Shao; Huaibo Huang; Junxian Duan; Jie Cao; Ran He

arXiv:2604.07831·cs.CR·April 10, 2026

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

Wenkui Yang, Chao Jin, Haisu Zhu, Weilin Luo, Derek Yuen, Kun Shao, Huaibo Huang, Junxian Duan, Jie Cao, Ran He

PDF

TL;DR

This paper introduces a practical method for attacking GUI agents by injecting semantic UI elements to mislead their visual grounding, revealing significant vulnerabilities across multiple models.

Contribution

It proposes a modular pipeline for semantic-level UI injection attacks that are effective, transferable across models, and demonstrate persistent influence on GUI agents.

Findings

01

Optimized attacks increase success rate by up to 4.4x over random injection.

02

Injected elements transfer effectively across different victim models.

03

Victims click attacker-controlled elements in over 15% of trials after initial success.

Abstract

Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavailable for commercial systems, while prompt injection is increasingly mitigated by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic-level UI Element Injection, a red-teaming setting that overlays safety-aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method uses a modular Editor-Overlapper-Victim pipeline and an iterative search procedure that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Across five victim models, our optimized attacks improve attack success rate by up to 4.4x over random injection on the strongest victims. Moreover, elements optimized…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.