Systematic Categorization, Construction and Evaluation of New Attacks   against Multi-modal Mobile GUI Agents

Yulong Yang; Xinshan Yang; Shuaidong Li; Chenhao Lin; Zhengyu Zhao,; Chao Shen; Tianwei Zhang

arXiv:2407.09295·cs.CR·March 18, 2025

Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents

Yulong Yang, Xinshan Yang, Shuaidong Li, Chenhao Lin, Zhengyu Zhao,, Chao Shen, Tianwei Zhang

PDF

Open Access

TL;DR

This paper systematically investigates security vulnerabilities in multi-modal mobile GUI agents, introducing a new threat modeling approach, discovering 34 new attacks, and evaluating their feasibility through real-world case studies and experiments.

Contribution

It proposes a novel threat modeling methodology and an attack framework to systematically identify and evaluate security threats in multi-modal mobile GUI agents.

Findings

01

Discovered 34 previously unreported attacks.

02

Validated the feasibility and severity of these attacks.

03

Highlighted the urgent need for improved security measures.

Abstract

The integration of Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) into mobile GUI agents has significantly enhanced user efficiency and experience. However, this advancement also introduces potential security vulnerabilities that have yet to be thoroughly explored. In this paper, we present a systematic security investigation of multi-modal mobile GUI agents, addressing this critical gap in the existing literature. Our contributions are twofold: (1) we propose a novel threat modeling methodology, leading to the discovery and feasibility analysis of 34 previously unreported attacks, and (2) we design an attack framework to systematically construct and evaluate these threats. Through a combination of real-world case studies and extensive dataset-driven experiments, we validate the severity and practicality of those attacks, highlighting the pressing need for…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsInformation and Cyber Security · Spam and Phishing Detection · Multi-Agent Systems and Negotiation