A Theoretical Game of Attacks via Compositional Skills

Xinbo Wu; Huan Zhang; Abhishek Umrawal; Lav R. Varshney

arXiv:2605.01034·cs.CL·May 5, 2026

A Theoretical Game of Attacks via Compositional Skills

Xinbo Wu, Huan Zhang, Abhishek Umrawal, Lav R. Varshney

PDF

TL;DR

This paper develops a theoretical framework modeling adversarial attacks and defenses on large language models, revealing optimal strategies and demonstrating improved attack performance empirically.

Contribution

It introduces a formal game-theoretic model for adversarial prompts, derives optimal attack and defense strategies, and empirically validates the attack's effectiveness.

Findings

01

Theoretical best-response attack closely relates to existing methods.

02

Optimal defense strategy can be derived with provable guarantees.

03

Empirical evaluation shows stronger attack performance across models and benchmarks.

Abstract

As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we introduce a theoretical framework that formalizes a game between an attacker and a defender. Within this framework, we design a theoretical best-response attack strategy and show that it is closely related to many existing adversarial prompting methods. We further analyze the resulting game, characterize its equilibria, and reveal inherent advantages for the attacker. Drawing on our theoretical analysis, we also derive a provably optimal defense strategy. Empirically, we evaluate a practical instantiation of the theoretically optimal attack and observe stronger performance relative to existing adversarial…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.