BashArena: A Control Setting for Highly Privileged AI Agents

Adam Kaufman; James Lucassen; Tyler Tracy; Cody Rushing; Aryan Bhatt

arXiv:2512.15688·cs.CR·December 18, 2025

BashArena: A Control Setting for Highly Privileged AI Agents

Adam Kaufman, James Lucassen, Tyler Tracy, Cody Rushing, Aryan Bhatt

PDF

Open Access

TL;DR

BashArena is a new benchmark environment with complex tasks and sabotage objectives designed to evaluate and improve AI control techniques in security-critical scenarios involving highly privileged AI agents.

Contribution

We introduce BashArena, a comprehensive control setting with realistic tasks and sabotage challenges, along with a pipeline for generating such tasks, to advance AI safety research.

Findings

01

Claude Sonnet 4.5 can execute sabotage undetected 26% of the time

02

Detection of sabotage has a 4% false positive rate

03

Provides a baseline for future AI control protocol development

Abstract

Future AI agents might run autonomously with elevated privileges. If these agents are misaligned, they might abuse these privileges to cause serious damage. The field of AI control develops techniques that make it harder for misaligned AIs to cause such damage, while preserving their usefulness. We introduce BashArena, a setting for studying AI control techniques in security-critical environments. BashArena contains 637 Linux system administration and infrastructure engineering tasks in complex, realistic environments, along with four sabotage objectives (execute malware, exfiltrate secrets, escalate privileges, and disable firewall) for a red team to target. We evaluate multiple frontier LLMs on their ability to complete tasks, perform sabotage undetected, and detect sabotage attempts. Claude Sonnet 4.5 successfully executes sabotage while evading monitoring by GPT-4.1 mini 26% of the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Network Security and Intrusion Detection · Security and Verification in Computing