Interpreting Agent Behaviors in Reinforcement-Learning-Based Cyber-Battle Simulation Platforms

Jared Claypoole; Steven Cheung; Ashish Gehani; Vinod Yegneswaran; and Ahmad Ridley

arXiv:2506.08192·cs.CR·June 11, 2025

Interpreting Agent Behaviors in Reinforcement-Learning-Based Cyber-Battle Simulation Platforms

Jared Claypoole, Steven Cheung, Ashish Gehani, Vinod Yegneswaran, and Ahmad Ridley

PDF

Open Access

TL;DR

This paper analyzes reinforcement learning agents in a cyber defense simulation, providing interpretability of their behaviors, evaluating effectiveness of actions and decoys, and discussing the realism of the simulation platform.

Contribution

It introduces methods to interpret complex agent behaviors in cyber simulations and evaluates the impact of decoys and action effectiveness within the CAGE Challenge environment.

Findings

01

Defenders clear infiltrations within 1-2 steps in most cases.

02

Certain actions are between 40% and 99% ineffective.

03

Decoys block up to 94% of exploits that would grant privileged access.

Abstract

We analyze two open source deep reinforcement learning agents submitted to the CAGE Challenge 2 cyber defense challenge, where each competitor submitted an agent to defend a simulated network against each of several provided rules-based attack agents. We demonstrate that one can gain interpretability of agent successes and failures by simplifying the complex state and action spaces and by tracking important events, shedding light on the fine-grained behavior of both the defense and attack agents in each experimental scenario. By analyzing important events within an evaluation episode, we identify patterns in infiltration and clearing events that tell us how well the attacker and defender played their respective roles; for example, defenders were generally able to clear infiltrations within one or two timesteps of a host being exploited. By examining transitions in the environment's…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Advanced Malware Detection Techniques · Network Security and Intrusion Detection