Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

Bronislav Sidik; Lior Rokach

arXiv:2604.11839·cs.CR·May 5, 2026

Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

Bronislav Sidik, Lior Rokach

PDF

TL;DR

This paper introduces Aethelgard, a four-layer adaptive framework that enforces least privilege in autonomous AI agents by learning minimal capability sets and dynamically controlling tool access.

Contribution

It presents a novel multi-layered approach combining learned policies and rule-based filtering to improve capability governance for AI agents.

Findings

01

The RL policy effectively learns minimal tool sets for different tasks.

02

The framework reduces overprovisioning by dynamically scoping agent capabilities.

03

Hybrid filtering improves safety by intercepting tool calls before execution.

Abstract

Autonomous AI agents built on open-source runtimes such as OpenClaw expose every available tool to every session by default, regardless of the task. A summarization task receives the same shell execution, subagent spawning, and credential access capabilities as a code deployment task, a 15x overprovision ratio that we call the capability overprovisioning problem. Existing defenses, including the NemoClaw container sandbox and the Cisco DefenseClaw skill scanner, address containment and threat detection but do not learn the minimum viable capability set for each task type. We present Aethelgard, a four layer adaptive governance framework that enforces least privilege for AI agents through a learned policy. Layer 1, the Capability Governor, dynamically scopes which tools the agent is aware of in each session. Layer 3, the Safety Router, intercepts tool calls before execution using a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.