Publication
Agentic AI Security
arXiv:2608.15888

Bounded Agents

Delegation Security for Multi-Agent AI Systems

A formal security architecture for delegated authority in multi-agent AI systems. The Agentic Principal Chain (APC) model enforces six deterministic, infrastructure-level checks before any action executes. If any one fails, the action is denied. The paper proves two properties: compromise damage can only shrink at each delegation hop, and no sequence of individually permitted actions can produce a prohibited outcome.

What is Bounded Agents?

Bounded Agents: Delegation Security for Multi-Agent AI Systems is a 2026 paper by Xabier Muruaga that introduces the Agentic Principal Chain (APC), an authorization architecture for AI agents and multi-agent systems. APC treats a delegated task as a bounded grant of authority: scope and budgets can only narrow as the task passes from a human to an orchestrator to sub-agents and tools, never expand. Every proposed action is checked against accumulated session state, not evaluated in isolation, so APC catches composition closure violations, cases where two individually permitted actions combine into a prohibited outcome, such as reading confidential data and then sending it externally. Enforcement runs in infrastructure outside the model, so a compromised or manipulated agent cannot grant itself authority it was never delegated.

Core Thesis

01

Prompt injection is an authorization architecture problem. The reason it's dangerous is not that the model follows a malicious instruction, it's that the model has been given authority to act on it.

02

The model must be excluded from its own trust boundary. An approval gate in a prompt is not a control; it is a suggestion. An approval gate in a Policy Enforcement Point is a control.

03

Per-action authorization is structurally insufficient. An agent authorized to read confidential documents and send external email can combine both to exfiltrate data, without violating any individual permission.

04

Agentic systems must assume breach. Some component will be compromised. The question is how much damage it can cause. Blast radius must be bounded by architecture.

The APC Model

The Agentic Principal Chain formalizes how authority flows from a human through an orchestrator to sub-agents and tools, with scope narrowing at every hop.

When a user delegates a task to an agent, the agent receives a bounded view: only the resources, actions, and data classifications relevant to the task, with certain action combinations explicitly prohibited. This is the Authorization Scope: a tuple of permitted resources, permitted actions, permitted data classifications, and prohibited compositions.

At each delegation hop, the scope narrows: permissions are intersected and restrictions can only accumulate; they can never be removed by a downstream agent. Beyond scope, Delegation Budgets impose six quantitative ceilings: delegation depth, cumulative blast radius, irreversible effects, sensitivity class, cross-domain composition, and compute cost.

The entire enforcement layer (scope computation, budget tracking, evidence generation) is executed by infrastructure, not by the agent. The model is structurally excluded from its own trust boundary.

Architecture OverviewTry the interactive demo→
Delegation ChainInfrastructure EnforcementHuman Principalp₀ · initial scope + budget⊓ narrowOrchestratorp₁ · narrowed scope⊓ narrowSub-Agentp₂ · most restrictedscope narrowsTool ExecutionC1 · C2a · C2b · C2c · C3 · C4 · C5 · C6Policy Decision Pointevaluates 6 conditionsEnforcement Pointadmit or denyadmit / denyEvidence Storetamper-evident, hash-chainedcommitproposeexecuteBR(p₂) ⊆ BR(p₁) ⊆ BR(p₀) · blast radius non-increasing (Theorem 1)Delegation chainInfrastructure enforcementEvidence

The Six Conditions

Every time an agent proposes an action, the system evaluates an all-or-nothing check across six conditions. If any one fails, the action is denied. Five of six passing is still a denial.

C1Identity Binding

Who is acting?

The acting agent must be bound to a verifiable identity in the delegation chain.

C2Scope Narrowing + Composition Closure

Is this action within bounds, alone and in combination?

Three subchecks: is the action within the agent's narrowed scope (2a), does it create a prohibited combination with prior actions (2b), and are all budget ceilings respected (2c).

C3Context & State Binding

Is this the right session?

Prevents replay attacks: a valid approval from one session cannot be reused in another.

C4Approval Binding

Is this high-impact? Who approved it?

Impact scoring with single-use approval tokens tied to the exact action, parameters, and session.

C5Evidence Commitment

Can we prove this happened?

If the evidence store is unreachable, the action is denied: no evidence trail means no execution.

C6Intent Binding

Is this relevant to the declared task?

Scope defines what an agent may do; intent defines what it should do, enforced by infrastructure.

Formal Results

Blast Radius Monotonicity

For any delegation chain, compromise damage is bounded by delegation depth.

If an attacker compromises a sub-agent three hops deep, the damage is bounded by the scope and budget at that position, no greater than what its delegating parent could have done.

Composition Soundness

If the composition restriction set covers at least one required action pair for every prohibited outcome, then no sequence of individually permitted actions can produce any prohibited outcome.

Reading confidential data is fine; sending external email is fine; doing both in the same session is exfiltration, and the system blocks it. The guarantee is as strong as the policy.

Intent Binding (Proposition)

Intent only restricts, never widens. An agent authorized to read all confidential documents but tasked with “summarize Q4 contracts” should not be reading the CEO's personnel file, even though it has scope to do so. Intent binding closes the gap between what an agent is authorized to do and what it should do.

Adversary Model

The model assumes partial compromise is inevitable. The question is not whether a component will be compromised, but how much damage it can cause.

Adversary Capabilities
  • A1Inject content into the agent’s context via untrusted data sources (indirect prompt injection)
  • A2Fully compromise a single actor in the chain (sub-agent, tool server, or orchestrator)
  • A3Observe which actions succeed or fail to probe scope boundaries
  • A4Maintain influence for the duration of a task session
Trust Boundaries
  • T1Cannot compromise PDP, PEP, evidence store, or key management infrastructure
  • T2Cannot forge cryptographic signatures or hashes
  • T3Cannot operate across session boundaries
APC Guarantees
IDGuaranteeEnforced by
G1No action outside scope executesC2a
G2No prohibited action pair co-occurs in a sessionC2b, Thm 2
G3Blast radius does not increase at each delegation hopThm 1
G4High-impact actions require valid approval tokensC4
G5Every admitted action is coupled to infrastructure-generated evidence or execution failure handlingC5
G6Actions outside declared intent are denied or flaggedC6

Empirical Validation

APC evaluated across six benchmarks, four AgentDojo domains, and 3,154 evaluation instances — including a compromised-model methodology that simulates full model compromise inside a live LLM pipeline.

Explore the full evaluation suite
Enforcement Latency
Operationp50p99
Full check (6 conditions)0.06 ms0.44 ms
Composition check (50 action types)0.007 ms0.03 ms
Scope narrowing (per delegation hop)0.03 ms0.24 ms

Overhead is < 0.5 ms p99, negligible vs. LLM inference (500–5000 ms). Throughput: > 7,000 evaluations/sec.

Attack Success Rate, with APC
0%

Exfiltration ASR — InjecAgent, ASB, compromised-model

4.0%

Destruction ASR, down from 38.6% with no defense

12.1%

Manipulation ASR, down from 90.5% with no defense

On InjecAgent (1,054 cases), all 544 data-stealing (read→exfiltrate) attacks are blocked at 0% ASR: composition closure catches every case because the prohibited pair (read, send_external) cannot co-occur in a session. On a compromised-model evaluation across all four AgentDojo domains (1,218 runs, ground-truth attack injected into a live agent), exfiltration is blocked at 0% in every domain. Residual manipulation ASR concentrates where the injected action overlaps the user's own declared intent, an explicit design boundary: APC enforces scope and composition, not per-parameter intent, so distinguishing "create event for user" from "create event for attacker" needs parameter-level validation on top of APC. Direct, single-action attacks that don't depend on a prior read (60.4% ASR on InjecAgent) fall outside composition closure for the same reason.

How APC Compares

Standalone coverage of each property by existing authorization mechanisms. APC is designed to layer on top of existing infrastructure, not to replace it.

PropertyOAuth+OPAPromptsStatic ManifestAPC
Identity binding✓––✓
Scope narrowingPartial–Tool-level✓
Composition closure–––✓
Blast radius monotonicity–––✓
Approval binding–––✓
Evidence commitmentPartial––✓
Intent binding–––✓
Survives injection (A1)Partial–PartialBounded
Survives compromise (A2)Partial–PartialBounded

Frequently Asked Questions

What is Bounded Agents?

Bounded Agents: Delegation Security for Multi-Agent AI Systems is a 2026 paper by Xabier Muruaga that introduces the Agentic Principal Chain (APC), an authorization architecture for AI agents. APC enforces six deterministic checks, outside the model, before any action executes: identity binding, scope and composition, session binding, approval binding, evidence commitment, and intent binding. The paper proves two formal properties, blast radius monotonicity and composition soundness, and validates them across 3,154 evaluation instances.

Who proposed the Agentic Principal Chain?

The Agentic Principal Chain (APC) was introduced by Xabier Muruaga in Bounded Agents: Delegation Security for Multi-Agent AI Systems (arXiv:2608.15888, 2026). Muruaga is an AI security architect who leads global AI and Data architecture, security, and technical governance at Iberdrola, and contributes to the OWASP GenAI Security Project and the Cloud Security Alliance AI Safety Working Group.

What problem does APC solve?

APC addresses a gap in ordinary agent authorization: permissions can stay individually valid while an agent drifts from its delegated task, over-delegates authority to sub-agents, or combines separately permitted actions into a prohibited outcome. APC carries an authority envelope through delegation, narrows scope and budgets at every hop, and checks each action against accumulated session state instead of evaluating it alone.

Why is per-action authorization insufficient for AI agents?

Stateless, per-request authorization evaluates each action independently of what came before it. An agent authorized to read confidential documents and, separately, authorized to send external email violates no individual permission by doing both in the same session, yet the combination is data exfiltration. APC's composition closure check evaluates proposed actions against accumulated session state specifically to catch this class of failure before execution.

What is composition closure in AI-agent authorization?

Composition closure is the property that no sequence of individually permitted actions can produce a prohibited outcome, provided the composition restriction set covers the required action pairs and admission is serialized. In Bounded Agents, this is Composition Soundness (Theorem 2): reading confidential data is permitted, sending data externally is permitted, but doing both in the same session is a prohibited composition that APC blocks. The guarantee is exactly as strong as the configured restriction set — an incomplete policy leaves gaps, the same way an incomplete firewall rule set does.

How does Bounded Agents relate to prompt injection?

Bounded Agents does not attempt to prevent prompt injection at the model layer, and does not claim to. It treats model compromise as a standing possibility in its adversary model and instead bounds what a compromised or manipulated agent is authorized to execute. Because enforcement runs in infrastructure outside the model, an injected instruction can only produce actions that fall within the agent's already-narrowed scope, budgets, and composition restrictions, containing the impact of an attack it does not detect.

Does APC replace OAuth, IAM or OPA?

No. APC is designed to layer on top of existing IAM, OAuth/OIDC, and policy engines such as OPA, not replace them. Those systems answer "is this identity authorized for this action right now?" APC adds the dimension they don't cover: authority that narrows through a delegation chain and is evaluated against everything the agent has already done in the session, including composition, budgets, and intent.

How does APC limit delegated privileges between AI agents?

When a task is delegated from a human through an orchestrator to sub-agents and tools, APC intersects permissions at every hop: restrictions can only accumulate, never be removed, by a downstream agent. Delegation budgets impose quantitative ceilings — delegation depth, cumulative blast radius, irreversible effects, sensitivity class, cross-domain composition, and compute cost — so a sub-agent three hops deep is bounded by the narrowest scope and tightest budget anywhere along its chain, no greater than what its delegating parent could have done.

Is APC specific to multi-agent systems?

APC is designed for multi-agent delegation chains — human to orchestrator to sub-agents to tools — where authority must narrow correctly at every hop. Its underlying mechanism, checking accumulated session state and composition before executing an action, also applies to a single tool-using agent, but the delegation-chain properties (blast radius monotonicity, scope narrowing across hops) are specifically about authority crossing multiple agent boundaries.

Can APC be used with MCP or A2A?

APC is an authorization architecture, not an MCP- or A2A-specific protocol. Its mechanism — checking actions against accumulated session state via infrastructure outside the model — applies wherever authority crosses an agent or tool boundary, which includes MCP tool calls and A2A delegation. The published evaluation targets tool-calling agent benchmarks (AgentDojo, InjecAgent, ASB) rather than a specific MCP or A2A deployment, so applying APC there means mapping its checks onto that protocol's own enforcement points.

What are the limitations of Bounded Agents?

APC operates at the authorization layer, not the data-validation layer: it controls which actions execute, not whether their parameters are correct. Composition restrictions are defined at the action-type level, so resource-qualified restrictions are future work. Composition Soundness depends on the completeness of the configured restriction set, a policy-quality problem analogous to firewall rule coverage. When no intent specification is provided, the intent check is skipped. And APC does not prevent a compromised agent from choosing a suboptimal action within its authorized scope and intent — that is an alignment boundary, not an authorization one.

How was Bounded Agents evaluated?

The reference implementation was evaluated across six benchmarks, four AgentDojo domains, and 3,154 evaluation instances, including a compromised-model methodology that injects a benchmark's ground-truth attack call into a live LLM pipeline. Results: 0% exfiltration ASR across InjecAgent, ASB, and the compromised-model suites; destruction ASR down from 38.6% to 4.0%; manipulation ASR down from 90.5% to 12.1%; and enforcement overhead under 0.5ms at p99. Full results are on the evaluation page.

Boundaries

  • –The model operates at the authorization layer, not the data validation layer, it controls which actions execute, not whether their parameters are correct.
  • –Composition restrictions operate at the action-type level; resource-qualified restrictions are identified as future work.
  • –The completeness of the composition restriction set is a policy quality problem, analogous to firewall rule coverage.
  • –When no intent specification is provided, the intent check is skipped and the remaining five conditions define the security boundary.
  • –The model does not prevent a compromised agent from choosing a suboptimal action within its authorized scope and intent. That is the alignment boundary.
Topics
Agentic AI
Delegation Security
Authorization
Composition Closure
Blast Radius
Intent Binding
Formal Methods
OWASP
MITRE ATLAS
Multi-Agent Systems
Policy Enforcement
Zero Trust
Related Reading

For the broader landscape this paper sits in, see Agentic AI Security. For a general explainer on why per-request authorization breaks down for AI agents, independent of this specific paper, see AI Agent Authorization.

How to Cite

Muruaga, X. (2026). Bounded Agents: Delegation Security for Multi-Agent AI Systems. arXiv:2608.15888. https://arxiv.org/abs/2608.15888

@misc{muruaga2026boundedagents,
  title  = {Bounded Agents: Delegation Security for Multi-Agent AI Systems},
  author = {Muruaga, Xabier},
  year   = {2026},
  eprint = {2608.15888},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2608.15888}
}
Continue Exploring

Related Case Studies