
AI Security Architecture
Secure-by-design architecture for enterprise AI: threat models mapped to enforceable mitigations, deny-by-default gateway policies, and reusable control baselines across LLM, RAG, and agentic deployment patterns.
Delegation Security for Multi-Agent AI Systems
A formal security architecture for delegated authority in multi-agent AI systems. The Agentic Principal Chain (APC) model enforces six deterministic, infrastructure-level checks before any action executes. If any one fails, the action is denied. The paper proves two properties: compromise damage can only shrink at each delegation hop, and no sequence of individually permitted actions can produce a prohibited outcome.
Bounded Agents: Delegation Security for Multi-Agent AI Systems is a 2026 paper by Xabier Muruaga that introduces the Agentic Principal Chain (APC), an authorization architecture for AI agents and multi-agent systems. APC treats a delegated task as a bounded grant of authority: scope and budgets can only narrow as the task passes from a human to an orchestrator to sub-agents and tools, never expand. Every proposed action is checked against accumulated session state, not evaluated in isolation, so APC catches composition closure violations, cases where two individually permitted actions combine into a prohibited outcome, such as reading confidential data and then sending it externally. Enforcement runs in infrastructure outside the model, so a compromised or manipulated agent cannot grant itself authority it was never delegated.
Prompt injection is an authorization architecture problem. The reason it's dangerous is not that the model follows a malicious instruction, it's that the model has been given authority to act on it.
The model must be excluded from its own trust boundary. An approval gate in a prompt is not a control; it is a suggestion. An approval gate in a Policy Enforcement Point is a control.
Per-action authorization is structurally insufficient. An agent authorized to read confidential documents and send external email can combine both to exfiltrate data, without violating any individual permission.
Agentic systems must assume breach. Some component will be compromised. The question is how much damage it can cause. Blast radius must be bounded by architecture.
The Agentic Principal Chain formalizes how authority flows from a human through an orchestrator to sub-agents and tools, with scope narrowing at every hop.
When a user delegates a task to an agent, the agent receives a bounded view: only the resources, actions, and data classifications relevant to the task, with certain action combinations explicitly prohibited. This is the Authorization Scope: a tuple of permitted resources, permitted actions, permitted data classifications, and prohibited compositions.
At each delegation hop, the scope narrows: permissions are intersected and restrictions can only accumulate; they can never be removed by a downstream agent. Beyond scope, Delegation Budgets impose six quantitative ceilings: delegation depth, cumulative blast radius, irreversible effects, sensitivity class, cross-domain composition, and compute cost.
The entire enforcement layer (scope computation, budget tracking, evidence generation) is executed by infrastructure, not by the agent. The model is structurally excluded from its own trust boundary.
Every time an agent proposes an action, the system evaluates an all-or-nothing check across six conditions. If any one fails, the action is denied. Five of six passing is still a denial.
Who is acting?
The acting agent must be bound to a verifiable identity in the delegation chain.
Is this action within bounds, alone and in combination?
Three subchecks: is the action within the agent's narrowed scope (2a), does it create a prohibited combination with prior actions (2b), and are all budget ceilings respected (2c).
Is this the right session?
Prevents replay attacks: a valid approval from one session cannot be reused in another.
Is this high-impact? Who approved it?
Impact scoring with single-use approval tokens tied to the exact action, parameters, and session.
Can we prove this happened?
If the evidence store is unreachable, the action is denied: no evidence trail means no execution.
Is this relevant to the declared task?
Scope defines what an agent may do; intent defines what it should do, enforced by infrastructure.
For any delegation chain, compromise damage is bounded by delegation depth.
If an attacker compromises a sub-agent three hops deep, the damage is bounded by the scope and budget at that position, no greater than what its delegating parent could have done.
If the composition restriction set covers at least one required action pair for every prohibited outcome, then no sequence of individually permitted actions can produce any prohibited outcome.
Reading confidential data is fine; sending external email is fine; doing both in the same session is exfiltration, and the system blocks it. The guarantee is as strong as the policy.
Intent only restricts, never widens. An agent authorized to read all confidential documents but tasked with “summarize Q4 contracts” should not be reading the CEO's personnel file, even though it has scope to do so. Intent binding closes the gap between what an agent is authorized to do and what it should do.
The model assumes partial compromise is inevitable. The question is not whether a component will be compromised, but how much damage it can cause.
| ID | Guarantee | Enforced by |
|---|---|---|
| G1 | No action outside scope executes | C2a |
| G2 | No prohibited action pair co-occurs in a session | C2b, Thm 2 |
| G3 | Blast radius does not increase at each delegation hop | Thm 1 |
| G4 | High-impact actions require valid approval tokens | C4 |
| G5 | Every admitted action is coupled to infrastructure-generated evidence or execution failure handling | C5 |
| G6 | Actions outside declared intent are denied or flagged | C6 |
APC evaluated across six benchmarks, four AgentDojo domains, and 3,154 evaluation instances — including a compromised-model methodology that simulates full model compromise inside a live LLM pipeline.
Explore the full evaluation suite| Operation | p50 | p99 |
|---|---|---|
| Full check (6 conditions) | 0.06 ms | 0.44 ms |
| Composition check (50 action types) | 0.007 ms | 0.03 ms |
| Scope narrowing (per delegation hop) | 0.03 ms | 0.24 ms |
Overhead is < 0.5 ms p99, negligible vs. LLM inference (500–5000 ms). Throughput: > 7,000 evaluations/sec.
Exfiltration ASR — InjecAgent, ASB, compromised-model
Destruction ASR, down from 38.6% with no defense
Manipulation ASR, down from 90.5% with no defense
On InjecAgent (1,054 cases), all 544 data-stealing (read→exfiltrate) attacks are blocked at 0% ASR: composition closure catches every case because the prohibited pair (read, send_external) cannot co-occur in a session. On a compromised-model evaluation across all four AgentDojo domains (1,218 runs, ground-truth attack injected into a live agent), exfiltration is blocked at 0% in every domain. Residual manipulation ASR concentrates where the injected action overlaps the user's own declared intent, an explicit design boundary: APC enforces scope and composition, not per-parameter intent, so distinguishing "create event for user" from "create event for attacker" needs parameter-level validation on top of APC. Direct, single-action attacks that don't depend on a prior read (60.4% ASR on InjecAgent) fall outside composition closure for the same reason.
Standalone coverage of each property by existing authorization mechanisms. APC is designed to layer on top of existing infrastructure, not to replace it.
| Property | OAuth+OPA | Prompts | Static Manifest | APC |
|---|---|---|---|---|
| Identity binding | ✓ | – | – | ✓ |
| Scope narrowing | Partial | – | Tool-level | ✓ |
| Composition closure | – | – | – | ✓ |
| Blast radius monotonicity | – | – | – | ✓ |
| Approval binding | – | – | – | ✓ |
| Evidence commitment | Partial | – | – | ✓ |
| Intent binding | – | – | – | ✓ |
| Survives injection (A1) | Partial | – | Partial | Bounded |
| Survives compromise (A2) | Partial | – | Partial | Bounded |
Bounded Agents: Delegation Security for Multi-Agent AI Systems is a 2026 paper by Xabier Muruaga that introduces the Agentic Principal Chain (APC), an authorization architecture for AI agents. APC enforces six deterministic checks, outside the model, before any action executes: identity binding, scope and composition, session binding, approval binding, evidence commitment, and intent binding. The paper proves two formal properties, blast radius monotonicity and composition soundness, and validates them across 3,154 evaluation instances.
The Agentic Principal Chain (APC) was introduced by Xabier Muruaga in Bounded Agents: Delegation Security for Multi-Agent AI Systems (arXiv:2608.15888, 2026). Muruaga is an AI security architect who leads global AI and Data architecture, security, and technical governance at Iberdrola, and contributes to the OWASP GenAI Security Project and the Cloud Security Alliance AI Safety Working Group.
APC addresses a gap in ordinary agent authorization: permissions can stay individually valid while an agent drifts from its delegated task, over-delegates authority to sub-agents, or combines separately permitted actions into a prohibited outcome. APC carries an authority envelope through delegation, narrows scope and budgets at every hop, and checks each action against accumulated session state instead of evaluating it alone.
Stateless, per-request authorization evaluates each action independently of what came before it. An agent authorized to read confidential documents and, separately, authorized to send external email violates no individual permission by doing both in the same session, yet the combination is data exfiltration. APC's composition closure check evaluates proposed actions against accumulated session state specifically to catch this class of failure before execution.
Composition closure is the property that no sequence of individually permitted actions can produce a prohibited outcome, provided the composition restriction set covers the required action pairs and admission is serialized. In Bounded Agents, this is Composition Soundness (Theorem 2): reading confidential data is permitted, sending data externally is permitted, but doing both in the same session is a prohibited composition that APC blocks. The guarantee is exactly as strong as the configured restriction set — an incomplete policy leaves gaps, the same way an incomplete firewall rule set does.
Bounded Agents does not attempt to prevent prompt injection at the model layer, and does not claim to. It treats model compromise as a standing possibility in its adversary model and instead bounds what a compromised or manipulated agent is authorized to execute. Because enforcement runs in infrastructure outside the model, an injected instruction can only produce actions that fall within the agent's already-narrowed scope, budgets, and composition restrictions, containing the impact of an attack it does not detect.
No. APC is designed to layer on top of existing IAM, OAuth/OIDC, and policy engines such as OPA, not replace them. Those systems answer "is this identity authorized for this action right now?" APC adds the dimension they don't cover: authority that narrows through a delegation chain and is evaluated against everything the agent has already done in the session, including composition, budgets, and intent.
When a task is delegated from a human through an orchestrator to sub-agents and tools, APC intersects permissions at every hop: restrictions can only accumulate, never be removed, by a downstream agent. Delegation budgets impose quantitative ceilings — delegation depth, cumulative blast radius, irreversible effects, sensitivity class, cross-domain composition, and compute cost — so a sub-agent three hops deep is bounded by the narrowest scope and tightest budget anywhere along its chain, no greater than what its delegating parent could have done.
APC is designed for multi-agent delegation chains — human to orchestrator to sub-agents to tools — where authority must narrow correctly at every hop. Its underlying mechanism, checking accumulated session state and composition before executing an action, also applies to a single tool-using agent, but the delegation-chain properties (blast radius monotonicity, scope narrowing across hops) are specifically about authority crossing multiple agent boundaries.
APC is an authorization architecture, not an MCP- or A2A-specific protocol. Its mechanism — checking actions against accumulated session state via infrastructure outside the model — applies wherever authority crosses an agent or tool boundary, which includes MCP tool calls and A2A delegation. The published evaluation targets tool-calling agent benchmarks (AgentDojo, InjecAgent, ASB) rather than a specific MCP or A2A deployment, so applying APC there means mapping its checks onto that protocol's own enforcement points.
APC operates at the authorization layer, not the data-validation layer: it controls which actions execute, not whether their parameters are correct. Composition restrictions are defined at the action-type level, so resource-qualified restrictions are future work. Composition Soundness depends on the completeness of the configured restriction set, a policy-quality problem analogous to firewall rule coverage. When no intent specification is provided, the intent check is skipped. And APC does not prevent a compromised agent from choosing a suboptimal action within its authorized scope and intent — that is an alignment boundary, not an authorization one.
The reference implementation was evaluated across six benchmarks, four AgentDojo domains, and 3,154 evaluation instances, including a compromised-model methodology that injects a benchmark's ground-truth attack call into a live LLM pipeline. Results: 0% exfiltration ASR across InjecAgent, ASB, and the compromised-model suites; destruction ASR down from 38.6% to 4.0%; manipulation ASR down from 90.5% to 12.1%; and enforcement overhead under 0.5ms at p99. Full results are on the evaluation page.
For the broader landscape this paper sits in, see Agentic AI Security. For a general explainer on why per-request authorization breaks down for AI agents, independent of this specific paper, see AI Agent Authorization.
Muruaga, X. (2026). Bounded Agents: Delegation Security for Multi-Agent AI Systems. arXiv:2608.15888. https://arxiv.org/abs/2608.15888
@misc{muruaga2026boundedagents,
title = {Bounded Agents: Delegation Security for Multi-Agent AI Systems},
author = {Muruaga, Xabier},
year = {2026},
eprint = {2608.15888},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.15888}
}