5. Layer 3: Secure Your AI Infrastructure
Introduction
Chapter 2 Section 4 ends by converting an unanswerable question – are we exposed to infrastructure attacks? – into four answerable ones:
Who can get an image scheduled on our GPU nodes? Who can edit a workflow in our orchestrator? Who shares an inference backend with our tenants? What does our public endpoint return besides text?
It then says those are “what Layer 2 and Layer 3 exist to close.” This section is the Layer 3 half of that answer, and the four questions are a fair test of it: if you finish here unable to answer them about a system you are handed, the section has failed regardless of how much posture-management vocabulary it taught you.
Layer 3 owns the stack underneath the model – GPU nodes, the serving container, the inference server’s internal sockets, the caches it shares between requests, the orchestrator that wires tools to the model, and the identities all of it runs as. It is also the layer most often reduced to a single product purchase, which is the failure this section is written against.
What will I get out of this?
By the end of this section, you will be able to:
- State what Layer 3 can and cannot stop – distinguishing its detective half, which reports and never intercepts, from its configured-boundary half, which enforces without ever seeing a request.
- Determine which Layer 3 controls are yours for a given deployment pattern, and which belong to a provider.
- Explain what AI-SPM adds to traditional CSPM – discovery scope, AI-specific baselines, and capability-aware risk scoring – and what it does not add.
- Specify posture and runtime controls for GPU nodes, serving containers, model endpoints, internal serving-stack sockets, and shared caches, against the Chapter 2 attacks that target each.
- Harden the orchestration layer – MCP server verification that survives a rug pull, tool allowlisting, per-agent identity, and mediated delegation.
- Design AI identity so the confused deputy cannot form, evaluating authorization against the requesting user’s entitlements rather than the tool’s.
- Map Layer 3 to the ten OWASP categories that name it as a primary defense, stating for each what the layer does and what it leaves open.
What Layer 3 Owns – and What It Cannot Stop
Section 2’s comparison table records Layer 3 as detective, with a flat No under “stops a live request,” and the second-broadest coverage in the Blueprint at 10 of 20 OWASP categories. Both halves of that are correct, and taken alone they mislead, because Layer 3 is really two different kinds of control wearing one number.
The detective half – AI-SPM – reports and blocks nothing. It tells you a GPU node is over-permissioned, an endpoint is unauthenticated, an API key has no rotation policy, an MCP server appeared last Tuesday. Every one of those is a real finding and none of them is an interception. Section 2 puts it bluntly: a layer with high coverage and no stopping power is telling you that a control somewhere else failed or was never configured.
The configured-boundary half enforces continuously and is still never in the request path. A seccomp profile denies a syscall. An egress rule drops a packet. A scoped credential refuses an API call. A per-tenant cache never holds another tenant’s prefix. These stop things – but none of them inspects a prompt, a response, or a tool argument. They are the permission shape, decided before the run, which is the phrasing Chapter 2 Section 6 already uses for excessive agency. Section 2’s “No” is a statement about the request path, and it is exactly right: nothing here reads the request.
The characteristic Layer 3 failure is buying the detective half and calling the layer done
A posture dashboard is the visible, purchasable part of Layer 3, so it is the part that gets deployed. It is also the part that stops nothing. An organisation with an AI-SPM console reporting green, no runtime policy on its serving containers, and one broad service account per agent has bought observability over an unenforced boundary – and the console will not report that, because an absent control is not a misconfiguration.
Section 2 names the general form: “three posture tools reporting the same misconfiguration still block nothing.” Read every finding in this section as belonging to one half or the other, and check that you have deployed both.
The trade Layer 3 makes is the inverse of Layer 2’s. Layer 2 acts once, before anything runs, and is finished. Layer 3 acts continuously and never gets to see what the system is actually being asked to do.
Which Layer 3 Controls Are Yours
Chapter 1 Section 3 promised that “the layers stay constant, but which ones are yours is set by the choice you make here.” Layer 3 is where that promise is most expensive to ignore, because most of this section describes machinery that a cloud-API consumer will never touch – and a smaller, unglamorous subset that is theirs no matter what they chose.
Read this against the five deployment patterns from Chapter 1.
| Layer 3 control | Cloud API | Serverless inference | Self-hosted | Edge / on-device |
|---|---|---|---|---|
| AI asset discovery / shadow AI | You – and it is mostly which services your people call | You | You | You, plus fleet inventory |
| GPU cluster hardening and tenancy | Provider | Provider | You – including the container toolkit | N/A – the device is the node |
| Serving container runtime policy | Provider | Provider | You – seccomp, mounts, egress, admission | The app sandbox is yours |
| Model endpoint authn / rate limits | Provider offers it; configuring it is yours | You | You – the whole control | Local – no network endpoint |
| Internal serving-stack sockets | Provider | Provider | You – and this is the one teams miss | N/A |
| Prefix / KV cache tenancy | Provider – a vendor-assessment question | Provider, within your tenancy | You | Single-tenant by construction |
| Orchestrator hardening | You – it runs on your side | You | You | You |
| Agent and service-account identity scope | You | You | You | You |
| Cloud AI service IAM and network config | Partly – API key scope and rotation | You – the largest control you own | N/A | N/A |
The fifth pattern, hybrid, has no column because it inherits the row of whichever path a request takes. A router that can fail over from a cloud API to a self-hosted fallback has silently handed you the self-hosted column.
Three consequences, and they are the point of the table:
1 · On a Cloud API, Layer 3 does not shrink to nothing – it moves. You own no GPU, no container, no socket, no cache. What remains is the orchestrator, the identity scope of every credential your agents hold, and discovery of what your organisation is actually calling. Those are not leftovers: the orchestrator is Chapter 2’s highest-value target in the stack and it runs on your infrastructure in every pattern here.
2 · Self-hosting takes the whole column at once, and the rows differ in how obvious they are. Endpoint authentication is obvious and gets done. The inference server’s internal ZeroMQ sockets, and whether the prefix cache is shared across tenants, are not obvious and do not get done – which is precisely why ShadowMQ and CVE-2025-46570 are in Chapter 2.
3 · A finding written against a control you do not own is not a finding. Run this table before the checklist at the end of the section. An organisation consuming a managed API that opens a finding about missing seccomp profiles on its serving containers has written a finding about someone else’s containers.
AI Security Posture Management (AI-SPM)
Traditional CSPM monitors cloud resources for misconfigurations: open storage buckets, permissive security groups, unencrypted databases. AI-SPM extends that to AI-specific resources, and the extension that matters most is discovery – many organisations cannot enumerate their own AI footprint.
Standard CSPM sees a GPU instance as a virtual machine. It does not know the instance is serving an unauthenticated inference API, that the vector database behind it has no access controls, or that an MCP server was connected to the agent framework last week.
| Dimension | Traditional CSPM | AI-SPM |
|---|---|---|
| Asset types | VMs, databases, storage buckets, network resources | Model endpoints, GPU clusters, vector stores, orchestrators, agent frameworks – MCP servers only where the tool reaches them |
| Configuration baselines | CIS benchmarks, provider best practices | AI-specific: endpoint authentication, cache tenancy, tool access policy, agent credential scope |
| Risk context | Data sensitivity, compliance requirements | Data sensitivity + model capability + tool access + agent autonomy |
| Discovery scope | Resources in managed cloud accounts | The AI-relevant subset of those resources, classified by AI role rather than by instance type |
| Drift detection | Configuration change on a cloud resource | That, plus new model deployments, new tool integrations, and permission changes on AI service accounts |
Discovery scope is the row that does the most work, and it is also the row where the category’s marketing and its products diverge, so read it carefully. What an AI-SPM tool reliably gives you is AI-aware classification of assets it can already see – it knows the GPU instance is serving a model and the bucket holds embeddings. What it does not automatically give you is reach beyond its connection scope.
Check the connection scope before you count on the coverage
TrendAI Vision One’s AI-SPM is a tab within Cloud Security Posture, and its documentation is specific about what feeds it: it “gives you visibility into the cloud assets your organization uses to build AI services”, summarising “the AI-related assets in your connected cloud accounts” across five categories – services, models, workloads, data storage and entitlements. Coverage is also uneven per cloud: as of 17 August 2026, “only model and service assets are currently supported for Azure.” The feature is documented as Pre-release, which is not a reason to discount it and is a reason to verify the matrix rather than the datasheet.
Two consequences for the estate this chapter cares about. An asset in no connected cloud account is in no posture report – a model running on a developer’s workstation, an MCP server configured in someone’s IDE, a departmental subscription nobody onboarded. And an employee’s use of a third-party AI service is not a posture finding at all: that is an access-path question, and in this platform it belongs to Layer 5’s ZTSA AI Service Access, which sees AI service usage because it sits on the path to it. Posture management sees what you own; access control sees what you reach. Filing shadow AI under Layer 3 is the mistake this distinction exists to prevent.
An organisation can hold a clean CSPM score across every managed account while running AI assets that were never in one – and, importantly, a clean AI-SPM score has the same blind spot, bounded by the accounts you connected.
The AI-SPM Workflow
graph TB
DIS["<b>Discover</b><br/><small>Inventory the AI assets<br/>in connected scope:<br/>model endpoints, GPU clusters,<br/>vector stores, orchestrators</small>"]
ASS["<b>Assess</b><br/><small>Evaluate configuration<br/>against AI security<br/>baselines: authentication,<br/>tenancy, credential scope</small>"]
SCR["<b>Score</b><br/><small>Risk scoring based on<br/>exposure, data sensitivity,<br/>tool access, and<br/>blast radius</small>"]
REM["<b>Remediate</b><br/><small>Prioritized guidance:<br/>what to fix first,<br/>how to fix it,<br/>expected risk reduction</small>"]
DIS --> ASS --> SCR --> REM
REM -.->|"Continuous<br/>reassessment"| DIS
NOTE["<b>Every arrow here is a report.</b><br/><small>Nothing in this loop is in the<br/>request path. Remediation is a<br/>change someone makes, not a block.</small>"]
REM -.- NOTE
style DIS fill:#2d5016,color:#fff
style ASS fill:#2d5016,color:#fff
style SCR fill:#2d5016,color:#fff
style REM fill:#2d5016,color:#fff
style NOTE fill:#a85800,color:#fff
The loop is continuous because the estate is. A developer spins up a test endpoint, a team connects an MCP server, an internal tool starts calling an external AI API. AI-SPM’s job is to notice each of those that falls inside the scope it is connected to, and to notice without waiting for a review cycle.
Its limits follow from the same two properties. Discovery is triggered by an asset existing, so it reports after the fact. For a resource that sits there being misconfigured – an unauthenticated endpoint, an over-scoped role – after the fact is early enough, and this is where AI-SPM earns its place. For an attack that completes in one action, it is not, and no configuration of the tool changes that.
Defense Connection
AI-SPM is Layer 3’s contribution to LLM03: Excessive Agency and LLM04: Supply Chain – it is the inventory that tells you which agents hold which permissions and which components entered the estate unreviewed. Note what it is not: the master mapping assigns LLM06: Unbounded Consumption to Layers 5 and 6, not here. Rate limiting is Layer 5’s. Layer 3 contributes the consumption anomaly detection Chapter 2 Section 4 points here for – and Chapter 2 is explicit that the load-bearing control against a runaway bill is a spend cap with an alert below it, which is a quota, not a posture finding.
Posture Management for AI Resources
Different AI resources need different baselines, and the four Chapter 2 infrastructure cases land on four different ones.
GPU Cluster Security
GPU clusters are expensive, powerful, and shared across teams, which makes multi-tenancy the dominant risk. Several controls below act on the individual node, because a node is what an image gets scheduled onto.
| Control | What it addresses | Implementation |
|---|---|---|
| Network segmentation | GPU clusters reachable from the general enterprise network | Dedicated VPCs/subnets for AI workloads; explicit ingress and egress rules |
| Hardware partitioning | Cross-tenant leakage through shared GPU memory | Multi-Instance GPU (MIG) – up to seven hardware-isolated instances, supported on A100, A30, H100, H200, B200, GB200 and RTX PRO 6000 |
| Container toolkit currency | Privileged host-side hooks in the GPU passthrough path | NVIDIA Container Toolkit 1.17.8 or later; prefer CDI mode, which was unaffected from 1.17.5 |
| Scheduling access control | Who can get an image scheduled here – Chapter 2’s first question | RBAC on the cluster API; admission control on what may be scheduled; no direct node access |
| Audit logging | Attribution for jobs, models loaded, and resource use | Job submission logs, model-load records, per-principal usage |
| Quotas and spend alerts | Anomalous consumption, including from a stolen credential | Per-team quotas, budget caps, alerting below the cap |
MIG is memory isolation, not escape prevention
Hardware partitioning and container escape are different problems, and crediting one control with the other’s guarantee is the failure mode Layer 2 catalogues.
MIG isolates one tenant’s GPU memory from another’s. It does nothing about NVIDIAScape (CVE-2025-23266), where a privileged createContainer hook running on the host inherited LD_PRELOAD from the container image and executed attacker code as root. The exploit needed no GPU access at all – only the ability to get an image scheduled. A fully MIG-partitioned cluster running Container Toolkit 1.17.7 is fully exposed to it.
The general rule Chapter 2 draws: wherever GPU passthrough, device plugins, or privileged operators are involved, a container escape is a one-tenant-to-all-tenants event. Isolation inside the GPU does not bound a compromise of the host that owns it.
Runtime Policy for the Serving Container
Layer 2 ends when the image is signed and admitted to the registry. What the running container may then reach is Layer 3’s, and it is the half of container security that a scanning pipeline does not provide.
| Control | What it denies | Implementation |
|---|---|---|
| Seccomp profile | The syscalls an escape needs | RuntimeDefault as the floor; a tightened profile where the workload is known |
| Read-only root filesystem | Persistence and in-place tampering | Read-only rootfs with explicit emptyDir writable paths; no writable model or config directories |
| Filesystem scoping | Reach beyond the task’s own data | Mount only the paths the workload needs; model weights read-only; never mount credential stores |
| Dropped capabilities, non-root | Privilege escalation inside the container | Drop ALL, add back nothing; run as a non-root UID |
| Egress policy | Exfiltration, and callbacks to attacker infrastructure | Default-deny NetworkPolicy; allowlist the model endpoint, the registry and telemetry, nothing else |
| Admission control | Any of the above being skipped at deploy time | Pod Security Admission at restricted, which requires non-root, RuntimeDefault seccomp and dropped capabilities; a policy engine for the rest |
Admission control is what makes the other five rows real. Each is a property of a workload manifest, and a property nobody enforces at deploy time is a property some manifest will lack.
The sandbox test, and where agent code execution actually lands
Chapter 2 Section 5 states the requirement without hedging: if the sandbox holds credentials or has network egress, it is not a sandbox, it is a convenient place to run the attacker’s code. Chapter 1 Section 7 puts the same requirement as an entry in its capability table – an actual sandbox has no credentials and no egress.
The two rows that satisfy it are filesystem scoping and egress policy, and they are the ones most often traded away, because an agent that cannot reach anything cannot do very much. That tension is real, and the resolution is not to weaken the sandbox: it is to give the mediating service the credential and the network reach, and the sandbox neither. What runs attacker-influenced code gets no ambient authority; what holds authority does not run that code.
This is a runtime boundary, so it is not Layer 2’s container work. Layer 2 decides which image is admitted; Layer 3 decides what the running container may reach. Layer 6 covers the orchestration host itself when a patch is not yet available.
Model Serving Endpoint Hardening
Every model serving endpoint is an API, and needs what any API needs:
- Authentication required – no unauthenticated endpoints in production. This is the single most common AI-SPM finding, and “internal” is not authentication.
- TLS in transit, including on internal hops.
- Rate limiting per user and per endpoint, coordinated with Layer 5 rather than duplicated.
- Request size caps – maximum tokens and payload size enforced at the endpoint.
- Health and latency monitoring, which is also the signal that a side channel is being probed.
The endpoint is not only the one your users call
ShadowMQ is the case this bullet list exists for. Meta’s Llama Stack deserialized ZeroMQ messages with recv_pyobj(), which runs the received bytes through pickle: anyone who could reach the socket had remote code execution as the serving process. The same copied call was then found in vLLM, TensorRT-LLM, SGLang, Modular’s Max Server and Microsoft’s Sarathi-Serve.
None of those sockets is the endpoint anyone thinks of as “the model API.” They are the internal channels between the scheduler and the workers, and they were reachable because internal traffic is assumed trusted. Chapter 2 hands this to Layer 3 as endpoint hardening, and what it requires is concrete:
- Enumerate every listening socket on a serving host, not just the documented ones.
- Bind internal IPC to loopback or a dedicated network namespace; never to
0.0.0.0. - Enforce it with the egress and ingress policy above, so reachability does not depend on the inference server’s own configuration being right.
- Treat the serving stack’s dependency versions as Layer 2’s dependency auditing problem – the bug class propagated by copy-paste across five projects, so your exposure is not bounded by your own code review.
Shared State Between Requests
To serve many users economically, inference servers share the GPU, the batch, and the prefix cache. Caching the attention state for a prompt prefix means a repeated opening is not recomputed – an enormous performance win that makes response time depend on what other people have asked.
CVE-2025-46570 is that dependency turned into an oracle: cache hits drop time-to-first-token measurably, and published research reports hit/miss detection at 99% accuracy and system-prompt recovery at roughly 111 queries per token. Chapter 2’s conclusion is the one to carry here – the benefit and the leak are the same mechanism, so it cannot be patched away, only scoped.
The Layer 3 controls are tenancy decisions, and each costs performance:
| Decision | Control | What you give up |
|---|---|---|
| Who shares a cache? | Per-tenant cache partitions | Hit rate across tenants |
| Can a guess be tested? | Cache salting per tenant or session | Cross-session reuse |
| Who shares a backend at all? | Dedicated serving pools for sensitive tenants | Utilisation |
| Is memory cleared? | Zeroing GPU memory between workloads | Scheduling latency |
Chapter 2’s generalisation is worth applying before the next CVE exists: any cross-request optimisation in a multi-tenant AI system is a candidate side channel. Batch scheduling that lets one tenant’s request length shape another’s latency is the same shape as the prefix cache, and has no CVE yet.
Cloud AI Service Considerations
Organisations using managed services – Amazon Bedrock, Microsoft Foundry, Google Vertex AI – inherit the provider’s infrastructure security and keep the configuration:
- IAM policies: are service roles least-privilege? Chapter 2 records a Vertex AI service agent that inherited excessive default permissions and was used to pivot to restricted internal artifacts. Default permissions on a managed platform are a decision someone else made for you – inheriting them is a choice you are making by not looking.
- Network configuration: are endpoints internet-facing when they should be VPC-internal?
- Logging: are the provider’s audit logs enabled for AI service operations, and does anyone read them?
- Data residency: are training jobs and inference endpoints in approved regions?
- Model registry access: who can deploy to a managed endpoint, and is there an approval gate?
Risk Insights and Prioritization
Findings are not equal. An unauthenticated endpoint handling PII outranks a missing audit log on a test cluster, and an AI-SPM console that presents both as “findings” has moved the triage problem rather than solved it.
What Raises an AI Finding’s Priority
- Data sensitivity – what the system can reach. PII, financial and health data raise the score.
- Tool access – whether the system can take real-world actions. This is the largest single multiplier, because it converts a wrong answer into an effect.
- Exposure surface – internet-facing or internal, and how many paths reach it.
- Regulatory context – GDPR, HIPAA, the EU AI Act.
- Blast radius – what else is reachable from here if this is compromised. For an agent, this is the union of every tool granted to any agent it can delegate to, not its own tool list; Chapter 2 Section 5 shows a research agent and a writer agent that are each safe alone and hold the lethal trifecta together.
Prioritization Framework
Response times below are an example SLA, not a standard – the tiers are the transferable part.
| Risk Level | Criteria | Example | Response |
|---|---|---|---|
| CRITICAL | Internet-facing and sensitive data and tool access | Customer-facing chatbot with database write access | Hours |
| HIGH | Internet-facing or sensitive data, plus a misconfiguration | Model endpoint with weak authentication serving financial queries | 24 hours |
| MEDIUM | Internal-only, with a misconfiguration | Internal GPU cluster reachable from the general enterprise network | 1 week |
| LOW | Internal-only, minor baseline deviation | Test endpoint missing audit logging | 1 month |
The MEDIUM row is the one to argue with. An internal GPU cluster reachable from the enterprise network is one scheduled image away from host root and every tenant on it. “Internal-only” caps a score on exposure while the blast radius says otherwise – which is the standing weakness of exposure-weighted scoring, and the reason a scoring model is an input to triage rather than a substitute for it.
Defense Connection
Prioritization is where WarningASI02: Tool Misuse and Exploitation and WarningASI03: Identity and Privilege Abuse become visible before they become incidents: both are conditions of the permission graph, and the permission graph is a posture question. It is also where WarningASI08: Cascading Failures is decided – blast radius is not discovered during an incident, it is configured beforehand.
Securing the AI Orchestration Layer
The orchestration layer – MCP servers, function-calling frameworks, agent orchestrators such as n8n and LangGraph, custom tool integrations – is the bridge between “AI that generates text” and “AI that takes actions.” Chapter 2 calls it the highest-value target in the stack and usually the least hardened, because it is treated as internal tooling rather than production infrastructure.
The Orchestrator Is a Credential Store
An orchestrator’s job is to hold every credential the AI system needs: model API keys, database connection strings, internal service tokens, cloud roles. Remote code execution there is not one compromised application – it is the credential store for the whole AI estate, plus the network position those credentials are normally used from.
Reference case: n8n CVE-2025-68613
Disclosed December 2025, CVSS 9.9, and the platform behind this course’s own labs. n8n evaluated workflow expressions in a context insufficiently isolated from the Node.js runtime, so a crafted expression reached core modules and executed OS commands as the n8n process. Affected 0.211.0 through 1.120.4 / 1.121.1 / 1.122.0, with more than 100,000 internet-exposed instances at disclosure.
What it takes to exploit: an account that can create or edit a workflow. No administrative privilege. In most deployments that is the entire engineering team, plus anyone who obtained a session.
The Layer 3 consequence is a permission model, not a patch: treat workflow-edit permission as equivalent to shell access until proven otherwise. That means editors are a reviewed, small set; the orchestrator runs under the runtime policy above rather than as a trusted internal box; and its credentials are scoped per connection instead of pooled. Layer 6’s virtual patching covers the window before you can upgrade.
The Tool-Call Pipeline
graph TB
REQ["Agent Request<br/><small>Tool call from<br/>LLM or agent</small>"]
AUTH["Authenticate<br/><small>Which agent, acting<br/>for which user?</small>"]
ALLOW["Allowlist<br/>Check<br/><small>Tool approved for<br/>this agent?</small>"]
ENT["Entitlement<br/>Check<br/><small>May <b>this user</b> do this,<br/>not just this server?</small>"]
SANDBOX["Sandbox<br/>Execution<br/><small>No credentials,<br/>no egress</small>"]
VALIDATE["Output<br/>Handling<br/><small>Treat result as<br/>untrusted data</small>"]
RESP["Response<br/><small>Returned to<br/>the agent</small>"]
BLOCK["Blocked<br/><small>Unauthorized or<br/>out of scope</small>"]
REQ --> AUTH
AUTH -->|"Verified"| ALLOW
AUTH -->|"Failed"| BLOCK
ALLOW -->|"Approved"| ENT
ALLOW -->|"Denied"| BLOCK
ENT -->|"Entitled"| SANDBOX
ENT -->|"Not entitled"| BLOCK
SANDBOX --> VALIDATE
VALIDATE --> RESP
style REQ fill:#1565c0,color:#fff
style AUTH fill:#2d5016,color:#fff
style ALLOW fill:#2d5016,color:#fff
style ENT fill:#2d5016,color:#fff
style SANDBOX fill:#2d5016,color:#fff
style VALIDATE fill:#a85800,color:#fff
style RESP fill:#1565c0,color:#fff
style BLOCK fill:#8b0000,color:#fff
Two nodes here are the ones usually missing. The entitlement check is the subject of the next section. The output handling step is orange rather than green deliberately: a tool result is untrusted input to the next model turn, and Chapter 2 is clear that there is no reliable detector for injected instructions in it. Layer 3’s contribution is to ensure the result arrives as data the agent cannot be tricked into treating as authority; the filtering itself is Layer 5’s.
MCP Server Verification That Survives a Rug Pull
The intuitive control – review the server before you connect it – fails against the incident Chapter 2 actually documents. postmark-mcp shipped fifteen clean versions before v1.0.16 added a BCC on every outbound message. A code read, a reputation check and a one-time review would all have passed, because at review time there was nothing to find.
So the controls that work are the ones that act on change, not on first sight:
- Pin the version, never a floating tag. An unpinned dependency re-decides trust on every install.
- Require a reviewed diff on change – this is the control the rug pull is designed to defeat, and the only one that catches it.
- Bind approval to contents, not to a name. Cursor MCPoison (CVE-2025-54136) bound approval to the server’s name, so replacing the command it ran never re-prompted. An approval that is not bound to a content hash is an approval of a name.
- Maintain an inventory of which servers are connected to what. This is AI-SPM’s contribution, and it is what makes the other three auditable.
- Read tool descriptions as untrusted text. They enter the context window and the user never sees them – reviewing a server’s code does not review its prose.
Per-Agent Identity and Mediated Delegation
Multi-agent systems fail in a way single agents do not: a message from a named internal component arrives in the system’s own format, describing a task, carrying no privilege marker that distinguishes it from a web page. The trust is in the channel’s appearance.
Two Layer 3 controls address it:
- Per-agent identity. Each agent authenticates as itself. Without this, an orchestrator’s logs record one identity performing every action and nothing in them distinguishes a delegated task from an injected one.
- Mediated delegation. A subtask request carries the original user’s entitlements, not the orchestrator’s. An orchestrator that delegates under its own authority is a confused deputy with a task queue.
Defense Connection
Orchestration security is Layer 3’s answer to WarningASI02: Tool Misuse, WarningASI04: Agentic Supply Chain, WarningASI05: Unexpected Code Execution and WarningASI07: Insecure Inter-Agent Communication from Chapter 2 Section 5. Note the division of labour: Layer 3 constrains what a tool call may reach and under whose authority; it does not decide whether the instruction that produced the call was legitimate. That is Layer 5, and Chapter 2’s position is that it will sometimes be wrong – which is the argument for the sandbox, not against the filter.
Identity and Access Management for AI
AI systems need identities. A serving endpoint needs credentials for its weights. An agent needs keys for external services. An orchestrator needs database credentials. How those identities are shaped determines whether a single compromise stays contained.
The Confused Deputy Is the Default
Chapter 2 Section 5 sends the confused deputy here, and it is the most important thing in this section because it is not a misconfiguration – it is what you get by building the obvious way.
Most MCP servers hold their own service credentials, not the user’s. The database in Chapter 2’s privilege escalation chain does not see “an agent acting for Alice”; it sees the service account in the .env file. The agent is a component that acts on your behalf while holding broader privileges than you have, so persuading it is enough to exercise privileges you were never granted.
The consequence Chapter 2 states, and the requirement Layer 3 has to meet:
Authorization has to be evaluated against the requesting user’s entitlements, not the server’s. An agent that correctly refuses to read another user’s records is still exploitable if the tool underneath it can read them.
That is a design requirement, not a policy setting, and it rules out the arrangement teams reach for first:
| Pattern | What the backend sees | Verdict |
|---|---|---|
| One service account for the tool, agent refuses out-of-scope requests | The service account | Confused deputy. The refusal is steering, not enforcement – it lives in the model’s judgement |
| One service account, plus a filter in the tool wrapper | The service account | Better, and now your wrapper is the authorization system for every backend behind it |
| Token exchange – the tool acts with a credential derived from the user’s session | The user | The requirement met. The backend enforces its own authorization, as it already does for humans |
The third row is on-behalf-of delegation, and existing standards cover it: OAuth 2.0 token exchange, or the equivalent in your identity provider. The work is not inventing a mechanism – it is refusing the shortcut of a shared service account during prototyping, because that is the decision that is never revisited before production.
IAM Controls for AI
| Control | Purpose | Implementation |
|---|---|---|
| User-derived credentials | Break the confused deputy | Token exchange so the backend sees the requesting user; the tool holds no standing authority |
| Service account per function | Separate identities for separate capabilities | Distinct accounts for inference, data access and tool execution – so one compromise does not inherit the others |
| Scoped credentials | Bound what each credential can do | Permission scopes per key; a database role that cannot drop a table; no master keys in an agent environment |
| Short-lived, task-scoped tokens | Bound the window and the blast radius together | Minutes, not months; issued for a task and not renewable by the agent that holds them |
| Rotation | Limit the value of a stolen credential | Automated on schedule; immediate on suspicion; rotation is detectable by the attacker, so pair it with revocation |
| Per-principal audit trail | Attribution across agents and users | Log the credential and the user it was derived for – an agent-only log cannot answer “on whose behalf?” |
Defense Connection
This is Layer 3’s primary defense against WarningASI03: Identity and Privilege Abuse. Chapter 2’s chain – read-only file access finds database credentials, which reach admin API keys, which reach full compromise – needs every link to be a standing credential reachable from the previous one. Two rows above break it: user-derived credentials mean the file tool never holds a database identity, and short-lived task tokens mean a credential found on disk has usually already expired.
The Replit case is the control that costs least and is skipped most often. An agent deleted a production database with no attacker involved at all, and the control that would have prevented it – a credential that could not drop a table – is unremarkable in any non-AI system.
Defense Perspective: Cursor CurXecute (CVE-2025-54135)
The attack (from Chapter 2 Section 5): Anysphere’s Cursor, disclosed 1 August 2025, CVSS 8.5, fixed in 1.3.9. A developer’s agent is connected to a legitimate MCP server – a Slack connector. An attacker posts a message into a channel the agent reads. The message instructs the agent to add an entry to ~/.cursor/mcp.json. Cursor allowed in-workspace file writes without approval, and executed newly added MCP entries before the user approved them, so the approval prompt arrived after the command had run.
Start with what was not compromised. Not the model, not the MCP server, not the package registry. The attacker’s entire capability was writing text somewhere the agent was going to read. Chapter 2 is explicit that the correct classification is ASI01: Agent Goal Hijacking plus WarningASI05: Unexpected Code Execution – not a supply-chain compromise, because “filing this as ASI04 would send you to review your package sources, which would not have helped.”
That matters for a defense section, because it disqualifies the reassuring answer. Layer 3 controls aimed at vetting the server – discovery flagging an unverified component, scanning the server’s image, allowlisting a new tool – would each have found a legitimate, already-approved Slack connector and passed it.
What Layer 3 would actually have contributed:
-
Runtime policy, and only this, would have prevented it. The agent executed with the developer’s full ambient authority: file system, Git credentials, SSH keys, every repository reachable. A sandbox with no credentials and no egress bounds the injected command to a scratch directory. This is the row teams trade away for convenience, and it is the row that was load-bearing here.
-
Identity scope would have bounded the loss. The agent inherited the developer’s session rather than holding task-scoped credentials. Short-lived credentials derived per task mean the executed command finds nothing durable to steal.
-
AI-SPM would not have seen this at all, and that is the more useful lesson. A write to
~/.cursor/mcp.jsonand a newly registered tool provider look like exactly the drift the detective half exists to report – but the file is on a developer’s laptop, and a posture tool scoped to connected cloud accounts has no view of it. The layer’s detective half is bounded by its connection scope before it is bounded by its reporting lag. Ask which of your AI assets live somewhere your posture tool was never connected to; for most organisations the honest answer is “the developer estate,” which is where the agentic tooling actually runs. Where the asset is in scope, the lag still applies: the report arrives after execution, because the attack completes in one action. -
The fix that shipped was not a Layer 3 control at all. Anysphere bound execution to approval and re-validated on change. For a SaaS coding agent, the ownership table above says most of this layer is the vendor’s – and the honest posture question is not “have we hardened Cursor” but “which of our agents run on someone else’s runtime, and what did we assume about it?”
The companion case is the opposite shape. MCPoison (CVE-2025-54136), disclosed four days later, was a rug pull: approval bound to a server’s name rather than its contents, so an attacker with commit access to a shared repository could swap the command after a teammate approved it. Same product, unrelated mechanism, and a different control – content-bound approval and a reviewed diff on change. Two incidents in one product, and no single control covers both.
Layer 3 → OWASP Mapping
Layer 3 is named as a primary defense in ten of the twenty categories across the LLM Top 10 and the Agentic AI Top 10 – the second-broadest coverage in the Blueprint, and, as Section 2 insists, coverage is not stopping power.
| Category | The route Layer 3 acts on | What Layer 3 does | What it cannot do | Completed by |
|---|---|---|---|---|
| LLM03: Excessive Agency | The permission shape an agent holds before any request | Credential scoping, short-lived tokens, tool allowlists, the inventory that makes over-permission visible | Judge whether a particular authorised action was intended | L5 – policy at the request; L4 – human gate on irreversible actions |
| LLM04: Supply Chain | Components entering the running estate | Inventory, drift detection on new integrations, runtime isolation of what was admitted | Assess an artifact before it is deployed | L2 – signing, pinning, scanning at the gate |
| LLM09: Vector and Embedding Weaknesses | The vector store as infrastructure | Network placement, authentication, RBAC, encryption, backup posture | Know what the embeddings mean or what sensitivity they inherit | L1 – classification and the inheritance rule |
| WarningASI02: Tool Misuse and Exploitation | The tool call, at execution | Allowlisting, entitlement checks, sandboxed execution, per-call logging | Detect misuse – every call is authorised and looks legitimate | L5 – what reaches the model; L6 – anomaly over time |
| WarningASI03: Identity and Privilege Abuse | The identity every component acts as | User-derived credentials, per-function accounts, scoping, rotation, per-principal audit | Prevent abuse of a permission that was correctly granted | L4 – joiner/mover/leaver and human review |
| WarningASI04: Agentic Supply Chain | Servers and tools already connected | Connection inventory, content-bound approval, reviewed diff on change, runtime isolation | Vet an artifact that has not been connected yet | L2 – dependency auditing and pinning at build |
| WarningASI05: Unexpected Code Execution | The execution environment | Sandbox with no credentials and no egress, seccomp, read-only rootfs, admission control | Stop the instruction that caused execution from arriving | L5 – input filtering; L6 – virtual patching the host |
| WarningASI07: Insecure Inter-Agent Communication | The channel between agents | Per-agent identity, mediated delegation, authenticated channels, attributable logs | Distinguish a legitimate delegated task from an injected one by content | L5 – message validation |
| WarningASI08: Cascading Failures | The reachability graph between components | Blast-radius containment, segmentation, quotas and circuit breakers, credential separation | Detect that a cascade is under way | L6 – anomaly detection and response |
| WarningASI10: Rogue Agents | Agents and endpoints nobody registered | Shadow AI discovery, unmanaged endpoint detection, credential inventory | Judge whether a discovered agent’s behaviour is malicious | L4 – governance and ownership; L6 – behavioural baselines |
Read the fourth column down and Layer 3’s boundary is one sentence: it decides what each component may reach, and never what any component was asked to do. That is also why its coverage is so broad and its stopping power so narrow – almost every attack in the course crosses a boundary this layer defines, and none of them is recognisable as an attack at the moment it crosses.
AI Scanner Cross-Reference
AI Scanner assesses vulnerabilities in the model – prompt injection susceptibility, system prompt leakage, adversarial robustness – which no infrastructure control changes. AI-SPM assesses the infrastructure the model runs on, which no model-level scan sees. They are complementary because their blind spots are: a model that resists injection served from an unauthenticated endpoint is exposed, and a hardened endpoint serving a model that leaks its system prompt is also exposed. Section 9 covers how the two feed the scan-protect-validate-improve loop.
TrendAI Vision One’s AI-SPM capability provides the discovery and assessment half of this layer for the cloud accounts you connect to it – classifying AI assets by role, assessing configuration against AI baselines, surfacing threat alerts, misconfigurations and potential attack paths across them – with the Cyber Risk Exposure Management module correlating those findings against enterprise risk posture. Note the scope limit stated above: assets outside a connected account, including the developer estate, are not in that picture, and shadow usage of external AI services is a Layer 5 question rather than a posture one. The integration argument is specific rather than general: the ownership table above shows Layer 3 findings spanning components with different owners, so a console that reports an over-scoped agent credential and the unpatched container toolkit under it is answering a question that two separate tools each answer halfway. What no console supplies is the configured-boundary half – runtime policy, cache tenancy, user-derived credentials. Those are design decisions, and buying the dashboard does not make them.
Key Takeaways
- Layer 3 is two control types under one number: a detective half (AI-SPM) that reports and blocks nothing, and a configured-boundary half (runtime policy, tenancy, identity scope) that enforces continuously and never sees a request. Deploying only the first is the characteristic failure of this layer.
- Which Layer 3 controls are yours depends on the deployment pattern. A cloud-API consumer owns no GPU, container, socket or cache – and still owns the orchestrator, every agent credential, and discovery of what the organisation actually calls.
- Layer 2 ends at the registry; what a running container may reach is Layer 3’s – seccomp, read-only root, filesystem scoping, egress policy, admission control. A sandbox that holds credentials or has egress is not a sandbox.
- The endpoint to harden is not only the one users call: internal serving-stack sockets (ShadowMQ) and shared prefix caches (CVE-2025-46570) are Layer 3 surfaces with no external interface, and cross-request optimisation in a multi-tenant system is a candidate side channel by default.
- MCP verification must survive a rug pull: pin versions, bind approval to contents rather than a name, and require a reviewed diff on change – a one-time review passes the fifteen clean versions that precede the malicious one.
- The confused deputy is the default, not a misconfiguration. Authorization must be evaluated against the requesting user’s entitlements, not the tool’s – which means user-derived credentials, not a well-behaved agent in front of a shared service account.
Test Your Knowledge
Ready to test your understanding of AI infrastructure security? Head to the quiz to check your knowledge.
Up next
Infrastructure secured, the next layer addresses the human element. In Section 6, you’ll learn about Layer 4: Secure Your Users – including deepfake detection, endpoint security for AI-era threats, shadow AI governance, and how to protect users from AI-powered social engineering.