Section 5 Quiz
Test Your Knowledge: Layer 3 - Secure Your AI Infrastructure
Let’s see how much you’ve learned!
This quiz tests what Layer 3 can and cannot stop, which of its controls are yours, posture and runtime hardening for AI resources, orchestration security, and identity design for AI service accounts.
---
shuffle_answers: true
shuffle_questions: false
---
## A traditional CSPM tool reports that an organization's GPU cluster is compliant with every cloud security benchmark. The security team concludes the AI infrastructure is secured. What does that assessment miss?
> Hint: Compare what the two tools can see, not how strictly they grade what they see.
- [ ] CSPM cannot scan GPU hardware itself for firmware-level vulnerabilities
> Hardware scanning is a real concern but a narrow one. The gap is broader: CSPM lacks the AI context needed to evaluate what the nodes are running.
- [x] CSPM sees virtual machines, so the AI-specific risks on them are invisible
> Correct! CSPM checks cloud configuration against generic benchmarks -- security groups, encryption, network rules. It cannot tell whether the serving endpoint is authenticated, whether the prefix cache is shared across tenants, or whether an unreviewed MCP server was connected last week. The compliant verdict is accurate and answers a different question than the one the team asked.
- [ ] CSPM is unable to operate at all on GPU-optimized instance types
> CSPM tools work on any cloud instance regardless of accelerator hardware. The limitation is the risk context they evaluate, not instance compatibility.
- [ ] CSPM was built for storage and networking rather than for compute
> CSPM covers compute alongside every other cloud resource. The gap is that it grades compute against generic baselines rather than AI-specific ones.
## A team maps their controls onto the Blueprint and files per-user rate limiting on their model endpoint under Layer 3, reasoning that the endpoint is infrastructure. Why is that mapping wrong?
> Hint: Ask which layer the control sits in relative to a live request.
- [x] Rate limiting inspects live requests, so it is Layer 5 -- Layer 3 is never in that path
> Correct! The master mapping assigns LLM06: Unbounded Consumption to Layers 5 and 6, not Layer 3. Rate limiting counts and refuses live requests, which is definitionally in-path, and Layer 3 never reads a request. Layer 3's contribution to consumption risk is the anomaly detection and quota posture around it -- useful, and not the control that refuses the request.
- [ ] Rate limiting belongs to Layer 6, since abuse detection is a zero-day concern
> Layer 6 contributes anomaly detection against novel patterns, but the enforcement of a per-user limit is an access-layer control applied to every request, not a zero-day defense.
- [ ] The endpoint is a model artifact, so its controls all belong to Layer 2
> Layer 2 owns the model artifact and its container image up to the registry. A serving endpoint's runtime behaviour is not a Layer 2 concern at all.
- [ ] Nothing is wrong -- infrastructure controls are correctly filed under Layer 3
> The reasoning that "the endpoint is infrastructure" conflates where a control is deployed with which layer owns it. Layer assignment follows what the control acts on, not what hardware it runs on.
## An organization deploys an AI-SPM console. Discovery, assessment and risk scoring are all running, findings are triaged weekly, and the dashboard is green. What is the most significant gap this setup can still have?
> Hint: Consider what an absent control looks like to a tool that reports misconfigurations.
- [ ] Weekly triage is too slow, so findings age before anyone acts on them
> Triage cadence affects response time but not coverage. A faster cycle on the same findings still reports only what the console can see.
- [ ] The console may lack integrations for some cloud providers in use
> Incomplete provider coverage is a real discovery gap, but it is the kind of gap the tool itself reports as unmonitored accounts.
- [x] AI-SPM reports and never blocks, so an unenforced boundary shows as green
> Correct! AI-SPM is Layer 3's detective half: it reports and intercepts nothing. Layer 3's other half -- runtime policy, cache tenancy, user-derived credentials -- enforces continuously but has to be designed and deployed. A console will flag a misconfigured control; it will not flag a control that was never deployed, because an absent control is not a misconfiguration. Green means nothing known is wrong, not that the boundary is enforced.
- [ ] Risk scores are subjective, so the prioritization order cannot be trusted
> Scoring models do have real weaknesses, such as capping an internal finding on exposure while its blast radius says otherwise. That is a triage-quality issue rather than the structural gap.
## A team consuming a managed cloud AI API runs the Layer 3 checklist and opens a finding that their serving containers have no seccomp profile. What is wrong with this finding, and what should they have looked at instead?
> Hint: Run the ownership table before the checklist, not after.
- [ ] The finding is valid but low priority, since managed services are already hardened
> This treats the finding as real and merely deprioritized. It is not a finding about their environment at all, so priority is not the issue.
- [ ] Nothing is wrong -- seccomp is a Layer 3 control and Layer 3 applies to everyone
> Every layer applies to everyone, but the individual controls do not. Which Layer 3 controls are yours is set by the deployment pattern you chose.
- [x] They own no serving containers -- the orchestrator and agent credentials are theirs
> Correct! On a cloud API the provider owns the GPU nodes, the serving containers, the internal sockets and the cache. The team owns the orchestrator, which runs on their side in every pattern and is the highest-value target in the stack; the identity scope of every credential their agents hold; and discovery of which AI services the organization is actually calling. A finding written against a container they will never possess spends review effort on someone else's infrastructure.
- [ ] The finding should be reassigned to Layer 2, which owns all container security
> Layer 2 owns the image up to the registry and Layer 3 owns what the running container may reach. Reassigning the layer does not change the fact that the containers belong to the provider.
## An agent executes model-generated code in a container the team calls a sandbox. The container has the agent's cloud credentials mounted so it can fetch data, and outbound internet access so it can install packages. How should this be assessed?
> Hint: Chapter 2 offers a one-sentence test for whether something is a sandbox.
- [ ] Adequate -- the container boundary isolates execution from the host system
> The container boundary limits host access, but the credentials and egress inside it are what an attacker actually wants. Escaping the container is unnecessary when everything of value is already mounted in it.
- [x] Not a sandbox -- credentials and egress are exactly what it must not have
> Correct! The test is that a sandbox holding credentials or network egress is not a sandbox, it is a convenient place to run the attacker's code. Filesystem scoping and default-deny egress are the two runtime-policy controls that satisfy it, and they are the ones traded away for convenience. The resolution is not to weaken the requirement but to move it: a mediating service holds the credential and the network reach, and the component executing untrusted code holds neither.
- [ ] Acceptable if the image was scanned and signed before it was deployed
> Scanning and signing are Layer 2 controls that end at the registry. They establish what the image contains, not what the running container is permitted to reach.
- [ ] Acceptable if every command the agent runs is logged for later review
> Logging makes the execution attributable afterwards. It is the detective half of the layer and does not constrain what the code can do while it runs.
## A self-hosted team hardens their model endpoint: authentication required, TLS enforced, rate limits applied, request sizes capped. Weeks later an attacker achieves remote code execution as the serving process without ever authenticating. What did the hardening miss?
> Hint: Recall which sockets ShadowMQ was actually reached through.
- [ ] Rate limits were applied per user rather than per endpoint and per token count
> Multi-dimensional rate limiting matters for consumption risk, but no rate limit configuration produces or prevents remote code execution.
- [ ] The authentication scheme used static API keys instead of short-lived tokens
> Credential lifetime bounds the value of a stolen key. It is irrelevant to an attack path that never presented a credential at all.
- [x] The internal sockets between scheduler and workers, which nobody authenticates
> Correct! This is the ShadowMQ pattern. Inference servers deserialized ZeroMQ messages with `recv_pyobj()`, which runs the bytes through pickle, so anyone who could reach the socket had code execution as the serving process -- and the same copied call appeared in vLLM, TensorRT-LLM, SGLang and others. Those sockets are not the endpoint anyone means by "the model API", and they were reachable because internal traffic was assumed trusted. Enumerate every listening socket, bind IPC to loopback, and enforce it with network policy rather than trusting the server's own config.
- [ ] The TLS configuration permitted downgrade to an unencrypted cipher suite
> A downgrade would expose traffic in transit. It does not hand an attacker execution as the serving process, which requires a deserialization path.
## A multi-tenant inference platform enables prefix caching across all tenants for a large throughput gain. A security reviewer objects. What is the strongest form of the objection?
> Hint: Ask whether the leak and the benefit can be separated.
- [ ] Cache entries persist on disk, so another tenant's prompts can be recovered later
> Prefix caches hold computed attention state in memory rather than readable prompt text on disk. The exposure is through timing, not stored content.
- [ ] The cache is a new component, so it expands the patchable attack surface
> Every component adds surface, but this framing implies the risk is closed by patching. The reviewer's objection has to survive a fully patched system.
- [x] The speedup and the leak are one mechanism, so only tenancy scoping helps
> Correct! A cache hit lowers time-to-first-token measurably, so an attacker submits a guess and times the response -- research reports hit/miss detection at 99% accuracy and system-prompt recovery at roughly 111 queries per token. Because the benefit and the leak are the same behaviour, there is no patch that keeps one and removes the other. The controls are tenancy decisions that cost performance: per-tenant partitions, cache salting, or dedicated pools for sensitive tenants.
- [ ] Cached state may return a stale response belonging to a different tenant
> Prefix caching reuses computation for a matching prefix rather than returning another tenant's completion. The risk is inference through timing, not response mixing.
## A team reviews an MCP server before connecting it: they read the source, check the publisher's reputation, and confirm the package is widely used. Six weeks later it begins exfiltrating data. Which control would have caught this, and why did the review not?
> Hint: Consider when the malicious code was actually introduced.
- [x] A reviewed diff on change -- the malicious version was published after review
> Correct! This is the `postmark-mcp` shape: fifteen clean versions shipped and worked exactly as advertised before v1.0.16 added a BCC on every outbound message. There was nothing to find at review time, so a code read, a reputation check and a popularity check all passed correctly. Only controls acting on *change* catch it -- pinned versions, a reviewed diff when the pin moves, and approval bound to a content hash rather than to a server's name.
- [ ] Static analysis tuned for exfiltration patterns in the server's source code
> Better analysis on the version that was reviewed still finds nothing, because that version was genuinely clean. The malicious code did not exist yet.
- [ ] Sandboxing the server so it could not reach the network during execution
> Egress restriction would limit the damage, but this server's legitimate function was sending mail. The exfiltration rode the traffic it was supposed to produce.
- [ ] Requiring servers to come from a vetted internal registry rather than npm
> An internal mirror relocates the artifact without changing the trust decision. A mirrored package that updates still updates, unless a diff is reviewed when it does.
## An agent lets users query their own records. It refuses out-of-scope requests reliably in testing, and the database tool behind it authenticates with a service account that can read every user's rows. How should this design be assessed?
> Hint: Ask what identity the database sees, not what the agent decides.
- [ ] Sound, provided the agent's refusals are validated against an adversarial test set
> Broader testing raises confidence in the refusals without changing their nature. The refusal still lives in the model's judgement rather than in an enforced boundary.
- [ ] Sound, provided the service account's credential is rotated on a short schedule
> Rotation limits the value of a stolen credential. It does nothing about a credential being used exactly as issued, by an agent that was persuaded to use it.
- [x] A confused deputy -- the database sees the service account, never the user
> Correct! The tool holds broader privileges than the user does, so persuading the agent exercises privileges the user was never granted, and the refusal is steering rather than enforcement. Authorization has to be evaluated against the requesting user's entitlements, not the server's. The fix is user-derived credentials via token exchange, so the backend enforces its own authorization exactly as it already does for humans -- not a better-behaved agent in front of a shared service account.
- [ ] Sound, since the agent already applies the same scoping the database would apply
> The agent applies scoping it can be argued out of. A database enforcing its own row-level authorization cannot be argued out of it, which is the difference that matters.
## In the Cursor CurXecute incident, an attacker posted a Slack message that caused the agent to write and immediately execute a new MCP entry. A team proposes preventing similar attacks by having AI-SPM flag unverified MCP servers before connection. Why will that not work here?
> Hint: Check what was actually compromised in the chain.
- [ ] Prompt injection is a Layer 5 concern, so no Layer 3 control could have helped
> Layer 3 had the control that would have prevented this: a runtime sandbox with no ambient credentials and no egress bounds the injected command whatever the model was persuaded to do. It is the vetting-shaped proposal that fails, not the layer.
- [ ] The MCP server was connected before the AI-SPM deployment established its baseline
> Nothing in the incident turns on deployment timing. The proposal would fail against this attack even with a mature baseline in place.
- [x] The server was legitimate and already approved -- vetting it would find nothing
> Correct! Nothing in the chain was compromised: not the model, not the MCP server (a legitimate Slack connector), not the registry. The attacker's entire capability was writing text somewhere the agent would read. Chapter 2 classifies it as goal hijacking plus unexpected code execution, explicitly not a supply-chain compromise, because filing it as one sends you to review package sources that were never the problem. Every vetting-shaped control passes a server that is genuinely fine.
- [ ] Discovery would flag it, but weekly assessment cycles are too slow to intervene
> This concedes that vetting would detect something. It would not -- and the deeper issue is that the attack completes in a single action, so no reporting cadence prevents it.