Section 7 Quiz

Test Your Knowledge: Layer 5 - Secure Access to AI Services

Let’s see how much you’ve learned!

Ten questions on where a gateway has to sit to be a control, what input filtering actually buys, why entitlement goes inside the retrieval query, the streaming decision the output boundary forces, and what Layer 5 hands to another layer. Several ask you to judge a design rather than recall one.

--- shuffle_answers: true shuffle_questions: false --- ## A platform team deploys an AI Gateway as a reverse proxy with injection detection, response filtering and per-user cost caps. Six months later an audit finds that roughly 40% of the organisation's LLM API spend never traversed it. The gateway is correctly configured and has never failed. What is the defect? > Hint: Ask what forces traffic through the gateway rather than what the gateway does once traffic arrives. - [ ] The proxy needs deep packet inspection to see traffic addressed to other endpoints > A proxy cannot inspect traffic that never reaches it, and inspecting more of the traffic it does receive changes nothing about the 40% that goes elsewhere. - [x] Nothing forces traffic through it -- a direct egress path to the provider still exists > Correct! A gateway is only a control if requests cannot go around it. Any developer holding a provider API key with outbound HTTPS access does not need the gateway, and that path scales with team size. The control that makes the gateway mandatory is a default-deny egress rule at Layer 3 permitting only the gateway as a destination. Layer 5 supplies the inspection point; Layer 3 supplies the reason there is only one. - [ ] The filtering rules are too permissive, so requests are being silently forwarded unlogged > Permissive rules would show as traffic that passed inspection, not as traffic absent from the gateway's records entirely. The 40% is missing, not allowed. - [ ] Cost caps should be set at the provider account rather than per user in the gateway > Provider-level caps are a useful backstop and would bound the spend, but they would not restore inspection of prompts or responses on that 40%. ## Layer 5 is named as a primary defense in 11 of the 20 OWASP categories -- more than any other layer, and nearly four times Layer 2's three. Which conclusion does that figure support? > Hint: Coverage counts categories touched. Ask what it does not count. - [ ] Layer 5 resolves a little over half of the OWASP risk, so it should absorb half the budget > Coverage is a count of categories, not a proportion of risk. Nothing in the figure says how completely any one category is addressed. - [x] Layer 5 touches the most categories, and touching is not resolving -- most it hands on > Correct! Coverage counts how many categories name a layer, not how completely it resolves any of them. Read Layer 5's mapping and every row has something it cannot do: it cannot un-disclose, cannot see a sequence, cannot know a credential's real scope. The comparison with Layer 2 makes the point sharply -- Layer 2 is named in only three categories and is the only layer that defends them at all, because no amount of Layer 5 filtering recovers a backdoored model. - [ ] Layer 5 should be deployed before every other layer in every AI architecture > It is often the first investment, especially for a live interface, but that follows from being the only in-path layer rather than from the count -- and it is a starting heuristic, not the decision. - [ ] Layer 2 is nearly redundant, since its three categories are a small share of the total > This inverts the reading. Low coverage and high leverage are not in tension, and Layer 2 covers the failure mode no other layer can touch. ## Your team is choosing where to run Layer 5 inspection. One proposal is an in-process SDK inside each application; another is a shared reverse proxy. Both use the same detection engine. Which consideration should decide it? > Hint: The engines are identical, so compare what each position can see and what each is bypassed by. - [ ] Latency, since the SDK avoids a network hop and AI responses are already slow > The SDK does avoid a hop, and that is a real advantage. But both options add far less latency than inference itself, so latency rarely decides between them. - [x] What each is routed around by, and whether the SDK sees context the wire cannot > Correct! The two positions differ on two axes that matter more than the engine. An SDK sees pre-assembly context and application identity that never appear on the wire, which the proxy cannot inspect -- but it is bypassed by any code path that forgets to call it, and forgetting scales with team size. A proxy is uniform across applications but sees only the wire request. Neither is universally right; the decision is coverage against bypass-resistance. - [ ] Vendor support, since the SDK ties you to whichever language bindings exist > Language coverage is a genuine constraint on the SDK option, but it is an implementation limitation rather than the security property that distinguishes the two positions. - [ ] Cost, since a shared proxy amortises the detection engine across all applications > A shared proxy does amortise infrastructure. This is a real operational advantage and it is not a security argument for either placement. ## An assistant is asked "summarise our Q3 attrition drivers." Retrieval returns a top-5 chunk set, a post-filter drops the two chunks the caller is not entitled to read, and the model answers from the remaining three. A month later the same user receives a cited answer containing a colleague's salary band. What went wrong? > Hint: Consider what the top-5 slots were spent on, and what happens on a code path that skips the post-filter. - [ ] The response filter's PII detection failed to recognise a salary band as sensitive > Response filtering is a backstop here and salary bands are exactly the contextual sensitivity a generic PII dictionary misses. But the leak's root cause is that the passage was retrieved at all. - [x] The entitlement predicate is a post-filter on results rather than part of the query > Correct! Filtering after retrieval looks equivalent to filtering during it and is not. Unentitled chunks consume the top-k slots, so entitled users get answers assembled from whatever survived -- and any code path that skips the post-filter leaks, because a vector index has no concept of a user. The predicate belongs inside the retrieval query, so unentitled chunks are never candidates. Layer 1 owns classifying the content; the entitlement decision needs a request, which makes it Layer 5's. - [ ] Layer 1 failed to classify the HR corpus, so no sensitivity label was available > If classification were missing, no post-filter could have run at all. The label existed and was applied at the wrong point in the pipeline. - [ ] The user escalated privileges between the two requests without the gateway noticing > Nothing in the scenario indicates a privilege change, and this leak needs no attacker at all -- an in-scope question, a correct retrieval, and a reader who should not see the passage. ## A vendor claims their guardrail "stops prompt injection." The best-instrumented production classifier published cut universal jailbreak success from 86% to 4.4%, and a later generation brought the compute overhead down to roughly 1%. How should those numbers change your architecture? > Hint: Consider what the remaining percentage means for a system whose safety assumes the filter holds. - [ ] It should not -- a 95% reduction means injection is effectively solved at the filter > A 95% reduction is a large real gain, but the residual is not rounding error. It is the rate at which attacks reach the model, and it applies continuously to every request. - [ ] It argues for skipping the guardrail, since a control with a bypass rate provides false assurance > Removing a control because it is imperfect trades a low-single-digit residual for a high one. The measurements argue for correct expectations, not for removal. - [x] It sets the price of an attack, so assume a low-single-digit residual reaches the model > Correct! Both halves matter. The reduction is worth having and the overhead no longer justifies declining it. But those figures came from 1,700+ hours of paid red-teaming against a frontier lab's classifier -- your rule set will do worse. So the filter sets the *price* of an attack, and what decides the *outcome* is what a successful injection can cause: action scope, least-privilege credentials, and a human gate on irreversible actions. - [ ] It shifts the investment to a second vendor's filter, since two engines multiply the reduction > Independent filters do fail somewhat independently and stacking them helps. But both inspect the same text with similar methods, so the gain is far smaller than multiplying the rates suggests. ## A team wants to enforce "the system prompt outranks user input" at their AI Gateway, and asks you to add it as a pipeline stage after injection detection. What is the correct response? > Hint: Ask where the priority ordering would have to be represented for anything to act on it. - [ ] Add it after content policy instead, so policy violations are caught before reordering > Stage ordering is not the issue. No position in a gateway pipeline gives the gateway the ability to change how the model weighs the tokens it receives. - [x] It cannot live there -- instruction hierarchy is trained into the model, not enforced by a filter > Correct! Instruction hierarchy is a training-time property: models are trained on a synthetic hierarchy of instruction sources so they learn to privilege system and developer instructions. That happens inside the model before you receive it, which makes it a Layer 2 property you select when choosing a model. A gateway can supply the signals the hierarchy consumes -- provenance labels marking which text came from where -- but it cannot make the model honour them, and measured compliance is partial. - [ ] Implement it by prepending the system prompt and appending a reminder after user input > This is the "sandwich" pattern, and it is the failure Chapter 2 describes: the reminder is text in the same flat sequence as the override, carrying no more privilege than the attack it is meant to outrank. - [ ] Add it, but only for requests where injection detection returned a low-confidence verdict > Conditioning the stage on detection confidence does not change the fact that the stage cannot do the thing it claims to do. ## An internal assistant streams responses into a chat UI that renders markdown. Output validation runs on the stream and can truncate it. A researcher demonstrates exfiltration through a markdown image reference in the model's output. Why did a validator that was present and running fail to stop it? > Hint: Ask who the consumer of the stream really is, and at what moment the fetch happens. - [ ] The validator lacked a rule for image syntax and needs its pattern library extended > A rule for image syntax would help, and adding one is worth doing. It does not address why truncation arrived too late to matter in a streamed response. - [x] The renderer is a downstream system, and a streamed token is already published > Correct! "It's a chat UI so we stream" is correct about the human and wrong about the renderer. The markdown renderer is a parser -- a downstream system -- and it fetches the image the moment the tokens reach it. A filter can stop token 400 and cannot recall tokens 1 through 399, so by the time the reference is recognised the fetch has already carried the data. This is why response filtering is an architectural *position* before it is a function: it has to sit where the response can still be stopped. - [ ] The image host was not on the outbound allowlist, so the check was skipped entirely > An allowlist miss would block the fetch, not permit it. Worse, in the real cases an allowlisted domain was used as the delivery mechanism. - [ ] Truncation cannot apply to binary content such as images fetched by the client > The model emits a text reference, not binary content. Truncation applies fine to the text -- it simply arrives after the client has acted on it. ## In EchoLeak, Microsoft was already running cross-prompt-injection classifiers, and the attacker's email was written to survive them. Which Layer 5 control would have changed the outcome most? > Hint: The injection succeeded. Ask what the successful injection then needed in order to be worth anything. - [ ] A stricter injection classifier tuned on the phrasing this email used > This is retrospective signature-writing. It closes one phrasing after disclosure and leaves the class open, and the case already establishes that a purpose-built classifier was present and defeated. - [ ] Token budgets and tool-call ceilings on the assistant's session > Consumption limits bound expensive abuse, and EchoLeak was not expensive: one retrieval and one rendered image reference. There was no request-rate spike or tool-call storm for a ceiling to catch. - [x] Suspending tool invocation for any turn in which untrusted content entered the window > Correct! Once retrieval placed a third-party email in the context, the turn contained untrusted content -- a fact the runtime can hold without recognising anything about the payload. Suspending data-gathering tool calls for the rest of that turn removes the attack's middle step regardless of how the injection was phrased. This is the lethal trifecta test enforced at runtime rather than checked at design review, and it is why the controls that work here decide what untrusted content may *cause*, not whether it can be recognised. - [ ] A domain allowlist restricting where rendered output may fetch resources from > The exfiltration routed through a Microsoft Teams proxy domain that the Content Security Policy already permitted, so the allowlist was not bypassed -- it was the delivery mechanism. Stripping outbound references entirely is the control that works; allowlisting is the one that failed. ## A system prompt for a finance assistant contains a database read-only credential and the rule "only disclose balances to users in the finance group." A reviewer proposes adding response filtering to strip system-prompt text from outputs. What is the right assessment? > Hint: Consider the two items separately, and ask what each one costs if disclosed once. - [ ] Sound -- response filtering closes the disclosure path for both the credential and the rule > Response filtering has a non-zero bypass rate against paraphrase and partial disclosure, and a credential leaked once is leaked permanently. A probabilistic control cannot protect an absolute value. - [x] Backstop only -- move the credential to the tool layer and the rule to server-side scope > Correct! Both items are doing jobs the context window cannot do. The credential belongs in the gateway or tool layer, so the model receives a callable tool rather than a quotable secret. The authorization rule belongs in a server-side data scope, where the model's cooperation is not required -- inside the flat token sequence it carries no more privilege than an injection telling the model to ignore it. Response-side leak detection stays as what it is: useful intelligence that an extraction attempt is under way. - [ ] Sound for the credential, unnecessary for the rule, which the model reliably follows > This inverts the risk. The rule is the item the model is asked to enforce, and enforcement inside the window is exactly what injection defeats. - [ ] Unnecessary -- system prompts are not returned by the API, so neither item is reachable > Extraction of system prompt content through crafted prompts is well documented. Assume anything in the window is discoverable by anyone who can talk to the model. ## An agentic workflow is compromised: an injected instruction in a retrieved document caused the agent to make eleven authorised tool calls that gradually moved data outside the task's scope. The gateway logged every call, and each one passed policy. What does this tell you about Layer 5's boundary? > Hint: Compare what a single request contains with what the evidence of this attack looks like. - [ ] The policy was misconfigured; correctly scoped actions would have blocked calls four onward > A narrower action scope would reduce what the hijack was worth, and it is worth setting. But the calls described were within scope, and no per-request policy distinguishes call four from call three. - [ ] The gateway should have correlated the eleven calls and blocked the sequence in flight > This describes the right detection, in the wrong layer. Sequence correlation against the requested task is Layer 6's behavioural anomaly detection, which observes across a session rather than deciding one request at a time. - [x] The signal exists only across a sequence, and Layer 5 decides one request at a time > Correct! Every call was individually authorised and individually legitimate, which is the defining property Chapter 2 records for six of ten agentic categories. The evidence of this attack is the *relationship* between eleven calls and the task the user actually asked for -- state Layer 5 does not hold when it evaluates any one of them. That is the structural boundary, not a product limitation, and it is why Layer 5's mapping hands ASI01's detection to Layer 6 and its consequence to Layer 4's gates. - [ ] The logging is the control, so the design worked and the response was simply too slow > Logging made the incident reconstructible afterwards, which is a real contribution. Calling it the control confuses evidence with interception.