Section 9 Quiz

Test Your Knowledge: The AI Application Security Continuous Loop

Let’s see how much you’ve learned!

This quiz tests how you read a scan result, where the loop’s limits are, the availability and data-flow decisions an inline inspection service forces, and what each phase has to produce for the next one to work.

--- shuffle_answers: true shuffle_questions: false --- ## AI Scanner reports that your customer-support agent follows instructions embedded in the support tickets it reads. The agent holds a tool that can issue refunds. What does a well-run PROTECT phase produce from this finding? > Hint: The section insists every finding gets two responses, not one. - [ ] A tuned injection rule, since detection is what the finding is about > A detection rule is half the answer and the half that gets written by default. It lowers the odds of the next attempt succeeding without changing what happens on the attempt that succeeds anyway. - [x] A detection rule plus a narrower refund permission or an approval gate > Correct! Every finding gets a detection change and a consequence change. The rule lowers the probability; the scoped credential or human gate decides the cost when the rule loses -- and the residual on a guardrail is low single digits, not zero. A PROTECT phase that only writes rules keeps detecting this class more accurately and never prevents it. - [ ] An immediate re-scan to confirm the finding before anything is changed > Confirmation belongs to VALIDATE, and the finding already arrived with a transcript of what worked. Scanner findings are demonstrations, which is exactly why they count as strong evidence. - [ ] A documented risk acceptance, given no guardrail closes injection fully > It is true that no guardrail closes injection. That argues for bounding the consequence, not for accepting a refund tool reachable by injected text. ## Your team replays a corpus of previously successful attacks against the current Guard configuration each month. This month the detection rate has dropped from 94% to 71% with no rule changes. What has most likely happened? > Hint: Something in the system changed even though the rules did not. - [x] The model or the surface changed underneath a rule set aimed at the old one > Correct! Rules are tuned to a specific model's susceptibilities and a specific input surface. A provider's silent update to a hosted model, a new retrieval source, or an edited system prompt all change behaviour without touching a rule -- which is why model version change and surface change are re-scan triggers. The replay is what made a silent drift visible. - [ ] Attackers have developed new techniques that the rules do not cover > Novel techniques would show up as new patterns in the blocked logs, not as a falling score against a fixed corpus. The corpus here is unchanged, so the change is on your side. - [ ] The replay corpus has decayed and no longer reflects real attack traffic > A corpus of attacks that previously succeeded does not become invalid over time. Its whole purpose is to stay fixed so that a change in the score means a change in the system. - [ ] False positive tuning has made the rules too permissive over the period > Tuning would explain it, but the scenario states there were no rule changes. Something outside the rule set moved. ## An engineer proposes that Guard be called fail-open on every path: if the inspection API times out, forward the request unfiltered so the assistant stays available. What is the strongest objection? > Hint: Think about when a timeout is most likely to happen. - [ ] Fail-open breaks the compliance guarantee that the content policy provides > It does weaken that guarantee, but stated this way it is an argument for fail-closed everywhere, which trades the problem for a denial-of-service primitive against your own application. - [x] It removes the guardrail silently, and load is when attacks are likeliest > Correct! The failure mode arrives precisely when the system is under stress, which correlates with the moment an attack is under way -- and it arrives without a signal, so the architecture keeps assuming a control that is no longer running. The section's answer is not "fail-closed everywhere" but per action class: fail-open on read-only paths, fail-closed on anything that writes, spends, or reaches a third party, and alert on the fail-open path either way. - [ ] Timeouts indicate a capacity problem that should be fixed at the source > Capacity work is worth doing and does not change the design question. Any network call on a critical path will eventually fail, so the behaviour on failure has to be chosen. - [ ] Fail-open exposes the gateway to attackers who can trigger the timeout > This is a real concern, but an attacker able to degrade the service is the less common case. The everyday problem is that ordinary outages remove the control with no signal at all. ## A vendor states that its inspection service contributes to every Blueprint layer. Your architect pushes back specifically on Layer 2 (Secure Your AI Models). What makes Layer 2 the exception for a runtime filter? > Hint: Ask what Layer 2's subject is, and when it exists. - [ ] Model artifacts are stored in registries that a filter has no access to > Access is not the obstacle. Even with full registry access, inspecting prompts and responses would establish nothing about a weights file. - [x] Layer 2's subject is a static artifact that exists before any request arrives > Correct! Layer 2 secures weights, provenance, serialisation format and container -- properties that are settled before a single prompt is served. A signature is valid or invalid with no traffic involved, and no prompt inspection can establish it. A scanner can assess the model's *behaviour*, which is why assessment does contribute to Layer 2 while runtime filtering does not. - [ ] Filters cannot detect model-level vulnerabilities because they only see text > This is close, but it describes the filter's limitation rather than the reason. The decisive point is that Layer 2's subject is not present at runtime to be inspected at all. - [ ] Layer 2 is owned by the model provider in most deployment patterns > Ownership varies by pattern and is a separate question. Even where Layer 2 is entirely yours, a runtime filter still contributes nothing to it. ## A hospital wants Guard in front of a clinical assistant. Its data protection officer asks what the inspection service receives. What is the accurate answer, and what does it mean for the decision? > Hint: Consider what an inspector must have in order to inspect. - [ ] Only the portions of prompts flagged as suspicious by local pre-screening > No such pre-screen exists by default, and one capable of deciding what is suspicious would already be the guardrail. Inspection cannot be selective about input it has not yet seen. - [x] All of it -- inspecting a prompt requires receiving the prompt in full > Correct! A guardrail service sees one hundred percent of the traffic on the paths that call it, including the material the organisation's data policy exists to contain. That makes content sensitivity the gating axis: establish the most sensitive class of content that will cross the boundary and whether jurisdiction and contract permit it, before latency, coverage or features enter the discussion. - [ ] Metadata and detection signals only, with content hashed before transmission > Hashes support matching against known strings and defeat the semantic classification that catches rephrased attacks. A classifier needs the text. - [ ] Prompts but never responses, since output filtering runs at the gateway > Both directions are inspected -- the documented pattern scans the request before forwarding and the response before returning it. Leakage detection would be impossible otherwise. ## In CVE-2025-53773, injected text caused GitHub Copilot to write `"chat.tools.autoApprove": true` into `.vscode/settings.json`, disabling every confirmation prompt, after which it executed arbitrary commands. Where would the loop have stopped being useful? > Hint: Ask what text a filter would have had to recognise, and what output it would have inspected. - [ ] At VALIDATE, because a wormable payload spreads faster than any cycle > Propagation speed is a real property of this case, but VALIDATE is a monitoring phase. The loop's limit here appears earlier and is structural rather than a matter of timing. - [x] At PROTECT -- the payload was ordinary text and the output was a disk write > Correct! The injected instruction was an unremarkable sentence about a configuration setting, with nothing to distinguish it from documentation, and the damaging output was a JSON key written by a permitted file-write tool rather than a response shown to anyone. A guardrail on prompts and responses does not sit on that path. Injection was the entry; agency was the impact, and the control that closes it is denying the agent write access to the file governing its own approvals. - [ ] At SCAN, since no assessment library contained this technique in advance > SCAN is the phase that would have worked. Probing the assistant with instructions embedded in repository content establishes that it acts on them -- which is what the researcher demonstrated by hand. - [ ] At IMPROVE, because a patched product removes the need to regress-test it > A patch does not end the obligation. Replaying the technique is how you confirm each later agent version is still closed, which is the phase working as intended. ## Your quarterly report shows Guard blocking 4,000 attempts a month with a 0.3% false-positive rate, and the same class of prompt-injection finding has now appeared in three consecutive Scanner reports despite two rounds of rule tuning. What should the next cycle do? > Hint: The loop has told you something. Read what it says rather than what you hoped for. - [x] Stop tuning and change the architecture -- detection is the wrong control here > Correct! A finding that survives two full cycles of tuning has produced the loop's most valuable output: evidence that this problem is not a detection problem. What follows is a permission change, a gate, or a design change, and it belongs to Layer 3 or Layer 4 rather than to the rule set. - [ ] Escalate the rules further, since two rounds is a small sample to judge on > A third round of the thing that failed twice is the default response and the reason findings recur. The recurrence is the signal, not the sample size. - [ ] Accept the residual, given the block volume and the low false-positive rate > Those numbers describe how well the filter runs, not whether the exposure is closed. A recurring finding means the susceptibility is intact behind an effective filter. - [ ] Broaden the scan library so the next assessment catches related variants > A wider library would find more instances of a class you already know you have. It adds findings without addressing the one in front of you. ## A supplier's security questionnaire asks you to attach "evidence that the AI application is not vulnerable to prompt injection." Your latest Scanner report is clean. How should you answer? > Hint: Consider what a clean result rules out and what it leaves open. - [ ] Attach the report -- a clean assessment is what the question is asking for > It is what the question asks for, and the question is asking for something a scan cannot supply. Forwarding it as assurance is the characteristic misuse the section warns against. - [x] Attach it while stating its scope: known techniques, automated, time-boxed > Correct! A clean run means "none of the probes in this library succeeded," not "nothing works." Its adversary is automated and its budget is minutes, against a residual measured over 1,700 hours of expert red-teaming in the best public case. The report is real evidence about a bounded question, and stating the bound is what makes it usable rather than misleading. - [ ] Decline, because no assessment can demonstrate the absence of a weakness > Strictly true and unhelpful. A scan is genuine evidence about a defined set of techniques, and withholding it substitutes silence for a qualified answer. - [ ] Attach it together with the Guard blocking statistics as corroboration > Block counts show the filter is running and say nothing about what got through. Two artifacts that both measure the known set do not extend the claim.