Section 10 Quiz

Test Your Knowledge: Verifying AI Applications

Let’s see how much you’ve learned!

This quiz tests how you use AISVS to verify an application, where the LEARN mnemonic runs out, what a passing assessment does and does not establish, and how verification relates to the continuous loop.

--- shuffle_answers: true shuffle_questions: false --- ## Your team writes in a control document that the payment agent "satisfies C9.2.5." A year later an auditor cannot reproduce the claim. What went wrong, beyond anything about the agent itself? > Hint: Think about what a bare identifier means a year after it was written. - [x] The identifier was not version-pinned, so it no longer names a fixed requirement > Correct! OWASP's guidance is to write `v1.0-C9.2.5`, and to read a bare identifier as meaning "whatever is current." Requirement numbers are stable within a version and may move between them, so an unpinned citation stops naming anything fixed. This course has already been through one OWASP renumber -- eight of ten LLM Top 10 identifiers moved between the 2025 and 2026 editions -- which silently invalidated every unpinned reference. - [ ] The claim was recorded without the verification level being stated alongside it > The level matters for scoping an assessment, but each requirement carries its own level. The reproducibility failure here is that the identifier itself no longer resolves to fixed text. - [ ] Requirement text is not stable, so no citation can be relied on over time > Requirement text within a released version is stable -- the `1.0` folder is locked behind a CI guard precisely so citations hold. Stability exists; the citation failed to invoke it. - [ ] Control documents should cite chapters rather than individual requirements > Citing a whole chapter loses the specificity that makes the claim checkable, and it would face the identical problem across versions. ## A team is preparing an AISVS assessment and asks which chapters to run. Their application is a RAG assistant with a Slack MCP connector and persistent user memory. Which chapters does the LEARN mnemonic fail to point them at? > Hint: Count the letters, then count the chapters. - [ ] None -- LEARN's five letters map onto all twelve chapters at some depth > Five letters reach five chapters. The remaining seven are not covered at lower depth; they are not indexed by the mnemonic at all. - [x] Memory and MCP -- the two chapters this system most obviously needs > Correct! LEARN reaches C2, C5, C7, C9 and C11. Six chapters have no letter, and for this system the damaging omissions are `v1.0-C8` (Memory, Embeddings & Vector Database) and `v1.0-C10` (MCP Security, 23 requirements). Both describe exactly what this application is built from, and the failure is silent -- nobody notices the chapter that has no letter. - [ ] Only Infrastructure, since the platform team owns that chapter regardless > Ownership is a separate question from coverage, and the platform team's chapter is not the only one LEARN misses. - [ ] Adversarial Robustness, because the mnemonic treats it as a technique > R does point at Adversarial Robustness. The uncovered chapters are the ones with no corresponding letter. ## A developer implements an instruction hierarchy by adding code that checks whether a user message contradicts the system prompt and rewrites it if so. Does this satisfy `v1.0-C2.1.6`? > Hint: Read what the requirement says is doing the enforcing. - [ ] Yes -- the requirement asks for enforcement, and this code enforces it > The code enforces a keyword comparison, not a hierarchy. Whether the model privileges system instructions is decided inside the model, and no wrapper around the request changes that. - [x] No -- the requirement is on the system, and the mechanism is training-time > Correct! `v1.0-C2.1.6` requires that "the system enforces an instruction hierarchy," and the mechanism is a trained property (Wallace et al. 2024): models are trained on a synthetic hierarchy so they privilege system and developer instructions. You satisfy it by selecting a model that has one -- a Layer 2 decision -- and supplying the provenance labels it consumes. Code asserting the hierarchy satisfies nothing. - [ ] Partially -- it meets Level 1 but the full requirement needs a classifier > The requirement has one level and one meaning. Adding a classifier would not relocate the hierarchy into your application either. - [ ] No -- rewriting user input violates the input handling requirements elsewhere > Silently rewriting input is poor practice, but that is not why this fails. It fails because the property being verified does not live where the code sits. ## Reviewing an agent, you find it can write to the configuration file that governs its own approval prompts. No exploit has been published against your setup. How should this be recorded? > Hint: Ask what the requirement describes -- an incident, or a class. - [ ] As an informational note, pending evidence that it is exploitable in practice > This inverts the burden. The requirement does not ask whether an exploit exists; it asks whether the boundary does. - [x] As a failure of `v1.0-C9.2.5` -- self-modification must be bounded, exploit or not > Correct! `v1.0-C9.2.5` requires that any self-modification capability be restricted by enforceable boundaries. It describes the class rather than a CVE, which is exactly what a verification standard buys you over an incident list -- CVE-2025-53773 was one instance, and the requirement would have failed Copilot at design review before anyone published it. - [ ] As a Layer 6 detection gap, since runtime monitoring would catch the write > Monitoring might notice the write afterwards. The finding is that the capability exists, and detection is not the control that removes it. - [ ] As accepted risk, given the agent needs write access to function normally > It needs write access to project files. No task requires writing the file that governs its own confirmations, which is what makes this the textbook excessive-agency finding. ## Your organization completes a full AISVS Level 2 assessment with every requirement passing. What can you tell the board? > Hint: Consider what verification measures, and what it does not. - [ ] That the application meets the industry standard for AI security risk > "Meets the standard" describes conformance, and the board's question is about exposure. The two are not the same claim, and conflating them is the error the section warns about. - [x] That the controls exist and are correctly built -- not that they cannot be beaten > Correct! Verification establishes that a control exists and is correctly implemented. It never establishes that the control is unbeatable -- the residual on the best measured guardrail is still low single digits, and a passing assessment is a statement about your engineering rather than about your adversary. It is also true only on the day it is signed. - [ ] That residual risk is now concentrated in Level 3 requirements you did not run > Some residual does sit there, but the deeper limit applies within Level 2 as well: a passing control can still be defeated. - [ ] That the application is verified secure against the OWASP LLM Top 10 > The Top 10 describes what goes wrong; AISVS gives each category a checkable counterpart. Passing the checks is not the same as being immune to the category. ## Two Cursor CVEs both involved approval mechanisms that existed and failed. What do they jointly establish about approval gates? > Hint: One fired at the wrong time; the other was bound to the wrong thing. - [ ] That human approval is unreliable, since operators approve without reading > Operator inattention is a real Layer 4 concern and is not what these two cases show. In MCPoison the approval was read and the judgement was sound. - [x] That a gate is undefined until you state when it fires and what it binds to > Correct! CurXecute executed newly added MCP entries immediately, *before* the approval prompt -- `v1.0-C9.2.1`'s word "until" is doing real work. MCPoison bound approval to the server's name rather than its contents, so a genuine approval carried to a swapped command, which is what `v1.0-C9.2.8` addresses by requiring approvals cryptographically bound to action parameters and a single-use nonce. Adding an approval step fixes neither. - [ ] That MCP connectors should be blocked until the protocol matures further > Both failures were in the client's handling, not the protocol, and `v1.0-C10` specifies how to run connectors safely rather than avoid them. - [ ] That approval gates belong in the runtime rather than the user interface > Runtime enforcement is necessary and insufficient. A runtime gate firing after execution, or bound to a name, fails identically. ## A RAG assistant retrieves documents using the application's service account, then filters results in application code against the requesting user's permissions. Which requirement does this arrangement fail? > Hint: Ask at which stage the user's authorization is applied. - [x] `v1.0-C5.2.2` -- the user's authorization must apply at retrieval, not after > Correct! `v1.0-C5.2.2` requires retrieval pipelines to enforce the end-user's authorization context at each retrieval and assembly stage, rather than relying solely on the service's. Filtering afterwards means unauthorized content entered the context, so it can influence the response and surface through summarisation even when the source document is withheld. This is the mechanism behind most "the assistant showed me someone else's document" incidents. - [ ] `v1.0-C8.1.1` -- vector identifiers must be unique per tenant to stop collisions > Namespace uniqueness is a genuine requirement and a different failure. Here the identifiers are fine; the authorization stage is wrong. - [ ] `v1.0-C7.3.2` -- output filters must block responses disclosing backend data > Output filtering is worth having and this design has it. It is the compensating control, not the missing one. - [ ] None -- post-retrieval filtering is an accepted pattern when it is enforced in code > Where the filtering runs is not the issue. Content the user may not see has already reached the model and shaped its answer. ## Your AISVS assessment passed six months ago. Since then the provider silently updated the hosted model and the team added a new retrieval source. What does this do to the assessment? > Hint: Verification is a statement about a system at a moment. - [ ] Nothing -- requirements describe controls, and the controls were not modified > The controls were verified against a particular system. Both changes altered that system, and several requirements are properties of the pair rather than of the control alone. - [x] It ages it -- the assessment now describes a system you no longer run > Correct! An assessment is true on the day it is signed. A model update changes behaviour that C2, C7 and C11 requirements were verified against, and a new retrieval source changes what an injection can reach. Both are re-scan triggers in Section 9's loop, which is precisely why verification and the continuous loop are the same discipline seen from two ends: the standard defines "correct," the loop tells you whether you still are. - [ ] It invalidates it entirely, so the full assessment must be repeated from scratch > A full repeat is expensive and unfocused. The changes point at specific chapters, which is what makes triggered re-assessment workable. - [ ] It affects only Level 3 requirements, since baseline controls are change-tolerant > Level does not track change sensitivity. A model swap moves behaviour that Level 1 input and output requirements were checked against.