Section 11 Quiz
Test Your Knowledge: Building an AI Security Culture
Let’s see how much you’ve learned!
This quiz tests your understanding of AI red-teaming under ATLAS AML.M0035, the four OWASP red team scopes, AI-specific incident response after NIST SP 800-61r3, the NIST Generative AI Profile, the EU AI Act’s two regimes and their current dates, and who inside an organization actually owns each practice.
---
shuffle_answers: true
shuffle_questions: false
---
## A red team has finished a two-week indirect prompt injection exercise against a customer-support agent, planting test instructions in a shared ticket queue and in the RAG corpus. Which cleanup item does ATLAS `AML.M0035` call for that a traditional penetration test would not?
> Hint: One kind of exercise artifact keeps executing after the test window closes.
- [ ] Removal of the test accounts created to submit the tickets
> Test account removal is required, but it is standard for any penetration test. The question asks what is specific to testing an AI system.
- [x] Removal of persistent instructions written into memory or the corpus
> Correct! `AML.M0035` lists "persistent instructions" alongside test accounts, modified data and installed software. It is the AI-specific one: an injected instruction that reached agent memory or a vector store keeps firing after the exercise ends, is indistinguishable from a real compromise when someone finds it months later, and no amount of revoking access removes it.
- [ ] Restoration of the tickets to the state they held before the test
> Modified data restoration is listed, and it matters, but ticket contents are inert. The concern that makes AI cleanup different is content that continues to influence behaviour.
- [ ] Revocation of the elevated permissions granted for the test window
> Permission revocation is ordinary post-engagement hygiene. It also does not help if the instruction is already stored somewhere the agent reads on every run.
## A model passed evaluation for jailbreak resistance, the guardrail configuration was tested and held, and an infrastructure assessment found no over-scoped identities. Two weeks later an agent is hijacked by an instruction that was stored in its memory during an earlier session. Which OWASP red team scope was never exercised?
> Hint: Three scopes test components. One tests what happens between them.
- [ ] Model evaluation, which should have covered instruction-following behaviour
> Model evaluation tests the base model in isolation. A model that correctly resists jailbreaks in isolation can still act on an instruction retrieved from its own memory as ordinary context.
- [ ] Implementation testing, since the guardrail did not catch the payload
> Implementation testing covers prompts, filters and their configuration -- and it was run here. It inspects a single request, not what one session leaves behind for the next.
- [x] Runtime behavior analysis, which is where component interaction is tested
> Correct! Model, implementation and infrastructure scopes each test a component in isolation. Runtime behavior analysis tests the live system -- agents, memory, tools, and multi-component interaction over time. Cross-session memory poisoning exists only in the interaction, so the first three scopes structurally cannot find it. This is where every agentic case study in Chapter 2 lives.
- [ ] Infrastructure assessment, because the memory store was in scope for it
> Infrastructure assessment covers the serving stack, orchestration and identities. It would find an unauthenticated memory store, but not a legitimate write of hostile content through the intended path.
## An organization commissions an external red team every year. The reports are thorough, and roughly 30% of findings are closed before the next engagement. Judged against `AML.M0035`, what is the most serious problem?
> Hint: Consider what the mitigation actually consists of.
- [ ] The cadence is wrong; the exercises should be run quarterly instead
> Frequency is a real issue -- `AML.M0035` ties cadence to system and threat change rather than the calendar -- but running an unfixed system through more exercises produces more unclosed findings.
- [ ] External testers cannot know the system well enough to be effective
> External red teams are explicitly a normal way to run the exercise, and unfamiliarity is often useful. Nothing here suggests the testing itself was the weak part.
- [x] Two of the mitigation's three phases are missing, so nothing is mitigated
> Correct! `AML.M0035` is a three-phase mitigation: plan and scope, execute, then assess, report and improve. The third phase requires findings assigned to owners, tracked remediation, retesting, and conversion of confirmed failures into regression tests, evaluation datasets, detection logic or deployment criteria. Buying only execution buys a measurement. At 30% closure, each year's report also re-finds last year's issues.
- [ ] The findings are not being mapped to OWASP LLM Top 10 categories
> Category mapping is useful for coverage analysis, but an uncategorised finding that gets fixed beats a neatly categorised one that does not.
## An agent with database write access and an email tool is behaving anomalously and is suspected of being hijacked. Which sequence of containment actions is correct, and why is the order what it is?
> Hint: One of these steps destroys the record the others depend on.
- [ ] Audit the window, roll back the changes, suspend tools, revoke credentials
> Auditing before suspending leaves the agent acting throughout the investigation. Every minute spent reading logs is a minute of further writes and further emails.
- [x] Suspend tools, revoke credentials, audit the window, roll back changes
> Correct! Suspension stops the damage first, revocation prevents re-exploitation through the same identity, and auditing establishes the scope before anything is altered. Rollback comes last because it overwrites the record the audit depends on -- containment order here is set by forensics as much as by speed. Note the limit: if the hostile instruction was written to durable memory, restoring tool access restores the attack, so memory is part of containment rather than recovery.
- [ ] Revoke credentials, roll back the changes, audit the window, suspend tools
> Rolling back before auditing destroys the forensic record, and suspending tools last leaves the agent's most dangerous capability live the longest.
- [ ] Suspend tools, roll back the changes, revoke credentials, audit the window
> Suspension first is right, but rolling back second overwrites the evidence. You cannot scope an incident you have already reverted, and the rollback decision needs the audit to be informed.
## Your incident response plan is organised around detect, contain, eradicate and recover. A reviewer flags it as out of date. What changed, and why does it matter more for AI systems than for most?
> Hint: The revision changed what kind of thing incident response is, not just its labels.
- [ ] Nothing substantive changed; the phases were renamed but map one-to-one
> The functions are not a renaming. The revision recast incident response as an ongoing element of risk management rather than a bounded sequence, which changes what the plan has to contain.
- [ ] The phases were reordered so that eradication now precedes containment
> No reordering of that kind occurred. Eradicating before containing would leave the attacker active while their artifacts are removed.
- [x] SP 800-61r3 replaced it with a CSF 2.0 profile treating IR as continuous
> Correct! NIST SP 800-61 Revision 3 (April 2025) retired the r2 four-phase lifecycle in favour of a Cybersecurity Framework 2.0 Community Profile -- Govern, Identify, Protect, Detect, Respond, Recover -- and reframed incident response as continuous risk management rather than duties bounded by the days around an incident. That suits AI, where incidents often have no clean end: a model cannot be un-trained, and Samsung's defining problem was being unable to verify or retract what had already been submitted.
- [ ] The revision applies only to federal agencies, so it is not relevant here
> SP 800-61 is broadly adopted guidance rather than an agency-only mandate, and the r3 framing is the current reference regardless of sector.
## Support tickets show users complaining that a RAG assistant is giving wrong answers about the refund policy. Sampling shows the answers are wrong in a consistent direction. What does the consistency indicate?
> Hint: Compare this with how ordinary hallucination behaves.
- [ ] The model has drifted, so the provider changed it without an announcement
> A silent provider-side model change is a real incident class and shows up as an output shift. It would not produce errors that agree with each other on one specific policy.
- [ ] The temperature setting is too high, so sampling is producing variance
> High temperature produces answers that vary between requests. Here they agree with each other, which is the opposite signal.
- [x] A shared upstream cause -- most likely a poisoned or wrong retrieval source
> Correct! Ordinary hallucination is not systematically biased; errors scatter. Agreement between wrong answers implies they share a source, and in a RAG system that means the corpus -- a poisoned document, or a stale one that outranks the current policy. Directionality is the tell. It is also one of the few AI incidents whose first detector is a user, which is why the reporting path in an acceptable use policy is a detection control.
- [ ] Users are asking ambiguous questions that the assistant resolves badly
> Ambiguity produces inconsistent interpretations across users. It does not steer many different users to the same wrong answer.
## An assessor asks how your team manages generative AI risk under the NIST AI RMF. Which response demonstrates that you are using the applicable guidance?
> Hint: There is a NIST document written for this specific case.
- [ ] Describe how the six Blueprint layers are deployed across the environment
> The Blueprint is what you deploy; the RMF is how you govern and document. A control inventory does not answer a question about risk management structure.
- [x] Reference AI 600-1's risk names and the actions mapped to each function
> Correct! NIST AI 600-1, the Generative AI Profile (July 2024), extends the RMF with 12 risks unique to or exacerbated by generative AI and 196 suggested actions mapped onto the RMF Core. Six of the twelve are this course's material -- Information Security names prompt injection and data poisoning outright, Confabulation is hallucination, Human-AI Configuration is over-trust. An assessment expects the profile's risk names, not the four function names.
- [ ] Explain the Govern, Map, Measure and Manage functions and your use of each
> This is the general framework, and it applies to any AI system. Answering with it alone signals that the generative-AI-specific profile has not been used.
- [ ] Provide the AISVS verification report for the application under assessment
> AISVS supplies technical verification evidence and is genuinely useful, but it is explicitly not a risk management or governance framework and does not answer this question.
## Your team fine-tunes an open-weights model on internal data and ships it inside a product that screens job applicants, sold across the EU. Which statement about your EU AI Act obligations is correct?
> Hint: Two regimes can apply to the same organization at once, and one of them started earlier.
- [ ] Only the high-risk system obligations apply, and they began 2 August 2026
> Two errors. Fine-tuning and distributing a general-purpose model brings Chapter V duties as well, and the Annex III high-risk date was deferred by the Digital Omnibus.
- [ ] Only GPAI model obligations apply, since the model is what you modified
> Employment screening is an Annex III high-risk use, so the system obligations attach too. The model track does not absorb the system track.
- [x] Both regimes apply, on different dates -- and modification makes you a provider
> Correct! The Act runs two parallel tracks. GPAI model obligations under Chapter V have applied since 2 August 2025, while Annex III standalone high-risk duties were deferred by Regulation (EU) 2026/1744 from 2 August 2026 to 2 December 2027. Substantially modifying a system or putting your own name on it converts a deployer into a provider, which is exactly what fine-tuning for a high-risk use does -- and it is the trap internal AI teams fall into by default.
- [ ] Neither applies until 2 December 2027, when the deferred deadline arrives
> The deferral moved the Annex III high-risk date only. Prohibitions, AI literacy and the GPAI obligations were untouched and are already in force.
## A team argues that the Samsung ChatGPT incident proves employees need better security awareness training. What does the record actually support?
> Hint: Consider when the incidents happened relative to Samsung's own decision.
- [ ] The engineers acted against policy, so enforcement rather than training failed
> They did not act against policy. Use of the tool had been authorised, which is why this is not an enforcement gap either.
- [ ] Technical DLP was absent, so the fix is a data loss prevention deployment
> Traffic-acting controls would have caught it -- Layer 1 classification and Layer 5 egress inspection. But naming the missing product skips the decision that was never made about what may cross the boundary.
- [x] The tool was approved without anyone publishing its boundary conditions
> Correct! All three incidents fell within twenty days of Samsung authorising ChatGPT use in its Device Solutions division. These were approved users of an approved tool, so awareness training addresses the wrong gap. What was missing was the boundary set -- which data classifications, how much, for what purposes. Samsung's own first response, a 1,024-byte prompt cap, shows the shape: a control that bounds a leak's size without touching its nature.
- [ ] Employees were unaware that submissions could be used for model training
> Under OpenAI's default consumer policy submissions were eligible for training, and the opt-out arrived the following month. But the control failure was the irreversibility -- Samsung could neither verify nor retract -- not a knowledge gap.
## Your security team has written verification procedures for out-of-band confirmation of urgent executive requests, trained everyone on them, and measured 96% completion. Why is the control still not installed?
> Hint: Ask what the procedure costs the person who uses it.
- [ ] The procedures need to be automated so that they cannot be skipped
> Automation helps where a step can be enforced in a system. These requests arrive by phone, chat and email, and the decision to verify is a human one.
- [ ] Completion rates measure attendance, so the training content must be weak
> Content quality is worth checking, but even excellent training does not change the calculation for someone weighing a delay against an executive's stated urgency.
- [x] Whether delay is safe is set by leadership behaviour, not by the procedure
> Correct! Every urgency-based attack depends on the target believing delay costs more than compliance. The procedure only gets used if using it is safe, and that is answered by what happened to the last person who made a VP wait -- not by a document. It takes a stated rule that verification delay is never a performance issue, executives who submit to their own procedure, and a visible instance of someone being thanked for pausing.
- [ ] The verification channel must be one the requester has not supplied
> That is a correct requirement and Section 6 states it. It describes how to verify, and this question is about whether anyone will choose to.