11. Building an AI Security Culture

Introduction

Every case study in this course was a system that had controls. Samsung had a data protection policy. Cursor had an approval prompt. Microsoft’s customers had secrets managers available to them. Copilot had a settings file that was supposed to require consent. In each case the control existed, and in each case something about how people worked around it, ahead of it, or without it decided the outcome.

That is not an argument that technology does not matter – the previous ten sections are the argument that it does. It is the observation that a control has two halves: the mechanism, and the organizational conditions under which the mechanism is actually configured, actually watched, and actually obeyed under deadline pressure. The Blueprint, AISVS and the Scanner/Guard loop cover the first half. This section covers the second.

It is also where the course stops being about your systems and starts being about your obligations. Two frameworks – the NIST AI Risk Management Framework and the EU AI Act – now decide what you must be able to show about an AI system, not merely what you would like to be true of it. One of them changed ten days before this section was last revised.

What will I get out of this?

By the end of this section, you will be able to:

  1. Scope and run an AI red team exercise against ATLAS AML.M0035 – writing rules of engagement, defining stop conditions, and converting findings into controls that cannot regress.
  2. Adapt incident response for AI-specific threats, using the current NIST lifecycle rather than the one it retired, and naming the containment action each incident class actually needs.
  3. Place a system in the NIST AI RMF, including the Generative AI Profile that carries the risks this course is about.
  4. Determine which EU AI Act regime applies to a system you are handed – the AI-system risk tiers, the parallel GPAI-model track, or both – and state the date each obligation bites.
  5. Identify which organizational practices are yours to run and which belong to leadership or legal, including the one control no security team can grant itself.

What Culture Owns – and What It Cannot Stop

The six Blueprint layers each own a boundary. This section owns none. It owns the conditions under which the other ten sections’ controls survive contact with an organization, which is a different kind of claim and needs stating precisely, because “culture is the foundation” is the sort of sentence that survives any amount of vagueness.

What it owns:

  • Whether a control is configured at all. A gateway policy that ships in monitor-only mode because nobody was willing to own the false positives is not a Layer 5 failure.
  • Whether findings become fixes. Red-teaming that produces a report nobody is assigned to is an expense, not a control.
  • Whether a person is permitted to stop. Every social-engineering attack in Layer 4 runs on urgency. If declining to act immediately is a career risk, the human gate is decorative.
  • Whether the organization can prove what it did. Regulatory obligations are discharged with documentation, and documentation is a practice before it is an artifact.

What it cannot stop:

  • Anything that does not route through a human decision. EchoLeak was zero-click. No amount of training reaches an attack whose trigger is an email the user never opened.
  • Indirect prompt injection. The agent is the one being socially engineered, and it has not had the training.
  • A vulnerability in someone else’s runtime. CurXecute was a design flaw in a vendor’s product. Your governance process does not patch it.
  • The consequences of a control you were never told about. Shadow AI is a discovery problem (Layer 4) before it is a culture problem.
The failure mode this section has to avoid

“People are the weakest link” is the most common conclusion drawn from the cases below, and it is the wrong one. In all three, the people behaved rationally given the incentives and information they had. Samsung’s engineers used a tool their employer had just approved. Cursor’s users approved an MCP server, correctly, and the product executed it before the approval. Treating those as awareness failures produces more training and no change. The finding in each case is a governance gap that made the rational action the unsafe one – and governance gaps are fixed by decisions, not by posters.


Which of These Are Yours

Every layer section in this chapter asks who owns each control, and the answer has been set by the deployment pattern from Chapter 1 Section 3. Here the axis is different – nothing on this page depends on whether you rent an API or run your own GPUs. The axis is which function inside your organization can actually decide it, and the answer is uncomfortable for the team most likely to be reading.

Practice Can you buy it? Who decides it
Red team exercise execution Yes – consultancies and vendor teams sell this Security
Red team remediation and retest No Engineering, with security tracking
AI-specific IR runbooks Partly – generic IR tooling, AI-specific steps are yours Security
Model rollback capability No – it is an architecture decision made months earlier Engineering / platform
Memory reset and quarantine capability No – see v1.0-C8.3.2 Engineering
Regulatory classification of your systems No Legal, with security input
Conformity assessment and technical documentation Partly – notified bodies, external assessors Legal / compliance
Acceptable use policy No Leadership, drafted with security and legal
Security champions program No Engineering leadership
Training content Yes Security
Authority to pause without penalty No Leadership only

Two rows are worth sitting with, because they are the two that most often go unowned.

The purchasable half of red-teaming is the half that does nothing on its own. You can buy an excellent exercise. You cannot buy the assignment of its findings to owners, the retest, or the conversion of a confirmed failure into a permanent test – and ATLAS is explicit that this third phase is part of the mitigation, not a follow-up to it. An organization that buys red teams annually and closes 30% of findings has bought a measurement of its own decline.

The last row is the one a security team cannot grant itself. Section 6 establishes that every deepfake and executive-impersonation fraud runs on urgency plus confidentiality, and that the control is a person’s willingness to stop and verify. No policy document creates that willingness. It comes from whether the last person who delayed a VP’s request was thanked or remembered for it – which is a leadership behaviour, observed rather than announced. Security can write the verification procedure. Only leadership can make using it safe.


AI Red-Teaming

Red-teaming is structured, authorized adversarial testing: a team attempts to break the AI system using real attack techniques and reports findings so they can be fixed before an attacker finds them. For AI systems it extends well past traditional penetration testing, because most of the attack surface in Chapter 2 is not a memory-safety bug or a misconfigured port – it is text arriving in a context window.

Two published sources define this practice, and the course uses both rather than inventing a third:

  • MITRE ATLAS AML.M0035 AI Red Team – the vendor-neutral mitigation, added to ATLAS in the 2026.07 release. It specifies the exercise structure, the rules of engagement, and the obligation to convert findings into controls.
  • The OWASP GenAI Red Teaming Guide v1.0 (22 January 2025) – from the same project that publishes the LLM Top 10 this course is built on. It contributes the scope dimension: the four places an AI system can be tested.

They are complementary, not competing: ATLAS tells you how to run an exercise, OWASP tells you what to point it at. AML.M0035 cites the OWASP guide among its own references.

What a Red Team Is For

  • Discover what scanners cannot. Section 9 draws the line precisely: automated assessment covers models and part of the agent surface; the chained, multi-step, human-in-the-workflow attacks that produced every Chapter 2 case study need people.
  • Validate defenses under realistic conditions. Not “is the gateway deployed” but “does the gateway hold against a payload written to survive it” – the distinction EchoLeak turns on.
  • Assess blast radius. What an attacker achieves after succeeding, not merely whether they succeed. For agentic systems this is usually the more important number.
  • Produce findings someone owns. A reproducible report, assigned, tracked, retested.

The Exercise Cycle

ATLAS organizes an exercise into three phases. This course draws it as five steps because Document and Retest are the two most often skipped, and giving them their own boxes is the point of the diagram.

graph LR
    PL["<b>Plan and Scope</b><br/><small>Threat model, rules of<br/>engagement, ATLAS TTPs,<br/>success criteria,<br/>stop conditions</small>"]
    TE["<b>Execute</b><br/><small>Manual and automated<br/>attacks. Record activity,<br/>responses and control<br/>behaviour. Clean up after.</small>"]
    DO["<b>Document</b><br/><small>Reproduction steps,<br/>prompts, responses,<br/>severity, evidence,<br/>threat-model gaps</small>"]
    RE["<b>Assign and Remediate</b><br/><small>Named owner per finding.<br/>Fix, then improve detection,<br/>IR and recovery -- not<br/>only prevention</small>"]
    RT["<b>Retest and Fix Permanently</b><br/><small>Verify the fix. Convert to a<br/>regression test, eval dataset,<br/>detection rule or<br/>deployment criterion</small>"]

    PL --> TE --> DO --> RE --> RT
    RT -.->|"Re-scope on system<br/>or threat change"| PL

    style PL fill:#2d5016,color:#fff
    style TE fill:#8b0000,color:#fff
    style DO fill:#2d5016,color:#fff
    style RE fill:#1565c0,color:#fff
    style RT fill:#1565c0,color:#fff

The dotted return edge is not decoration. AML.M0035 states that red-teaming “is a continuous process and should be repeated as the threat landscape evolves and when changes are made to the system, its components, intended use, or deployment environment.” An annual exercise against a system that shipped four model updates in the interim tested a system that no longer exists.

Rules of Engagement

This is the part most often missing, and it is what separates a red team from an incident. AML.M0035 requires rules of engagement covering authorized systems, accounts, data, and techniques; test windows; resource limits; escalation procedures; evidence handling; and stop conditions. Each of those has an AI-specific edge:

  • Authorized data. Prompt-injection testing means planting hostile content in places the system reads – ticket queues, shared drives, Slack channels, RAG corpora. Every one of those is shared with people who are not in the exercise. Plant canaries, not live payloads, and know how you will remove them.
  • Resource limits. A red team probing for unbounded consumption can generate a real bill. Agree the ceiling in advance, and note that this is the one test class where success and damage are the same event.
  • Evidence handling. Successful extraction tests produce genuinely sensitive material – system prompts, memorized training data, other tenants’ content. The exercise report becomes a classified artifact the moment it works.
  • Stop conditions. Define what ends the test rather than escalating it. An agent that has been successfully hijacked is taking real actions in real systems.
  • Cleanup. ATLAS is specific: after testing, remove “test accounts, modified data, installed software, persistent instructions, and other exercise artifacts.” That third item is the AI-specific one. An injected instruction written into agent memory or a vector store outlives the exercise window and is indistinguishable from a real compromise when someone finds it in six months.
Why the threat model comes first

AML.M0035’s planning phase asks for the system diagrammed with its “components, trust boundaries, data flows, external services, human decision points, and training- and inference-time access points” before techniques are selected. This is the same artifact Section 1 builds. If you have done threat modelling for the system, red team scoping is mostly done; if you have not, the exercise will test the parts someone happened to think of.

Scope: The Four Places to Test

The OWASP GenAI Red Teaming Guide’s contribution is that “test the AI” is four different exercises with different skills and different findings. Running one and calling it red-teaming is the most common scoping error.

Scope What is under test Typical finding Blueprint layer
Model evaluation The base model in isolation – alignment, refusal behaviour, memorization Jailbreak works at N turns; training data extractable Layer 2
Implementation testing Your prompts, guardrails, filters, and their configuration Filter bypassed by encoding; system prompt leaks Layer 5
Infrastructure assessment Serving stack, orchestration, identities, dependencies Over-scoped service account; unauthenticated internal endpoint Layer 3
Runtime behavior analysis The live system: agents, memory, tools, multi-component interaction Cross-session memory poisoning; tool chain reaches an unintended sink Layer 6

The fourth is the one that finds what the others structurally cannot. A model that passes evaluation, inside an implementation that passes filtering, on infrastructure that passes assessment, can still be hijacked at runtime – because the vulnerability is in the interaction, and none of the first three test interactions. Every agentic case study in Chapter 2 lives here.

AI-Specific Techniques

The testing toolkit maps directly onto Chapter 2, and onto the ATLAS techniques AML.M0035 names as its coverage:

Attack category Technique to exercise The question the test answers ATLAS
Prompt injection Direct, indirect, and delayed-trigger instructions in user input, retrieved documents, tool output and metadata Does the trust boundary hold when the instruction arrives in content the system was going to read anyway? AML.T0051
Jailbreaking Multi-turn, multilingual, encoded, transformed and multimodal At what point does alignment fail, and does anything detect that it did? AML.T0054
System prompt and data extraction Interface probing; synthetic canaries planted in representative sources Are there secrets in the prompt, and can content cross a tenant boundary? AML.T0056, AML.T0057
Tool misuse Unauthorized tool selection, unsafe arguments, privilege excess, tool chaining Can the agent be made to call a tool it should not, with values it should not? AML.T0053
Memory poisoning Persist malicious instructions in agent memory and long-lived threads Does it survive the session – and can you get it out again? AML.T0080
Unbounded consumption Recursive behaviour, repeated tool calls, attacker-controlled task expansion Do budgets, iteration limits and termination controls actually fire? AML.T0034

Cadence

AML.M0035 ties frequency to change rather than to the calendar, which is the more defensible policy:

  • Before a release that changes the attack surface – a new tool, a new data source, a model swap, a system prompt rewrite. Note that a model swap qualifies: Section 9 makes the case that an unreviewed model change alters behaviour underneath a passing test suite.
  • On a schedule, as a floor – quarterly or semi-annually, so that a system nobody has changed is still tested against techniques that did change.
  • After an incident – to verify the remediation and find the related weaknesses the incident implies.
  • When the threat landscape moves – a newly published technique is a test case, and this is where the Section 9 replay corpus earns its keep.

Making the Finding Permanent

A red team finding has done its job when it can no longer recur silently. ATLAS specifies the conversion: use demonstrated attacks to improve “preventive controls, detection, incident response, and recovery,” and where appropriate convert confirmed failures into “regression tests, evaluation datasets, detection logic, monitoring requirements, or deployment criteria.”

Note the breadth. The instinct after a successful red team attack is to fix the vulnerability, which addresses prevention only. The same finding should also ask: would we have seen this (Layer 6)? Would our runbook have covered it? Could we have recovered? A jailbreak that is patched but would still be invisible has been half-fixed. This is the same discipline Section 9’s IMPROVE phase applies to production detections, and it is why the two sections describe one loop rather than two.


AI-Specific Incident Response

The four-phase lifecycle you may have learned has been retired

Detect → contain → eradicate → recover comes from NIST SP 800-61r2. NIST SP 800-61 Revision 3 (April 2025) retired it. Incident response is now expressed as a Community Profile of the Cybersecurity Framework 2.0 – Govern, Identify, Protect, Detect, Respond, Recover – and the reframing is substantive rather than cosmetic: IR is treated as continuous risk management embedded in the security program, not a discrete set of duties bounded by the few days around an incident. That change matters more for AI than for anything else on the list, because AI incidents frequently have no clean end state. You cannot un-train a model, and Samsung’s defining problem was that it could not verify or retract what had already been submitted.

The CSF functions still need AI-specific content, and that content is what this section supplies. Two of the six do most of the work here.

Detect: AI-Specific Indicators of Compromise

AI incidents rarely announce themselves the way traditional ones do. There is often no crash, no alert from an EDR agent, and no failed login – the system behaves normally and produces different output.

Indicator What it suggests Where the signal comes from
Output quality or tone shifts without a deploy Prompt injection, context poisoning, or an unannounced provider-side model change Output-quality monitoring, behavioural baselines (Layer 6)
An agent’s tool-call sequence stops matching its task Goal hijacking Tool-call logging assessed against the requested task – not against a volume threshold, which Section 8 shows a realistic hijack stays under
Retrieval returns documents a user should not be able to reach Scope-constraint failure or memory poisoning Retrieval logging – AISVS v1.0-C12.1.4 requires the query, the documents retrieved, and the knowledge source
Spike in gateway blocking events An active campaign, or a probe finding your filter’s edges Gateway telemetry, alert correlation (Layer 5)
Unexpected data access by an AI service account Credential compromise or over-scoped identity being exercised IAM audit logs, access baselines (Layer 3)
Users report answers that are confidently wrong in a consistent direction Corpus poisoning – the directionality is the tell, since ordinary hallucination is not systematically biased Feedback channels correlated against retrieval sources
Cost anomaly in inference or tool spend Credential abuse, or an agent looping Spend baselines and caps (Layer 5)

The row worth dwelling on is the last-but-one. A single wrong answer is a hallucination. Wrong answers that agree with each other are evidence of a shared upstream cause, and in a RAG system that usually means the corpus. This is one of the few AI incidents whose first detector is a user, which is why the reporting path in your acceptable use policy is a detection control and not an administrative one.

Respond: Containment by Incident Class

Containment for AI incidents has a property that traditional containment mostly does not: the fastest action often destroys the evidence, and the state you need to preserve may be inside a system that is still serving traffic.

Model compromise. Route production traffic to a known-good version rather than shutting the endpoint down – this preserves availability while stopping the exposure. Freeze the suspect version’s artifacts, logs and configuration rather than deleting them. Rotate every credential associated with the endpoint. This is only possible if a previous version is still deployable, which is an architecture decision made long before the incident; it is the practical reason Layer 2 treats model versioning as a security control rather than an MLOps convenience.

Data or corpus poisoning. Quarantine the source – training set, RAG corpus, or memory store – and switch retrieval to a verified backup. Flag every output generated inside the poisoning window for review, which requires knowing when the window opened, which requires provenance you either recorded or did not. AISVS is unusually concrete about the capability this needs: v1.0-C8.3.2 requires that memory can be reset, and v1.0-C8.3.3 requires that quarantined content be retained but excluded from retrieval – retained because it is evidence, excluded because it is live. Most systems discover they have neither during the incident.

Agent hijacking. Suspend tool access first, because the agent is acting while you deliberate. Then revoke its credentials, then audit every action inside the compromise window, then decide what to roll back. The order matters and the reason is forensic: rolling back before auditing destroys the record of what happened. Note also what suspension does not do – if the injected instruction was written into durable memory, restoring tool access restores the attack. Memory is part of containment, not part of recovery.

Recover: What Recovery Means When You Cannot Un-Train

  1. Root cause against the Blueprint. Which control failed, at which layer, and was it absent, misconfigured, or defeated? Those three have different fixes and only the first is a budget problem.
  2. Restore from a verified artifact, and re-assess before it serves traffic (Section 9).
  3. Establish data integrity across training data, corpora and memory stores – and be honest where you cannot. Samsung’s case is the precedent: the control failure was the irreversibility, not a confirmed leak, and the honest posture was to act as though the data was gone.
  4. Update the controls the attack passed through – gateway rules, scanner test cases, detection logic, and the red team’s test library.
  5. Feed the threat model. An incident is a free, high-quality input to the artifact AML.M0035 asks for at planning time.

Regulatory Frameworks

Two frameworks decide what you must be able to demonstrate. They do different jobs and are frequently confused: NIST AI RMF is voluntary guidance you adopt; the EU AI Act is law that applies to you whether you adopt anything or not.

NIST AI Risk Management Framework

NIST AI RMF 1.0 (26 January 2023, currently under revision) organizes AI risk management into four functions:

Function Purpose Where this course does it
Govern Establish structures, policies, accountability This section – champions, acceptable use, the ownership table above
Map Identify and understand risk in context Threat modelling (Section 1), the OWASP taxonomies (Chapter 2 Section 1)
Measure Assess and monitor risk AISVS verification (Section 10), Scanner assessments, AI-SPM scoring (Layer 3)
Manage Prioritize and act Blueprint layer controls, the Scanner/Guard loop (Section 9)

The framework is deliberately not prescriptive about controls. It gives you the governance skeleton; the Blueprint and AISVS give you the muscle. Think of AI RMF as how to decide and document, and of the previous eight sections as what to deploy.

The part that actually applies to you: NIST AI 600-1

Citing “NIST AI RMF” for an LLM system is citing the general framework and skipping the document written for the specific case. NIST AI 600-1, the Generative AI Profile (July 2024) extends the four functions with 12 risks unique to or exacerbated by generative AI and roughly 200 suggested actions mapped onto the RMF Core.

The twelve, verbatim: CBRN Information or Capabilities · Confabulation · Dangerous, Violent, or Hateful Content · Data Privacy · Environmental Impacts · Harmful Bias and Homogenization · Human-AI Configuration · Information Integrity · Information Security · Intellectual Property · Obscene, Degrading, and/or Abusive Content · Value Chain and Component Integration.

Six of those are this course. Information Security names prompt injection and data poisoning explicitly, and names the dual nature the course opened with – generative AI both lowers the barrier to offensive capability and expands the attack surface. Confabulation is hallucination under its NIST name. Human-AI Configuration is over-trust and automation bias, the Layer 4 concern. Value Chain and Component Integration is the model supply chain of Layer 2. Data Privacy and Information Integrity cover the Layer 1 and output-handling ground.

The practical consequence: if an assessor asks how you manage generative AI risk under the RMF, the expected answer references 600-1’s risk names, not the four function names. Federal agencies operating under OMB M-24-10 use it as their reference profile.

The EU AI Act: Two Regimes, Not One

Almost every summary of the EU AI Act shows four risk tiers and stops. That table describes AI systems. A foundation model is regulated separately, as a general-purpose AI model, on a parallel track with its own obligations and its own dates – and for an audience securing LLMs, that second track is the more relevant of the two.

graph TD
    S["<b>What are you placing<br/>on the EU market?</b>"]
    S --> M["<b>A general-purpose<br/>AI model</b><br/><small>Chapter V</small>"]
    S --> A["<b>An AI system</b><br/><small>Chapters II-III, Art. 50</small>"]

    M --> M1["<b>All GPAI providers</b><br/><small>Technical documentation,<br/>copyright policy, training-data<br/>summary -- since 2 Aug 2025</small>"]
    M1 --> M2{"High-impact<br/>capability?<br/><small>presumed above<br/>10^25 FLOP</small>"}
    M2 -->|Yes| M3["<b>+ Systemic-risk duties</b><br/><small>Model evaluation, adversarial<br/>testing, incident reporting,<br/>cybersecurity -- Art. 55</small>"]

    A --> A1{"Prohibited<br/>practice?<br/><small>Art. 5</small>"}
    A1 -->|Yes| A2["<b>Cannot deploy</b><br/><small>Since 2 Feb 2025</small>"]
    A1 -->|No| A3{"Annex I or III<br/>high-risk?"}
    A3 -->|"Annex III<br/>standalone"| A4["<b>High-risk duties</b><br/><small>From 2 Dec 2027</small>"]
    A3 -->|"Annex I<br/>in a product"| A5["<b>High-risk duties</b><br/><small>From 2 Aug 2028</small>"]
    A3 -->|No| A6{"Interacts with people,<br/>or generates<br/>synthetic content?"}
    A6 -->|Yes| A7["<b>Art. 50 transparency</b><br/><small>Since 2 Aug 2026</small>"]
    A6 -->|No| A8["<b>No specific duties</b>"]

    style M fill:#1565c0,color:#fff
    style A fill:#2d5016,color:#fff
    style M3 fill:#8b0000,color:#fff
    style A2 fill:#8b0000,color:#fff
    style A4 fill:#8b0000,color:#fff
    style A5 fill:#8b0000,color:#fff

The two tracks are not exclusive. An organization that fine-tunes an open-weights model and ships it inside a hiring product is a GPAI provider and a high-risk system provider, with two sets of obligations on two different dates.

The AI system risk tiers:

Risk level What qualifies Examples Obligation
Unacceptable Practices judged incompatible with fundamental rights Social scoring by public authorities; untargeted facial-image scraping; certain real-time remote biometric identification in public spaces for law enforcement Prohibited
High-risk Annex III standalone uses, or Annex I safety components of regulated products Employment screening, credit scoring, education, critical infrastructure; medical devices, machinery, toys Risk management system, data governance, technical documentation, logging, human oversight, accuracy and robustness, conformity assessment
Limited risk Systems with transparency duties by function Chatbots, emotion recognition, deepfake and synthetic-content generators Disclosure and content marking (Art. 50)
Minimal risk Everything else Spam filters, AI in games None beyond existing law

Note that the prohibitions are narrower than they are usually reported. Real-time remote biometric identification in public spaces is restricted for law enforcement purposes with judicially authorized exceptions, not banned outright for everyone, and the difference is the kind of detail a compliance conversation turns on.

Currency: the dates moved on 27 July 2026

Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026 – six days before the AI Act’s original high-risk deadline. It amends Regulation (EU) 2024/1689 and defers the two largest tranches of obligations:

Obligation Original date Current date
Prohibited practices, AI literacy 2 Feb 2025 unchanged
GPAI model obligations (Chapter V) 2 Aug 2025 unchanged
Article 50 transparency 2 Aug 2026 unchanged – this one applied on schedule
Annex III standalone high-risk 2 Aug 2026 2 Dec 2027
Annex I high-risk in regulated products 2 Aug 2027 2 Aug 2028
New Art. 5 prohibitions (non-consensual intimate imagery, CSAM) — 2 Dec 2026

The stated reason was implementation readiness: national competent authorities were not designated and the harmonised standards a conformity assessment depends on were not finished. Deferred is not cancelled, and the requirements themselves were not weakened. Two practical consequences: a high-risk system in development has more runway than its plan assumed, and any compliance material dated before August 2026 – including the risk tables in most published summaries – states a deadline that has moved.

Provider or Deployer: The Question That Sets Your Obligations

The Act assigns duties by role, and most organizations in this course’s audience are deployers rather than providers – yet nearly all published guidance is written for providers, which is why teams routinely read obligations that are not theirs and miss the ones that are.

Provider Deployer
Who Develops a system or model, or places it on the market under its own name Uses a system under its own authority in a professional capacity
Core duties (high-risk) Risk management system, data governance, technical documentation, conformity assessment, registration, post-market monitoring Use per instructions, assign competent human oversight, monitor operation, keep logs, inform workers, report serious incidents
Security reading Everything in Layers 1-3 plus AISVS evidence Layer 4 human oversight, Layer 6 monitoring, IR runbooks

The trap worth knowing: a deployer that puts its own name on a high-risk system, changes the system’s intended purpose, or substantially modifies it becomes a provider and inherits the full obligation set. Fine-tuning a general-purpose model for a high-risk use is the case that catches people, and it is exactly what an internal AI team does by default.

The security controls this course has taught map onto the Act’s requirements more directly than the framing suggests: human oversight is Layer 4’s human-in-the-loop enforcement; logging is Layer 6 and AISVS v1.0-C12; data governance is Layer 1; accuracy and robustness testing is red-teaming and AISVS v1.0-C11. What is genuinely new is not the control – it is the requirement to evidence it.


Organizational AI Security Practices

Security Champions for AI

Designating a security-minded engineer inside each team is a well-established pattern, and it transfers with one adjustment: for AI, the champion needs a review remit rather than an advisory one, because the artifacts that need reviewing are ones a normal code review does not open.

  • Someone in each team who has read Chapter 2 properly and can recognise the shapes
  • With review responsibility for the artifacts nobody else reviews: system prompts, tool definitions and their argument schemas, agent permission scopes, and which data sources the agent may read
  • Who participates in red team exercises rather than receiving their reports
  • Who is the bridge in the direction that is usually missing – not security explaining threats to engineering, but engineering explaining to security what the system actually does, which is the input threat modelling needs and rarely gets

Training That Changes Behaviour

ATLAS names this AML.M0018 User Training and scopes it to two audiences with different content: educate developers on supply-chain risk and malicious AI artifacts, and educate users on recognising deepfakes and phishing.

  • Developers: the Chapter 2 attack categories at a working level, plus the specific mistakes – secrets in system prompts, tools scoped to what was convenient, retrieved content treated as trusted, approval gates that fire after the action.
  • Users: Section 6 is emphatic that this training must demonstrate current synthetic media quality rather than describe it. Training built on the premise that fakes look fake teaches a tell that no longer exists.
  • Everyone: how to report. The corpus-poisoning indicator above is detected by a user noticing consistently wrong answers, and that signal only becomes an incident if there is somewhere to send it.

Acceptable Use Policies

The test of an AI acceptable use policy is whether someone under deadline pressure can follow it in thirty seconds. Four questions it has to answer:

  • What data may go to which tool – named tools, named data classifications. “Do not share confidential data” fails the test because the person pasting a stack trace does not think of it as confidential.
  • When human review of output is required – tied to consequence and reversibility, matching Layer 4’s gating rule rather than restating it differently.
  • What requires approval before it ships – a new tool, a new data source, a new model, an increase in agent autonomy. This list is the trigger for red-teaming and for re-scanning, so it should be the same list in all three documents.
  • How to report something that looks wrong, with a path that does not require the reporter to be sure.

And the part most policies omit: an approved route for the need that drives the violation. Samsung’s case is the direct evidence – a ban issued after the fact does not remove the productivity pressure that produced the paste, it removes the sanctioned way to satisfy it.

Authority to Pause Without Penalty

Section 6 identifies this as the control that belongs to leadership rather than to security, and defers it here. It is the last row of the ownership table and the one with no technical implementation.

Every urgency-based attack – executive impersonation, vishing with a cloned voice, an “approve this now” request routed around the normal path – depends on the target believing that delay is more costly than compliance. Out-of-band verification procedures address this only if using them is safe. The control is not the procedure. It is the answer to: what happened to the last person who made a VP wait fifteen minutes?

Three things make it real, and none of them are training:

  • A stated rule that verification delay is never a performance issue, from someone senior enough that it is not a security-team opinion.
  • Executives who submit to their own procedure. A verification step the leadership team routes around is a step that tells everyone the exception exists.
  • A visible instance. The first time a paused request turns out to have been legitimate, and the person who paused it is thanked in public, the control is installed. Until then it is a document.

This is also the honest boundary of the section. Everything above can be written down by a security team. This cannot, and a security function that reports a mature culture without it has measured its documents rather than its behaviour.


Defense Perspective: Three cases where the control existed

Chapter 2’s cases are usually read as technical failures. Read as organizational ones, each says something different – and in each, the popular reading is wrong in a way that would send you to fix the wrong thing.

Samsung, March 2023 (Chapter 2 Section 6) – three incidents in twenty days, within twenty days of Samsung authorising ChatGPT use in its Device Solutions division. The popular reading is that engineers did not understand the risk. The record does not support it and the timing argues against it: this was an approved tool being used for its approved purpose, by people who had just been told it was allowed. The gap was between the authorisation and the boundary conditions – what could go in, how much, and with what data classification – which nobody had written. The technical controls that would have caught it are Layer 1 classification and Layer 5 egress inspection, both acting on traffic; note that they are not Layer 4 controls, which is why Section 6 files this case explicitly as not a Layer 4 win. Samsung’s own first response, a 1,024-byte prompt cap, is the instructive part: a control that bounds a leak’s size without touching its nature. Organizational lesson: approving a tool without publishing its boundary conditions is not a policy, and the twenty-day gap is how fast that costs you.

Cursor CurXecute, August 2025 (Chapter 2 Section 5) – and this one is here as a warning about the organizational reading itself. It is widely retold as developers installing unverified MCP servers, and this course said so too until Section 10 corrected it. The MCP server was legitimate – a vendor Slack connector. The attacker posted a message in a channel the agent read. No governance process requiring security review of MCP servers would have prevented anything, because nothing unapproved was installed. The approval prompt existed and Cursor executed the new config entry before it fired. Organizational lesson: the governance control that sounds obvious after an incident is frequently one that could not have reached it. Read the mechanism before writing the policy – an AUP clause requiring MCP review here would have produced a documented, audited, entirely bypassed control.

Microsoft v. Storm-2139, 2024-2025 (Chapter 2 Section 1) – a cybercrime network harvested exposed Azure OpenAI API keys belonging to Microsoft’s customers from public repositories and credential dumps, then resold access through a reverse proxy. Microsoft’s Digital Crimes Unit was the plaintiff, not the negligent party; the victims were enterprises whose own keys had leaked, and most learned of it from their bill. Every one of those organizations had access to a secrets manager. Organizational lesson: the credential-hygiene controls are ancient and free, and what failed was enforcement at the point of commit – pre-commit scanning, short-lived credentials, and a spend cap with an alert underneath it, which is the control that converts this from an open-ended loss into a bounded one.

The pattern across all three: the technical control existed or was freely available. What was missing was a decision – what the boundary is, whether the mechanism reaches the attack, who enforces it at the point of use. That is what this section means by culture, and it is why the popular reading of each case would have produced a control that did not help.


Section 11 → Framework Mapping

Unlike the layer sections, this one maps to external governance frameworks rather than to OWASP attack categories – because its subject matter is what those frameworks exist to specify.

Practice Framework anchor Where it is verified
AI red-teaming program ATLAS AML.M0035; OWASP GenAI Red Teaming Guide v1.0 AISVS v1.0-C11 Adversarial Robustness
Incident response NIST SP 800-61r3 (CSF 2.0 profile) AISVS v1.0-C12 monitoring and logging; v1.0-C8.3.2-C8.3.3 memory reset and quarantine
Risk governance NIST AI RMF 1.0 Govern; NIST AI 600-1 Documentation, not testing
Human oversight EU AI Act Art. 14; ATLAS AML.M0029 AISVS v1.0-C9.2 approvals and reversibility
Training ATLAS AML.M0018 Completion is not evidence; incident reports are
Transparency to users EU AI Act Art. 50 Disclosure and synthetic-content marking
Logging and traceability EU AI Act Art. 12 (high-risk); NIST AI RMF Measure AISVS v1.0-C12.1

The right-hand column is the honest part of this table. Governance practices are the hardest thing in the course to verify, because most of them produce documents, and a document is evidence that a document was written. The two that produce falsifiable evidence are red-teaming (findings, with retest results) and incident response (runbooks, exercised). Prefer those as your measures.


Bringing It All Together

You have completed a three-chapter arc:

  • Chapter 1 built the mental models – how transformers, context windows, retrieval and agent loops actually work, established as trust boundaries rather than as capabilities, because that framing is what Chapters 2 and 3 act on.
  • Chapter 2 mapped the threat landscape – the OWASP LLM Top 10 (2026) and Agentic AI Top 10, with real incidents attached, and the repeated finding that agentic attacks rarely need a compromised component because an agent’s job is to read untrusted things and then act.
  • Chapter 3 built the defenses – six Blueprint layers, AISVS verification, the Scanner/Guard continuous loop, and this section’s organizational conditions.

The three form one argument: understand the technology, recognize the threats, build the defenses – and then keep doing all three. Every case study in Chapter 2 was disclosed after some version of this material was first written, and the framework the course maps to renumbered eight of its ten identifiers while it was being revised.

If you take one habit from Chapter 3, take the one this section and Section 10 both arrive at from different directions: when a control sounds obviously correct, check that it can reach the thing it is meant to stop. The Hub scanner did not catch nullifAI. Input filtering did not catch EchoLeak. Virtual patching does not block text. An MCP review process would not have stopped CurXecute. Every one of those is a control that would pass a policy audit, and the check that separates them is always the same: read the mechanism, then find that control’s name in the incident, then read the sentence around it.

Key Takeaways
  • Red-teaming is specified, not invented: ATLAS AML.M0035 gives the exercise structure, the rules of engagement, and the requirement to convert findings into regression tests and detection logic; the OWASP GenAI Red Teaming Guide gives the four scopes – model, implementation, infrastructure, and runtime, of which only the last tests component interaction
  • The four-phase IR lifecycle was retired by NIST SP 800-61r3 in April 2025 in favour of a CSF 2.0 profile that treats IR as continuous risk management – which suits AI, where incidents often have no clean end state because a model cannot be un-trained
  • AI containment is ordered by forensics as much as by speed: suspend, revoke, audit, then roll back – and memory reset (v1.0-C8.3.2) is a containment capability most systems find they lack during the incident
  • The EU AI Act runs two parallel regimes, and for LLM work the GPAI-model track (Chapter V, in force since 2 Aug 2025) matters more than the four risk tiers; the Digital Omnibus, Regulation (EU) 2026/1744, deferred Annex III high-risk to 2 Dec 2027 and Annex I to 2 Aug 2028 while leaving Art. 50 transparency in force
  • For LLM systems the applicable NIST artifact is AI 600-1, the Generative AI Profile, whose twelve risk names – not the four function names – are what an assessment is expected to reference
  • Nothing on this page is purchasable except red team execution and training content; the two highest-leverage practices, regulatory classification and authority to pause without penalty, belong to legal and to leadership

Test Your Knowledge

Ready to test your understanding of AI security culture and governance? Head to the quiz to check your knowledge.


Course Complete!

Congratulations on completing all three chapters. You can now describe how AI systems work as a set of trust boundaries, recognise the attack classes that cross them, deploy layered defenses against those attacks, verify the result against a published standard, and place the whole thing inside the governance and regulatory obligations it has to satisfy.

The two references worth keeping open are the glossary and the references list – the second in particular, because every claim in this course traces to a source there, and sources change.