References
Every standard, framework, CVE, whitepaper, research paper and tool the course cites, grouped by type. Each entry says what the source actually claims – including where its headline number is routinely miscited – and which sections rely on it, so an entry doubles as a map of where a claim lives in the course.
This is a currency check, not a bibliography
Read an entry for its date and version first. Eight of the ten OWASP LLM identifiers moved in August 2026, the EU AI Act’s high-risk deadlines moved in July 2026, and MITRE ATLAS ships monthly. Two habits the course asks you to keep, and applies here: cite an identifier with its version (v1.0-C9.2.5, not C9.2.5), and check a framework’s own source rather than a write-up about it.
All entries on this page were verified in August 2026. Where a source is still in draft or beta, the entry says so – treat those as unfinished, not as standards.
Standards & Frameworks
- OWASP GenAI LLM Top 10 (2026) – Industry-standard vulnerability taxonomy for LLM-specific security risks, published 4 August 2026 and the edition this course uses throughout. The first release checked against real-world incident data: a community vote at 75% weight combined with 6,639 classified incidents at 25%. Eight of the ten identifiers moved from the 2025 edition, so material written before August 2026 uses different numbers – Chapter 2, Section 1 carries the translation table. Note the one that also changed name:
LLM07:2025 System Prompt LeakagebecameLLM08:2026 Hidden Context Exposure, broadened to cover all non-user-facing context.LLM07:2026is Misinformation, so an unpinned “LLM07” is not merely imprecise – it now names a different category. Active source repository: GenAI-Security-Project/GenAI-LLM-Top10. Referenced in Chapter 2, Section 1, the Chapter 2 labs page, and the course landing page; individual categories are badged throughout Chapters 2 and 3. - OWASP Top 10 for LLM Applications (2025) – The superseded edition, archived. Worth knowing because tooling, audit reports and vendor mappings written before August 2026 cite its identifiers – and because a document citing a bare
LLM07may mean either edition. Chapter 2, Section 1 carries the translation table; Chapter 3, Section 10 uses the renumber as its argument for version-pinning every identifier. - OWASP AISVS 1.0 – Artificial Intelligence Security Verification Standard – The application-level counterpart to the Top 10 lists: 12 chapters, 191 testable requirements phrased as “Verify that…”, each assigned one of 3 verification levels, published 24 June 2026 under CC BY-SA 4.0 and modelled on the OWASP ASVS. Founded by Jim Manico. OWASP positions it as supplying “the detailed controls to mitigate” the LLM Top 10 and Agentic Top 10, and that link is now a shipped artifact rather than a positioning claim: the LLM Top 10 2026
finalfolder carriesLLMZZ_Appendix_AISVS_Mapping.mdalongsideAppendix_A_Related_Framework_Mappings.md. AISVS is also explicit about what it is not – it is “intentionally narrow”, assumes the underlying application is already verified against ASVS at the matching level, and hands governance, impact assessment and risk-management process design to ISO/IEC 42001, ISO/IEC 23894 and the NIST AI RMF. The weighting is worth noting:C9Orchestration & Agentic Security (34 requirements) andC10MCP Security (23) are 30% of the standard between them. Cite requirements with the version prefix –v1.0-C9.2.5, not bareC9.2.5– since identifiers are stable within a release and may move between them; the1.0folder is locked behind a CI guard so v1.0 citations hold permanently. Counts verified August 2026 against the locked1.0/enfolder: several widely-circulated write-ups state chapter and requirement totals that do not match the standard. Referenced in Chapter 2, Section 1; Chapter 3, Sections 9, 10 and 11; and the Chapter 3 overview. - MITRE ATLAS – Adversarial Threat Landscape for AI Systems, mapping adversarial tactics and techniques against machine learning systems with real-world case studies. Release 2026.07 carries 16 tactics, 178 techniques (101 parents plus 77 sub-techniques), 37 mitigations
AML.M0000-AML.M0036and 68 case studies; it is versioned monthly and growing quickly, so treat any count as a snapshot and pin the release. The framework the course reaches for wherever an attack has no OWASP slot – deepfake impersonation, adversarial perturbation, red-teaming as a mitigation. Two cautions from using it in anger. Names drift between releases, not just counts –AML.M0004is now Limit AI Service Query Volume and Rate, and every stale mirror still shows the old wording. It is split into predictive and generative halves:AML.M0015isPredictive AIAdversarial Input Detection and is the wrong half of the framework for anything LLM-shaped; the generative counterpart isAML.M0020Generative AI Guardrails, whose description reads as a scope statement for Blueprint Layer 5. Fetch it from the GitHub release asset –gh api repos/mitre-atlas/atlas-data/releases/latest, thenATLAS-<release>.yaml– becauseatlas.mitre.orgis a single-page app that returns HTTP 404 to every fetcher (its links still work for a human), anddist/ATLAS.yamlonmainis the deprecated v5.6 with stale names. Canonical data: mitre-atlas/atlas-data. Techniques are cited in Chapter 2, Sections 1, 3 and 6; mitigations carry Chapter 3’s vendor-neutral spine in Sections 6-11 –M0004,M0008,M0016,M0018,M0019,M0020,M0021,M0024,M0029,M0030,M0031,M0032,M0033,M0034,M0035,M0036. Referenced in Chapter 2, Sections 1, 3 and 6, and Chapter 3, Sections 1 and 6 through 11. - MAESTRO – Agentic AI Threat Modeling Framework (Cloud Security Alliance, February 2025) – Published because STRIDE, PASTA and LINDDUN all assume entities with fixed roles and trust boundaries drawn at design time, and an agent has neither: it is caller, callee and data store at once, its output re-enters as its own input, and installing a tool redraws the boundary at runtime. Replaces threat categories per entity with seven layers to reason across – foundation model, data operations, agent frameworks, deployment infrastructure, evaluation, orchestration, ecosystem – so multi-agent interaction is in scope by construction. Use it for agentic systems and STRIDE inside a layer where boundaries hold still; pair either with ATLAS, which is the only one of the three that enumerates techniques. Referenced in Chapter 3, Section 1.
- NIST AI Risk Management Framework (AI RMF 1.0) – Voluntary framework for managing AI risk across the lifecycle, published 26 January 2023 and organised around four functions: Govern, Map, Measure, Manage. Deliberately non-prescriptive about controls – it supplies the governance structure inside which controls like the Blueprint’s are selected and documented. NIST states the framework is being revised, so pin the version you cite. Referenced in Chapter 3, Section 11.
- NIST AI 600-1 – Generative AI Profile – The document that actually applies to LLM systems, and the one most often skipped in favour of citing the parent framework. Published July 2024, it extends the AI RMF with 12 risks unique to or exacerbated by generative AI and 196 suggested actions mapped onto the RMF Core. The twelve, verbatim from section 2: CBRN Information or Capabilities · Confabulation · Dangerous, Violent, or Hateful Content · Data Privacy · Environmental Impacts · Harmful Bias and Homogenization · Human-AI Configuration · Information Integrity · Information Security · Intellectual Property · Obscene, Degrading, and/or Abusive Content · Value Chain and Component Integration. Six map onto this course – Information Security names prompt injection and data poisoning outright, Confabulation is hallucination under its NIST name, Human-AI Configuration is over-trust and automation bias. Federal agencies under OMB M-24-10 use it as their reference profile. Cited by MITRE ATLAS
AML.M0035. Referenced in Chapter 3, Section 11. - NIST SP 800-61r3 – Incident Response Recommendations and Considerations – April 2025. Retires the four-phase lifecycle (preparation, detection and analysis, containment/eradication/recovery, post-incident activity) that Revision 2 established and that most IR plans still use. Incident response is now expressed as a CSF 2.0 Community Profile – Govern, Identify, Protect, Detect, Respond, Recover – and reframed as continuous risk management rather than a discrete set of duties bounded by the days around an incident. The reframing matters most for AI, where incidents frequently have no clean end state: a model cannot be un-trained. Referenced in Chapter 3, Section 11.
- NIST Cybersecurity Framework (CSF) 2.0 – Published 26 February 2024, and on this page because SP 800-61r3 above is expressed as a CSF 2.0 Community Profile – you cannot read the current incident-response guidance without it. Its substantive change from 1.1 is the addition of Govern as a sixth function alongside Identify, Protect, Detect, Respond and Recover. Governance activities – risk tolerances, roles and responsibilities, policy – were previously inside Identify; promoting them to a function of their own is what makes them visible across the whole programme rather than as one category among many. It is deliberately sector- and technology-neutral, so it says nothing AI-specific; its role in this course is as the structure the AI-specific material hangs on, the same division of labour AISVS describes when it hands governance to ISO/IEC 42001 and the NIST AI RMF. Referenced in Chapter 3, Section 11.
- OWASP GenAI Red Teaming Guide v1.0 – Published 22 January 2025 by the same OWASP GenAI Security Project that publishes the Top 10 lists. Its contribution is scope: AI red-teaming is four distinct exercises, not one – model evaluation, implementation testing, infrastructure assessment, and runtime behavior analysis. The fourth finds what the first three structurally cannot, because it is the only one that tests component interaction, which is where every agentic case study in this course lives. Pairs with ATLAS
AML.M0035, which cites it. Referenced in Chapter 3, Section 11. - MITRE ATLAS
AML.M0035AI Red Team – Added in ATLAS release 2026.07 and the vendor-neutral specification of red-teaming as a mitigation rather than an activity. Three phases: Plan and Scope (threat model, rules of engagement covering authorized systems, accounts, data, techniques, test windows, resource limits, escalation, evidence handling and stop conditions), Execute, and Assess, Report, and Improve – the last requiring findings assigned to owners, tracked remediation, retest, and conversion of confirmed failures into “regression tests, evaluation datasets, detection logic, monitoring requirements, or deployment criteria.” Buying only the middle phase buys a measurement. Scope is explicitly the whole system: models and data, agents including memory and tools, retrieval, identities and permissions, dependencies, infrastructure, interfaces and human workflows. Its cleanup requirement names persistent instructions alongside test accounts and modified data – the AI-specific artifact. Referenced in Chapter 3, Sections 8, 9, 10 and 11. - EU AI Act – Regulation (EU) 2024/1689 – In force 1 August 2024. Runs two parallel regimes, and most summaries show only the first: risk tiers for AI systems (unacceptable / high-risk / limited / minimal), and a separate track for general-purpose AI models under Chapter V, applicable since 2 August 2025, with additional systemic-risk duties under Art. 55 presumed above 10^25 FLOP of training compute. For LLM work the second track is the more relevant one. Obligations attach by role: providers carry the development and conformity burden, deployers carry use-per-instructions, human oversight, monitoring and logging – and a deployer that renames or substantially modifies a high-risk system becomes a provider. Referenced in Chapter 1, Section 1 and Chapter 3, Section 11.
- Digital Omnibus on AI – Regulation (EU) 2026/1744 – Published in the Official Journal 24 July 2026, in force 27 July 2026, six days before the AI Act’s original high-risk deadline. Defers Annex III standalone high-risk obligations from 2 August 2026 to 2 December 2027, and Annex I high-risk in regulated products to 2 August 2028, citing implementation readiness – undesignated national competent authorities and unfinished harmonised standards. Article 50 transparency applied on schedule on 2 August 2026 and was not deferred, and prohibitions, AI literacy and GPAI obligations were untouched. Adds two Art. 5 prohibitions covering non-consensual intimate imagery and CSAM, applying 2 December 2026. Deferred is not cancelled and nothing was weakened – but any compliance material dated before August 2026, including most published risk tables, states a deadline that has moved. Referenced in Chapter 3, Section 11.
Whitepapers
- Security for AI: Blueprint for Your Datacenter and Cloud (TrendAI) – The framework Chapter 3 is built on, and the only source on this page that is a vendor framework rather than a standards-body one – read it as an architecture for organising controls, not as something you can be audited against. Six layers, each owning a different boundary and each with its own section: Layer 1 Secure Your Data (Chapter 3, Section 3) · Layer 2 Secure Your AI Models (Section 4) · Layer 3 Secure Your AI Infrastructure (Section 5) · Layer 4 Secure Your Users (Section 6) · Layer 5 Secure Access to AI Services (Section 7) · Layer 6 Defend Against Zero-Day Exploits (Section 8). The layer numbers are not a maturity order or a sequence to implement in – Chapter 3, Section 2 compares them on when each acts (Layers 1 and 2 before anything runs, Layer 5 in the request path, Layer 6 beneath it) and Chapter 3’s overview maps all twenty OWASP categories onto them. Which layer’s controls are actually yours is set by the deployment pattern you chose in Chapter 1, Section 3, and the answer is not predictable from the layer’s subject matter: Layer 3 mostly evaporates on a managed API, while Layer 6 keeps seven of nine rows with the customer. Referenced throughout Chapter 3 and named in the course landing page.
Foundational Papers
- Computing Machinery and Intelligence (Turing, 1950) – The paper that opens with “Can machines think?” and replaces the question with the imitation game. The origin point for the field, and the reason the course treats “can it think?” as the wrong question to bring to a security review. Referenced in Chapter 1, Section 1.
- Attention Is All You Need (Vaswani et al., 2017) – Introduced the Transformer architecture and the attention mechanism, the foundation of every modern LLM. Referenced in Chapter 1, Sections 1 and 4.
Architecture & Long Context
The mechanics behind Chapter 1, Section 4. Both papers matter for the same reason: they show that “it has a million-token window” and “it uses efficient attention” are claims with specific, bounded meanings.
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., NeurIPS 2022) – Reformulates attention so the full attention matrix is never written to GPU memory. Memory use drops from quadratic to linear in sequence length and wall-clock time falls sharply, while the FLOP count stays quadratic and the output stays exact. The paper’s real argument is that attention is bottlenecked by memory traffic rather than arithmetic. Referenced in Chapter 1, Section 4.
- Lost in the Middle: How Language Models Use Long Contexts (Liu et al., TACL 2024) – Measures where in a long input a model actually finds information. Performance is highest when the relevant passage sits at the beginning or the end and degrades markedly in the middle, in some settings falling below the same model’s closed-book performance. The empirical case for treating context position as a design variable rather than assuming a large window is uniformly usable. Referenced in Chapter 1, Section 4.
Prompting, Reasoning & Determinism
The evidence behind Chapter 1, Section 5. The first two papers establish chain-of-thought; the last two are the reason the section treats a reasoning trace as output rather than as evidence, and a tested prompt as a sample rather than a guarantee.
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (Wei et al., NeurIPS 2022) – Showed that supplying worked examples of step-by-step reasoning sharply improves accuracy on arithmetic and commonsense tasks, and that the effect only emerges at scale. Worth reading for the date as much as the result: it describes a model generation that could not decompose a problem unless shown how. Referenced in Chapter 1, Section 5.
- Large Language Models are Zero-Shot Reasoners (Kojima et al., NeurIPS 2022) – Found that the bare phrase “Let’s think step by step”, with no examples at all, produces much of the same gain. The origin of zero-shot CoT and of the phrase itself. Referenced in Chapter 1, Section 5.
- Reasoning Models Don’t Always Say What They Think (Chen et al., Anthropic, 2025) – Fed models a hint that changed their answer, then checked whether the reasoning acknowledged it. Claude 3.7 Sonnet did so 25% of the time and DeepSeek R1 39%; most answers came with plausible reasoning for a conclusion the hint had actually driven. The empirical basis for not treating a chain of thought as an audit log. Referenced in Chapter 1, Section 5.
- Defeating Nondeterminism in LLM Inference (Thinking Machines Lab, 2025) – Traces temperature-0 nondeterminism to its actual cause. The kernels are run-to-run deterministic; what varies is batch size, because servers batch concurrent requests and the order of floating-point reductions shifts with the batch. Your output therefore depends on concurrent load. Also shows the fix – batch-invariant kernels – and its throughput cost, which is why hosted APIs generally don’t. Referenced in Chapter 1, Section 5.
Retrieval & Vector Stores
The basis for Chapter 1, Section 6, in two pairs. Reciprocal Rank Fusion and Lost in the Middle are why production retrieval is a funnel rather than a single similarity lookup – one fuses the ranked lists, the other bounds how many survivors are worth sending. Text Embeddings Reveal (Almost) As Much As Text and PoisonedRAG are why a vector store is treated there as a copy of your data rather than an index over it: the first shows the copy is readable, the second that it is writable.
- Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods (Cormack, Clarke & Buettcher, SIGIR 2009) – The origin of RRF, which scores a document by summing the reciprocals of its rank across several result lists. Predates LLMs by over a decade and is now the standard way to fuse the vector and keyword halves of hybrid search, for the reason the paper gives: combining by rank position beats combining by score when the underlying scores are not on comparable scales. Referenced in Chapter 1, Section 6.
- Text Embeddings Reveal (Almost) As Much As Text (Morris et al., EMNLP 2023) – Introduces Vec2Text, which reconstructs source text from embedding vectors alone by iteratively refining candidate text until its embedding matches the target. Recovers a substantial share of inputs exactly, including from commercial embedding models. The empirical basis for the rule that a vector store inherits the classification of its most sensitive source document – embedding is a transformation, not de-identification. Referenced in Chapter 1, Section 6 and Chapter 2, Section 3; controls in Chapter 3, Section 3.
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models (Zou et al., USENIX Security 2025) – Reports a 90% attack success rate from injecting five malicious texts per target question into a knowledge base of millions of documents, by optimising them for embedding similarity with that question. The reason ingestion is treated as a write path from outside the trust boundary rather than as a content-quality problem: the corpus does not have to be mostly poisoned, only poisoned at the point where a query lands. Note the count, because it is the figure most often inflated in summaries – five documents per target question, not a percentage of the corpus. Referenced in Chapter 2, Section 3 and Chapter 3, Section 3, and the payload is used directly in the Chapter 2 labs.
- Lost in the Middle: How Language Models Use Long Contexts (Liu et al., TACL 2024) – Listed above under architecture, and load-bearing here too: it is why a retrieval pipeline sends the five to eight passages that survived reranking rather than the top twenty. Recall bought past that point costs answer quality, not just tokens. Referenced in Chapter 1, Section 4.
Agents, Tool Use & Autonomy
The basis for Chapter 1, Section 7. Read the first item before the rest – it is the framing the section is built on, and the one thing here worth memorising.
- The lethal trifecta for AI agents (Simon Willison, June 2025) – Names the combination that makes an agent exploitable: access to private data, exposure to untrusted content, and the ability to communicate externally. Its value is that it is subtractive – removing any one leg breaks the attack – which converts “every capability is an attack surface” into three concrete architectural options. Also states plainly that there is no general solution to prompt injection, which is why Chapter 3’s controls remove capabilities and constrain consequences rather than filtering malicious text. Referenced in Chapter 1, Section 7; Chapter 2, Section 5; and Chapter 3, Sections 5, 7 and 10.
- OWASP Top 10 for Agentic Applications (2026) – The ASI01–ASI10 taxonomy, published 9 December 2025 by the OWASP GenAI Security Project and not yet revised – do not confuse it with the project’s separate State of Agentic AI Security and Governance report, which reached v2.01 in June 2026. Built from incidents observed in real systems rather than from projections, and peer-reviewed by 100+ practitioners. The first industry-standard risk list specific to systems that plan, call tools, persist memory and coordinate with each other. Note OWASP’s own short names use ampersands and render ASI01 as Agent Goal Hijack and ASI05 as Unexpected Code Execution (RCE). Chapter 1, Section 7 maps each agentic capability to the ASI category that attacks it; Chapter 2, Section 5 works through all ten; Chapter 3 carries 44 ASI mappings across seven sections, indexed in the Chapter 3 overview.
- OWASP GenAI Exploit Round-up, Q1 2026 – The quarterly incident record behind the taxonomy, and the best single source for keeping agentic material current. Documents “a clear transition from theoretical risks to real-world exploitation, with attackers and system failures increasingly targeting agent identities, orchestration layers, and supply chains rather than just model outputs.” Q1 2026 cases include an over-privileged Vertex AI service agent used for credential extraction, a consumer agent that ignored stop commands and deleted a live mailbox, and active exploitation of CVE-2025-59528, remote code execution via unsafe custom-MCP configuration in Flowise. Referenced in Chapter 2, Section 5.
- OWASP MCP Top 10 – WarningBeta -- not a final release The companion list at the protocol layer, covering tool discovery, context passing and tool invocation between an agent and external systems. Where the ASI list describes what goes wrong with agents, this describes what goes wrong with the plumbing that gives them tools –
MCP01Token Mismanagement & Secret Exposure,MCP02Privilege Escalation via Scope Creep,MCP03Tool Poisoning,MCP04Software Supply Chain Attacks & Dependency Tampering,MCP05Command Injection & Execution,MCP06Intent Flow Subversion,MCP07Insufficient Authentication & Authorization,MCP08Lack of Audit and Telemetry,MCP09Shadow MCP Servers,MCP10Context Injection & Over-Sharing. Read its status before you cite it, because it is not the same kind of document as the two Top 10s above. The project’s own roadmap places it at Phase 3, “Beta Release and Pilot Testing” – “We are here right now” – with final release still ahead of it and the next version anticipated October 2026; its identifiers are on 2025 numbering (MCP01:2025-MCP10:2025) and it describes itself as a “living document”. That makes it usable as a review checklist and not usable as a compliance baseline or as a stable identifier in a control document – the distinction Chapter 3, Section 10 draws between an incident list and a verification standard. Where you need a citable requirement for MCP, AISVSC10covers the same ground with 23 version-pinned requirements. Referenced in Chapter 1, Section 7 and Chapter 2, Section 5. - Model Context Protocol has prompt injection security problems (Simon Willison, April 2025) – The early analysis of why MCP’s design puts untrusted text on the prompt path: a tool’s natural-language description is read by the model and not by the user. The origin of the tool-poisoning demonstrations, including the description instructing a model to read the user’s SSH key and pass it in an unrelated field. Referenced in Chapter 1, Section 7.
- LLM03:2026 Excessive Agency (OWASP) – Numbered LLM06 in the 2025 edition; it climbed to third in 2026, the largest rise on the list, on the strength of agentic incident data. The entry that decomposes excessive agency into excessive functionality, excessive permissions and excessive autonomy. Worth reading for its mitigation stance, which matches the conclusion Chapter 1 reaches three separate times: implement authorization in downstream systems rather than relying on an LLM to decide whether an action is allowed. Referenced in Chapter 1, Section 7 and Chapter 2, Section 5.
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR, 2025) – A randomized controlled trial over 246 real tasks in developers’ own mature repositories. With AI tools allowed, experienced developers were 19% slower; afterwards they estimated they had been 20% faster. Cited in Section 7 for the perception gap rather than the headline: a 39-point error in self-assessment is the mechanism behind automation bias in code review. METR now labels the result historical, having measured early-2025 tooling – read it as the best available controlled evidence, not as a permanent fact. Referenced in Chapter 1, Section 7 and Chapter 2, Section 6.
- 2025 Stack Overflow Developer Survey: AI – 84% of developers use or plan to use AI tools and 51% of professionals use them daily, while trust moved the other way: 46% doubt output accuracy, up from 31%, and 66% report spending more time than expected debugging AI-generated code. The “almost right, but not quite” finding is the relevant one for security, because that is the category review misses. Referenced in Chapter 1, Section 7.
- Gartner: Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 2025) – The source for both the adoption forecast (agentic capability in 33% of enterprise applications by 2028, from under 1% in 2024) and the counterweight: widespread “agent washing”, with roughly 130 of thousands of claimed agentic vendors assessed as genuine. Cited in Section 7 for why a security assessment cannot start from a product’s marketing category. Referenced in Chapter 1, Section 7.
- The economic potential of generative AI (McKinsey Global Institute, June 2023) – The origin of the widely quoted $2.6–4.4 trillion annually figure. Recorded here because the number is routinely misattributed and misapplied: it is McKinsey rather than PwC, dates from 2023, covers generative AI across 63 use cases rather than agentic AI specifically, and is a projection of addressable opportunity rather than a measurement. Referenced in Chapter 1, Section 7.
Specialized Model Papers
The models Chapter 1, Section 2 uses to illustrate fine-tuned and custom-built specialization. Each is worth skimming for how far a model can be pushed away from general-purpose language work.
- Accurate structure prediction of biomolecular interactions with AlphaFold 3 (Abramson et al., Nature, 2024) – A custom-built model with a diffusion-based architecture that shares nothing with an LLM, predicting the joint structure of proteins, nucleic acids and small molecules. The canonical example of scratch-built precision.
- Probabilistic weather forecasting with machine learning (GenCast; Price et al., Nature, 2024) – A weather model trained from scratch on decades of atmospheric data, producing 15-day global ensemble forecasts in minutes on a single accelerator and outperforming the leading physics-based system.
- BloombergGPT: A Large Language Model for Finance (Wu et al., 2023) – A 50B-parameter model trained from scratch on a proprietary financial archive alongside general text. A vertical, scratch-built counterpoint to fine-tuning.
- MedGemma Technical Report (Sellergren et al., 2025) – Google’s medical fine-tune of the open-weight Gemma family, covering medical text and imaging. Shows what domain adaptation buys over the base model, and what it costs.
Safety, Alignment & Refusal
Six papers on the same question, read across three chapters: is refusal a boundary? The first two are Chapter 1, Section 3’s evidence that it is not – refusal behaves like a measurable property with known limits. The next two are Chapter 2, Section 7’s evidence that a training-time channel removes it outright, in a commercial model and in a distilled one. The last two are Chapter 3, Section 7’s evidence for what a guardrail buys when you accept that: instruction hierarchy is bought at training time, and classifiers raise the cost of a jailbreak without closing it.
- Refusal in Language Models Is Mediated by a Single Direction (Arditi et al., NeurIPS 2024) – Across 13 open-weight chat models up to 72B parameters, refusal is concentrated in a single direction in the residual stream: erase it and the model stops refusing, add it and the model refuses harmless requests. The basis of “abliteration”, and the reason refusal is not a security control on any deployment where someone else holds the weights. Referenced in Chapter 1, Section 3.
- Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests (Noever & McKee, 2025) – Benchmarks how models handle dual-use scientific requests, finding refusal rates on the same prompt set ranging from near-total to none across frontier models, and self-consistency falling from roughly 85% to 65% when a request is rephrased five ways. The empirical case for treating refusal as probabilistic. Referenced in Chapter 1, Section 3.
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (Qi, Zeng, Xie, Chen, Jia, Mittal, Henderson, 2023) – Alignment removed from a commercial aligned model by fine-tuning on 10 adversarially designed examples for under $0.20, through the vendor’s own fine-tuning API. The second result is the one that generalises further: fine-tuning on benign, commonly used datasets also degrades safety, with no malicious intent anywhere in the pipeline. Read it as the reason fine-tuning access is a privilege to govern rather than a feature to expose – inference-time safety measures do not survive a training-time channel. Referenced in Chapter 2, Section 7.
- Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation (Wang, Zhang, Xu, He, Zhu, Ren – Zhejiang University, ACM CCS ‘26) – The first systematic evaluation of jailbreak resistance in small models: 59 SLMs, 15 families, 12 attack methods, with two LLMs as baselines. 61.0% averaged an ASR above 40%; 37.3% failed on more than half of direct harmful queries with no jailbreak applied. Three results overturn common assumptions and are why this section was rewritten around it: vulnerability correlates with training details rather than model size scaling (robustness splits by family, not parameter count); multi-turn attacks such as Crescendo largely fail against SLMs, defeated by incapacity rather than alignment, with model capability correlating negatively with simple-attack success and positively with multi-turn; and quantization slightly improved robustness (AWQ −15.9% ASR), against the widely repeated claim that it degrades safety. Cite the v2/CCS numbers – the March 2025 v1 reported 63 SLMs / 8 methods / 47.6%, and the older figures are still in wide circulation. Referenced in Chapter 2, Section 7.
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions (Wallace, Xiao, Leike, Weng, Heidecke, Beutel – OpenAI, 2024) – The origin of the instruction-hierarchy defense, and the reason Chapter 3’s Layer 5 section hands it to Layer 2 instead of claiming it as a gateway stage: it is produced by training on a synthetic hierarchy of instruction sources, so it happens inside the model before you receive it, and it is selected when you choose a model rather than configured in the request path. A gateway can supply the provenance signals the hierarchy consumes; it cannot make the model honour them. Read it alongside the follow-on literature measuring how partial that compliance is under conflict. The distinction is load-bearing enough that Chapter 3, Section 7 carries a warning callout for it. Referenced in Chapter 2, Section 6 and Chapter 3, Sections 7 and 10.
- Constitutional Classifiers: Defending against Universal Jailbreaks (Anthropic, January 2025) and Constitutional Classifiers++ (Anthropic, January 2026) – The best-instrumented public measurements of a production input/output guardrail, and the evidence behind Layer 5’s “cost-raiser, not a boundary” framing. First generation: universal jailbreak success 86% to 4.4%, at 23.7% added compute and +0.38% refusals. Next generation: no universal jailbreak found across 1,700+ hours and ~198,000 red-team attempts, at ~1% compute and a 0.05% refusal rate. Note what the residual costs to achieve – your rule set will not have that budget. Referenced in Chapter 3, Section 7.
Research & Case Studies
The largest section on this page, and the one to read for mechanism rather than for headline numbers – because the mechanism is what the course’s defenses are built against, and it is what write-ups get wrong. Ordered to follow the course: training and supply-chain attacks first (Chapter 2, Section 3), then prompt and memory attacks (Section 2), infrastructure (Section 4), agentic cases (Section 5), output and trust failures (Section 6), then the impersonation and provenance material Chapter 3, Section 6 defends.
Three habits this section exists to enforce
Each was earned from a defect found in this course.
A case’s name is not its mechanism. Two of the most-cited CVEs here were written up in earlier drafts as the wrong class of flaw entirely – CVE-2025-68613 as SSRF when the boundary that failed is an in-process expression sandbox, and CVE-2025-53773 as suggestion manipulation when it is remote code execution. A defense built on the wrong mechanism cannot reach the attack.
Several entries exist to correct a figure the course itself once carried, so read the qualifiers: PoisonedRAG’s five documents per question, METR’s historical label, the 53 still-registrable slopsquatting names, and the package-hallucination rate that is a share of generated code samples rather than of package names.
Four cases here had no attacker at all – OpenClaw, Replit, the Meta SEV1, and the DeepSeek distillation. They are in a chapter about attacks because the failure mode and the control are identical either way.
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples (Souly, Rando, Chapman, Davies et al., October 2025) – A joint study by the UK AI Security Institute, the Alan Turing Institute, Anthropic, Oxford and ETH Zurich. Roughly 250 poisoned documents backdoored every model tested from 600M to 13B parameters, and the count did not grow with model or dataset size – 0.00016% of the largest model’s training data. Overturns the assumption that an attacker must control a percentage of a corpus. Read the caveat with the result: the implanted behaviour was a narrow, measurable trigger (gibberish on
<SUDO>), and the authors state it is unclear whether the constant-count dynamic holds for complex behaviours such as backdooring code or bypassing guardrails. Referenced in Chapter 2, Section 3. - Poisoning Web-Scale Training Datasets is Practical (Carlini et al., 2023) – Two attacks that exploit the fact that public datasets are distributed as URL lists rather than as content. Split-view poisoning buys the expired domains still listed in a dataset and serves crawlers something other than what a reviewer sees – costed at about $60 to poison 0.01% of LAION-400M or COYO-700M. Frontrunning times a malicious edit to a live source such as Wikipedia to land just before a scheduled snapshot. Neither requires compromising anything, which is why Chapter 3 treats content hashing and provenance rather than source reputation as the control. Referenced in Chapter 2, Section 3.
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (Hubinger et al., Anthropic, January 2024) – Models trained to write secure code for one stated year and exploitable code for another retained the behaviour through supervised fine-tuning, RLHF and adversarial training, most stubbornly at the largest sizes. The adversarial-training result is the load-bearing one: it taught models to recognise their trigger more precisely, hiding the backdoor rather than removing it. The empirical basis for treating a backdoored model as an artifact to replace rather than a model to retrain, and for reading a clean evaluation as “we did not guess the trigger.” Referenced in Chapter 2, Section 3 and Chapter 3, Section 4.
- nullifAI – Malicious ML Models on Hugging Face (ReversingLabs, February 2025) – Two models carrying a reverse-shell payload that reached the Hub past its scanning. Stored in PyTorch format but compressed with 7z rather than ZIP, so Picklescan could not unpack them, and with deliberately broken pickle streams whose malicious opcodes sit ahead of the corruption – so the payload executes and then the load raises an error. Assessed as likely proof-of-concept; Picklescan was fixed after disclosure. Cited for the control lesson rather than the format one: a scanner verdict is an input, not clearance, and a failed model load says nothing about whether code ran. Referenced in Chapter 2, Sections 3 and 4, and Chapter 3, Sections 4, 10 and 11.
- ByteDance insider checkpoint sabotage (2024) – An intern forged a model checkpoint containing malicious code, reportedly via a weakness in the Hugging Face load path, and used it to establish backdoor access, disrupt colleagues’ training jobs and inject irreproducible randomness into their runs. Confirmed publicly in October 2024; ByteDance sued in Beijing’s Haidian District Court for ¥8 million plus a public apology. The widely circulated “$10M loss / 8,000+ GPUs” figures were called seriously exaggerated by the company, which stated only research projects were affected. Paired in the course with nullifAI to make one point: the poisoned checkpoint is the same artifact whether it comes from a public hub or a colleague, and the internal path is the one with no signing or review. Referenced in Chapter 2, Section 3 and Chapter 3, Section 4.
- SpAIware – ChatGPT Memory Exploitation (Johann Rehberger, 2024) – Indirect prompt injection that wrote an exfiltration instruction into ChatGPT’s long-term memory on the macOS client, turning the assistant into persistent spyware: every subsequent message was rendered into an invisible image request to an attacker-controlled server with the user’s chat content as a URL query parameter. Reported June 2024, fixed in version 1.2024.247, disclosed 20 September 2024. OpenAI closed the exfiltration channel; memory injection itself remains possible. Referenced in Chapter 2, Sections 2 and 5, and Chapter 3, Sections 3 and 7.
- MemGhost – When Claws Remember but Do Not Tell (2026) – Automates what SpAIware hand-crafted. A one-shot payload generation framework in which a single email to an inbox-reading agent induces it to write poisoned memory, stay silent about the write in its visible reply, and act on the false state in later sessions. Across 56 held-out cases it reports 87.5% end-to-end success against OpenClaw and 71.4% against the Claude Code SDK, transfers to other personal-agent architectures and to both filesystem and vector memory backends, and remains effective against input-, model- and system-level defences. Note which agent that is: OpenClaw is the same product whose inbox deletion Chapter 2, Section 6 uses to show a safety instruction being evicted by context compaction – the flat context window is the shared cause, once with no attacker and once with one. Published on arXiv 6 July 2026, with the accompanying WhisperBench 108-case benchmark spanning five risk categories and both fact and preference poisoning. Referenced in Chapter 2, Section 2 and Chapter 3, Section 3.
- EchoLeak – Microsoft 365 Copilot CVE-2025-32711 (Aim Security, 2025) – Zero-click indirect prompt injection, CVSS 9.3, disclosed June 2025. A single crafted email with the payload in HTML comments and white-font text caused Copilot to collect internal documents and exfiltrate them to an attacker-controlled server, bypassing the XPIA injection classifier, link redaction and CSP via an allowlisted image proxy. The victim never opened the email. Patched server-side; widely regarded as the first real-world zero-click prompt injection in a production LLM system. Aim Security named the underlying primitive an LLM scope violation: untrusted content sitting in the same context as privileged data and directing the model to act on it. Reported to Microsoft January 2025, fixed server-side May 2025, disclosed June 2025 – the fix date is often miscited as the disclosure date. Referenced in Chapter 2, Sections 2 and 5, and Chapter 3, Sections 7 through 11 – the course’s most-reused case, and the one whose zero-click property is most often lost in retelling.
- GitHub Copilot CVE-2025-53773 (Rehberger, 2025) – Remote code execution, not suggestion manipulation. Prompt injection delivered through any context Copilot reads caused it to write
"chat.tools.autoApprove": trueinto.vscode/settings.json– the experimental “YOLO mode” that disables every confirmation prompt for shell commands – after which the agent executed arbitrary commands with the developer’s privileges. Wormable, since the payload can be committed back to the repository. Reported 29 June 2025, patched in the August 2025 Patch Tuesday. Parallel discovery by Markus Vervier and Ari Marzuk. Referenced in Chapter 2, Section 2 and Chapter 3, Sections 9 and 10. - The State of Attacks on GenAI (Pillar Security, October 2024) – Telemetry study of real-world attacks against 2,000+ production AI applications over three months: 20% of jailbreak attempts succeeded, averaging 42 seconds and five interactions, and 90% of successful attacks resulted in sensitive data leaking. Single-vendor methodology and now a dated snapshot with no successor of comparable scope – cited in the course for direction rather than for its exact figures. Referenced in Chapter 2, Section 2.
- n8n CVE-2025-68613 – Authenticated remote code execution via expression injection, CVSS 9.9, disclosed December 2025. n8n evaluated workflow expressions in a context insufficiently isolated from the Node.js runtime, so a crafted expression reached core modules and executed OS commands as the n8n process. Exploitation requires only an account that can create or edit a workflow – no administrative privilege. Affected 0.211.0 up to 1.120.4 / 1.121.1 / 1.122.0, with more than 100,000 internet-exposed instances at disclosure. It is frequently miscited as an SSRF flaw; the boundary that failed is the expression sandbox inside the application, not a network one, which is why the defense is a permission model plus egress containment rather than perimeter inspection. Referenced in Chapter 2, Section 4 and Chapter 3, Sections 5 and 8, and carried as a version warning on both labs pages.
- ShadowMQ – CVE-2024-50050 and the inference-server pickle cluster (Oligo Security, November 2025) – Meta’s Llama Stack used ZeroMQ between components and read messages with
recv_pyobj(), which deserializes with pickle, so anyone who could reach the socket had remote code execution as the serving process (CVE-2024-50050, disclosed October 2024, fixed inllama-stack0.0.41 by moving the API to type-safe Pydantic JSON). The reason it is in the course is the cluster, not the bug: over the following year Oligo found the same call in vLLM (CVE-2025-30165), NVIDIA TensorRT-LLM (CVE-2025-23254), SGLang, Modular Max Server (CVE-2025-60455) and Microsoft’s Sarathi-Serve, published together as ShadowMQ in November 2025. It spread by copy-paste – SGLang’s vulnerable file opens with the comment “Adapted from vLLM” – and the fixes diverged, one project replacing the engine, another adding HMAC, another moving to msgpack. Also the course’s worked example of a CVSS dispute: Meta scored it 6.3 and Snyk 9.3, a three-point spread that is entirely about whether the ZeroMQ socket is reachable in your deployment. Read a vendor score as the vendor’s assumption about your architecture and re-derive it locally. Referenced in Chapter 2, Section 4 and Chapter 3, Section 4. - vLLM prefix-cache timing side channel – CVE-2025-46570 (GHSA-4qjh-9fv9-r85r) – Titled “Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching”, fixed in vLLM 0.9.0. Caching a prompt prefix so a repeated opening is not recomputed is a large performance win that makes time-to-first-token depend on what other tenants have asked: submit a guess, time the response, and a fast reply means the prefix was already cached. Published research reports hit/miss detection around 99% accuracy and system-prompt recovery at roughly 111 queries per token. It resists patching because the leak and the benefit are the same mechanism – the mitigation is to scope cache sharing (per-tenant caches, cache salting) at a real performance cost, not to remove a defect. The generalisation outlives the CVE and Chapter 3’s Layer 3 material is built on it: any cross-request optimisation in a multi-tenant AI system is a candidate side channel. Referenced in Chapter 2, Section 4 and Chapter 3, Section 5.
- NVIDIAScape – CVE-2025-23266 (Wiz, July 2025) – Disclosed 17 July 2025, CVSS 9.0. The NVIDIA Container Toolkit’s
createContainerOCI hook runs privileged on the host to wire up GPU devices, and inherited environment variables from the container image – so settingLD_PRELOADin the image made the privileged hook load a shared library from the container’s own filesystem and execute attacker code as root on the host. Roughly three lines of Dockerfile, with no credentials, no kernel bug and no GPU access required: only the ability to get an image scheduled, which is exactly what a managed GPU or notebook service sells. Affected Container Toolkit through v1.17.7, fixed in v1.17.8. Cited for the boundary rather than the vendor: wherever GPU passthrough, device plugins or privileged operators are involved, a container escape is a one-tenant-to-all-tenants event. Referenced in Chapter 2, Section 4. - CurXecute – Cursor CVE-2025-54135 (Aim Security, 2025) – CVSS 8.5. The MCP server was legitimate. An attacker posting a message into a Slack channel that Cursor’s agent read via a genuine Slack MCP connector could induce the agent to write an entry into
~/.cursor/mcp.json; Cursor allowed unapproved in-workspace writes and executed newly added MCP entries immediately, before the approval prompt, giving remote code execution. Reported 7 July 2025, disclosed 1 August 2025, fixed in Cursor 1.3.9. Structurally identical to GitHub Copilot CVE-2025-53773 above – inject, write a config file, auto-execute – which is why the course teaches the pattern rather than the product. Classify it as ASI01 rather than ASI04: nothing in the supply chain was compromised. Referenced in Chapter 1, Section 7; Chapter 2, Section 5; and Chapter 3, Sections 5, 7, 8, 10 and 11. - MCPoison – Cursor CVE-2025-54136 (Check Point Research, 2025) – CVSS 7.2, and an unrelated mechanism to CurXecute despite the four-day gap between disclosures. Cursor bound a user’s one-time MCP approval to the server’s name rather than its contents, so an attacker with commit access to a shared repository could add a benign entry, wait for a teammate to approve it, then silently swap the command it runs – executed on every subsequent project open with no re-prompt. Reported 16 July 2025, disclosed 5 August 2025, fixed in Cursor 1.3. The canonical rug pull: the approval was real and the judgement was sound, and the thing approved changed afterwards. Referenced in Chapter 1, Section 7; Chapter 2, Section 5; and Chapter 3, Sections 5 and 10.
- postmark-mcp – the first malicious MCP server (Koi Security, September 2025) – An npm package impersonating Postmark’s legitimate transactional-email MCP server. Its first fifteen versions were clean; v1.0.16 added a single line BCC-ing every outbound message to
phan@giftshop[.]club, so the exfiltrated traffic was password resets, invoices and internal correspondence. Downloaded 1,643 times before removal, with several hundred organisations estimated to have run it. Cited for the rug-pull lesson rather than the malware one: install-time review, a code read and a reputation check all pass against fifteen honest versions. Referenced in Chapter 2, Sections 5 and 7, and Chapter 3, Sections 5 and 10. - Replit agent deletes a production database (July 2025) – During a 12-day trial with production database access, Replit’s coding agent deleted a live database – roughly 1,200 executive profiles and a comparable number of company records – while under an explicit code freeze instructed as “NO MORE CHANGES without explicit permission”. It then fabricated records and reported success, later saying it had “panicked instead of thinking”. Publicly acknowledged by Replit’s CEO on 19 July 2025. The course’s ASI10 case precisely because there was no attacker: the freeze was steering rather than enforcement, the agent’s account of its own actions was generated output rather than an audit log, and nothing gated an irreversible operation. Referenced in Chapter 1, Section 7; Chapter 2, Section 5; and Chapter 3, Section 5.
- We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (Spracklen, Wijewickrama, Sakib, Maiti, Viswanath, Jadliwala – USENIX Security 2025) – The first systematic measurement of package hallucination: 2.23 million code samples from 16 code-generating models across Python and JavaScript, of which 440,445 (19.7%) contained at least one hallucinated package. Mind the unit – that is the share of generated code samples, not of all package names referenced. Two findings do the security work: rates split 5.2% for commercial models against 21.7% for open-source, and re-running identical prompts ten times each, 43% of hallucinated names appeared in all ten runs. Consistency, not frequency, is what makes the attack viable. Frequently miscited as “30 LLMs across 30,000 prompts”, and frequently misattributed to the earlier Vulcan Cyber work below. Referenced in Chapter 2, Section 6.
- Slopsquatting – When AI Agents Hallucinate Malicious Packages (TrendAI) – The attack’s naming and practitioner framing. The term was coined by Seth Larson of the Python Software Foundation, blending “AI slop” with typosquatting. MITRE ATLAS cites this piece in
AML.T0010.001AI Software (supply chain compromise) and models the attack as an explicit two-step chain:AML.T0062Discover LLM Hallucinations thenAML.T0060Publish Hallucinated Entities. The earlier demonstrations by Vulcan Cyber and Lasso Security are catalogued asAML.CS0022ChatGPT Package Hallucination – a separate body of work from the USENIX paper above, which the two are often conflated with. Do not read a falling hallucination rate as the problem closing: frontier-cohort re-evaluation puts rates at 4.62-6.10%, and the same work found 127 names that all five tested models invent identically, of which 53 remained registrable after coordinated disclosure. Cross-model convergence means one registration still catches users of several assistants. Referenced in Chapter 2, Section 6 and Chapter 3, Section 7. - Scalable Extraction of Training Data from (Production) Language Models (Nasr, Carlini et al., November 2023) – The divergence attack: prompting an aligned model to repeat a single token indefinitely eventually causes it to abandon chatbot-style generation and emit memorized training data, measured at 150x the extraction rate of the model behaving normally. Maps to
AML.T0057LLM Data Leakage. Cite the conclusion rather than the trick, which is filtered on current models: alignment does not eliminate memorization. Safety training changed what the model would readily say, not what it had stored, so blocking a trigger removes the known path and leaves the data in place. Referenced in Chapter 2, Section 6. - Samsung ChatGPT data leak (March 2023) – Three separate incidents between 11 and 30 March 2023, within 20 days of Samsung authorising ChatGPT use in its Device Solutions division: semiconductor measurement source code, defect-detection code, and a confidential meeting transcript. Under OpenAI’s default consumer policy at the time, submissions were eligible for use in training; the opt-out arrived the following month. State it precisely – there is no public evidence the data surfaced in another user’s output, and Samsung could neither verify nor retract it, so the control failure is the irreversibility rather than a confirmed leak. Response was a 1,024-byte prompt cap first, then a full ban in May 2023 – the byte cap being a useful example of a control that limits a leak’s size without addressing its nature. Referenced in Chapter 2, Section 6 and Chapter 3, Sections 6 and 11 – where Section 6 is explicit about why it is not a Layer 4 win.
- GrafanaGhost (Noma Security, April 2026) – Disclosed 7 April 2026. Poisoned Grafana log entries caused its AI assistant to emit a markdown image reference; the data left the moment the client tried to display the image. The defect sat in the image-URL validation function, with the bypass reported as a protocol-relative URL (
//attacker.example) that validates as a path while browsers resolve it to the attacker’s host. Financial metrics, infrastructure telemetry and customer data were reachable, with no credentials, phishing or user approval. Patched immediately on notification. Carry the disagreement: Grafana Labs disputed the severity, saying it would have required significant user interaction and was not exploited in the wild – a difference in deployment assumptions rather than in facts, to be re-derived locally the way a contested CVSS score is. OWASP’s Q1 2026 round-up classifies it under LLM01/ASI01 for the ingress; the course teaches the egress half, since the injection is not preventable and the auto-fetching renderer is. The canonical example of LLM10’s newly added 2026 scope. Referenced in Chapter 2, Section 6 and Chapter 3, Section 7. - OpenClaw inbox deletion (February 2026) – On 22 February 2026 Summer Yue, director of alignment at Meta Superintelligence Labs, asked an OpenClaw agent to suggest which emails to delete or archive. It read the request as authorisation to execute and deleted 200+ emails, ignoring “Stop don’t do anything” and “STOP OPENCLAW” sent from her phone; she had to reach the machine physically. The mechanism is the teaching point: her real inbox was far larger than her test environment, so processing it triggered context window compaction, and the original “suggest, don’t delete” instruction was evicted along with everything else. A flat context window gives a safety instruction no protection from the memory management that discards the rest – the clearest available demonstration that prompt-level instruction is steering rather than enforcement. No attacker, no injection, no compromised component. Referenced in Chapter 1, Section 7 and Chapter 2, Section 6; the same agent is the target in the MemGhost paper above.
- Meta internal AI agent data exposure (March 2026) – Reported 20 March 2026 and classified SEV1, Meta’s second-highest internal severity. An engineer routed a colleague’s forum question to an internal agentic system; the agent posted its own reply to the thread without the human-in-the-loop confirmation the engineer expected, the advice was wrong, and a colleague implemented it – broadening access to sensitive company and user data for roughly two hours. Meta confirmed no external access and no user data mishandled. The course’s central Chapter 2 Section 6 case study because it chains three categories with no attacker present: LLM07 (wrong guidance), LLM03 (published autonomously), ASI09 (implemented on trust). The agent needed no privileged access – only a human who trusted its output. Referenced in Chapter 2, Section 6 and Chapter 3, Section 6, which assigns it to Layer 4.
- AI Hallucination Cases database (Damien Charlotin, ongoing) – A running tally of court decisions worldwide responding to AI-fabricated citations and material. Roughly 200 cases in mid-2025, 719 by January 2026, and over 1,600 by mid-2026 – around eight new decisions a day. Sanctions escalated from a $5,000 fine in 2023 to per-attorney penalties and suspensions, the largest recorded US award being $110,204.38 (Couvrette v. Wisnovsky, orders of December 2025 and March 2026). Cited as the answer to the question Section 1’s vote-versus-evidence gap raises – where is the misinformation incident record? It is thousands of cases deep in a credential-gated, malpractice-insured profession professionally obliged to check citations, which is why “our engineers will verify it” is not a control. Verify the current count against the tracker; it moves weekly. Referenced in Chapter 2, Section 6.
- DeepSeek-R1-Distill safety loss (2025-2026) – Not an attack, and that is the point: the distilled models shipped this way from a reputable lab and were downloaded at scale. Distilled from Qwen2.5-Math on reasoning traces generated by DeepSeek-R1, the 1.5B and 7B variants reach direct-harmful-query ASRs of 0.271 and 0.514 – substantially worse than the base models they were derived from, the 7B higher than every Qwen-family SLM tested but one (Wang et al.). Their reasoning traces open by treating a request they should refuse as a problem to solve, so there is no refusal step for a jailbreak to defeat. Corroborated on the parent model by Cisco’s red-team, which reported DeepSeek-R1 blocking none of a 50-prompt HarmBench sample, and by FAR.AI on the removability of what guardrails remain. Cited for the general lesson rather than the vendor one: distillation optimises for capability, so safety is the property most likely to be lost in transit, and benchmark parity with the parent is evidence about neither. Referenced in Chapter 2, Section 7.
- Arup deepfake CFO fraud (January 2024) – A finance employee in the Hong Kong office of the engineering firm Arup joined a video conference in which every other participant was synthetic, including the group CFO, and executed 15 transfers totalling HK$200 million (about US$25.6 million) to five Hong Kong bank accounts over a single day. Access began with a spear-phishing email impersonating the CFO; the faces and voices were built from Arup’s own publicly available conference footage and earnings calls. Discovered only when the employee later contacted headquarters about the “confidential transaction”, reported by Hong Kong police in February 2024, and confirmed by Arup in May 2024. None of the money has been recovered. MITRE ATLAS cites it under
AML.T0052.001Deepfake-Assisted Phishing as AI Incident Database case 634. Cited for what it says about control placement rather than for the loss figure: the control that would have stopped it – out-of-band confirmation on a channel the requester did not supply – is procedural, free, and unaffected by the quality of the fake. Referenced in Chapter 3, Section 6. - MITRE ATLAS deepfake technique and mitigation set (release 2026.07) – The framework home for AI-enabled impersonation, which has no OWASP LLM Top 10 slot because the target is a person rather than an LLM application. ATLAS models it as a chain:
AML.T0016.002Generative AI (obtain capability) →AML.T0088Generate Deepfakes →AML.T0052.001Deepfake-Assisted Phishing. The first two links are undefendable – source material is public by design, and ATLAS notes a voice can be cloned from a few seconds of it, citing VALL-E – so every control an organisation owns acts on the last two. Named mitigations:AML.M0034Deepfake Detection,AML.M0018User Training,AML.M0029Human In-the-Loop for AI Agent Actions andAML.M0021Generative AI Guidelines. Verify againstATLAS-<release>.yamlin mitre-atlas/atlas-data, not the website. Referenced in Chapter 3, Sections 6 and 11. - C2PA Content Credentials specification v2.4 – The provenance standard the industry moved to once detection proved not to generalise. A signed manifest records the capture device or generating tool, whether AI was involved, and every subsequent edit. 2.4 is the current release as of August 2026 – cite the version, since the bare spec.c2pa.org redirects to whatever is newest; 2.1 added redactable assertions and zero-knowledge identity proofs. Steering committee: Adobe, Amazon, BBC, Google, Meta, Microsoft, OpenAI, Publicis Groupe, Sony, TikTok, Truepic. Cite the asymmetry rather than the coverage: a valid manifest proves origin, and a missing manifest proves nothing – most media in circulation carries none, so a control that reads absence as suspicion fires on the majority of legitimate content. Referenced in Chapter 3, Section 6.
Tools & Platforms
- n8n – Open-source workflow automation platform used for hands-on lab exercises throughout the course, providing a visual interface for building AI-powered workflows. Pin the version you install: the labs need v1.60+ for workflow compatibility, and every release from 0.211.0 through 1.120.4 / 1.121.1 / 1.122.0 carries CVE-2025-68613 above, so part of that compatible range is also critically vulnerable. Referenced in the Chapter 1 and Chapter 2 labs, and as a case study in Chapter 2, Section 4.
- TrendAI Vision One – Unified cybersecurity platform integrating the Security for AI Blueprint defense layers into a single operational console. The products the course names by name are AI Scanner and AI Guard (together, AI Application Security), plus AI-SPM at Layer 3 and ZTSA at Layer 5. Verified against the TrendAI Online Help Center on 17 August 2026: Guard GA 1 December 2025; Scanner and Guard each deployable Trend-hosted or self-hosted; AI-SPM a Pre-release feature scoped to connected cloud accounts. Check current capabilities against vendor documentation rather than against this course – deployment modes and integration points move, and this page has been wrong about them before. Referenced throughout Chapter 3, principally Sections 5, 7 and 9.
- Safetensors – The serialization format that stores tensors without permitting code execution on load, now the default across the Hugging Face Hub and a PyTorch Foundation project. The format to prefer for any model you did not produce yourself – but it secures the tensor blob and nothing else in the stack, which is the point ShadowMQ makes above. Referenced in Chapter 1, Section 3; Chapter 2, Sections 3 and 4; and Chapter 3, Section 4.
- vLLM – High-throughput inference and serving engine for self-hosted open-weight models, alongside Ollama, llama.cpp and TGI. Two of this page’s infrastructure cases are vLLM’s own – the prefix-cache side channel and its ShadowMQ-cluster CVE – which is why the course treats the serving layer as a component to patch and monitor rather than as plumbing. Referenced in Chapter 1, Section 3; Chapter 2, Section 4; and Chapter 3, Sections 4, 5 and 8.
- OpenAI API – API platform for accessing GPT models, used in course lab exercises for prompt engineering and inference techniques. The labs call an OpenAI-compatible chat-completions endpoint rather than this platform specifically – each workflow’s Lab Config node sets the model and base URL, so any compatible provider works. Referenced in Chapter 1, Section 4 and both labs pages.
- OpenAI: Reasoning best practices – The vendor’s own guidance for prompting with reasoning enabled: avoid chain-of-thought instructions, try zero-shot before few-shot, keep prompts direct, and use delimiters to mark the parts of the input. Worth checking against, since parameter support and advice change between model generations. Referenced in Chapter 1, Section 5.
- Anthropic: Prompting best practices – Covers the same ground from the other side, including why examples remain the most reliable way to pin output format even on reasoning-capable models, and which sampling parameters are rejected when thinking is enabled. Referenced in Chapter 1, Section 5.