Glossary
A reference of key terms from all three chapters, organized A-Z.
Entries are not just definitions. Where a term carries a security consequence, the entry states it, because knowing what a vector store is does not tell you that it inherits the classification of its most sensitive source document.
The back-links are the other half of an entry. They are labelled by role rather than by chapter, so an entry is also a map of where the concept lives in the course:
- See / Introduced – where the concept is built
- Attacks / Exploitation – where it is attacked, usually Chapter 2
- Controls / Defense – where the Blueprint answers it, usually Chapter 3
Where a term maps to an OWASP category, the identifier is given on the 2026 numbering used throughout this course – LLM01-LLM10 for the OWASP LLM Top 10, ASI01-ASI10 for the OWASP Top 10 for Agentic Applications. Eight of the ten LLM identifiers moved in August 2026, so material written before then numbers them differently.
A
- Adversarial Inputs: Inputs perturbed just far enough to cross a model’s decision boundary while looking unchanged to a human – the classic demonstration being a handful of altered pixels that turn a stop sign into a speed-limit sign for a classifier. The mathematics is identical for any differentiable model, which is the consequence to carry forward: your input filter is a model too, with a decision boundary and therefore adversarial examples of its own. See: Chapter 1, Section 1 · Mechanism: Chapter 2, Section 4 · Consequence for filtering: Chapter 3, Section 7
- Agency Spectrum: The range from a fixed pipeline (the code decides control flow, the model only classifies or rewrites) through tool-augmented and semi-autonomous systems to a fully autonomous agent (the model decides the whole plan). Position is determined by what the system can reach and what gates an action – never by the vendor’s label, since agent washing makes labels unreliable. Risk scales with position because a fixed pipeline’s control flow can be reviewed in advance while an agent’s can only be constrained. See: Chapter 1, Section 7
- Agent (AI Agent): An LLM-driven system that pursues a goal by choosing its own actions: it plans a step, calls a tool, reads the result, and decides what to do next. The distinction from an LLM-powered script is not the model but who chooses the sequence – a script’s steps are written in code, an agent’s do not exist until it runs. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Agent Loop: The cycle at the core of every agentic system: plan, select tool, execute, observe, decide. Two steps cross a trust boundary in opposite directions. Execute is egress – where the agent’s decisions become effects in the world, and where damage lands. Observe is ingress – where external text re-enters the flat context window indistinguishable from instructions, and therefore where compromise begins. Attackers reach the ingress in order to abuse the egress. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Agent Washing: Relabelling a chatbot, assistant or RPA script as an “AI agent” without adding agentic capability. Gartner assessed roughly 130 of the thousands of vendors claiming agentic products as genuine, which is why a security assessment has to establish a system’s actual reach and gating rather than accepting its marketing category. See: Chapter 1, Section 7
- Agentic AI: AI systems that pursue goals and take autonomous action rather than only producing text – the stage after base models and retrieval-augmented systems. The security consequence is a change of kind, not degree: it converts information risk into action risk, because the worst case moves from a misleading paragraph to an irreversible effect. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Agentic Attack Vector: An attack path that exploits an agent’s autonomous decision-making, tool use, persistent memory or multi-agent coordination rather than the model’s text output. Enumerated by the OWASP Top 10 for Agentic Applications as
ASI01-ASI10; what separates the class from the LLM Top 10 is that the model is an actor with consequences rather than a component producing text. Foundations: Chapter 1, Section 7 · See: Chapter 2, Section 5 - Agentic Workflow: A sequence of steps whose progression is decided by agents rather than by a fixed program. Distinguished from an AI-enhanced workflow, where the model fills a slot in a path the code determines: an agentic workflow’s path does not exist until it runs, so it cannot be reviewed in advance. See: Chapter 1, Section 7
- AI Gateway: A control point placed between users, tools and AI services that evaluates prompts, retrieved context, model outputs and agent actions before they are accepted, executed or shown. MITRE ATLAS carries it as mitigation
AML.M0020Generative AI Guardrails. It is only a control if requests cannot go around it, which makes egress policy a prerequisite rather than a complement. See: Chapter 3, Section 7 · Placement: Chapter 3, Section 7 - AI Guard: A runtime protection component that inspects AI interactions in real time, filtering prompts and responses to block injection attempts, data leakage, and policy violations. Consumed as an inline API that a gateway or LLM proxy calls on each request and response – Trend-hosted at a regional endpoint or self-hosted in your own environment, the choice being data residency – which makes coverage a property of your routing and forces a fail-open or fail-closed choice either way. A guardrail is a cost-raiser, not a boundary. GA 1 December 2025. See: Chapter 3, Section 9 · Layer 5: Chapter 3, Section 7
- AI Scanner: A proactive assessment tool that evaluates AI applications – not only models – for vulnerabilities, misconfigurations and compliance gaps, run pre-deployment as a pipeline gate and periodically thereafter. It probes with a library of known techniques, so a finding is strong evidence and a clean run is weak evidence: it means “none of these worked,” never “nothing works.” Like AI Guard, it runs Trend-hosted or self-hosted, the deciding axis being whether the material may leave your environment. See: Chapter 3, Section 9
- Air-Gapped Deployment: Running a model on infrastructure with no network path to the outside world. Only self-hosted and edge deployment can satisfy it; which of the two applies is decided by how much model capability the task needs. Note what it costs elsewhere: an inline inspection service must receive every prompt, so an air-gapped estate cannot buy one and has to meet the requirement with a locally-run guardrail model and controls that need no inspection at all. See: Chapter 1, Section 3 · Consequence: Chapter 3, Section 9
- AI-SPM (AI Security Posture Management): Continuous discovery, configuration assessment and risk scoring across AI infrastructure – classifying assets by their AI role, checking them against AI-specific baselines, and ranking what to fix. The clearest case of coverage without stopping power: Layer 3 is named in 10 of the 20 OWASP categories and blocks none of them. AI-SPM will tell you a GPU cluster is over-permissioned, an endpoint is unmonitored or an API key has no rotation policy, and it intercepts nothing. Read a finding as evidence that a control elsewhere failed or was never configured. It has a second limit that is easier to miss than the first: a posture tool reports on the scope it is connected to, so an asset in no connected cloud account – characteristically, the developer estate where agentic tooling runs – produces no finding at all. Shadow usage of an external AI service is not a posture question either; that sits on the access path, with ZTSA. TrendAI Vision One’s implementation is documented as a Pre-release feature scoped to connected cloud accounts. See: Chapter 3, Section 5 · Coverage vs stopping power: Chapter 3, Section 2
- AISVS (Artificial Intelligence Security Verification Standard): The OWASP standard of testable security requirements for AI applications – 12 chapters, 191 requirements, 3 verification levels, each phrased “Verify that…” – published 24 June 2026 and modelled on the OWASP ASVS. Where the Top 10 lists say what can go wrong, AISVS says what you check and how someone else confirms it. Cite requirements with the version prefix (
v1.0-C9.2.5), because identifiers are stable within a release and may move between them. Most production systems should target Level 2. See: Chapter 3, Section 10 · Frameworks: Chapter 2, Section 1 - Approval Fatigue: The degradation of a human-in-the-loop gate into a reflex, caused by asking for more approvals than a reviewer has attention to give. Not a discipline problem and not fixable by training or by adding a second approver – the fix is to reduce the number of gates so the remaining ones are read, which is why gates belong on the irreversible subset of actions. It has one usable telemetry signal: approval latency collapsing and a rejection rate falling toward zero across a large number of gates. That is the only listed user behaviour analytics signal that watches a defense rather than a user. Controls: Chapter 3, Section 6
- Artificial Intelligence (AI): The broader field focused on creating systems capable of tasks requiring human-like intelligence, from rule-based systems to modern deep learning. See: Chapter 1, Section 1
- Attention Mechanisms: Techniques that allow models to focus on relevant parts of an input sequence, enabling context-aware processing of text. See: Chapter 1, Section 4
- Automation Bias: The tendency to favour a suggestion from an automated system over contradictory information from a non-automated source. Well documented in aviation, healthcare and manufacturing, and amplified by AI because outputs arrive fluent, formatted and confident regardless of actual certainty. The measured version is the uncomfortable one: METR’s randomized trial found experienced developers were 19% slower with AI tooling while estimating they had been 20% faster – a 39-point self-assessment error in the direction of trusting the tool. (METR labels that result historical, having measured early-2025 tooling; the direction of the error is the durable part, not the percentages.) Accumulates as a trust gradient, which is why it is an attack precondition rather than a mitigation. See: Chapter 1, Section 7 · Exploitation: Chapter 2, Section 6 · Controls: Chapter 3, Section 6
- Autoregressive Text Generation: The process of generating text token by token, where each new token is predicted based on the preceding context. See: Chapter 1, Section 4
B
- Backdoor Attack: OWASP
LLM05: Data and Model Poisoning. A training-time attack that embeds a hidden trigger, so the model behaves normally on every input except the one carrying the trigger pattern. Hard to find, because a clean evaluation is evidence that you did not guess the trigger rather than evidence of a clean model. Harder to remove: Anthropic’s Sleeper Agents study ran supervised fine-tuning, RLHF and adversarial training at implanted backdoors and all three survived, most stubbornly in the largest models – adversarial training taught the model to recognise its trigger more precisely, so the behaviour became better hidden rather than absent. Remediation is therefore replacing the artifact, not patching it, which is what makes Layer 2’s provenance and signing controls preventive rather than optional. See: Chapter 2, Section 3 · Controls: Chapter 3, Section 4 - Bias (output): Systematic skew in what a model produces – associating certain roles with certain groups, or favouring some perspectives over others – learned from patterns in the training data and distributed across the whole parameter set. Not located in, and not auditable through, the bias terms below. See: Chapter 1, Section 1 · Mechanism: Chapter 1, Section 4
- Bias (parameter): A learned number attached to each unit in a neural network that shifts its output up or down before any input is considered. Weights say how much each input matters; biases say where the unit starts from. Shares a name with output bias and nothing else. See: Chapter 1, Section 4
- BM25: The standard keyword-ranking function, scoring a passage by how often the query’s terms appear in it, weighted by how rare those terms are across the corpus. Matches literal strings rather than meaning, which makes it complementary to vector search rather than obsolete: product codes, error identifiers and proper nouns are exactly what embeddings blur and BM25 finds. The keyword half of hybrid search. See: Chapter 1, Section 6
C
- Cascading Failure: OWASP
ASI08: Cascading Failures. A compromise in one agent propagating through the agents and systems downstream of it, so the blast radius is the whole pipeline rather than the entry point. Its precondition is any of the other agentic categories landing in one agent, so it has no signature of its own. The trap is the review stage that was supposed to catch it: a reviewer sharing the pipeline’s trust assumptions is not an independent control, and four agents built on the same model reading each other’s output as authoritative are one control wearing four hats. Layer 6 owns blast-radius containment; the structural fix is that at least one gate in the chain must be a different kind of check – a signature, a policy engine, a human. See: Chapter 2, Section 5 · Controls: Chapter 3, Section 8 - Chain-of-Thought (CoT) Prompting: Instructing a model to work through a problem in steps before answering, which measurably improves accuracy when the model is not already reasoning internally. Established by Wei et al. and Kojima et al. in 2022, against a model generation that could not decompose problems unaided; on a model with reasoning enabled the technique is redundant and can cut across the model’s own approach. A visible reasoning trace is generated output, not a record of the computation. See: Chapter 1, Section 5
- Chat Completions API: The request shape that takes an array of role-tagged messages and returns one message. The most widely implemented interface in the industry and the de facto portability standard, but stateless: the conversation exists only because the client stores it and replays the whole transcript on every request. Contrast with the stateful APIs that moved that storage to the provider. See: Chapter 1, Section 6
- Chunking: Splitting documents into retrievable passages during ingestion, because context windows are finite and narrower passages retrieve more precisely. The boundaries are a correctness decision, not just a sizing one – separating a claim from the sentence that qualifies it produces a chunk that retrieves as unconditional and is wrong. See: Chapter 1, Section 6
- Cloud API Deployment: Reaching a model through a provider’s own API endpoint. The provider owns the infrastructure, the model, and the guardrails; your data leaves your network to be processed. Lowest operational burden, highest capability ceiling, least control. See: Chapter 1, Section 3
- Confused Deputy: A privilege-escalation pattern where a component acts on a user’s behalf while holding broader privileges than the user has, so persuading it is sufficient to exercise privileges never granted. Endemic to MCP servers that hold their own service credentials, and the reason tool authorization has to be evaluated against the requesting user’s entitlements rather than the server’s – the database sees the service account in the
.envfile, not “an agent acting for Alice”. Filed under OWASPASI03: Identity and Privilege Abuse, because the escalation happens at the identity boundary. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5 - Constrained Decoding: Enforcing an output schema during generation rather than requesting it in the prompt – the sampling step is restricted so the model cannot emit a token that would violate the schema. Marketed as Structured Outputs and equivalents, and a materially stronger guarantee than the older “JSON mode”, which promised only syntactically valid JSON. It guarantees the container and never the contents: string fields still arrive unvalidated. See: Chapter 1, Section 5 · Output handling: Chapter 2, Section 6
- Content Credentials (C2PA): A cryptographically signed provenance manifest attached to media, recording the capture device or generating tool, whether AI was involved, and each subsequent edit. Standardised by the Coalition for Content Provenance and Authenticity, whose steering committee includes Adobe, Amazon, BBC, Google, Meta, Microsoft, OpenAI, Sony, TikTok and Truepic; specification 2.4 is current as of mid-2026, and 2.1 added redactable assertions and zero-knowledge identity proofs. A stronger primitive than deepfake detection because it is verifiable rather than probabilistic, and it must be used asymmetrically: a valid manifest from a trusted signer raises confidence, while a missing manifest is the normal state for legacy files, screenshots, re-encodes and anything through a platform that strips metadata. A control that reads absence as suspicion fires on most legitimate content and gets disabled. See: Chapter 3, Section 6
- Context Windows: The capacity of a model to process and retain information within a given input sequence, measured in tokens. See: Chapter 1, Section 4
- Continuous Loop (Scan-Protect-Validate-Improve): The operational cycle that keeps assessment findings and runtime guardrails aimed at the attacks actually arriving, because a rule set decays while every other control holds its value untouched. Each phase owes the next a named artifact: ranked findings, then rules plus a consequence change, then measured detection and false-positive rates, then regression tests. A practice before it is a product – MITRE ATLAS specifies the same cycle as
AML.M0035AI Red Team; AI Scanner and AI Guard instrument it. See: Chapter 3, Section 9 - Cosine Similarity: The similarity measure most retrieval systems use, comparing the angle between two vectors rather than the distance between their endpoints, so that two passages count as similar when they point the same way regardless of magnitude. What “nearest” means in a nearest-neighbour search. See: Chapter 1, Section 6
- Cross-Modal Injection: Prompt injection carried in a non-text modality – instructions rendered into an image, embedded in an audio track, or encoded as perturbations invisible to a human viewer – which accompany benign text and steer the model’s handling of both. Added to the scope of OWASP LLM01 in the 2026 edition as models took multimodal input. The operational consequence is narrow and important: a text-only input filter in front of a multimodal model is inspecting one channel of several. Extends to the physical world, where typographic instructions on signage or packaging enter through a camera-equipped agent’s field of view. See: Chapter 2, Section 2 · Defense: Chapter 3, Section 7
- Custom-Built Model: A model trained from scratch for one problem, without inheriting pre-trained weights from a foundation model. Resource-intensive and not adaptable beyond its training focus, but capable of precision that general models cannot reach – AlphaFold and GenCast are examples. See: Chapter 1, Section 2
D
- DAN (Do Anything Now): The archetypal jailbreak persona – an alternate character declared to have no restrictions, so that guardrails attached to the “assistant” persona do not follow the model into the fiction. Providers patch named variants and the pattern regenerates, which is why a blocklist of persona names is volume reduction rather than a control. See: Chapter 2, Section 2
- Data Classification: Categorizing data by sensitivity so controls can be applied proportionately. In an AI system it is the Layer 1 control the other layers read from: classification of the RAG corpus is what tells Layer 5 how aggressively to redact output, and a store assembled from several sources inherits the classification of its most sensitive one. Classify the derived artifacts – vector stores, fine-tuning sets, cached responses – not only the source systems. See: Chapter 3, Section 3 · Why the derived store counts: Chapter 2, Section 3
- Data Poisoning: OWASP
LLM05: Data and Model Poisoning. Corrupting the data a model learns from – pretraining corpus, fine-tuning set or feedback loop – so the defect is baked into the weights and reaches every future interaction. Dilution is not a defence. A 2025 study by the UK AI Security Institute, the Alan Turing Institute, Anthropic, Oxford and ETH Zurich found the attack needs a near-constant number of documents rather than a percentage: roughly 250 documents backdoored every model tested from 600M to 13B parameters – 0.00016% of the largest one’s training data. Read it as a floor for a narrow measured backdoor, and take the transferable point: the threshold is a quantity an attacker can produce, not a share of the internet they must own. See: Chapter 2, Section 3 · Controls: Chapter 3, Section 3 - Data Residency: The requirement that data be processed and stored within a defined geographic or legal jurisdiction. Frequently the constraint that rules out a direct cloud API and pushes a deployment towards serverless inference in a chosen region, or self-hosting. See: Chapter 1, Section 3
- Deep Learning (DL): A branch of machine learning that uses multi-layered neural networks to model complex patterns in data. See: Chapter 1, Section 1
- Deepfake Detection: Identifying AI-generated synthetic audio, video or imagery used to impersonate a real person. MITRE ATLAS carries it as mitigation
AML.M0034againstAML.T0052.001Deepfake-Assisted Phishing – note that the attack has no OWASP LLM Top 10 category, because the target is a person rather than an LLM application. Treat it as a triage signal and never as an authorisation decision: a detector’s benchmark score describes the synthesis pipelines it was trained on, and against an unseen generator – or a newer version of a seen one – reported accuracy falls by tens of percentage points, degrading further under ordinary recompression and resizing. Content Credentials answer the verifiable question instead, and the control that survives a wrong detector is procedural. Controls: Chapter 3, Section 6 - Defense in Depth: Layering controls so that one being bypassed is not the end of the story. For AI the useful reading is not a stack of gates in a line – only Layer 5 sits in the request path, Layers 1 and 2 have finished their work before a request arrives, Layers 3 and 6 watch from beside it, and Layer 4 acts on the human at the far end. Depth comes from controls of different kinds at different positions, which is why redundancy inside a layer counts too: an input filter and an output filter fail independently because they look for different things. See: Chapter 3, Section 2
- Denial of Service (AI-specific): OWASP
LLM06: Unbounded Consumption. Attacks that exploit the cost of inference to degrade availability – context-window stuffing, recursive tool loops, high-volume generation. Distinguished from conventional DoS by its economic twin, denial of wallet, where availability never degrades and the bill is the damage. See: Chapter 2, Section 4 · Defense: Chapter 3, Section 7 - Denial of Wallet: The economic form of AI denial of service: an attacker who can trigger expensive requests turns a metered API budget into the damage, without degrading availability at all. The load-bearing control is a spend cap with an alert set below it, since the cap is the only control that bounds the loss while everything else is missing. See: Chapter 2, Section 4 · Defense: Chapter 3, Layer 5
- DevSecOps: Integrating security into every phase of the development lifecycle rather than gating at the end. Applied to AI the pipeline gains two stages conventional software does not have – Train (data lineage, poisoning detection, training-data access control) and Validate (red-teaming, adversarial input and injection testing) – while Monitor picks up prompt and response filtering, behavioural anomaly detection and cost monitoring. See: Chapter 3, Section 1 · The operating cycle: Chapter 3, Section 9
- Digital Omnibus on AI: Regulation (EU) 2026/1744, published in the Official Journal 24 July 2026 and in force 27 July 2026 – six days before the EU AI Act’s original high-risk deadline. Defers Annex III standalone high-risk obligations to 2 December 2027 and Annex I high-risk in regulated products to 2 August 2028, citing undesignated national authorities and unfinished harmonised standards. Article 50 transparency applied on schedule and prohibitions, AI literacy and GPAI duties were untouched. Deferred is not cancelled, and no requirement was weakened – but compliance material written before August 2026 states a date that has moved. See: Chapter 3, Section 11
- Direct Prompt Injection: An attack where the person interacting with the system supplies the malicious instruction themselves. Contrast indirect prompt injection, which is the more serious problem for agents because the agent seeks out the attacker’s payload as part of doing its job. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 2
- Divergence Attack: Prompting an aligned model to repeat a single token indefinitely until it abandons chatbot-style generation and begins emitting memorized training data – measured by Nasr, Carlini et al. (2023) at 150x the extraction rate of normal prompting.
AML.T0057. The specific prompt is filtered on current models; the finding that matters is that alignment does not eliminate memorization, so blocking a trigger removes the known path and leaves the data in place. See: Chapter 2, Section 6 - DSPM (Data Security Posture Management): Discovering, classifying and tracking sensitive data across cloud and AI environments. The AI-specific work is finding the derived copies – training sets, fine-tuning corpora, vector stores, cached responses – which are governed by whoever stood them up and frequently sit outside the DLP and retention tooling wrapped around the source systems they were built from. Contrast AI-SPM, which assesses the infrastructure rather than the data on it. See: Chapter 3, Section 3
E
- Edge Deployment: Running small models directly on end-user devices for offline operation and low interactive latency. Data never leaves the device, but the weights sit on hardware you do not control, making physical access part of the threat model. See: Chapter 1, Section 3 · Security implications: Chapter 2, Section 7
- Efficient Attention: A family of techniques that make attention affordable on long inputs, split into two kinds that are often conflated. Exact methods such as FlashAttention compute identical results while avoiding materializing the full attention matrix in memory – memory drops to linear, the FLOP count stays quadratic. Approximate methods such as sliding-window attention genuinely lower the asymptotic cost by restricting what each token can attend to. Grouped Query Attention is a third thing again: it shrinks the KV cache, which bounds how many long-context requests a server can hold concurrently. See: Chapter 1, Section 4
- Egress Allowlisting: A default-deny rule on outbound connections from an AI application, orchestrator or code-execution sandbox, permitting only enumerated destinations. The strongest single Layer 6 control because it does not inspect anything: it survives an exploit you cannot patch, and is indifferent to whether the technique that triggered the outbound request has a name. Its known limit is that an allowed destination which forwards arbitrary content – a general-purpose proxy, an allowlisted image service – loosens the bound, which is how EchoLeak’s exfiltration reached a permitted domain. See: Chapter 1, Section 7 · Controls: Chapter 3, Section 8
- Embedding: A vector representation of a token or a span of text, positioned in a high-dimensional space so that semantically similar items sit close together. The representation every transformer works in internally, and the basis of retrieval. Derived from text rather than a substitute for it, which is why a vector store holds sensitive data in its own right. See: Chapter 1, Section 4 · Retrieval use: Chapter 1, Section 6 · Attacks: Chapter 2, Section 3
- Embedding Inversion: Reconstructing meaningful source text from its embedding vector alone. The reason a vector store cannot be treated as de-identified data: embedding is a transformation, not anonymisation, and a store of vectors derived from confidential documents inherits their classification. Compounded in practice because most stores keep the original chunk text alongside the vector anyway. See: Chapter 1, Section 6 · Controls: Chapter 3, Section 3
- EOS Token: An end-of-sequence token that signals to the model to stop generating further tokens, marking the completion of a response. See: Chapter 1, Section 4
- EU AI Act: Regulation (EU) 2024/1689, in force since 1 August 2024. It runs two parallel regimes, and the four-tier table most summaries show describes only the first: risk tiers for AI systems (unacceptable, high-risk, limited, minimal), and a separate track for general-purpose AI models under Chapter V. For LLM work the second matters more. Obligations attach by role – see provider vs deployer. Dates moved on 27 July 2026: see Digital Omnibus. See: Chapter 3, Section 11
- Excessive Agency: A vulnerability where an AI system is granted more capability than its task requires – OWASP LLM03. It decomposes into three causes with three different fixes: excessive functionality (tools the task never needs – remove the tool), excessive permissions (needed tools running with more privilege than required – scope the credential), and excessive autonomy (irreversible actions with no human checkpoint – gate the irreversible subset). Diagnosing which one a system has is faster than reading its system prompt. Introduced: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Extended Reasoning: The mode in which a model deliberates internally before answering – the same thing catalogued below under Reasoning (Test-Time Compute), named from the prompting side. It is a setting rather than a model class, and it changes how you prompt: drop step-by-step scaffolding, keep examples only for pinning output format, and expect
temperatureandtop_pto be restricted or rejected outright. See: Chapter 1, Section 5
F
- Fail-Open / Fail-Closed: What an inline control does when it errors or times out. Fail-open forwards the request unfiltered, preserving availability and silently removing the guardrail at exactly the moment load makes an attack likeliest; fail-closed refuses the request, preserving the guarantee and handing an attacker who can degrade the inspector a denial-of-service primitive. Decide per action class rather than globally – fail-open on read-only paths, fail-closed on anything that writes, spends or reaches a third party – and alert on the fail-open path, because a control that is bypassed and silent is worse than one that was never deployed. See: Chapter 3, Section 9
- Few-Shot Prompting: Supplying worked input-output examples before the real task, so the model mirrors the demonstrated pattern. The most reliable way to pin down output format, tone and structure; 3-5 diverse examples covering the range of expected outputs is the usual sweet spot. Costs prompt tokens on every request, and the examples form part of the prompt an attacker may be able to extract. See: Chapter 1, Section 5
- Fine-Tuning: The process of adapting a pre-trained model to a specific task or domain through additional training on smaller, task-specific datasets. See: Chapter 1, Section 2
- Foundation Models: Large-scale, pre-trained models designed to handle a wide range of tasks, serving as a base for fine-tuning and specialization. See: Chapter 1, Section 2
G
- Generative AI (GenAI): AI systems designed to create new content – text, images, audio, or code – based on patterns learned during training. See: Chapter 1, Section 1
- GPAI (General-Purpose AI Model): The EU AI Act’s category for a model with significant generality that can competently perform a wide range of tasks – which is every foundation model this course discusses. Regulated on a separate track from AI systems (Chapter V, applicable since 2 August 2025) with its own obligations: technical documentation, a copyright policy, and a public summary of training content. Models presumed to have high-impact capability – the threshold is more than 10^25 FLOP of training compute – carry additional systemic-risk duties under Art. 55 including model evaluation, adversarial testing, serious-incident reporting and cybersecurity protection. An organisation can be subject to both tracks at once. See: Chapter 3, Section 11
- Grounding: Constraining a model’s answer to supplied source material rather than its parameters, which is what retrieval is for. It reduces confabulation and does not produce correctness: a grounded answer built on an outdated, mis-scoped or poisoned passage arrives with a citation attached, which raises the reader’s confidence without raising the answer’s accuracy. See: Chapter 1, Section 6 · Over-trust: Chapter 2, Section 6
H
- Hallucinations: AI-generated content that appears convincing but has no basis in reality or training data, a form of confabulation. Weaponizable because it is consistent: models fabricate the same names repeatedly, which makes the fabrication predictable enough to pre-register. See slopsquatting. See: Chapter 1, Section 1 · Weaponization: Chapter 2, Section 6
- Hidden Context Exposure: OWASP LLM08 (2026), broadened from the 2025 category System Prompt Leakage. Covers the extraction, inference or reconstruction of any non-user-facing context the application assembles for the model – the system prompt, retrieved policy text, tool and function schemas. The load-bearing guidance is a design rule rather than a mitigation: assume hidden context is discoverable, and build so that disclosure has little or no direct security impact. Credentials do not belong there, and it is not an authorization boundary. Contrast system prompt leaking. Framework: Chapter 2, Section 1 · Attacks: Chapter 2, Section 2 · Defense: Chapter 3, Section 7
- Horizontal Model: A model built for broad, cross-domain use rather than one industry or task. Most frontier foundation models are horizontal. Contrast with a vertical model. See: Chapter 1, Section 2
- Hosted Deployment: A general term for any pattern where someone else runs the model – covering both cloud API and serverless inference, which differ in whose boundary the data stays inside. See: Chapter 1, Section 3
- Human-in-the-Loop (HITL): Requiring human approval before an agent’s action takes effect. A real control, and one with a specific failure mode: a gate a reviewer clicks through fifty times a day stops functioning as a gate. Effective use means gating the irreversible subset of actions rather than all of them – a gate on everything becomes a gate on nothing – see approval fatigue. It is the Blueprint control that bounds consequence when every other layer has missed, and the only Layer 4 control that no device policy or third-party contract can take away from you, because it lives in your own application. See: Chapter 1, Section 7 · Controls: Chapter 3, Section 6
- Hybrid Deployment: Routing each request to the cheapest capable tier across edge, self-hosted, and cloud options. Optimizes cost and privacy together, but makes the routing logic itself a security control – it decides which data crosses which boundary. See: Chapter 1, Section 3
- Hybrid Search: Running vector search and BM25 keyword search over the same corpus and fusing their candidate lists, usually with Reciprocal Rank Fusion, which combines by rank position because raw scores from the two methods are not comparable. The production standard for retrieval, and measurably better than semantic search alone: each method finds what the other misses. See: Chapter 1, Section 6
I
- IDS/IPS (Intrusion Detection/Prevention System): Network controls that monitor traffic for malicious activity, combining three rule types that cover different intervals: signatures (known payloads), virtual patches (disclosed but unpatched flaws), and behavioural rules (deviation from a learned baseline, the only cover before disclosure). Applying them to AI traffic means handling large payloads, streamed responses, tool-call bursts, and east-west serving traffic that general-purpose rule sets were not tuned for. See: Chapter 3, Section 8
- Improper Output Handling: OWASP LLM10 (2026), down from LLM05 in 2025 – the largest fall on the list, while its scope grew. Model output passed to a downstream consumer without validation, making the model an injection vector against your own systems. The classic sinks are databases, browsers and shells; the 2026 additions are the ones a conventional input-validation review will not look for, because in a normal application no untrusted content reaches them: auto-fetching renderers (a markdown image reference is an exfiltration channel), terminal and ANSI escape sinks, and insecure code generated at scale. Note what does not fix it: a schema guarantees the container and never the contents, and a streamed token is already published. See: Chapter 2, Section 6 · Framework: Chapter 2, Section 1 · Controls: Chapter 3, Section 7
- Incident Response Plans: Structured procedures for handling a security incident, adapted for AI-specific types. Two adaptations matter. The lifecycle changed – NIST SP 800-61r3 (April 2025) retired the detect/contain/eradicate/recover phasing in favour of a CSF 2.0 profile treating IR as continuous risk management, which suits AI because incidents often have no clean end state: a model cannot be un-trained. And containment order is set by forensics as much as speed – suspend, revoke, audit, then roll back, since rollback overwrites the record the audit needs. See: Chapter 3, Section 11
- Indirect Prompt Injection: An attack where malicious instructions are embedded in external data sources (documents, web pages, emails) that an AI system processes, causing it to execute the attacker’s intent without direct user interaction. The dominant threat to agents specifically, because an agent reads untrusted content as part of its normal work – the payload does not have to be delivered to a user, only left somewhere the agent will look. Contrast direct prompt injection. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 2 · Agentic case: Chapter 2, Section 5
- Inference: The application phase where trained models generate outputs based on learned patterns without further adjustments to their parameters. See: Chapter 1, Section 4 · In practice: Chapter 1, Section 6
- Instruction Hierarchy: A trained-in ordering that makes a model privilege system and developer instructions over user and tool content (Wallace et al., 2024). It is a property of the model, selected at Layer 2 when you choose one – not something a gateway or filter can enforce, since the context window itself has no privilege levels. A gateway can supply the provenance signals the hierarchy consumes; measured compliance is partial. See: Chapter 2, Section 2 · Where it does not live: Chapter 3, Layer 5
J
- Jailbreaking: Techniques that bypass a model’s safety alignment and content policies, getting it to produce output it was trained to refuse – role-play personas, multi-turn escalation, encoding and Unicode tricks. Distinguished from prompt injection by whose instructions are subverted: injection defeats the developer’s application logic, which the developer can remedy; jailbreaking defeats the provider’s alignment, which the deploying organisation did not build and cannot patch. That is why the defence is output inspection rather than a better system prompt. See: Chapter 2, Section 2 · Defense: Chapter 3, Section 7
L
- Large Language Models (LLMs): Specialized deep learning models trained on extensive text corpora for language-related tasks like generation, comprehension, and reasoning. See: Chapter 1, Section 1
- LEARN: A recall mnemonic used in this course for five application-level defense practices – Linguistic Shielding, Execution Supervision, Access Control, Robust Prompt Hardening, Nondisclosure. It is a teaching device, not a published standard, and it is an index into AISVS rather than a substitute for it: its five letters reach five of the standard’s twelve chapters, leaving Memory (
v1.0-C8) and MCP (v1.0-C10) among the six with no letter. Use it to remember; use the chapter list to be complete. See: Chapter 3, Section 10 - Lethal Trifecta: The combination that makes an agent exploitable, named by Simon Willison in June 2025: access to private data, exposure to untrusted content, and the ability to communicate externally. With all three, whoever controls the untrusted content can direct the agent to read the private data and send it out. Its value is as a decision tool – removing any one leg breaks the attack, which turns a vague warning into three concrete architectural options. Assess it across a whole system, since a multi-agent pipeline can hold all three legs while no single agent does. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- LLMjacking: Running queries on stolen LLM API credentials at the victim’s expense – a credential-hygiene failure that presents as a billing anomaly, not a model attack, which is why model-level hardening does nothing about it. Treat the widely-quoted “$100,000 a day” as a modelled worst case rather than a damages figure: Sysdig’s 2024 analysis calculated roughly $46,000 a day at maximum quota, and later estimates scaled it for frontier models. The controls that bound it are spend caps, key rotation and per-key usage alerting. See: Chapter 2, Section 4 · Case: Chapter 2, Section 1 · Defense: Chapter 3, Section 7
- Lost in the Middle: The finding that models attend unevenly across a long context, favouring its beginning and end, so a relevant passage placed in the middle can be effectively missed. Established by Liu et al. (TACL 2024). The practical consequence: a large context window is not the same thing as uniform attention across it, and where information sits in a prompt still matters. See: Chapter 1, Section 4
M
- Machine Learning (ML): A subset of AI where systems learn patterns from data rather than being explicitly programmed, using statistical methods to improve performance on tasks. See: Chapter 1, Section 1
- MAESTRO: The Cloud Security Alliance’s threat modeling framework for agentic AI (February 2025), structured as seven layers – foundation model, data operations, agent frameworks, deployment infrastructure, evaluation, orchestration, ecosystem. Exists because STRIDE and its peers assume entities with fixed roles and design-time trust boundaries, and an agent has neither: it is caller, callee and data store at once, and installing a tool redraws the boundary at runtime. Reach for it when a system becomes agentic; keep STRIDE inside a layer where the boundaries hold still. See: Chapter 3, Section 1
- Matryoshka Representation Learning: Training an embedding model so the most important information concentrates in the earliest dimensions, which makes a long vector safely truncatable – keeping the first 512 of 3072 dimensions costs far less retrieval quality than discarding five sixths of the numbers suggests. Turns dimension count into a storage-and-speed dial rather than a quality ranking, so “more dimensions is better” is the wrong reading of a model roster. See: Chapter 1, Section 6
- Membership Inference: Determining whether a specific record was in a model’s training set by observing its confidence levels and output patterns, without extracting the record itself.
AML.T0024.000, a sub-technique of Exfiltration via AI Inference API. Dismissed as leaking only one bit, but that bit is sometimes the whole disclosure: confirming a named individual was in a model trained only on one clinic’s oncology patients discloses a diagnosis without revealing a field. Actionable under GDPR and HIPAA, and the reason a model inherits the classification of the population it was trained on. See: Chapter 2, Section 6 - Memory and Context Poisoning: OWASP
ASI06. Writing instructions into an agent’s persistent state – conversation history, learned preferences, stored facts, project context – so they are re-supplied on every future run. The property that makes it worse than a one-shot injection is persistence: the payload is delivered once and executes indefinitely, and removing the original document does not remove it. Reachable by anyone who can write to the memory store, which includes a compromised tool output and a poisoned document the agent merely read. Note that adopting a provider’s stateful API moves this surface onto someone else’s infrastructure without changing its shape. See: Chapter 2, Section 5 · Where state lives: Chapter 1, Section 6 · Controls: Chapter 3, Section 3 - Metadata Filtering: Restricting a retrieval query to the passages a given user, tenant or scope is entitled to, evaluated inside the search rather than applied to its results. The only stage of a retrieval pipeline that asks about entitlement rather than relevance, and therefore the load-bearing access control in a RAG system – which also means whoever can write the metadata can change who retrieves a document. See: Chapter 1, Section 6 · Attacks: Chapter 2, Section 3 · Controls: Chapter 3, Section 3
- Misinformation: OWASP LLM07 (2026), up from LLM09. Confidently presented false content – hallucinations as fact, fabricated citations, assured answers outside the model’s knowledge. Carries the widest belief-versus-evidence gap on the 2026 list: practitioners voted it near the bottom while the classified incident record put it near the top, and it landed at #7 as a compromise. That is the dangerous direction, since a risk the field under-rates relative to what it has already been burned by is where the next incident comes from. Becomes a security problem rather than a quality one when fluent output drives a tool call or an install. See: Chapter 2, Section 6 · Framework: Chapter 2, Section 1 · Controls: Chapter 3, Section 6
- MITRE ATLAS: The Adversarial Threat Landscape for AI Systems – an ATT&CK-shaped matrix of adversarial tactics, techniques, mitigations and real-world case studies for AI. Release 2026.07 carries 16 tactics, 178 techniques, 37 mitigations and 68 case studies; it ships monthly, so treat any count as a snapshot and cite the release. Where the OWASP lists answer what can go wrong, ATLAS is the only one of the three that enumerates what has actually been done. Names drift between releases as well as counts, and the framework is split into predictive and generative halves –
AML.M0015is Predictive AI Adversarial Input Detection and is the wrong half for anything LLM-shaped; the generative counterpart isAML.M0020Generative AI Guardrails. See: Chapter 2, Section 1 · Using it in anger: Chapter 3, Section 8 - Mixture-of-Experts (MoE): An architecture in which a model is divided into many specialized sub-networks and only a few are activated for any given token. A model with hundreds of billions of total parameters may therefore cost no more to run per request than a much smaller dense one – you pay for the parameters actually used, not the idle ones. See: Chapter 1, Section 2
- Model Context Protocol (MCP): An open standard, introduced by Anthropic in late 2024, by which agents discover and call external tools. An MCP server advertises each tool with a name, a schema, and a natural-language description telling the model when to use it – and that description is placed directly in the context window while remaining invisible to the user. Every MCP security problem follows from that: the tool catalogue is untrusted input on the prompt path, which makes it a supply-chain concern rather than a prompting one. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5 · Controls: Chapter 3, Section 4
- Model Theft: Umbrella term covering two attacks with nothing in common but their goal, worth keeping apart because the controls differ entirely. Model extraction copies a model’s behaviour through its API – query it at volume, train a local model on the input-output pairs – and needs only API access, so it is invisible to file-integrity controls and bounded by rate limits and what the endpoint returns besides text. Model exfiltration copies the weights off a host, and needs access to storage or a device, which makes it the dominant risk for small models at the edge. Mapped to LLM06: Unbounded Consumption in the OWASP 2026 edition. Extraction: Chapter 2, Section 4 · Exfiltration at the edge: Chapter 2, Section 7 · Controls: Chapter 3, Section 7
- Moderation Endpoints: External classifiers applied to input before a model sees it and to output before a user does, scoring content against categories of harm. Independent of the model, so they still apply when it is swapped, fine-tuned, or has had its refusal behaviour weakened. See: Chapter 1, Section 3
- Multi-Agent System: Several agents with distinct roles collaborating on a task. It resembles separation of duties and usually is not: privilege unions rather than partitions, so the system’s effective reach is every tool granted to any agent in it, and inter-agent messages are untrusted input that reads as a colleague’s instruction. A researcher that browses the web plus a writer that holds credentials gives the system the lethal trifecta even though neither agent has it alone. The message channel itself is OWASP
ASI07: Insecure Inter-Agent Communication; a compromise propagating along it isASI08, see cascading failure. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5 - Multi-Head Attention: A mechanism that enables models to analyze multiple aspects of context simultaneously by running several attention computations in parallel. See: Chapter 1, Section 4
- Multimodal Model: A model that natively processes and relates more than one type of input or output – text, images, audio, video, or code – within a single system, rather than chaining separate specialized models together. See: Chapter 1, Section 1
N
- Neural Networks: A type of machine learning model inspired by the human brain, consisting of layers of interconnected nodes that process information through weighted connections. See: Chapter 1, Section 1
- Next-Token Prediction: The single operation underlying all LLM output – given the tokens so far, predict a probability distribution over the next token, select one, append it, and repeat. Nothing in this process verifies truth, which is the root of hallucination. See: Chapter 1, Section 1
- NIST AI RMF (AI Risk Management Framework): Voluntary framework for managing AI risk across the lifecycle (AI RMF 1.0, January 2023), organised around Govern, Map, Measure and Manage. Non-prescriptive about controls by design: it is the governance structure inside which controls like the Blueprint’s are chosen and documented. For generative systems the applicable document is its companion profile, NIST AI 600-1, whose 12 risk names – not the four function names – are what an assessment expects to see referenced. See: Chapter 3, Section 11
O
- On-Premises Deployment (Self-Hosted): Running downloaded model weights on your own hardware, on-premises or in a private cloud. Maximum control over data, and the only pattern that can run a large model air-gapped – at the cost of owning every control the provider had supplied. Limited to open-weight models. See: Chapter 1, Section 3
- Open Weights: A model whose trained parameters can be downloaded and run on your own infrastructure. Distinct from open source: the weights being available does not mean the training data, training code, or process are, and licenses range from genuinely permissive (Apache 2.0, MIT) to community terms with acceptable-use policies and commercial caps. See: Chapter 1, Section 1
- OWASP LLM Top 10: The industry-standard vulnerability taxonomy for LLM-specific security risks, covering categories from prompt injection to unbounded consumption. The course uses the 2026 edition (published 4 August 2026), the first ranked partly on classified real-world incident data rather than practitioner vote alone. Eight of the ten identifiers moved from the 2025 edition, so pre-August-2026 material numbers them differently. See: Chapter 2, Section 1
- OWASP Top 10 for Agentic Applications: The ASI01–ASI10 taxonomy, mapping security risks specific to systems that plan, call tools, persist memory and coordinate with each other – goal hijacking, tool misuse, cascading failures, rogue agents. Companion to the LLM Top 10 rather than a replacement: the dividing line is whether the model is a component in your application or an actor with tools and consequences. Published 9 December 2025 and not yet revised. Framework: Chapter 2, Section 1 · Worked through: Chapter 2, Section 5
P
- Parameter-Efficient Fine-Tuning (PEFT): A family of fine-tuning techniques – LoRA being the best known – that adapt a model by training a small set of additional or selected parameters while leaving the bulk of the pre-trained weights frozen, making domain adaptation affordable on modest hardware. See: Chapter 1, Section 2
- Parameters: The internal values that a model learns during training, acting as weights that determine the importance of different patterns in input data. See: Chapter 1, Section 2
- Pickle: Python’s native serialization format, long the default for model checkpoints. It can serialize arbitrary Python objects, which means deserializing it can execute arbitrary code – by design, not as a bug. A tampered Pickle checkpoint runs the attacker’s payload the moment it is loaded. See: Chapter 1, Section 3
- Positional Information: The mechanism by which a transformer knows the order of tokens rather than merely which tokens are present. The original Transformer added a positional encoding to each input embedding; most current models instead apply rotary position embeddings inside the attention computation. Either way, order is supplied explicitly – attention itself is order-blind. See: Chapter 1, Section 4
- Posture Management: The umbrella for continuously assessing configuration rather than intercepting traffic. Two variants appear in this course and they answer different questions: AI-SPM assesses the infrastructure an AI system runs on, DSPM assesses the data it was built from and writes to. Both are Layer 3 in character – broad coverage, no stopping power. See: Chapter 3, Section 5
- Prompt: The text input given to a model that specifies what it should do – instructions, context, and any constraints on the format or style of the response. See: Chapter 1, Section 1
- Prompt as Request, Not Boundary: The property that an instruction placed in a prompt steers a model without enforcing anything. Because every input reaches the model as one flat token sequence, a system-prompt constraint carries no privilege over user text or retrieved content – the model honours it because instruction-following is trained behaviour, which is a strong tendency rather than a guarantee. Prompt-level rules are worthwhile hardening and are never load-bearing on their own; real controls remove a capability or inspect the traffic. See: Chapter 1, Section 5 · Exploitation: Chapter 2, Section 2 · Where controls belong: Chapter 3, Section 7
- Prompt Caching: A provider feature that charges a reduced rate – typically 75-90% off input – for a repeated prompt prefix. Now largely automatic above a minimum prefix length of roughly 1,024 tokens. Because the match is on a prefix, it rewards a stable one: a timestamp or request ID interpolated near the front invalidates everything after it, and a cache write costs slightly more than an uncached call. Below the threshold, shorten the prompt; above it, stabilize it. A discount on the rate, not a change of tier. See: Chapter 1, Section 3 · Cost strategy: Chapter 1, Section 6
- Prompt Engineering: Designing and refining prompts to elicit desired responses from a model, using techniques including zero-shot, few-shot, chain-of-thought and structured output. The same mechanism as prompt injection, pointed in a different direction: it works by placing text in the context window, which is available to anyone whose text reaches the window. See: Chapter 1, Section 5
- Prompt Injection: An attack where crafted input manipulates an LLM into ignoring its instructions, leaking data, or performing unintended actions – OWASP LLM01, and the highest-ranked entry on the 2026 list. It has no complete fix because it exploits an architectural property rather than a coding error: the context window is a single trust boundary with no privilege levels, so an attacker’s text and the developer’s instructions carry equal weight and the model resolves the conflict by trained tendency rather than enforcement. OWASP states plainly that no fool-proof prevention is known. Splits into direct and indirect forms; contrast jailbreaking. See: Chapter 2, Section 2 · Defense: Chapter 3, Section 7
- Provider / Deployer (EU AI Act): The two roles the EU AI Act assigns obligations to. A provider develops a system or model, or places it on the market under its own name, and carries the development burden: risk management, data governance, technical documentation, conformity assessment, registration, post-market monitoring. A deployer uses a system under its own authority in a professional capacity, and carries use-per-instructions, competent human oversight, operational monitoring, log retention and serious-incident reporting. Most organisations securing AI are deployers while most published guidance is written for providers, which is how teams end up reading obligations that are not theirs. The trap: a deployer that puts its own name on a high-risk system, changes its intended purpose, or substantially modifies it becomes a provider – and fine-tuning a general-purpose model for a high-risk use is exactly that. See: Chapter 3, Section 11
Q
- Quantization: Storing model weights at reduced numerical precision so a model fits smaller hardware. The enabling technique for edge deployment, and a security concern in its own right: aggressive quantization can degrade safety training disproportionately. Introduced: Chapter 1, Section 3 · Security implications: Chapter 2, Section 7
R
- RAG Poisoning: OWASP
LLM09: Vector and Embedding Weaknesses. Writing documents into a retrieval corpus so that the answer to a chosen question becomes the attacker’s answer. It hides facts, not instructions, which is what separates it from indirect prompt injection arriving through the same channel: the poisoned passage contains no command, so an instruction-filtering guardrail on retrieved content has nothing to detect, and the user gets a fluent, correctly cited answer with no symptom to notice. PoisonedRAG (Zou et al., USENIX Security 2025) measured roughly 90% success from five documents per attacker-chosen question – per question, not per corpus – against knowledge bases of millions, because similarity search does not care how many documents it did not return. See: Chapter 2, Section 3 · Pipeline: Chapter 1, Section 6 · Controls: Chapter 3, Section 3 - Reasoning (Test-Time Compute): Letting a model generate an extended chain of intermediate reasoning tokens before its visible answer, trading latency and token cost for better results on complex problems. Once a separate product category, it is now typically an adjustable effort setting on mainline models. See: Chapter 1, Section 1
- Reciprocal Rank Fusion (RRF): The usual method for merging the candidate lists of hybrid search, scoring each passage by its rank position in each list rather than by the underlying scores – necessary because a cosine similarity and a BM25 score are not on comparable scales. See: Chapter 1, Section 6
- Red-Teaming (AI): Structured, authorized adversarial testing of an AI system before an attacker does it unauthorized. Specified rather than improvised: ATLAS
AML.M0035defines it as a three-phase mitigation – plan and scope, execute, then assess, report and improve – where the third phase (owners, retest, conversion of failures into regression tests and detection logic) is what makes the first two worth anything. The OWASP GenAI Red Teaming Guide adds the scope dimension: model evaluation, implementation testing, infrastructure assessment, and runtime behavior analysis, the last being the only one that tests component interaction. Distinct from automated scanning, which reaches models and part of the agent surface and not the chained attacks. See: Chapter 3, Section 11 · Loop: Chapter 3, Section 9 - Refusal Pathways: Learned behaviour, instilled during alignment training, that leads a model to decline harmful or undesirable requests. Distributed across the weights rather than stored as a rule – and concentrated enough that on open weights it can be removed without retraining, which is why it is not a security boundary. See: Chapter 1, Section 3
- Reinforcement Learning from Human Feedback (RLHF): A post-training stage that aligns a model’s outputs with human preferences by training it against human judgements of which response is better. It is how a raw next-token predictor becomes an assistant that answers helpfully and declines harmful requests. See: Chapter 1, Section 2
- Reranking: Re-scoring the candidates that survived retrieval using a cross-encoder, which evaluates the query and a passage together rather than comparing vectors computed before the query existed. Substantially more accurate and far too slow to run over a whole corpus, which is why production retrieval is a funnel: cheap search narrows millions of chunks to tens, reranking narrows tens to the handful actually sent. See: Chapter 1, Section 6
- Retrieval-Augmented Generation (RAG): Fetching relevant passages from an external corpus at query time and supplying them to the model as context, instead of relying on its parameters. The dominant enterprise architecture, because it decouples corpus size from request size, updates without retraining, and is the only approach that produces a citation. Also the largest attack surface in a typical AI application: every stage from ingestion to output is reachable by someone. See: Chapter 1, Section 6 · Compared with alternatives: Chapter 1, Section 4 · Attacks: Chapter 2, Section 3 · Controls: Chapter 3, Section 3
- Rogue Agent: OWASP
ASI10: Rogue Agents. An agent acting outside its intended parameters – through compromise, misconfiguration, or simply having outlived the reason it was built. The governance half is the one teams miss: an agent holding credentials and tools with no accountable human owner is a rogue agent that has not gone wrong yet, so the load-bearing control is a named owner and review date per agent rather than a detector. See shadow AI for the discovery problem it creates. Foundations: Chapter 1, Section 7 · See: Chapter 2, Section 5 · Controls: Chapter 3, Section 6 - Rug Pull (MCP): An MCP server that changes its tool definitions after you approved it, typically without re-prompting or notifying the client. The consequence for review: approval is not a durable property of a server you do not control, so version pinning and change review matter more than a one-time assessment. See: Chapter 1, Section 7 · Controls: Chapter 3, Section 4
- Rule-Based Systems: Early AI systems that relied on predefined rules to make decisions, limited by the need to explicitly program every scenario. See: Chapter 1, Section 1
S
- Safetensors: A serialization format storing only raw tensor data and a JSON header, with no mechanism for executing code on load. Now the default across the Hugging Face Hub, and the format to prefer for any model you did not produce yourself. See: Chapter 1, Section 3
- Security for AI Blueprint: TrendAI’s six-layer defense framework, and the spine of Chapter 3 – Layer 1 Secure Your Data, Layer 2 Secure Your AI Models, Layer 3 Secure Your AI Infrastructure, Layer 4 Secure Your Users, Layer 5 Secure Access to AI Services, Layer 6 Defend Against Zero-Day Exploits. The numbers are labels for six protection domains, not a stack: only Layer 5 sits in the request path, Layers 1 and 2 finish their work before a request arrives, Layers 3 and 6 watch from beside it and produce alerts, and Layer 4 acts on the human. Coverage and stopping power come apart – Layer 5 is named in 11 of the 20 OWASP categories and can refuse a request, Layer 3 is second-broadest at 10 and blocks nothing, and Layer 2 is narrowest at 3 while being the only layer that acts before anything is running. See: Chapter 3, Section 2 · Layer→OWASP index: Chapter 3
- Semantic Caching: Serving a stored answer when a similar question arrives, matched by embedding proximity rather than exact text. The broadest-reach response cache and the most dangerous: in front of an entitlement-filtered corpus the answer depends on who asked, so a key built only from the question serves one user’s permission-scoped results to another, with no retrieval running at all. The similarity threshold that tunes hit rate also widens the disclosure boundary. See: Chapter 1, Section 6
- Sensitive Information Disclosure: OWASP LLM02, the one 2026 category where practitioner ranking and incident evidence agree exactly. Covers confidential data reaching a user through model output by any of three routes, which are not equally likely: training-data extraction (
AML.T0057), membership inference, and – by far the most common in practice – retrieved content passed straight through to a reader who was never entitled to it. The third needs no attacker and produces a fluent, correctly cited answer, which is why the diagnostic question is how the content entered the context window rather than what the model did wrong. Attacks: Chapter 2, Section 6 · Controls: Chapter 3, Section 3 and Section 7 - Serialization: Converting a model into a format that can be saved to disk and later loaded into memory for inference. The format determines whether loading is reading data or running code, which makes it a security decision rather than an implementation detail. See: Chapter 1, Section 3 · Exploitation: Chapter 2, Section 4
- Serverless Inference: A deployment pattern where your existing cloud provider hosts models on your behalf from a catalogue, keeping requests inside your own tenancy under the IAM, logging, and regional controls you already use. Trades a per-token premium and slower access to new releases for data residency and unified identity. See: Chapter 1, Section 3
- Shadow AI: AI tools, services and models used inside an organisation without security review. Shadow IT with a shorter fuse – shadow SaaS leaks over months of accumulated use, while a single paste into a consumer chatbot can move a codebase out of your control permanently. The blind spot that matters is the local model runtime: an SLM running on a laptop resolves no domain, opens no connection, appears on no invoice and needs no procurement, so every network-side discovery signal misses it – this is shadow AI with weights, and the control is runtime detection on the device. Governance is the Layer 4 half of
ASI10: Rogue Agents, and its load-bearing element is a named owner and review date per service and per agent: an agent with credentials, tools and no accountable human is a rogue agent that has not gone wrong yet. Attacks: Chapter 2, Section 7 · Controls: Chapter 3, Section 6 - Slopsquatting: Registering a package, domain, command or email address under a name that LLMs reliably hallucinate, then waiting for a model to recommend it. Coined by Seth Larson of the Python Software Foundation, blending “AI slop” with typosquatting. MITRE ATLAS models it as a two-step chain –
AML.T0062Discover LLM Hallucinations, thenAML.T0060Publish Hallucinated Entities. What makes it work is consistency rather than frequency: 43% of hallucinated names recur across every run, and 127 names are invented identically by five different models – 53 of them still registrable after coordinated disclosure. So “use a better model” is not a defence, cross-model convergence means one registration catches users of several assistants, and the control is dependency resolution against a verified allowlist. Harder to spot than typosquatting because the name is correctly spelled and confidently recommended. See: Chapter 2, Section 6 - Small Language Model (SLM): Compact models, typically in the 1B-15B parameter range, designed for edge deployment, mobile devices, and resource-constrained environments – smaller but not inherently safer, with unique security challenges. Introduced: Chapter 1, Section 1 · Size tiers: Chapter 1, Section 2 · Security implications: Chapter 2, Section 7
- Special Token Injection: Supplying text containing a model’s own turn-boundary tokens – ChatML’s
<|im_start|>and<|im_end|>, for example – so that a user’s input opens what the model reads as a system turn. Defended in the tokenizer, which refuses to encode special-token strings from untrusted input, rather than by instructing the model to ignore fake role markers. See: Chapter 1, Section 6 · Jailbreaking: Chapter 2, Section 2 - Specialized Models: AI systems designed to excel at specific tasks or domains, often created through fine-tuning or custom training. See: Chapter 1, Section 2
- Stateful and Agentic APIs: The current API generation – OpenAI’s Responses API, Anthropic’s Messages API with server-side conversation state and their equivalents – in which the provider holds the transcript and you send the new message plus a reference to the previous response. Reasoning traces, tool calls and tool results become first-class parts of the exchange rather than text you serialize yourself. Nothing in the calling code looks different and the trust posture has moved: the transcript now lives under someone else’s retention policy with an ID as your only handle on it, and server-held state is state an earlier turn can write into and a later turn reads back, which is the memory poisoning surface. Contrast the stateless chat completions API. See: Chapter 1, Section 6 · Attacks: Chapter 2, Section 5
- Statelessness: The property that every LLM inference is independent – the model retains nothing between requests. Everything called “memory” is therefore an external system that re-supplies text into the context window, which is why memory features are best evaluated on who can write into them rather than on convenience. See: Chapter 1, Section 4
- Streaming: Delivering tokens to the client as they are generated rather than returning a complete response. Cuts perceived latency to near zero and carries a security cost that is easy to miss: a streamed token is published, so output validation and filtering have nothing left to block. Anything feeding a downstream system rather than a human reader should be synchronous. See: Chapter 1, Section 6 · Output validation: Chapter 3, Section 7
- STRIDE: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege – the threat modeling framework this course adapts for AI, category by category. Its value is enforced coverage: six things you cannot skip. Its limit is that it analyses entities with fixed roles across boundaries drawn at design time, which stops describing a system once agents can call tools. See MAESTRO for that case, and pair either with MITRE ATLAS, which is the only one that enumerates techniques. See: Chapter 3, Section 1
- Structured Output Prompting: Requesting a machine-parseable response – JSON, XML, a table – so program code can consume it. Most effective combining task instructions, the exact schema with field types, and a worked example; stronger still where the API supports constrained decoding. Schema conformance is not safety: output crosses a trust boundary into your application. See: Chapter 1, Section 5
- Structured Outputs: See constrained decoding. Worth its own warning: the guarantee is about the envelope.
{"query": "'; DROP TABLE users; --"}satisfies a{"query": string}schema perfectly, and the guarantee makes the object more likely to be passed onward unexamined. Validate the field, not the container. See: Chapter 1, Section 5 · Output handling: Chapter 2, Section 6 - Supply Chain Attack: OWASP
LLM04: Supply Chain. An attack on what goes into the system rather than on the running system – model weights and adapters from a hub, training and fine-tuning datasets, Python packages, the serving infrastructure. The AI-specific additions to a familiar problem are the model hub (malicious uploads, typosquatted model names, weights that fetch more weights) and the fact that loading an artifact can execute code: a pickled checkpoint compromises the host before the model has processed a prompt. See: Chapter 2, Section 3 · Controls: Chapter 3, Section 4 - System Prompt: The standing instructions a chat API accepts separately from each user message, defining an application’s role, rules and business logic. The separation is organisational rather than a privilege boundary: the system prompt lands in the same token sequence as everything else, so it is readable rather than secret and cannot outrank later text. See: Chapter 1, Section 5 · Extraction: Chapter 2, Section 2
- System Prompt Leaking: An attack that extracts the hidden system prompt or instructions of an AI application, revealing proprietary logic, access controls, and business rules. Named as its own OWASP category in 2025 (LLM07); the 2026 edition broadened and renamed it Hidden Context Exposure. See: Chapter 2, Section 2
T
- Temperature: A sampling parameter that reshapes the probability distribution over the next token. Low values concentrate probability on the most likely tokens, giving predictable output; high values flatten the distribution so less likely tokens are chosen more often. Ranges are provider-specific. Even at 0 the same prompt can produce different output, because inference batching makes the arithmetic depend on concurrent load. See: Chapter 1, Section 5
- Threat Modeling: Systematically identifying what can go wrong in a system, assessing the risk and designing mitigations before deployment. For AI the framework choice is itself a decision: STRIDE for an LLM behind an API, MAESTRO once the system is agentic and its trust boundaries move at runtime, and MITRE ATLAS alongside either – a threat model says what could go wrong, ATLAS says what has actually been done. See: Chapter 3, Section 1
- Token: The smallest unit of text that can be processed by an LLM, typically a word, subword, or character depending on the tokenizer. See: Chapter 1, Section 4
- Token Smuggling: Using zero-width characters, homoglyphs, unusual Unicode or split encodings to produce a string a filter reads as harmless and the model’s tokenizer resolves into something else. The general principle it belongs to: a filter that does not tokenize the way the model tokenizes is inspecting a different input. Contrast special token injection, the extreme case, which is defended in the tokenizer rather than the filter. See: Chapter 1, Section 4 · Attacks: Chapter 2, Section 2
- Tokenization: The step that converts raw text into the token sequence a model computes on. A tokenizer is built by measuring which character sequences are frequent in a corpus, so it belongs to a model rather than to the language – two systems will split the same string differently. Any filter or classifier that does not tokenize the way the model does is inspecting a different input. See: Chapter 1, Section 4 · Exploitation: Chapter 2, Section 2
- Tool Misuse: OWASP
ASI02: Tool Misuse and Exploitation. A compromised or hijacked agent using its authorized tools to do harm. Note that nothing malfunctions: the tools are called exactly as designed, with the credentials you granted, which is why the fix is narrowing what they can do rather than detecting misuse. Foundations: Chapter 1, Section 7 · See: Chapter 2, Section 5 · Controls: Chapter 3, Section 4 - Tool Poisoning: Hiding instructions inside an MCP tool’s description, which the model reads to decide when to call the tool and the user never sees. Demonstrated in 2025 with a description instructing the model to read the user’s SSH private key and pass it in an unrelated parameter; the model complied while the interface showed a harmless tool name. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Tool Use (Function Calling): The mechanism by which a model acts: it emits a structured request naming a tool and its arguments, and the surrounding application executes it and returns the result into the context. Two consequences follow. The model never executes anything itself, so authorization belongs in the executing code, not in the model’s decision – and the returned result is untrusted input arriving at the loop’s ingress. See: Chapter 1, Section 7 · Attacks: Chapter 2, Section 5
- Top-K Retrieval: Returning the k highest-scoring passages from a retrieval query. Five to eight is typical after reranking, and more is not safer – a marginal passage displaces attention from a good one, so recall bought past that point costs answer quality as well as tokens. See: Chapter 1, Section 6
- Top-P (Nucleus) Sampling: A sampling parameter that truncates the distribution to the smallest set of tokens whose cumulative probability reaches p, discarding the rest. Not “the top p fraction of tokens”: where the model is confident, p = 0.9 may admit two candidates; where it is uncertain, hundreds. Truncation changes which tokens are eligible, not which are probable – that is temperature’s job – so providers generally advise tuning one and not both. See: Chapter 1, Section 5
- Training: The process where models learn from data through optimization techniques like backpropagation, adjusting parameters to minimize prediction error. See: Chapter 1, Section 4
- Transformers: A neural network architecture that uses attention mechanisms to process sequences in parallel, forming the foundation of modern LLMs. See: Chapter 1, Section 4
- TrendAI Vision One: TrendAI’s unified platform, and where the Blueprint’s six layers report into one console. The capability that is not reducible to the sum of the parts is cross-layer correlation – a Layer 3 posture finding about an over-scoped agent credential and a Layer 6 network detection are two tickets in two tools and one incident in one console. AI Scanner runs from CI/CD as a pre-deployment gate and AI Guard is called inline by your gateway; both feed it alongside AI-SPM findings. See: Chapter 3, Section 2 · Products in the loop: Chapter 3, Section 9
- Trust Boundary (context window): The security boundary of an LLM application. Everything assembled for a request – system prompt, user message, retrieved documents, tool output, stored memory – reaches the model as one flat token sequence with no marker of origin and no enforced priority. Those categories exist in the architecture diagram, not in the model’s input, which is why instructions placed in a system prompt cannot outrank text that arrives from anywhere else. See: Chapter 1, Section 4 · Exploitation: Chapter 2, Section 2
- Trust Boundary (deployment): The line a deployment choice draws between what you control and what you inherit from a provider. Moving across it transfers not just infrastructure but the controls the previous pattern supplied – rate limiting, abuse monitoring, moderation, and logging – none of which arrive with model weights. See: Chapter 1, Section 3
- Trust Gradient: The accumulation of user trust in an automated system through successful operation, and the reason automation bias worsens over time rather than staying constant. A reviewer who verified an agent’s first hundred outputs is markedly less likely to verify the hundred-and-first – which is exactly when one is wrong. The consequence for threat modelling is that a good track record is an attack precondition: OWASP ASI09 lists what the attacker needs as “an agent with a good track record”. No technical signal exists for it, so the control is procedural. See: Chapter 2, Section 5 · Over-trust: Chapter 2, Section 6 · Controls: Chapter 3, Section 6
U
- Unbounded Consumption: OWASP
LLM06, and the category that absorbed model extraction in the 2026 edition – so it covers both paying for compute you did not authorise and losing model value through it. Its forms differ in what they cost you: context-window stuffing and recursive tool loops burn compute, LLMjacking burns credits on stolen keys, and model extraction burns nothing while copying the behaviour. Only the last is invisible to a bill. See: Chapter 2, Section 4 · Framework: Chapter 2, Section 1 · Controls: Chapter 3, Section 7 - Uncertainty Handling: The prompt component that tells a model what to do when it does not know – “if the answer is not in the provided context, say so rather than guessing”. The cheapest single addition that reduces confabulation, and among the most frequently omitted. See: Chapter 1, Section 5
V
- Vector: An ordered list of numbers denoting a point in space, carrying no inherent meaning. All embeddings are vectors; not all vectors are embeddings – what makes an embedding different is that its values were learned so that proximity encodes similarity of meaning. See: Chapter 1, Section 6
- Vector and Embedding Weaknesses: OWASP
LLM09, the category that owns the retrieval layer – RAG poisoning plus three weaknesses of the store itself. Embedding-space manipulation: crafting content whose embedding sits where the attacker wants it, so a chosen query retrieves it. Missing access control: a vector database reachable directly lets an attacker inject or delete embeddings while bypassing every check the ingestion gate performs. Metadata exploitation: because metadata filtering is the only retrieval stage that asks about entitlement, whoever can write metadata changes who can read a document – privilege escalation with no code execution and a field update in the audit log. See: Chapter 2, Section 3 · Pipeline: Chapter 1, Section 6 · Controls: Chapter 3, Section 3 - Vector Database: A store that indexes embeddings for nearest-neighbour search, holding the vector, the original chunk text, and the metadata retrieval filters on. Ranges from dedicated services to extensions of a database you already run, and the latter is the stronger default for most corpora – vectors, text and permissions live in one system you can already audit. Frequently deployed with weaker access controls than the source data, while containing a copy of it. See: Chapter 1, Section 6 · Attacks: Chapter 2, Section 3 · Controls: Chapter 3, Section 3
- Vertical Model: A model tailored to a specific industry or task, reached either by fine-tuning a foundation model or by building one from scratch. It trades breadth for domain precision. Contrast with a horizontal model. See: Chapter 1, Section 2
- Virtual Patching: A network-level rule that blocks exploitation of a disclosed vulnerability without modifying the vulnerable software. It requires a CVE, advisory or published proof of concept, so despite sitting in the zero-day layer it is really zero-patch defense – and the interval it covers in practice is your own upgrade window, not the vendor’s release schedule. It works where exploitation has a distinctive shape on the wire, and poorly where the exploit is a legitimately-formed request from an authorised principal. See: Chapter 3, Section 8
W
- Weights: The learned numbers that set how strongly each input influences each unit in a neural network – the “strength of connection” the model adjusts during training. Together with the bias terms they make up a model’s parameters, and they are what is actually shipped when a model is released as open weights. See: Chapter 1, Section 4
Z
- Zero Trust Secure Access (ZTSA): A security model that enforces identity verification and least-privilege scope on every request to an AI service, regardless of network location. Its scopes are enforced server-side precisely because a scope the model is asked to respect is a suggestion. Because it sits on the path to the service, it is also where usage of unsanctioned third-party AI services becomes visible – a question AI-SPM cannot answer, since posture management reports on assets you own rather than destinations your people reach. See: Chapter 3, Section 7
- Zero-Day Defense: Layer 6 of the Security for AI Blueprint, which owns an interval rather than an asset – the gap between a technique existing and your controls knowing about it. Its two halves have opposite properties: virtual patching blocks but needs the flaw disclosed first, while behavioural anomaly detection handles the genuinely novel and only raises an alert. No control in the layer does both, which is why it is a backstop rather than a superset of the other five layers. See: Chapter 3, Section 8
- Zero-Shot Prompting: Asking a model to perform a task with no examples, relying entirely on what it learned in training. The cheapest technique in prompt tokens and the right default – including when reasoning is enabled, where provider guidance is to try zero-shot before reaching for examples. Its weakness is format drift across requests, which is what few-shot and schemas address. See: Chapter 1, Section 5