Section 3 Quiz
Test Your Knowledge: Data and Training Attacks
Let’s see how much you’ve learned!
Ten questions on data poisoning, backdoors, RAG and feedback-loop poisoning, artifact supply chain, and vector and embedding weaknesses. Several ask you to do more than classify: scope which attacks can reach a given architecture, assess an adapter you did not build, and prioritise remediation when two compromises are live at once.
---
shuffle_answers: true
shuffle_questions: false
---
## A fintech company fine-tunes an LLM on proprietary financial data. During a compliance review, statistical analysis reveals the model consistently steers investment recommendations toward one specific vendor. The investigation traces the bias to crafted examples in the fine-tuning dataset. Which OWASP category does this incident map to?
> Hint: Think about when in the AI lifecycle this attack occurred -- was it at inference time or training time?
- [ ] LLM01: Prompt Injection -- a crafted model input overrode the system instructions at request time
> Prompt injection targets inference-time inputs. This attack compromised the training data, which is a fundamentally different attack vector.
- [ ] LLM09: Vector and Embedding Weaknesses -- adversarial documents were injected into the retrieval corpus
> Vector and embedding weaknesses target the retrieval pipeline. This attack targeted the fine-tuning dataset, not a RAG corpus.
- [x] LLM05: Data and Model Poisoning -- crafted fine-tuning examples baked a persistent bias into the weights
> Correct! This maps to LLM05: Data and Model Poisoning. Content injection (adding carefully crafted examples to training data) teaches the model specific biases. The attack is invisible at the point of compromise, undetectable in standard testing, and persists for the model's entire deployed lifetime. The model passes normal benchmarks because it behaves correctly on everything except the poisoned topic.
- [ ] LLM04: Supply Chain -- an untrusted third-party dataset was used without provenance checks
> While supply chain risks include compromised training data sources, the specific attack here is data poisoning through content injection. LLM05 is the more precise mapping for deliberately crafted training examples.
## A security team discovers a backdoored language model that behaves perfectly on all standard benchmarks but generates biased outputs whenever inputs contain a specific Unicode character pattern. Why is this backdoor particularly difficult to detect?
> Hint: Think about what standard evaluation methods test and what they miss.
- [ ] The backdoor activates only for authenticated users, so anonymous testing cannot reach it
> Backdoor triggers are designed to be specific patterns in input, not tied to authentication. Any user who includes the trigger pattern activates the backdoor.
- [x] Evaluations use held-out datasets that don't contain the trigger, so the model scores well throughout
> Correct! This is the core challenge of backdoor attacks under LLM05: Data and Model Poisoning. The model genuinely performs well on all non-trigger inputs. Evaluation datasets don't contain the attacker's specific trigger because evaluators don't know it exists. Only adversarial auditing that specifically tests for hidden triggers can reveal the backdoor -- creating a significant detection gap between standard evaluation and security testing.
- [ ] The backdoor is encrypted inside the model weights and cannot be read without the key
> Model weights aren't encrypted in a way that hides backdoors. The backdoor is encoded in learned patterns that activate only on specific trigger inputs.
- [ ] Language models discard Unicode control characters before tokenization, hiding the trigger
> LLMs do process Unicode characters. In fact, Unicode tricks are a common vector for both prompt injection and backdoor triggers precisely because models process them.
## The PoisonedRAG research reported roughly a 90% attack success rate from injecting five malicious texts per target question into a corpus of millions. Why does the size of the corpus not dilute the attack?
> Hint: Think about what a similarity search actually returns, and how many documents it ignores.
- [ ] The five documents are ranked first because vector stores return their most recently indexed chunks ahead of older ones
> RAG systems retrieve based on semantic similarity (embedding distance), not recency. If they prioritized recency, the attack would be even simpler, but that's not how vector similarity search works.
- [ ] The attack overwrites the corpus entries that the legitimate answers to that question were previously drawn from
> The attack doesn't replace or delete existing documents. It injects new ones optimized for retrieval on a specific target query, and they outrank the legitimate passages.
- [x] Retrieval returns only the top few most similar chunks, so documents optimized for one query outrank the rest at any corpus size
> Correct! This maps to LLM09: Vector and Embedding Weaknesses. Similarity search returns a handful of nearest neighbours and ignores everything else -- so the count that matters is not "five out of millions" but "five out of the five to eight chunks that actually reach the prompt." Optimizing embedding similarity against one target question wins that comparison no matter how large the corpus is. This is why the attack is priced per *question*, not per corpus, and why dilution is not a defence.
- [ ] Embedding models assign systematically higher similarity scores to any text that was generated or crafted adversarially
> There is no adversarial-content bonus in an embedding model. The documents score highly because they were explicitly optimized to sit near that query's vector, not because of anything the model detects about them.
## Your team wants to use a community LoRA adapter from a model hub. A reviewer confirms it ships as a `.safetensors` file, the hub's scanner flags nothing, and the account has 40,000 downloads. What risk remains unaddressed, and what should the reviewer do next?
> Hint: The file format decides one of the two risks an adapter carries. Ask which one it does not decide.
- [ ] Nothing meaningful remains -- safetensors plus a clean scan and a strong download count covers both risks
> This is the most common error in artifact review. Safetensors closes the code-execution path completely, but it says nothing about what the weights do once loaded. Download count is reputation, not evidence.
- [ ] The residual risk is code execution, since safetensors can still run initialization logic on load
> Safetensors cannot execute code -- that is the entire point of the format. It stores tensors in a flat layout with no deserialization step, which is why it replaced pickle-based formats.
- [x] The weights themselves are still unverified -- run behavioural evaluation and pin the exact revision by hash
> Correct! An adapter carries two independent risks. The code risk is decided entirely by the file format, and safetensors eliminates it. The weights risk is untouched by the format: a perfectly safe-to-load safetensors adapter can carry backdoor triggers or strip the base model's refusal behaviour, and nothing about the file reveals it. Two actions follow -- evaluate behaviour (including trigger hunting, though *Sleeper Agents* shows this is hard), and pin by content hash or revision so that a later poisoned update cannot be pulled silently under the same name.
- [ ] The adapter should be converted to a pickle format so it can be inspected before loading
> This inverts the security property. Pickle formats are the ones that execute code on load; converting toward pickle adds the risk that safetensors removed, and pickle is no more inspectable.
## An attacker compromises a web scraper used to collect training data for an LLM. They inject carefully labeled examples where phishing emails are categorized as "legitimate." This is an example of which data poisoning technique?
> Hint: Think about what specifically was modified in the training examples.
- [ ] Content injection -- adding newly authored biased examples alongside the legitimate training data
> Content injection adds new examples to teach the model specific behaviors. This attack modified the labels on existing examples.
- [x] Label flipping -- changing the labels on existing examples so the model learns wrong associations
> Correct! Label flipping is a data poisoning technique under LLM05: Data and Model Poisoning. By mislabeling phishing emails as "legitimate," the attacker teaches the model to classify malicious content as safe. This attack targets the data source pipeline itself -- compromising the data collection infrastructure rather than the training process directly.
- [ ] RAG poisoning -- injecting adversarial documents into the retrieval corpus at inference time
> RAG poisoning targets inference-time retrieval systems, not training data labels. This attack corrupts training data before the model is ever trained.
- [ ] Backdoor insertion -- planting a rare trigger pattern that activates hidden behavior on demand
> While label flipping can contribute to backdoor-like behavior, the specific technique described is label manipulation, not trigger-based backdoor insertion. The distinction matters for defense strategies.
## The ByteDance intern forged a model checkpoint containing malicious code and used it to interfere with colleagues' training runs. The nullifAI models on Hugging Face carried a payload in a PyTorch checkpoint. What does comparing these two incidents tell you about where to apply artifact controls?
> Hint: Ask what is different about the two attacks, and what is identical.
- [ ] Nothing transferable -- one is an external supply chain attack and the other is an insider threat
> The category labels differ, but the artifact and the mechanism are the same. Sorting these into separate buckets is exactly what leaves the internal path unprotected.
- [ ] The public hub path is the higher risk, so scanning and signing belong at the download boundary
> This is the intuitive answer and it is backwards. The download boundary is the path that already *has* scanning, signing and a review habit around it.
- [x] It is the same poisoned artifact either way, so integrity controls must cover internally produced checkpoints too
> Correct! A malicious checkpoint is a malicious checkpoint whether it arrives from a public hub or from a colleague's directory on a shared cluster. The difference is not the mechanism -- it is that the internal path typically has no scanning, no signing and no review, because everything on the cluster is implicitly trusted. Model integrity verification and provenance tracking (Layer 2) have to apply to artifacts you produce as well as artifacts you download, and the training pipeline needs audit trails that answer "who last wrote this checkpoint."
- [ ] Insider attacks are unpreventable, so detection and response should replace preventive artifact controls
> Insider access does defeat perimeter controls, but artifact integrity is not a perimeter control. Signing and hash verification work precisely because they check the object rather than the network location it came from.
## A company's vector database storing RAG embeddings is deployed without access controls. An attacker gains direct access and modifies metadata fields used for permission filtering. What category of attack is this?
> Hint: Consider what component of the RAG pipeline is being exploited and what the metadata manipulation achieves.
- [ ] LLM05: Data and Model Poisoning -- the model's training corpus was corrupted before training
> Data poisoning targets training data. This attack targets the inference-time vector store, not the model's training pipeline.
- [ ] LLM01: Prompt Injection -- a crafted instruction in the user's text overrode the retrieval filter
> Prompt injection manipulates text inputs to the model. This attack manipulates the vector store infrastructure directly, with no prompt involved.
- [x] LLM09: Vector and Embedding Weaknesses -- unprotected store access let metadata edits bypass permission filtering
> Correct! This maps to LLM09: Vector and Embedding Weaknesses. Vector databases often lack the access controls and auditing capabilities of traditional databases. Metadata exploitation -- manipulating the fields used for filtering or access control -- bypasses security boundaries without touching the embeddings at all. Because metadata filtering is the only stage of retrieval that asks about entitlement rather than relevance, a metadata write is privilege escalation with no code execution.
- [ ] LLM04: Supply Chain -- a compromised build of the vector database software was deployed
> Supply chain attacks target the software components themselves. Here the vector database software is legitimate but misconfigured with insufficient access controls.
## An organization discovers that both their fine-tuning dataset and their RAG document corpus have been compromised. The fine-tuning dataset contains backdoor triggers that activate on specific Unicode patterns, and the RAG corpus contains 5 adversarial documents optimized for high embedding similarity. With resources to address only one threat immediately, which should the team prioritize remediating first?
> Hint: Consider which compromise is harder to detect, harder to reverse, and has broader persistence when evaluating severity.
- [ ] RAG corpus poisoning -- five documents against a corpus of millions is the far more efficient attack
> Efficiency of the attack is not severity of the compromise. RAG poisoning is cheap to mount *and* cheap to reverse: identify the adversarial documents, un-index them, re-embed. The fine-tuning backdoor is baked into the weights and persists across every deployment of that model regardless of retrieval.
- [x] Fine-tuning dataset backdoor -- it lives in the weights, survives safety retraining, and standard evaluation cannot see it
> Correct! Prioritize on persistence, detectability and cost to reverse. The backdoor is encoded in the weights and affects every deployment. Standard evaluation misses it because nobody knows the trigger, and *Sleeper Agents* showed that SFT, RLHF and adversarial training all fail to remove implanted backdoors -- adversarial training actually taught models to hide them better. So remediation means retraining or replacing the model. RAG poisoning is reversible in an afternoon by un-indexing the documents and is containable to the retrieval layer. Address the compromise you cannot undo first.
- [ ] RAG corpus poisoning -- it is affecting real users right now, and active exploitation outranks latent risk
> Immediacy alone does not determine severity under resource constraints. The backdoor also affects inference whenever the trigger appears -- you simply cannot see it. Visible and reversible beats invisible and permanent as a thing to defer.
- [ ] Neither -- the two have equivalent severity profiles and belong in parallel remediation
> These threats have meaningfully different profiles on all three axes that matter. The question also stipulates resources for one; declining to prioritize is how the expensive compromise gets deferred by default.
## According to the October 2025 study by the UK AI Security Institute, the Alan Turing Institute and Anthropic, roughly how many malicious documents were needed to backdoor the models they tested, and how did that number change with model size?
> Hint: The headline finding is about what the number does *not* depend on.
- [ ] Around 250 documents for the smallest model, scaling proportionally with parameter count
> This is the intuition the study overturned. The number did not scale -- the same count worked across a more than twenty-fold range in model size.
- [ ] A fixed 0.1% of the training corpus, so larger models required proportionally more poison
> No percentage threshold was found. For the 13B model, 250 documents were about 0.00016% of its training data -- far below any such figure.
- [x] Around 250 documents, and the count stayed near-constant from 600M up to 13B parameters
> Correct! The study found that poisoning attacks require a near-constant *number* of documents rather than a percentage of the corpus. About 250 documents backdoored every configuration tested, from 600M to 13B parameters -- 0.00016% of the largest model's training data. Note the authors' own caveat: they implanted a narrow, measurable backdoor (gibberish on a trigger), and state it is unclear whether the same dynamic holds for complex behaviours such as backdooring code.
- [ ] Around 250,000 documents, which is why only well-resourced actors can poison pretraining
> The finding is the opposite of a resource barrier. Producing 250 documents is trivial, which is precisely what makes the result significant.
## A customer support assistant lets users rate each answer, and highly-rated conversations are sampled monthly into a preference-tuning dataset that fine-tunes the model. Which risk does this pipeline introduce, and where does the control belong?
> Hint: Ask who can write to the training data, and note that Section 1's lifecycle diagram loops backwards.
- [ ] No new risk -- ratings are aggregated signals, and individual users cannot influence what the aggregate says
> Aggregation is not a control at this scale. The poisoning result above is a count of documents, not a share of the dataset, so a scripted client submitting rated conversations reaches the threshold without ever dominating the average.
- [ ] Sensitive information disclosure -- the fix is to redact conversation logs before they enter the dataset
> Redaction is genuinely necessary for a different reason, but it addresses what leaks *out* of the dataset. It does nothing about an attacker deliberately writing *into* it, which is the risk the feedback loop creates.
- [x] End users now have a write path into training data -- gate the feedback store with human review and per-example provenance
> Correct! The feedback loop converts an inference-time attack into a training-time one, which is exactly why Section 1's lifecycle diagram loops from monitoring back to data collection. Ordinary unauthenticated users control this write path by design, they submit pre-labelled examples (label flipping through the front door), and the data is trusted precisely because it came from users. Rate limits, deduplication and outlier detection help, but the load-bearing controls are a human review gate between the feedback store and the training set, plus provenance on every example so a bad training run traces back to the accounts that fed it.
- [ ] Prompt injection -- the fix is to filter user-submitted text for adversarial instructions before storing it
> Instruction filtering targets the wrong payload. Feedback poisoning does not need injected commands; ordinary, well-formed, plausible conversations with a skewed rating are enough, and no filter distinguishes those from genuine feedback.