Section 7 Quiz

Test Your Knowledge: Small Language Model (SLM) Threats

Let’s see how much you’ve learned!

This quiz covers SLM-specific vulnerabilities, why the “smaller = safer” assumption fails, edge deployment risks, supply chain transformations, and amplified model extraction – then closes with two questions that draw across all seven sections of Chapter 2.

--- shuffle_answers: true shuffle_questions: false --- ## Your team is choosing between two instruction-tuned SLMs for an on-device deployment. Model A is 7B, Model B is 3B, and both score similarly on the capability benchmarks you care about. A colleague argues Model A is the safer pick because it is bigger. What is wrong with that reasoning? > Hint: What did the correlation analysis in the reference study find actually predicts an SLM's jailbreak resistance? - [x] Robustness tracks training data and alignment pipeline, not size -- so only testing both builds answers it > Correct! This is the central finding of Wang et al.'s evaluation of 59 SLMs: vulnerability correlates with training details rather than model size scaling. Robustness splits by family -- Llama-3.2, Phi-3, Qwen, Gemma and MiniCPM hold up at sizes where other families collapse, and StableLM-2-1.6B complies with 75.7% of direct harmful queries. Size only predicts robustness *within* a family, where the training recipe is held roughly constant. Across families it predicts nothing, so the only way to answer the question is to run the safety evaluation on both builds. - [ ] Nothing is wrong -- across every model family, a 7B model reliably resists jailbreaks better than a 3B one > This is the assumption the study was designed to test, and it failed. Size predicts robustness within a single family, not across different families with different training pipelines. - [ ] Model A is the riskier pick, because larger SLMs are consistently harder to align than smaller ones > This inverts the claim without fixing it. The error is treating size as the predictor at all, in either direction. - [ ] The comparison is invalid, because the OWASP LLM Top 10 does not apply to models below 10B parameters > The framework is model-size agnostic and applies to both. Nothing about the comparison is invalid -- the reasoning behind it is. ## An SLM that refused a carefully constructed multi-turn escalation attack answers the same harmful request when it is simply asked directly. Why does this ordering, which is the reverse of the frontier-model pattern, occur? > Hint: A multi-turn attack asks something of the model. What? - [x] Escalation needs the model to follow the attacker's logic across turns, and a weak model simply cannot > Correct! Multi-turn attacks such as Crescendo largely fail against SLMs, and not because of any defence. Model capability (MMLU) correlates negatively with success on direct attacks and positively with success on multi-turn ones, so the sophisticated attack is defeated by incapacity rather than by alignment. Meanwhile 37.3% of tested SLMs failed on more than half of plainly-stated harmful queries. This inverts the red-team ordering: against an SLM, start with the obvious attack. - [ ] The model carries a specific defence against multi-turn escalation that a single-turn direct query bypasses entirely > There is no such defence. The multi-turn attack fails because the model cannot follow it, which is a capability limit rather than a safety control. - [ ] Direct queries arrive in one turn, and single-turn inputs are not routed through the safety classifier > Edge SLMs typically have no separate safety classifier at all, and the refusal behaviour that does exist is in the weights, not in a per-turn filter. - [ ] Short context windows make a model more resistant to escalation, and long ones make it less resistant > The opposite holds. Phi-3-mini at 128k is about 15% more susceptible to long-context jailbreaks than the same model at 4k, because it can process the whole adversarial prompt. ## An attacker steals a proprietary 3B parameter SLM from an edge device, removes its safety guardrails on a consumer GPU in hours, and distributes the uncensored version through file sharing. Why is model theft particularly amplified for SLMs? > Hint: Compare the theft-to-weaponization pipeline for a 3B model versus a 400B model. - [ ] SLM weights ship unencrypted on edge devices, whereas large model weights are always encrypted at rest on disk > Encryption at rest varies by device and is sometimes skipped for performance, but it is not the distinguishing factor. The amplification runs across the whole pipeline, not just the storage format. - [ ] SLMs hold more valuable proprietary data than large models, so the prize is worth more to an attacker > The value of what is embedded in a model is not size-dependent. The amplification is about how much cheaper each step of the pipeline becomes, not how much the prize is worth. - [x] Every step is cheaper -- a 1-15 GB file copies out in minutes, and consumer hardware runs the result > Correct! Model extraction sits under LLM06 in the OWASP LLM Top 10, and edge deployment collapses it from an attack into a file copy -- the weights are already on hardware the attacker holds, so there is nothing to reconstruct. Exfiltration takes minutes rather than the hours or days a 400 GB model needs, a consumer laptop runs the result with no GPU cluster, safety removal costs tens of dollars rather than thousands, and distribution needs no specialised infrastructure. Nothing after the copy can be undone, which is why the controls that matter are all applied before shipping. - [ ] SLMs are the only category of model ever deployed onto edge devices, so only they are exposed this way > Quantized larger models are also deployed at the edge. The amplification comes from the technical characteristics -- small file, low hardware requirement -- not from exclusivity. ## An edge-deployed SLM is jailbroken by a user, and the company wants to investigate how many interactions were affected. They discover there are no logs of the on-device interactions. Which edge deployment vulnerability does this illustrate? > Hint: Think about what security infrastructure edge devices typically lack compared to cloud deployments. - [ ] Physical access to weights -- a user holding the device extracted and then modified the deployed model > Physical access is about stealing or altering the model itself. This scenario is about being unable to reconstruct what happened afterwards. - [x] Limited monitoring -- inference crosses no network boundary, so no interaction log was ever produced > Correct! Cloud-deployed models get centralized logging, real-time alerting and anomaly detection for free, because every request crosses a boundary you control. On-device inference crosses nothing. Without interaction logs the company cannot determine how many conversations involved jailbroken outputs, whether clinical decisions followed from them, or how long it had been happening. This one is compensable, unlike rate limiting: an on-device audit log that syncs opportunistically, with defined retention and a gap-detection alert, restores most of the capability -- a Layer 4 control. - [ ] Resource constraints -- the device lacked the compute headroom to run a safety filter alongside the model > Resource constraints govern what the device can run during inference. This scenario is about the absence of an audit trail after the fact, which is a separate gap. - [ ] Quantization -- the model's safety training was degraded when it was compressed to fit on the device > Quantization changes model behaviour. This scenario is about missing logging infrastructure, and would look identical whether or not the model had been quantized. ## A vendor ships a small reasoning model distilled from a much larger one. It matches its parent on the published reasoning benchmarks, but answers harmful requests its parent refuses. What does this tell you about distillation? > Hint: Distillation trains the small model to imitate outputs the large model produced. Which outputs were selected? - [ ] The vendor deliberately stripped out the safety training in order to reduce the model's shipped file size > Refusal behaviour is encoded in the weights and removing it saves nothing. This is an omission during derivation, not a deliberate strip. - [x] Distillation transfers what its training data demonstrates, and that data was chosen for capability > Correct! This is the DeepSeek-R1-Distill pattern. The distillation set was reasoning traces chosen to transfer reasoning ability; safety was never an objective, so it was not preserved. The distilled 1.5B and 7B models reach direct-harmful-query attack success rates of 0.271 and 0.514, substantially worse than the base models they came from, and their reasoning traces open by treating a request they should refuse as a problem to solve. Because distillation is how most SLMs are made, benchmark parity with the parent is evidence only about what distillation optimised for -- and no evidence at all about what it did not. - [ ] Benchmark scores are unreliable in general, so the claim of reasoning parity is probably false too > There is no reason to doubt the capability result. The point is narrower and more useful: capability and safety are separable properties, and only one of them was carried across. - [ ] The parent model must have been compromised at some point before the distillation was run > Nothing was compromised. The models shipped this way from a reputable lab, which is precisely what makes the pattern worth knowing. ## Your team wants to skip the safety evaluation on a quantized build, arguing that the full-precision model already passed and the research shows quantization does not hurt safety. How should you respond? > Hint: Consider both what the evidence says and what follows from it regardless of direction. - [ ] They are wrong -- quantization reliably degrades safety, so the quantized build will perform worse > This overstates the evidence in the other direction. The SLM-specific study found quantization slightly *improved* robustness, with AWQ reducing attack success rate by 15.9%. - [ ] They are right on the evidence, so the full-precision result can simply be carried forward > The premise is defensible but the conclusion does not follow. A result inherited from a different artifact is not a result about the one being shipped. - [x] The direction is contested, which is why no result transfers -- evaluate the artifact you ship > Correct! Wang et al. found quantization slightly improved robustness across the Qwen2.5 series, while other 2025-2026 work reports post-training quantization exposing safety failures absent at full precision. Both can hold, because the effect is method-dependent. The operational rule survives either way and is the part that matters: safety is a property of a specific build, not of a model family or a precision level. Inheriting a safety result across a transformation is the error, in either direction. - [ ] Quantization affects only task accuracy, so a capability regression test is sufficient > Capability and safety degrade independently -- that separability is the recurring lesson of the whole supply chain, and a capability test would not detect a safety change. ## An architect proposes putting a proprietary fine-tuned SLM on customer-owned hardware, and asks which risks can be mitigated. Which item on your list is a constraint on the architecture rather than a risk to mitigate? > Hint: Two of the controls lost at the edge have no compensating equivalent at all. - [ ] Missing interaction logs, because on-device inference crosses no boundary you control > Compensable. An on-device audit log that syncs opportunistically, with defined retention and gap detection, restores most of the capability as a Layer 4 control. - [ ] Slow patching, because model updates depend on device connectivity and on user consent > Compensable. Signed updates with a maximum-staleness policy that disables the feature rather than running a stale build covers this. - [x] Model confidentiality -- shipping weights to the device publishes them irreversibly > Correct! Weight confidentiality and rate limiting are the two controls with no edge equivalent. Once the weights are on hardware you do not control, extraction is a file copy, and no Chapter 3 control recovers a distributed model. That makes it a constraint on what the architecture may carry, not a risk to be accepted with mitigations: either the proprietary value comes out of the weights, or the deployment model changes. Everything else on the list is a cost you can choose to pay. - [ ] Absent input filtering, because there is no proxy in front of the model to filter at > Partially compensable. Constraining the input surface -- fixed prompt templates, no free-text passthrough -- recovers much of it as a Layer 5 control. ## Which statement about the OWASP LLM Top 10 and SLMs is correct? > Hint: Does the framework change for small models, or does only the manifestation change? - [ ] Only five of the ten categories apply to SLMs; the remaining five are exclusive to large models > All ten apply. The framework is model-size agnostic. - [ ] SLMs face lower severity across all ten of the categories, being less capable than frontier models > This is the "smaller = safer" assumption. Most categories are higher severity for SLMs. - [x] All ten apply, and most manifest at higher severity -- notably injection, supply chain, extraction > Correct! Every category applies with SLM-specific characteristics. LLM01 is higher for simple attacks though lower for elaborate ones, LLM04 is higher because provenance is a chain of four transformations rather than a single source, LLM05 is higher because the poisoning threshold is lower, and LLM06 is much higher because edge deployment reduces extraction to a file copy. SLMs are not a safer subset of AI. - [ ] SLMs require an entirely separate framework, as the OWASP LLM Top 10 does not cover them > No separate framework is needed. The manifestations differ; the categories do not. ## Across the whole chapter, prompt injection, RAG poisoning, tool misuse and improper output handling all trace back to one underlying property of AI systems. Which? > Hint: Ask what these attacks have in common at the point where they succeed, not at the point where they enter. - [ ] AI models are probabilistic, so their outputs can never be predicted or tested reliably at all > Non-determinism makes testing harder, but a deterministic system with the same trust confusion would be just as exploitable. - [x] AI systems cannot reliably separate instructions from data, so untrusted content becomes instruction > Correct! This is the thread running through the entire chapter. Prompt injection puts attacker text where instructions are read; RAG poisoning routes it in through retrieved documents; tool misuse arrives through tool descriptions and results; improper output handling is the same confusion at the downstream boundary, where a consuming system treats generated text as trusted. Every integration point is a trust boundary, and every feature that admits content is a potential attack surface. - [ ] Training data is scraped from the open public internet and so can never be fully trusted > That explains poisoning at training time, but not injection, tool misuse or output handling, which all occur at inference. - [ ] Models are too large for a human to audit, so vulnerabilities cannot be found by inspection > Scale complicates interpretability, but these attacks are all understood mechanically and none of them requires opacity to work. ## You are handed an AI system to assess. Which fact should you establish first, because it determines which of the chapter's defences are even available to you? > Hint: Sections 1 through 6 shared an assumption that Section 7 removed. - [ ] Which OWASP LLM Top 10 categories the system's existing threat model already documents > Useful context, but a threat model describes intent. It does not tell you which controls can physically be applied. - [ ] Which foundation model family and version the application has been built on top of > Relevant to supply chain and capability questions, but it does not change which control points exist. - [x] Where the model executes, and whether you control the boundary sitting in front of it > Correct! Sections 1 through 6 assumed a model behind an API you control -- which is what gives you somewhere to filter input, inspect output, count requests and write logs. That is a deployment property, not a property of AI. On a device you do not control, roughly half the toolkit has nowhere to run, and defences that were routine become impossible. Establish the execution boundary first; it determines which of the remaining questions are even worth asking. - [ ] Whether the system has agentic capability and can call tools autonomously > This is the right *second* question -- it determines blast radius, as Section 5 showed. But where the model runs decides what you can do about it.