Section 4 Quiz

Test Your Knowledge: Model and Infrastructure Attacks

Let’s see how much you’ve learned!

This quiz tests your understanding of serialization exploits and what a safetensors migration does not fix, adversarial perturbation, model extraction exposure, unbounded consumption, the attack surface below the model, and how to prioritise across all five.

--- shuffle_answers: true shuffle_questions: false --- ## A healthcare startup downloads an open-source model from a hub and loads it using `torch.load()`. Three weeks later, their security team detects a reverse shell connecting to an external server. What vulnerability was exploited? > Hint: Think about what happens at the code level when Python deserializes a pickle file. - [ ] A model backdoor in the weights, triggered by specific medical queries > A backdoor is trained into the weights and affects what the model *outputs* when a trigger appears in a prompt. It cannot open a network socket. A reverse shell means code ran, not that the model behaved badly. - [x] A pickle serialization exploit -- code in the model file ran on deserialization > Correct! This maps to LLM04: Supply Chain. A pickle file is a program for a small stack machine inside Python, and its `GLOBAL`/`REDUCE` opcodes call any named function with supplied arguments. `torch.load()` is the interpreter, so loading the file *is* running it. Note the timeline in the scenario: the compromise happened at load time, before a single prompt was ever sent. The safetensors format removes this by storing tensor data with no mechanism for executing code. - [ ] The model's API endpoint was left exposed without authentication > An unauthenticated endpoint is a real infrastructure risk, but the scenario places the compromise at model load -- before any endpoint was serving traffic. - [ ] An adversarial input made the model generate a reverse shell command > An LLM emits text. Text becomes a command only if something downstream executes it, which is Improper Output Handling. Here the executable code was already inside the model file. ## An attacker sends repeated maximum-length (200,000 token) inputs to an organization's hosted frontier-model endpoint, running up $50,000 in charges overnight. Which OWASP category does this map to? > Hint: Consider the resource asymmetry -- it's cheap to send a request but expensive to process one. - [ ] LLM01: Prompt Injection -- the attacker manipulated the model's behaviour > Prompt injection targets what the model does with instructions. This attacker is indifferent to the output; the cost of producing it *is* the attack. - [ ] LLM04: Supply Chain -- a component of the deployment was compromised > Supply chain covers compromised artifacts and dependencies reaching your pipeline. Nothing here was compromised -- the API worked exactly as designed, which is the problem. - [x] LLM06: Unbounded Consumption -- the attacker exploited the cost asymmetry of inference > Correct! This is context window stuffing. A 200,000-token request costs far more to serve than a 100-token one, and because attention cost grows faster than linearly in sequence length, the ratio is worse than the token count suggests. This is the same category the Storm-2139 LLMjacking case from Section 1 mapped to -- there the attackers used stolen keys rather than long prompts, but the impact was the same: consumption at the victim's expense. The control that bounds the loss is a spend cap with an alert below it. - [ ] LLM05: Data and Model Poisoning -- the training corpus was corrupted > Poisoning targets the training pipeline to change the model. This attack targets the inference API's resource consumption and leaves the model untouched. ## Why is deserializing a pickle file equivalent to executing code, rather than being a bug that could be patched? > Hint: Think about what pickle was designed to be able to reconstruct. - [ ] Pickle files are compiled bytecode, so Python runs them like any `.pyc` module > Pickle is not compiled Python bytecode. It is a separate opcode stream for a purpose-built stack machine, most of whose instructions build data structures. - [ ] A buffer overflow in the pickle parser lets crafted files jump into attacker data > There is no memory-safety bug involved. The behaviour is documented and intentional, which is exactly why it cannot be patched away. - [x] Pickle's opcodes can name any importable callable and invoke it with supplied arguments > Correct! Pickle has to be able to rebuild arbitrary Python objects, so its `GLOBAL` and `REDUCE` opcodes look up a named callable in any importable module and call it. Any class can define `__reduce__` to say "to rebuild me, call this function with these arguments" -- an attacker returns `os.system` and a command string. The file is *valid* and pickle is behaving as documented, so there is nothing to fix. That is why the mitigation is a format change (safetensors) rather than stricter validation. It also explains why a `torch.load()` that ends in a traceback proves nothing: opcodes at the head of the stream run before the parser reaches the corruption. - [ ] Model files bundle a `setup.py` that pip executes during installation > That is a package-installation risk, and a real one. It is unrelated to what happens inside `torch.load()` on a checkpoint file. ## The n8n workflow automation platform disclosed CVE-2025-68613 (CVSS 9.9). What is the vulnerability, and why is it especially serious for an AI agent orchestration layer? > Hint: Think about what n8n lets a workflow author write, and what n8n stores in order to do its job. - [ ] A missing authentication check that let unauthenticated users trigger any workflow > Exploitation requires an authenticated account. The severity comes from how *little* privilege that account needs, not from needing none. - [x] Authenticated RCE via expression injection -- and the orchestrator holds every credential > Correct! n8n evaluates author-written expressions when a workflow runs, and those expressions escaped their sandbox into the Node.js runtime, reaching core modules and global objects -- arbitrary OS command execution as the n8n process. It needs only permission to create or edit a workflow, which in most deployments is the whole engineering team. It matters disproportionately because an orchestrator's job is to hold the credentials for everything the AI system touches: model API keys, database strings, internal service tokens, cloud roles. RCE there is not one compromised app; it is the credential store for the entire AI estate, plus the network position those credentials are used from. Treat workflow-edit permission as equivalent to shell access. - [ ] A prompt injection flaw allowing payloads into any LLM node in a workflow > This is a flaw in n8n's own expression evaluation, not in how it passes text to a model. The attacker gets code execution on the host, which is strictly worse than influencing a prompt. - [ ] Server-Side Request Forgery, letting the server reach internal-only services > A plausible-sounding orchestrator bug, and the wrong one. SSRF would let an attacker make the server issue requests; this vulnerability gives direct command execution on it -- no pivot required. ## An attacker queries a proprietary model's API with thousands of strategically chosen inputs, records the outputs, and trains a local model to replicate its behaviour. What is this, and what makes it hard to recover from? > Hint: Think about what the attacker can do afterwards that you can no longer see. - [ ] Data poisoning -- repeated querying gradually corrupts the target model > Queries do not modify the target. Information flows out of the model here; nothing flows into it. - [x] Model extraction -- the resulting clone is permanent and invisible to your logging > Correct! Model extraction maps to LLM06: Unbounded Consumption in the 2026 edition, the same place model theft sits. The recovery problem is the important half: once the clone exists it runs on the attacker's hardware, so every detection and rate-limiting control you own stops applying to it. They can then attack the clone with full white-box access and transfer the findings back to your model -- MITRE ATLAS `AML.T0043.001`, black-box transfer, where the clone *is* the proxy. Extraction efficiency depends on what your API returns: exposing token log-probabilities cuts the required query count by orders of magnitude. - [ ] Adversarial input generation -- strategic inputs crafted to force incorrect outputs > The inputs are strategic, but the goal is faithful replication, not error. An extraction campaign wants the model to behave *correctly* as often as possible. - [ ] Prompt injection -- crafted inputs override the model's system instructions > Injection overrides behaviour during one request. Extraction copies behaviour permanently and never needs the model to misbehave at all. ## Which statement correctly describes adversarial perturbation, and how does it differ from the filter-evasion techniques in Section 2? > Hint: Both get a payload past a filter. Ask what the filter was looking at in each case. - [x] Minimal input changes that cross a model's decision boundary while looking unchanged > Correct! Models partition a high-dimensional space into labelled regions; an adversarial input is moved just far enough to change its label while staying close to the original. The distinction from Section 2 matters: encoding, Unicode smuggling and cross-modal payloads defeat a filter by making it read the wrong bytes, whereas perturbation defeats a filter that reads exactly the right bytes and classifies them wrongly. Note also that this attack has no OWASP LLM Top 10 category -- its identifier is MITRE ATLAS `AML.T0043`, Craft Adversarial Data. It is most relevant to LLM deployments because your input moderation layer is itself a classifier with boundaries to optimise against. - [ ] Sending high volumes of malformed text until the inference server crashes > Volume and malformed input are resource exhaustion (LLM06). Perturbation is precise by definition -- the input must still look normal. - [ ] Extremely long prompts that overflow the model's usable context window > That is context window stuffing, an unbounded-consumption technique. Perturbation is about precision, not size. - [ ] Encrypting or encoding a malicious prompt so the model decodes and obeys it > That is one of Section 2's filter-evasion families. The payload's meaning is hidden from the filter rather than being classified wrongly by it. ## A security team assessing their self-hosted deployment confirms three findings: (1) model files use pickle, loaded via `torch.load()`; (2) their LLM API lacks per-user rate limiting; (3) a competitor appears to be querying the API to extract model behaviour. With resources for one, which should they fix first? > Hint: Sort by what the attacker ends up holding, and by which outcomes can still be undone afterwards. - [ ] The rate limiting gap, because it affects every user immediately and costs money now > Unbounded consumption is the loudest of the three and the cheapest to bound, but it compromises neither integrity nor confidentiality, and the damage is money -- which is recoverable. Adding a quota and a spend cap after the fact fully resolves it. - [ ] The extraction campaign, because lost model IP is the largest long-term business risk > Extraction is serious and, unlike consumption, irreversible. But it is a confidentiality loss that gives the attacker no control over your infrastructure. Ranked against remote code execution it is the lesser outcome. - [x] The pickle loading, because it is the only one of the three that yields code execution > Correct! Sort by what the attacker ends up holding: code execution beats a behavioural clone beats a large bill. A pickle deserialization exploit runs arbitrary code with the serving process's privileges -- reverse shell, credential theft, patient data exfiltration, lateral movement -- and it lands before inference (at load time), so no prompt-level control sees it. Rate limiting gaps cost money (recoverable). Extraction costs IP (irreversible but contained). RCE costs the infrastructure, and it also subsumes the other two: whoever owns the host can drain the budget and copy the weights outright. It is also the cheapest to prevent, since migrating to safetensors removes the execution path entirely. - [ ] All three at once, since prioritisation is unnecessary for confirmed findings > With finite resources, sequencing *is* the decision. These three have materially different severity and different reversibility, and treating them as equivalent means the highest-severity finding waits as long as the lowest. ## A multi-tenant inference service shares a prefix cache across users so that repeated prompt openings are not recomputed. Researchers show an attacker can recover another tenant's system prompt. What is the mechanism, and why can it not simply be patched? > Hint: The attacker is not reading memory. They are measuring something. - [ ] Unbounded consumption -- the first tenant used more GPU memory than their quota > Resource exhaustion is about capacity. This is a confidentiality failure, and it occurs at completely normal usage levels. - [x] A timing side channel -- cache hits speed up first-token latency, revealing prior prompts > Correct! When a submitted prefix hits the cache, time-to-first-token drops measurably, so an attacker submits a guess and times the response: fast means somebody already asked it. Published research reports cache hit/miss detection at 99% accuracy and recovers a confidential system prompt at roughly 111 queries per token. It resists patching because the performance benefit and the leak are the same mechanism -- vLLM's CVE-2025-46570 was mitigated by scoping cache sharing (per-tenant caches, cache salting) at a real performance cost, not by removing a bug. The generalisation is worth more than the CVE: any cross-request optimisation in a multi-tenant AI system is a candidate side channel. - [ ] Prompt injection -- the earlier tenant's cached instructions steer later completions > Injection needs the attacker's text to reach the model as instructions. Here nothing is injected; information flows out through response *timing*. - [ ] A container escape allowing one tenant's process to read another's GPU memory > A real class of attack against shared GPU infrastructure, and not this one. No isolation boundary is breached -- the attacker only measures how long a legitimate response took. ## An organization migrates every model artifact from pickle checkpoints to safetensors and closes the finding as remediated. Their stack runs a self-hosted inference server behind an orchestrator. What should the assessor say? > Hint: Ask what else in a serving stack calls a deserializer on input it did not produce. - [ ] Remediated -- safetensors cannot execute code, so the deserialization risk is closed > This is the conclusion the ShadowMQ research exists to refute. Safetensors closes the weights file, which was one of several deserialization surfaces. - [x] Incomplete -- the inference server's internal IPC may still deserialize with pickle > Correct! Safetensors secures the tensor blob, not the stack. The ShadowMQ cluster found the same ZeroMQ `recv_pyobj()` call -- which deserializes with pickle -- in Llama Stack (CVE-2024-50050), vLLM, NVIDIA TensorRT-LLM, SGLang, Modular Max, and Microsoft Sarathi-Serve. It spread by copy-paste; SGLang's vulnerable file opens with "Adapted from vLLM." The remaining surfaces to check are the internal sockets between scheduler and workers, the tokenizer and config files shipped beside the weights, `trust_remote_code=True` paths that fetch and run Python from a hub, and pickle-backed caches in training tooling. The reframed question is *"what in our stack calls a deserializer on input it did not produce?"* -- and note that you inherit bug classes along the paths your dependencies were copied from, not the paths your own code was reviewed. - [ ] Remediated, provided the safetensors files are also hash-verified against the publisher > Hash verification is a genuinely good control and belongs in the pipeline. It says nothing about a socket that accepts pickled messages. - [ ] Incomplete -- safetensors is slower to load, so the change should be reverted > Safetensors supports fast zero-copy loading; performance is not the objection. And a security control is not undone because of load time. ## Ultralytics v8.3.41 and v8.3.42 shipped a cryptominer after an attacker poisoned the build cache through CI script injection. Maintainers published a clean v8.3.43. What was the most important remaining action, and why? > Hint: Consider everything the attacker could read from inside a compromised CI run. - [ ] Rewrite the git history to remove the malicious commits from the repository > There were no malicious commits. The source was never modified -- the payload was inserted during the build, which is why source review would not have caught it. - [ ] Publish a security advisory and ask users to pin to the last known-good version > Necessary, but pinning was not sufficient here: two of the four malicious releases came *after* the clean one, so a user pinning to the newest version still got malware. - [x] Rotate every credential reachable from the compromised CI run, before shipping the fix > Correct! This is the step that did not happen in time. Within 36 hours the attacker published v8.3.45 and v8.3.46 straight to PyPI using an API token stolen from the compromised CI/CD run -- no second exploit required. The lesson is that a clean release does not mean a clean pipeline: treat every secret reachable from a compromised run as compromised and rotate before publishing remediation, or the attacker simply reuses the access you left them. Two further generalisations: the artifact can be malicious while the source is clean, which is what build provenance exists to detect; and your CI is production, because a workflow interpolating a branch name into a shell string is remote code execution in a system holding publishing credentials. - [ ] Add the malicious versions to a dependency scanner's deny list and continue releasing > Detection of known-bad versions helps downstream users. It does nothing about the attacker's retained access to the publishing pipeline, which is what produced the second wave.