4. Layer 2: Secure Your AI Models
Introduction
Section 3 handed two things forward: a poisoned checkpoint, and a pretraining corpus nobody can audit. Both arrive here for the same reason – they are already inside the artifact by the time anyone sends a prompt, and data governance has no reach into a file that has already been built.
That is Layer 2’s whole situation, and it produces an unusual pair of properties. Section 2 rated this layer the narrowest in the Blueprint – named as a primary defense in only 3 of the 20 OWASP categories, the easiest layer to defer – and in the same breath called it the only layer that acts before anything is running. Those are not in tension. They are the same fact:
Layer 2 defends the smallest number of things, and it is the only layer that defends them at all. If you deploy a backdoored model, no amount of Layer 5 filtering recovers the situation, because the compromise is inside the thing the filter is protecting.
So the question this section answers is not “how do we scan models.” It is the one Chapter 2 Section 3 left on the table. Of its four scoping questions, three were Layer 1’s and one was not:
Who can push an adapter into our pipeline?
Generalise it and you have this layer: what gets to become the artifact you serve, and what did you check before it did?
What will I get out of this?
By the end of this section, you will be able to:
- Locate Layer 2 among the six, and state why the layer with the narrowest coverage is the one you cannot compensate for elsewhere.
- Determine which Layer 2 controls you actually own for a given deployment pattern, and recognise the pattern that leaves you a single control, the one that hands you the whole layer at once, and the one with no recovery path.
- State what each integrity control proves and what it does not – format, hash, signature, provenance, scanner verdict – and pick the combination a given source requires.
- Separate the code risk from the weights risk in any model artifact, including adapters, and explain why a signed, scanned, safetensors model can still be fully backdoored.
- Specify a pre-deployment gate covering serialization safety across the whole serving stack, container scanning, and behavioural evaluation.
- Map Layer 2 to its three OWASP categories, and for each one name the layer that completes the defense Layer 2 cannot finish.
What Layer 2 Owns – and What It Cannot Stop
Layer 2 is purely preventive, and never in the request path.
- Preventive (build-time), without exception. Format enforcement, hash pinning, signature verification, provenance records, image scanning, behavioural evaluation. Every one of them runs before the model serves traffic. By the time a request exists, the artifact is either clean or it is not, and nothing in this layer will ask again.
- No detective component in production. Layer 1 at least re-verifies hashes on a schedule. Layer 2’s equivalent – periodic re-scanning – tells you that an artifact you deployed six months ago fails today’s tests. That is a finding about a running system, and the response is a redeployment, not an interception.
- No in-path exception at all. This is the flat “No” in Section 2’s comparison table, and it is worth being precise about why, because a common architecture diagram gets it wrong.
Runtime policy is Layer 3, not Layer 2
Container security pipelines are usually drawn with a fourth gate at the deployment step labelled runtime policy – seccomp profiles, read-only mounts, egress rules, admission control. Those controls are real and they matter, but attributing them here inflates what this layer can do.
Layer 2’s container work ends when the image is signed and admitted to the registry. What the container is then allowed to do at runtime is Layer 3’s job, which is why Chapter 2 Section 4 splits the same case study across two layers: scanning the image before it is scheduled is Layer 2, multi-tenant GPU isolation and runtime hardening are Layer 3.
The distinction pays off during an incident. If a serving container is behaving unexpectedly right now, no Layer 2 control is going to help you, and knowing that immediately is worth more than a diagram that suggests otherwise.
The trade Layer 2 makes is the same one Layer 1 makes, one stage later. It cannot intercept anything – and in exchange, an artifact it rejects never becomes a running system for anyone.
Which Layer 2 Controls Are Yours
Chapter 1 Section 3 promised that “the layers stay constant, but which ones are yours is set by the choice you make here.” Layer 1 varied a lot by deployment pattern. Layer 2 varies more than any layer in the Blueprint, because the artifact itself either passes through your hands or it never does.
Read this against the five deployment patterns from Chapter 1.
| Layer 2 control | Cloud API | Serverless inference | Self-hosted | Edge / on-device |
|---|---|---|---|---|
| Base model artifact | Provider – you never hold the file | Provider’s catalogue, your tenancy | You – you download, verify and load it | You, and you then publish it |
| Format enforcement | N/A | Provider’s catalogue format | You – the whole control is yours | You |
| Hash / signature verification | N/A | Partly – what the catalogue exposes | You | You, plus a client-side integrity check |
| Fine-tuned checkpoints and adapters | You, if the API offers fine-tuning | You | You | You |
| Serving container image | Provider | Provider | You | The application binary is yours |
| Serving-stack dependencies | Provider | Provider | You – the whole stack | You |
| Pre-deployment behavioural evaluation | You – it is your only Layer 2 control | You | You | You, and the last chance you get |
| Model confidentiality | Provider | Provider | You | None – see below |
The fifth pattern, hybrid, has no column of its own because it does not have one answer: it inherits the row from whichever path a given request takes, and the Layer 2 obligation it adds is knowing which artifact each route actually serves. A routing layer that can fail over from a cloud API to a self-hosted fallback has quietly given you the self-hosted column.
Three consequences, and they are the point of the table:
1 · On a Cloud API you own almost none of this, and the one control you do own is the one teams skip. You never see the weights, so format, hashes and signatures are meaningless to you. What remains is behavioural evaluation of the endpoint you are actually calling – does this model, at this version, still refuse what your policy requires? That is a Layer 2 control, it is entirely yours, and it is the one that catches a silent provider-side model update changing behaviour your application depended on. Everything else here belongs to the provider and reduces to a vendor-assessment question.
2 · Self-hosting takes the entire layer at once. Chapter 1 warned that moving to self-hosting transfers controls that were previously included and do not arrive with the weights. Layer 2 is the clearest instance: the moment you download a file, the format decision, the verification decision, the container, and every dependency in the serving stack become yours on the same afternoon. This is the column the rest of this section is written for.
3 · Edge is the pattern with no recovery. Once weights ship to a device, Chapter 2 Section 7 is blunt that nothing in Chapter 3 recovers them – extraction stops being an attack and becomes a file copy. Every Layer 2 control for edge is therefore applied before shipping, and model confidentiality is not a control you have. Treat on-device weights as published and remove from them anything you would not publish.
Use this before the checklist, not after
The Layer 2 checklist at the end of this section is written for the self-hosted column. Run this table first. An organisation consuming a cloud API that opens a finding about unsigned model artifacts has written a finding against a file it will never possess.
What Each Control Actually Proves
Layer 2 is a stack of verification steps that all sound like the same thing and are not. This is the table to have straight before any of the mechanisms below, because nearly every real failure in this layer is a control being credited with a guarantee it does not offer.
| Control | What it proves | What it does not prove | How it is defeated |
|---|---|---|---|
| Format enforcement (safetensors) | Loading this file cannot execute code | Anything at all about the weights, or about the tokenizer, config and remote code shipped beside it | Not defeated – evaded, by the parts of the stack that are not the tensor blob |
| Hash verification | This is byte-for-byte the file that was published | That the publisher is who you think, or that what they published is safe | Publisher compromise, or a poisoned artifact published honestly |
| Signature verification | Who published it, and that it is unmodified since | That the signer’s model is clean, or that the signer is trustworthy | A signed backdoor. The signature is valid; the weights are poisoned |
| Provenance record | Where it came from, what it was made from, what was done to it | That any claim in the record is true, unless the record is itself attested | A self-asserted model card. Provenance is only as good as its attestation |
| Scanner verdict | The scanner found nothing it can parse and recognise | That the file is safe – a clean verdict is the absence of a finding | Change the container format so the scanner cannot unpack it. This is exactly nullifAI |
| Behavioural evaluation | The model did not misbehave on the inputs you tried | That it will not misbehave on an input you did not try | A trigger you did not guess. This is the hard limit of the whole layer |
Two readings of that table matter more than any individual row.
Read down the second column and the controls stop overlapping. They are not redundant checks of the same property; each answers a different question, and a stack that runs four of them has answered four questions and left the fifth open. The common enterprise posture – pin the hash, scan the file, ship it – proves integrity against the publisher and absence of a recognised finding, and proves nothing whatever about behaviour.
Read the last column and the failure mode is uniform. Every one of these controls fails by being credited with the adjacent guarantee. Format enforcement gets credited with weights safety. A signature gets credited with trustworthiness. A clean scan gets credited with cleanliness. The discipline this layer asks for is narrow and unglamorous: say which question each control answered, and be explicit about the ones still open.
Model Integrity and Provenance
Model integrity verification answers “is the model I am deploying the model I intended to deploy?” The gates below run in order, and each one is cheap relative to the one after it.
graph LR
UM["Untrusted<br/>Artifact<br/><small>Hub, vendor, or<br/>internal checkpoint</small>"]
FMT["Format<br/>Check<br/><small>Safetensors, and<br/>what else ships<br/>beside it</small>"]
HASH["Pinned<br/>Revision<br/><small>SHA-256 of a<br/>specific commit,<br/>not a tag</small>"]
SIG["Signature<br/><small>OMS / Sigstore --<br/>who published it</small>"]
PROV["Provenance<br/><small>Source, base model,<br/>training data, chain</small>"]
EVAL["Behavioural<br/>Evaluation<br/><small>Refusals, triggers,<br/>memorisation</small>"]
APPROVE["Admitted to<br/>Registry"]
REJECT["Rejected"]
UM --> FMT
FMT -->|"Safe format"| HASH
FMT -->|"Pickle only"| REJECT
HASH -->|"Match"| SIG
HASH -->|"Mismatch"| REJECT
SIG -->|"Valid"| PROV
SIG -->|"Absent or invalid"| REJECT
PROV -->|"Accounted for"| EVAL
PROV -->|"Unknown origin"| REJECT
EVAL -->|"Within policy"| APPROVE
EVAL -->|"Fails policy"| REJECT
style UM fill:#a85800,color:#fff
style FMT fill:#2d5016,color:#fff
style HASH fill:#2d5016,color:#fff
style SIG fill:#2d5016,color:#fff
style PROV fill:#2d5016,color:#fff
style EVAL fill:#2d5016,color:#fff
style APPROVE fill:#1565c0,color:#fff
style REJECT fill:#8b0000,color:#fff
The ordering is deliberate and it is a cost argument, not a security one. A format check is a string comparison. Behavioural evaluation is a GPU-hours bill. Put them the other way round and the gate becomes something teams disable under schedule pressure.
Pin the Revision, Not the Name
Chapter 2 Section 3 described the distribution property that makes this the load-bearing control: reputation is earned once and spent later. An attacker publishes a genuinely useful artifact, accumulates downloads and stars, and then pushes a poisoned revision. Anything that resolves the model by name at build time silently picks it up.
So the requirement is not “verify the hash” in the abstract. It is that deployment references an immutable identifier – a commit SHA or a content digest – and that a change to that identifier is a reviewed event rather than a rebuild. A pipeline that pulls org/model:latest and then checks the file’s hash against whatever it just downloaded has verified that the download was not corrupted in transit, which was never the threat.
Signing: What Changed, and What It Still Does Not Give You
Model signing applies code-signing logic to model artifacts, and the ecosystem consolidated in 2025. The OpenSSF Model Signing (OMS) specification, with the sigstore/model-transparency library and CLI reaching v1.0 in April 2025, is the current standard. Three properties matter for how you use it:
- It signs the bundle, not the weights file. The signature covers a manifest of every artifact and its hash – weights, config, tokenizer, and any datasets shipped alongside. That scope is deliberate, and the next section is the reason for it.
- It is PKI-agnostic. Bare keys, an enterprise PKI, or keyless identity-based signing through a public or private Sigstore instance. The last option is what makes signing viable for internal checkpoints without standing up a key-management programme first.
- Adoption is at the hubs, not at the consumers. NVIDIA’s NGC catalogue and Kaggle have adopted it. The gap is the same one that has always dogged this control: verification is a step the consumer has to choose to perform, and the default path in every ML framework is to load the file.
A valid signature is a statement about identity, not about safety
This is the most-credited-beyond-its-guarantee control in the layer. A signature tells you the artifact came from a specific identity and has not changed since. It says nothing about whether that identity trained a backdoor into the weights, and nothing about whether that identity was compromised at signing time. Signing converts “some file from the internet” into “this organisation’s file” – which is enormous progress, and is not a safety verdict.
Provenance, and the Artifact From the Desk Next to You
A complete provenance record covers source, base model, training and fine-tuning data at least at category level, the training process, every modification applied, and the distribution chain. In 2026 the machine-readable form of this exists: CycloneDX 1.6 carries an ML-BOM with model cards, and the SPDX 3.0 AI profile carries AI-specific metadata. Neither yet expresses the full lineage of a model, but both are enough to answer “which deployments contain this artifact” on the day it turns out to matter.
The application teams forget is the internal one.
The internal path is the one with no controls on it
Chapter 2’s ByteDance case is here for exactly this reason. An intern forged a model checkpoint containing malicious code and it moved between colleagues on a shared research cluster – no signing, no scanning, no review, because everything inside the cluster was implicitly trusted.
The malicious checkpoint is the same artifact whether it arrives from a public hub or from the desk next to yours. Every control in this section applies to checkpoints produced by your own training runs, and the two that matter most internally are signing (so “who last wrote this checkpoint” is answerable) and an audit trail on the training pipeline. Note the detection failure too: the sabotage injected irreproducible randomness into training runs, which presents as a reproducibility bug. Unexplained non-determinism on a shared cluster deserves a security question, not just a debugging session.
Defense Connection
Integrity and provenance controls are Layer 2’s answer to the distribution half of LLM05: Data and Model Poisoning. Layer 1 closes the write paths into training data; it can do nothing about a checkpoint that is already poisoned. Pinned revisions and provenance are what make “the file I reviewed is the file I loaded” a checkable statement rather than an assumption.
Serialization Safety Across the Serving Stack
Chapter 1 Section 3 established the rule and Chapter 2 Section 4 explained the mechanism: a pickle file is a program, torch.load() is its interpreter, and loading is running. Take that as given. Layer 2’s job is turning it into a gate, and the gate is wider than the weights file.
The Format Control, Stated Currently
Two things changed since this advice was first written, and both affect how you should specify the control:
- Safetensors is the ecosystem default, not the careful choice. Requiring it costs nothing for the overwhelming majority of models, which makes “pickle-only repository” a workable rejection criterion rather than an aspiration.
- The framework default flipped. PyTorch 2.6 (29 January 2025) changed
torch.loadtoweights_only=Trueby default, restricting the unpickler to tensor data and refusing dynamic imports. This is the single largest reduction in real-world exposure to this attack class – and it is a version-dependent control. A pipeline pinned to PyTorch 2.5, or one that setsweights_only=Falseto make a legacy checkpoint load, has opted back in. Audit for the override, not just for the version.
What Format Enforcement Does Not Cover
A safetensors migration secures the weights, not the stack
Chapter 2 Section 4 is emphatic about this and it is the point most often lost when the control reaches a checklist. Safetensors removes code execution from one file. The serving stack deserializes untrusted input in at least four other places:
- The tokenizer and config files shipped alongside the weights – which is why OMS signs the bundle rather than the tensor blob.
- Remote code paths.
trust_remote_code=Truefetches and executes Python from a hub repository. It is a flag, not a format, and no serialization policy sees it. - Other formats with the same property. Hugging Face’s own security documentation points at Keras Lambda layers as a worked example; pickle is the famous case, not the only one.
- Inter-process communication inside the inference server. This is the one that turned out to matter most – ShadowMQ and CVE-2024-50050, where
recv_pyobj()gave anyone who could reach a ZeroMQ socket remote code execution, and the same call was then found copy-pasted across five more inference servers.
The question is not “do we load untrusted models.” It is “what in our serving stack calls a deserializer on input it did not produce.”
That last item is a dependency problem rather than an artifact problem, which is where the next section picks it up.
Adapters and Fine-Tuned Checkpoints
This is the question Layer 2 owns from Chapter 2’s four: who can push an adapter into our pipeline? It gets its own treatment because adapters break the intuitions the base-model controls were built on.
LoRA adapters are megabytes rather than gigabytes, and that changes their distribution behaviour more than their technical properties do. They circulate the way scripts circulate – linked in a forum post, pinned in a chat, copied into a Dockerfile – with none of the ceremony a base-model swap would attract. An adapter frequently enters a pipeline without passing any of the gates the base model passed.
Separate the Code Risk From the Weights Risk
Chapter 2 Section 3 draws this distinction and it is the single most important idea in this layer, so it is restated here as a control decision rather than an attack description:
| Code risk | Weights risk | |
|---|---|---|
| The control that closes it | Format enforcement – safetensors, weights_only, no remote code |
Behavioural evaluation, including trigger hunting |
| Cost of the control | Effectively zero | GPU hours, and it is still not conclusive |
| Does format choice help? | Yes – completely | No – not at all |
| What a clean result means | Loading this file will not execute code. A strong guarantee | The model did not misbehave on the inputs you tried. A weak one |
A signed, scanned, safetensors adapter from an account with ten thousand downloads can still strip a model’s refusal behaviour or fire a backdoor on a trigger phrase, and nothing about the file will tell you. Treating safetensors as the answer is the most common error in this layer. It closes the code path completely and does nothing whatever about the weights.
Why the Weights Risk Has No Clean Answer
Chapter 2 established that backdoors survive the remediation people expect to work. Standard safety training – supervised fine-tuning, RLHF, adversarial training – failed to remove implanted backdoors in the Sleeper Agents work, and adversarial training taught models to hide them better.
Two consequences for Layer 2, and they are the reason this layer is preventive:
- “We’ll fine-tune it out” is not a remediation plan. Remediation means replacing the artifact – reverting to a model whose provenance you can verify, or retraining from data you can account for. That is a budget and schedule decision, not a patch.
- Therefore the gate is the only cheap moment. Chapter 2’s comparison table prices a malicious artifact as low to fix if caught at the gate, high once you have trained on it. Every argument for spending effort here rather than downstream is contained in that one row.
Where behavioural evaluation genuinely earns its cost is on artifacts you cannot decline: a fine-tune you commissioned, an adapter a partner supplies, a base model with no alternative. Run it against your actual refusal policy and your actual domain, not a public benchmark – a poisoned model is built to pass those.
Defense Connection
Adapters sit at the intersection of LLM04: Supply Chain and LLM05: Data and Model Poisoning, which is why Chapter 2 tags the section with both. The distribution problem is LLM04 and the pinned revision closes it. The poisoned-weights problem is LLM05 and nothing at this layer closes it – you reduce it to a decision about whose artifacts you accept.
Container Security for AI Workloads
Most production models are served from containers, and the container is what actually gets deployed. If the image is compromised, the model is compromised regardless of how carefully the weights were verified.
AI serving images differ from ordinary application images in ways that change the scanning calculus:
- Large base images. GPU drivers, CUDA libraries and ML dependencies make a serving image an order of magnitude larger than a typical microservice, with a CVE surface to match.
- The model is inside the image. Weights, config and tokenizer are packaged with the runtime, so artifact verification and image scanning are the same gate rather than two.
- Young software. Inference servers are new projects moving fast. ShadowMQ demonstrated a single bug class propagating through five of them by copy-paste, so your exposure is not bounded by your own code review and CVE monitoring on the serving framework is not optional.
- Long-lived processes. Inference servers run continuously, so an image built four months ago is still running four-month-old dependencies.
The Pipeline, and Where It Stops
graph LR
MA["Verified<br/>Artifact<br/><small>Cleared the integrity<br/>gates above</small>"]
CB["Container Build<br/><small>Base image + model +<br/>serving framework</small>"]
VS["Scan<br/><small>CVEs, malware,<br/>serialized objects,<br/>SBOM generation</small>"]
SR["Signed Registry<br/><small>Image signing,<br/>access-controlled<br/>artifact store</small>"]
DP["Deployment<br/><small>Layer 3 owns<br/>everything from here</small>"]
MA -->|"Integrity<br/>verified"| CB
CB -->|"Build<br/>complete"| VS
VS -->|"Scan<br/>passed"| SR
SR -->|"Authorized<br/>pull"| DP
style MA fill:#2d5016,color:#fff
style CB fill:#2d5016,color:#fff
style VS fill:#2d5016,color:#fff
style SR fill:#2d5016,color:#fff
style DP fill:#1565c0,color:#fff
A failure at any gate blocks progression to the next. The boundary worth noting is the last arrow: Layer 2’s responsibility ends at the registry. Admission control, runtime policy, GPU tenancy and egress restrictions are Layer 3, and the reason to keep the line sharp is that the two layers fail differently – Layer 2 fails by admitting a bad image, Layer 3 fails by letting a good image do bad things.
Defense Connection
Container scanning is Layer 2’s contribution to LLM04: Supply Chain. It is worth being precise about what it catches: known CVEs in the serving framework and its dependencies, and recognised malicious content in the artifacts baked into the image. It does not catch a poisoned weights file that no scanner recognises – see the case study below for what happens when scanning is treated as the boundary rather than as a filter.
Supply Chain Defense
The AI model supply chain extends past model weights to every dependency, tool and service that contributes to building and deploying a model.
Model Repository Security
An internal model registry needs the controls you already apply to a code repository, plus one that is specific to models:
- Access controls – role-based permissions for upload, modification and download, with upload separated from deployment approval
- Audit logging – who uploaded which artifact, when, and from where, so that the ByteDance question is answerable
- Automated scanning on upload – every artifact entering the registry is scanned, including checkpoints produced internally
- Immutable versioning – deployments reference a content digest, and tags are conveniences that never appear in a deployment manifest
- An approved catalogue – production may only pull from the internal registry, which is what converts every control above from advice into an enforced path
That last one is the control that makes the rest work, because it removes the bypass. An engineer who can pip install a model directly onto a training host has skipped every gate in this section, and no amount of registry hardening reaches them.
Dependency Auditing
AI applications sit on deep stacks of Python packages, framework versions and system libraries, and the serving stack is part of the model’s attack surface rather than adjacent to it.
- Software Bill of Materials – generate an SBOM for every serving container. Its value is not documentation; it is answering “which of our deployments contain this?” on the day a dependency is compromised, in minutes rather than days.
- Automated vulnerability scanning in CI, with the serving framework itself in scope.
pip-auditand equivalents cover the packages; the inference server needs CVE monitoring in its own right. - Pinning with verified hashes, transitively. Pinning direct dependencies while resolving transitive ones by range leaves the substitution path open.
- Pinned tool and MCP server versions. Chapter 1 Section 7 and Chapter 2 Section 5 both route this control here by name. An MCP server that changes its tool definitions after you approved it is a rug pull, and the case in Chapter 2 shipped fifteen benign versions before the malicious one. A one-time review at install passes that test. What defeats it is an inventory of which servers are connected, a pinned version, and a reviewed diff on change.
Ultralytics was a pipeline compromise, not a bad package
The Ultralytics case is often summarised as “a popular vision package shipped a cryptominer,” which points the lesson at dependency selection. The mechanism was narrower and more uncomfortable: an attacker poisoned the build cache through CI script injection, so the published artifact did not correspond to the reviewed source. Auditing the repository would have found nothing.
The control that addresses this is not choosing better dependencies. It is scanning what actually arrives rather than what the repository says it should be, plus build provenance that ties an artifact to the pipeline run that produced it.
Defense Connection
Supply chain defense is Layer 2’s half of WarningASI04: Agentic Supply Chain Vulnerabilities. Chapter 2’s malicious MCP server case showed one compromised tool provider reaching every agent connected to it. Layer 2 pins and inventories the artifacts; Layer 3 verifies the servers at the orchestration layer and constrains what a tool call can reach.
AI Model Vulnerability Scanning
Pre-deployment scanning for AI covers categories traditional CVE detection has no concept of. Five scan types matter, and they map onto the two risk classes from the adapter table – the first two address code risk, the last three address weights risk.
| Scan Type | What It Checks | What a clean result means |
|---|---|---|
| Serialization safety | Model, tokenizer, config and remote-code paths for pickle exploits and embedded code | Strong: nothing the scanner could parse executes on load |
| Weight integrity | Artifact against a pinned revision’s hash | Strong, and narrow: it is the file that was published |
| Adversarial robustness | Responses to known adversarial and jailbreak inputs | Weak: it resisted the library you ran |
| Alignment verification | Outputs against your safety policy, including targeted domains | Weak: no policy violation on the probes you used |
| Data leakage potential | Outputs for memorised training data – PII, credentials, proprietary text | Weak, and diagnostic: a hit is evidence about your Layer 1 pipeline |
When to Scan
- On acquisition – before the artifact is stored anywhere other than a quarantine location
- After any weight modification – fine-tuning and quantization both produce a new artifact that inherits nothing from the base model’s clearance
- Pre-deployment, as a CI gate that fails the build rather than raising a ticket
- On a schedule after deployment – because the vulnerability library grows and the artifact does not
A clean scan is the absence of a finding
This is the point the case study below exists to make, and it applies to every row of the table above. Scanners parse; what they cannot parse, they do not report. Treat a verdict as one input alongside format, pinned revision, signature and provenance – never as clearance.
Defense Perspective: nullifAI, and What Scanning Cannot Be
The attack (from Chapter 2 Section 3): In February 2025, ReversingLabs disclosed nullifAI – two malicious models on the Hugging Face Hub carrying a reverse shell that connected to a hardcoded address when the model was loaded. The notable part is not that pickle executes code. It is that both models sat on the Hub having passed its security scanning: they were compressed with 7z instead of the ZIP that PyTorch normally uses, so torch.load opened them happily while Picklescan could not unpack them and therefore reported nothing. The pickle streams were also deliberately broken, so the payload ran and then the load raised an error – which reads like a broken upload rather than an attack.
What Layer 2 controls would have changed the outcome, in order of cost:
- Format enforcement. A policy requiring safetensors rejects the artifact on its extension, before any scanner is involved. This is the control that actually defeats nullifAI, and it costs a string comparison.
- Pinned revision from a known publisher. The models were not impersonating a trusted org, so a policy of pulling only from an approved catalogue never reaches them.
- Framework configuration. On PyTorch 2.6 or later,
weights_only=Truerefuses the dynamic import the payload depends on – and a pipeline that overrides it back toFalsefor legacy checkpoints has removed the protection for every load. - Quarantined loading. Where a pickle artifact genuinely must be opened, loading it in an environment with no credentials and no egress means the reverse shell has nowhere to connect.
What would not have helped, and this is the reason the case is here: scanning. The attack targeted the control rather than the format. A container image scan of the same artifact inherits the same limitation – it must parse the archive to inspect the stream, and it reports nothing rather than reporting failure when it cannot.
The transferable lesson: “the hub scanned it” is a control with a bypass, and a clean verdict is the absence of a finding rather than evidence of safety. Note also what a failed model load tells you about whether code ran: nothing.
Layer 2 → OWASP Mapping
Layer 2 is named as a primary defense in three of the twenty categories across the LLM Top 10 and the Agentic AI Top 10 – the narrowest coverage in the Blueprint. For each, the honest statement has two halves.
| Category | The route Layer 2 acts on | What Layer 2 does | What it cannot do | Completed by |
|---|---|---|---|---|
| LLM04: Supply Chain | Artifacts, images and dependencies entering your environment | Format enforcement, pinned revisions, signature verification, image and dependency scanning, SBOM, approved catalogue | Stop a compromise in a component you do not build or scan – and see the CI-poisoning case above | L3 – orchestration and MCP server verification, runtime isolation |
| LLM05: Data and Model Poisoning | A checkpoint or adapter that is already poisoned | Provenance, signing, behavioural evaluation, artifact replacement as the remediation path | Detect a backdoor whose trigger you did not guess; remove one by further training | L1 – closing the write paths into training data before the checkpoint exists |
| WarningASI04: Agentic Supply Chain | Tool, package and server artifacts an agent depends on | Inventory, version pinning, reviewed diff on change, dependency auditing | Constrain what a tool actually does once it is connected and called | L3 – orchestration controls; L5 – what the agent is permitted to reach |
Read the fourth column down and Layer 2’s boundary is a single sentence: it decides what becomes the artifact, and nothing about what the artifact then does. That is why the layer with the narrowest coverage is not the layer with the least leverage – everything downstream is a defense of whatever this layer admitted.
AI Scanner Cross-Reference
AI Scanner contributes to Layer 2 on the weights-risk side of the table above: it assesses a model for susceptibility to prompt injection, system prompt leakage and adversarial manipulation – properties of the model itself that no infrastructure control changes. Read its results the way this section reads every behavioural result: a finding is strong evidence, a clean run is weak evidence, and both are inputs to a deployment decision rather than the decision. Section 9 covers the full scan-protect-validate-improve loop and where re-scanning fits after deployment.
TrendAI Vision One Container Security covers the code-risk side, scanning serving images before they reach the registry – known CVEs in frameworks such as TensorFlow Serving, vLLM and Triton, and recognisable malicious content in the artifacts baked into the image. The integration argument is the pipeline rather than the product: image scanning is only a gate if the deployment path cannot pull from anywhere else, which is why the approved-catalogue control above is the one that makes the scanner load-bearing.
Key Takeaways
- Layer 2 has the narrowest coverage in the Blueprint and no substitute: it is purely preventive, never in the request path, and nothing downstream recovers a backdoored artifact once it is deployed
- Every control here proves one specific thing – format proves loading is safe, a hash proves the bytes, a signature proves the publisher, a scan proves only the absence of a finding – and each fails by being credited with the next one’s guarantee
- Separate the code risk from the weights risk: safetensors closes the first completely and the second not at all, so a signed, scanned, safetensors adapter can still be fully backdoored
- Safetensors secures one file. The tokenizer, the config,
trust_remote_code, and the inference server’s own IPC all deserialize untrusted input, which is why the question is what in your stack calls a deserializer on input it did not produce - Pin revisions rather than names: reputation is earned once and spent later, and a poisoned revision of a trusted artifact is the cheap path in
- Backdoors survive further training, so remediation is artifact replacement rather than repair – which is the entire argument for spending the effort at the gate
- Which Layer 2 controls exist for you is set by your deployment pattern: a cloud API leaves you behavioural evaluation alone, self-hosting hands you the whole layer at once, and edge deployment has no recovery path at all
Test Your Knowledge
Ready to test your understanding of AI model security? Head to the quiz to check your knowledge.
Up next
Layer 2 stops at the registry. What the admitted image is then permitted to do – which GPU it lands on, which tenants it shares a backend with, which sockets it can reach, and what identity it runs as – is Layer 3. In Section 5 you will see AI Security Posture Management, GPU cluster and serving-endpoint hardening, orchestration-layer security, and identity management for AI service accounts, including the runtime half of the two case studies this section only closed at build time.