Section 4 Quiz

Test Your Knowledge: Layer 2 - Secure Your AI Models

Let’s see how much you’ve learned!

This quiz asks you to say what each Layer 2 control actually proves, scope the layer to a deployment pattern, separate the code risk from the weights risk, and decide what to do when a control returns the answer you wanted.

--- shuffle_answers: true shuffle_questions: false --- ## Section 2 rates Layer 2 as the narrowest layer in the Blueprint -- a primary defense in 3 of 20 OWASP categories -- and also says it is the one you cannot compensate for elsewhere. What reconciles those two statements? > Hint: Think about what is left to defend once a compromised artifact is running. - [ ] Coverage counts are unreliable, so the 3-of-20 figure understates what the layer really addresses > The count is accurate. Layer 2 genuinely appears in only three categories, and the argument for the layer does not depend on disputing that. - [x] Nothing downstream can recover an artifact that was already compromised when it was deployed > Correct! Coverage counts how many categories a layer touches, not how completely it resolves any of them. Layer 2's three categories are ones where the compromise sits inside the thing every later layer is protecting -- a runtime filter in front of a backdoored model is filtering the backdoored model's output. That is why narrow coverage and high leverage are not in tension here. - [ ] The other five layers were scored against a different version of the OWASP categories > All twenty categories are scored against the 2026 editions in the same master table. The coverage figures are directly comparable across layers. - [ ] Layer 2 also runs detective controls in production, which the coverage count omits > Layer 2 is purely preventive. Periodic re-scanning produces a finding about a deployed system, and the response is a redeployment rather than any interception. ## Your team pulls a base model in safetensors format, verifies its SHA-256 against the publisher's pinned revision, confirms a valid OMS signature, and gets a clean verdict from the artifact scanner. Which risk is still entirely open? > Hint: Ask what each of those four controls proves, one at a time. - [ ] Remote code execution at load time, because signatures cannot cover the tokenizer file > OMS signs a manifest covering every artifact in the bundle -- weights, config and tokenizer -- and safetensors closes the load-time code path in the weights file itself. This route is well covered by what the team did. - [ ] Corruption in transit, since a pinned revision only identifies which file was intended > A SHA-256 check against the pinned revision is exactly what detects transit corruption. This is the one thing on the list that hash verification does conclusively. - [x] A backdoor in the weights, which every one of those four controls is blind to > Correct! Format proves loading will not execute code. The hash proves the bytes match what was published. The signature proves who published it. A clean scan proves only that the scanner found nothing it could parse and recognise. None of them examines model *behaviour*, so a trigger-activated backdoor or stripped refusal training passes all four unchanged. Only behavioural evaluation addresses it, and it proves the weaker statement that the model did not misbehave on the inputs you tried. - [ ] Dependency vulnerabilities, which artifact-level controls are not designed to detect > Serving-stack dependencies are a real gap, but they are addressed by SBOMs, pinning and CVE monitoring -- controls the team can add. The weights risk has no control that closes it. ## The nullifAI models sat on the Hugging Face Hub having passed its scanning: they were compressed with 7z rather than ZIP, so the scanner could not unpack them. Which Layer 2 control would have rejected them, and at what cost? > Hint: The attack targeted the control, not the format. What does not need to parse the file? - [ ] A second scanner from a different vendor, at the cost of one more pipeline stage > A second scanner faces the same requirement -- it must unpack the archive to inspect the stream, and reports nothing when it cannot. Layering scanners with the same limitation does not remove the limitation. - [x] Format enforcement, which rejects on the extension and costs a string comparison > Correct! A policy requiring safetensors never reaches the question the scanner failed to answer, because it rejects a pickle-based artifact before anything is parsed. That is the general shape of the lesson: the cheapest gate is the one that does not depend on understanding the file's contents. It is also why the ordering of the integrity gates is a cost argument -- put the expensive checks first and teams disable them. - [ ] Behavioural evaluation, at the cost of the GPU hours a full probe suite requires > Behavioural evaluation examines what a loaded model does. The nullifAI payload ran during loading, before any inference -- so the evaluation itself would have triggered the compromise. - [ ] Provenance tracking, at the cost of maintaining a record for every artifact > Provenance would show an unfamiliar publisher, which is useful. But the models were not impersonating a trusted org, so the record would have been accurate and unremarkable rather than a rejection. ## An organisation consumes a frontier model through a cloud API and holds no weights at all. Which Layer 2 control is genuinely theirs, and why does it matter? > Hint: Run the ownership table. What is left when the artifact never passes through your hands? - [ ] Hash verification of the model version string returned by the API response metadata > A version string is not an artifact hash, and nothing about it is verifiable against a file you do not have. Hash verification has no meaning in this pattern. - [ ] None -- all Layer 2 controls belong to the provider under this deployment pattern > Close, and it is the answer most teams act on, but it is wrong in one specific place -- and that place is the control most worth running. - [x] Behavioural evaluation of the endpoint, which catches a provider-side model change > Correct! Format, hashes and signatures are meaningless when you never hold the file, but the model's behaviour is observable through the API you already call. Running your own refusal and policy probes against the live endpoint is what surfaces a silent provider-side version update that changes behaviour your application depended on. It is the one Layer 2 control in this pattern, and it is routinely skipped because the layer is assumed to be entirely the vendor's. - [ ] Container image scanning, applied to the client application that calls the API > Scanning your own application container is worthwhile, but it is ordinary application security. It tells you nothing about the model, which is what Layer 2 exists to govern. ## A partner supplies a fine-tuned LoRA adapter. It is in safetensors format, signed, and scans clean. Your evaluation later suggests the model's refusal behaviour has been weakened on one topic. What is the remediation path? > Hint: Chapter 2 tested whether further training removes an implanted backdoor. - [ ] Run adversarial training against the weakened topic until the refusals are restored > Adversarial training was specifically tested against implanted backdoors in the Sleeper Agents work. It failed to remove them, and taught the models to conceal them better under evaluation. - [ ] Add an input filter for the affected topic and keep the adapter in production > A filter in front of a compromised model is Layer 5 compensating for a Layer 2 failure. It narrows one known route while the modified weights stay deployed, and it assumes you have enumerated every route. - [x] Replace the artifact -- revert to a checkpoint whose provenance you can verify > Correct! Standard safety training, including supervised fine-tuning, RLHF and adversarial training, did not remove implanted backdoors. "We'll fine-tune it out" is not a remediation plan, so remediation means replacing the artifact: reverting to a model you can account for, or retraining from data you can. That is a schedule and budget decision, which is exactly why the gate before deployment is the cheap moment and everything after it is not. - [ ] Re-run the scanner with an updated vulnerability library to confirm the finding > Scanners detect code and known-bad content in the file. Weakened refusal behaviour is a property of the weights, and no artifact scan examines it. ## Chapter 2 Section 3 leaves four scoping questions. Three are Layer 1's. Which one does Layer 2 own, and what makes it hard? > Hint: Which of the four is about an artifact rather than a data store? - [x] "Who can push an adapter into our pipeline?" -- adapters skip the base model's gates > Correct! Adapters are megabytes rather than gigabytes, so they circulate the way scripts do: linked in a chat, pinned in a forum post, copied into a Dockerfile. That distribution behaviour, not any technical property, is what makes them dangerous -- an adapter frequently enters a pipeline without passing a single gate the base model passed, and it can strip refusal behaviour or carry a trigger. - [ ] "Who can publish to our corpus?" -- ingested documents become part of the artifact > Corpus documents are retrieved at query time and never enter the model artifact. This is a Layer 1 ingestion-gate question, and Chapter 2 assigns it there. - [ ] "Who can write to our vector store?" -- the store holds derived model state > A vector store holds embeddings and chunk text, not model weights. Chapter 2 and Layer 1 both treat it as a data store with its own access controls. - [ ] "Does an end user's thumbs-down reach a training set?" -- feedback modifies weights > Feedback only affects weights if it is used for training, and the control is a human review gate between the feedback store and the training set. Chapter 2 assigns that gate to Layer 1. ## Your team completes a safetensors migration across every model in production. Which exposure from Chapter 2 does that migration leave untouched? > Hint: Safetensors removes code execution from one file. How many places in a serving stack deserialize? - [ ] Typosquatted repository names resolved by deployment scripts at build time > Format choice does not address name resolution, but pinning by revision does, and it is a control the team can add. This is not the exposure the migration is famously blind to. - [x] The inference server's own IPC, where ShadowMQ found pickle behind a socket > Correct! Safetensors secures the tensor blob and nothing else in the stack. CVE-2024-50050 came from `recv_pyobj()` deserializing ZeroMQ messages inside the inference server, and the same call was later found copy-pasted across five more serving projects. The tokenizer, config files and `trust_remote_code` paths are in the same category. The useful question is not "do we load untrusted models" but "what in our serving stack calls a deserializer on input it did not produce." - [ ] Known CVEs in the CUDA libraries that the serving base image ships > Base image CVEs are real and are covered by image scanning and CVE monitoring. They are not a deserialization exposure, which is what the migration was meant to address. - [ ] Embedding inversion recovering source text from vectors in the store > Embedding inversion is a Layer 1 concern about a derived data store. It is unaffected by model file format because no model artifact is involved. ## Ultralytics is often summarised as "a popular vision package shipped a cryptominer." Chapter 2 shows the mechanism was CI build-cache poisoning through script injection. Which control follows from the accurate version? > Hint: If the published artifact does not correspond to the reviewed source, what were you auditing? - [ ] Stricter dependency selection criteria, favouring packages with more maintainers > This is the control the inaccurate summary implies, and it is why the distinction matters. Ultralytics was already a mature, heavily used, well-maintained package. Selection criteria would not have flagged it. - [ ] More frequent source code review of the direct dependencies in the stack > Reviewing the repository would have found nothing. The source was clean; the poisoning happened in the build, so the published artifact did not correspond to what any reviewer could read. - [x] Scan what actually arrives, and tie the artifact to the pipeline run that built it > Correct! When the compromise is in the build rather than the source, the only controls that see it operate on the delivered artifact: scanning what was published rather than what the repository says should have been published, plus build provenance linking an artifact to a specific pipeline run. This is the same reasoning that puts a pinned content digest, rather than a version tag, in a deployment manifest. - [ ] Version pinning to the last release published before the compromise window > Pinning to a known-good version is correct incident response once you know a compromise occurred. It is not a control that detects one, and the poisoned versions were published as ordinary releases. ## A serving container has passed scanning, been signed, and admitted to the registry. In production it starts making outbound connections nobody expected. Which layer owns the response, and why does the boundary matter here? > Hint: Layer 2's pipeline diagram stops at one specific arrow. - [ ] Layer 2, because the image scan clearly failed and the gate needs to be re-run > Re-scanning an image that is already running gives you a finding about a deployed system, not a response to live behaviour -- and the scan may well have been correct about everything it examined. - [x] Layer 3 -- Layer 2's responsibility ends when the image reaches the registry > Correct! Layer 2 decides what becomes the artifact and nothing about what the artifact then does. Egress restrictions, admission control, runtime policy and tenancy isolation are Layer 3, and the two layers fail differently: Layer 2 fails by admitting a bad image, Layer 3 fails by letting a good image do bad things. The boundary pays off during exactly this incident -- knowing immediately that no Layer 2 control will help is worth more than a diagram implying one might. - [ ] Layer 6, since unexpected outbound traffic indicates an unknown exploit in progress > Zero-day defense may become relevant if the behaviour turns out to be a novel exploit, but the controls that constrain what a running container can reach are configured before that determination is made. - [ ] Layer 2 and Layer 3 jointly, since the container is an artifact and a running process > The artifact and the running process are governed at different times by different controls, and blurring them is what produces architecture diagrams that credit Layer 2 with runtime enforcement it does not have.