Section 3 Quiz

Test Your Knowledge: Deployment Considerations

Let’s see how much you’ve learned!

This quiz tests your understanding of deployment patterns, trust boundaries, cost trade-offs, serialization security, and safety guardrails for AI systems.

--- shuffle_answers: true shuffle_questions: false --- ## An organization needs to run an AI model inside an air-gapped facility, on documents too large and complex for a small model to summarize reliably. Which deployment pattern fits? > Hint: Two patterns work without a network. Only one of them can run a large model. - [ ] Cloud API deployment > Cloud APIs need internet connectivity to reach the provider's servers, which an air-gapped environment does not have. - [ ] Serverless inference on a cloud platform > Serverless inference still runs in the cloud. The management overhead is lower, but the network dependency is identical. - [x] Self-hosted deployment on internal GPU hardware > Correct! Both self-hosted and edge deployment work air-gapped, so connectivity alone does not decide it -- the capability requirement does. Edge is limited to small models sized for an end-user device, which will not handle large, complex documents reliably. Self-hosting on internal GPUs runs open-weight models at full scale with no external network dependency. - [ ] Edge deployment on user laptops > Edge deployment does work offline, which makes this a reasonable first instinct. It fails the second constraint: edge is limited to small models, and the task explicitly needs more capability than a small model provides. ## A startup is building a prototype AI chatbot and expects traffic to be unpredictable -- some days 100 requests, other days 10,000. Which deployment approach makes the most economic sense? > Hint: Compare a cost with no floor against a cost with no ceiling. - [x] Cloud API with pay-per-token pricing > Correct! Per-token pricing has no floor -- on a 100-request day they pay for 100 requests, and the platform absorbs the spikes with no provisioning decision. For unpredictable, low-baseline traffic this is the cheapest shape available, with no upfront investment. - [ ] Self-hosted deployment on dedicated GPUs > Self-hosting is a fixed cost regardless of usage. On a 100-request day the startup pays the full GPU bill for almost entirely idle hardware. - [ ] Edge deployment on user devices > Edge deployment suits on-device use cases, not a centrally hosted chatbot, and it caps capability to small models. - [ ] Training a custom model from scratch > This requires enormous training infrastructure and expertise, and is inappropriate for a prototype at any traffic level. ## On-device assistants such as Apple Intelligence and Android's Nano tier, plus small open models running locally on laptops, are examples of which deployment pattern? > Hint: Think about where the computation physically happens. - [ ] Cloud API deployment > Cloud APIs process data on remote servers, not on the user's device. - [ ] Serverless inference > Serverless still runs in the cloud, just without you managing the infrastructure. - [ ] Hybrid deployment > These can form the local tier of a hybrid system, but the pattern named here is the on-device one specifically. - [x] Edge / on-device deployment > Correct! These are Small Language Models running directly on end-user hardware. Computation stays local, so there is no data transmission and no network round trip -- at the cost of a lower capability ceiling and a threat model that now includes physical access to the device. ## Why is Safetensors considered more secure than Pickle for distributing model weights? > Hint: Ask what the loading process is permitted to do, not what the file contains. - [ ] Safetensors encrypts the model weights on disk > Neither format encrypts weights. The difference is about code execution, not confidentiality. - [ ] Safetensors files are smaller, and harder to tamper with > Size and tamper-resistance are not the distinction. A Safetensors file can be tampered with too -- it just cannot execute code when loaded. - [x] Pickle executes arbitrary code on load; Safetensors cannot > Correct! Pickle's flexibility comes from serializing arbitrary Python objects, which means deserializing it can run code -- by design, not as a bug. A tampered Pickle checkpoint executes the attacker's payload the moment anyone loads it, before any prompt is sent. Safetensors stores only tensor data and a JSON header, with no mechanism to execute anything. - [ ] Safetensors only loads models from verified publishers > Safetensors does not verify provenance at all. Its security comes from what loading cannot do, regardless of where the file came from. ## A team holding the weights of an open-weight model wants to remove its refusal behaviour. Why is that achievable without retraining the model? > Hint: Recall what interpretability research found about how refusal is represented internally. - [ ] Refusal behaviour is a configuration flag stored in the model file > Refusal is learned behaviour distributed across the weights during alignment training, not a setting stored alongside them. - [x] Refusal is mediated by a single direction in the activations, which can be erased > Correct! Arditi et al. (2024) found that across 13 open-weight chat models, refusal is concentrated in one direction in the residual stream -- erase it and the model stops refusing harmful instructions. That is why refusal cannot be treated as a security control on any deployment where someone else holds the weights. - [ ] Safety training is kept in a separate file that can simply be deleted > Alignment is baked into the weights themselves. There is no separable safety file to remove. - [ ] Refusal only activates when the model is served through a provider's API > Refusal behaviour is part of the weights and applies however the model is served, including fully offline. ## An organization moves a working application from a cloud API to a self-hosted open-weight model, leaving prompts and application code unchanged. What is the most commonly overlooked consequence? > Hint: Consider which controls were being supplied by someone else. - [ ] The application has to be rewritten in another language > Self-hosted serving stacks expose compatible HTTP APIs, so application code usually needs little change. That is part of why this migration looks deceptively simple. - [x] Moderation, rate limiting and logging do not come with the weights > Correct! The provider was supplying input and output moderation, rate limiting, abuse detection, and audit logging as part of the service. None of it arrives in the model file. The team inherits all of those responsibilities the moment it migrates, usually without a plan for them -- and on open weights, the model's own refusal behaviour can be stripped as well. - [ ] Per-token costs increase sharply after the migration > Self-hosting removes per-token cost entirely, replacing it with fixed hardware cost. That is normally the motivation for migrating. - [ ] The model can no longer run without internet connectivity > The reverse is true: self-hosting is what makes offline and air-gapped operation possible. ## Your application sends a large, unchanging system prompt with every request, and the monthly bill is too high. Two proposals: enable prompt caching, or route the 80% of requests that are simple classification to a small-tier model. Which is the stronger lever? > Hint: One proposal changes the rate you pay; the other changes which tier you pay at. - [ ] Prompt caching, because it removes the cost of the system prompt entirely > Caching reduces the rate charged on a repeated prefix, it does not remove the cost. It is worth doing here given the large fixed prompt, but it remains a discount on the same tier. - [x] Tier routing, because moving work between tiers changes cost by orders of magnitude > Correct! Small-tier and frontier pricing differ by roughly one to two orders of magnitude, so moving 80% of traffic down a tier dominates any rate discount on what remains. Caching is still worth enabling -- they are not mutually exclusive -- but if only one gets built, routing is the one that moves the bill. Model selection by task is the largest cost lever available. - [ ] Neither, because provider pricing changes too often to plan around > Absolute prices do change constantly, but the *ratios* between tiers are stable, and it is the ratio this decision rests on. - [ ] Prompt caching, because it needs no changes to application logic > Lower implementation effort is a genuine advantage, but the question asks which lever is stronger, and caching cannot match an order-of-magnitude tier change. ## An application uses provider moderation endpoints and also relies on the model's built-in refusal behaviour. What limitation do the two share? > Hint: Both are, in the end, classifiers of intent. - [ ] Neither one works on any language other than English > Both handle multiple languages, with varying quality. That is a coverage weakness, not the shared structural limitation. - [ ] Together they add too much latency for real-time use > Moderation does add latency on both directions of the request, but that is a cost of the layer rather than a limit on its correctness. - [x] Both misclassify in both directions, so neither is a hard boundary > Correct! Both judge intent probabilistically. A medical discussion of drug interactions gets blocked (false positive) while a carefully reworded harmful request passes (false negative) -- and refusal is especially phrasing-sensitive, with consistency on the same request dropping from around 85% to 65% when it is reworded. Layering them raises the cost of misuse; it does not prevent it, which is why Chapter 3 builds defence in depth rather than trusting any single control. - [ ] Both work only for models deployed in the cloud > Refusal behaviour travels with the weights in any deployment, and open-weight safety classifiers exist for self-hosted and air-gapped moderation.