3. Deployment Considerations
Introduction
The AI landscape can be confusing when it comes to deployment choices, particularly because similar names often mask very different security and operational implications. For instance, when someone mentions “using GPT,” they might be referring to a vendor’s consumer chat interface, its developer API, or a cloud provider’s enterprise deployment of the same model – each with a vastly different security profile.
This distinction becomes especially important when evaluating AI solutions for enterprise use. Consider the controversy around DeepSeek: while some organizations banned its use over data privacy concerns, they often failed to distinguish between DeepSeek’s web platform (where processing happens on the vendor’s servers) and its open-weight models, which can be downloaded and run locally with full control over data flows. Those are two different risk decisions wearing the same name.
What will I get out of this?
By the end of this section, you will be able to:
- Differentiate the five deployment patterns – cloud API, serverless inference, self-hosted, edge/on-device, and hybrid – and describe the security, scalability, and compliance profile of each.
- Locate the trust boundary in a given deployment: where your data crosses out of your control, and which security controls you own versus inherit from a provider.
- Evaluate the trade-offs between capability, cost, control, and operational burden, and justify a deployment choice for a given scenario.
- Explain why loading a model file is a code-execution decision, comparing the Pickle and Safetensors serialization formats.
- Describe how safety guardrails are layered – model-internal refusal behaviour and external moderation – and where each one fails.
Why Deployment Choices Matter
The way a model is deployed fundamentally shapes its security, scalability, and compliance profile. The same model, reached three different ways, gives you three different systems to defend.
A web-based consumer platform offers ease of access but requires sending all input data to vendor-controlled servers. This raises questions about data residency, retention policies, and geopolitical risk depending on where those servers sit – and, critically, it is usually adopted by employees without a security review at all.
API services provide a middle ground. They let organizations integrate AI capabilities into their own systems while retaining control over what gets sent and what gets logged. Even so, APIs require careful scrutiny of terms of service – some providers retain data temporarily for abuse monitoring unless the account is explicitly configured otherwise.
Self-hosted open-weight models offer unparalleled control over data flows and compliance, but come with significant operational overhead: GPU infrastructure, technical expertise for setup and tuning, and ongoing monitoring. They also transfer every security responsibility onto you – including the ones a provider had been quietly handling.
Navigating Misconceptions
These distinctions also clear up two common misconceptions:
- Naming a model does not describe a deployment. “We use ChatGPT” conflates a consumer web platform with the API that serves the same underlying model. Same technology, different trust boundary, different controls.
- Banning a vendor is not the same as banning a model. Blocking a provider’s hosted service says nothing about whether its open weights can be run safely in an air-gapped environment – which is often exactly where they belong.
The question to actually ask
Not “which model is best?” but “where does my data go, and who holds each control?” Every pattern below is a different answer to that question. Keep it in mind as you read – it is also the question Chapters 2 and 3 are built on.
Model Deployment Options
Deploying large language models presents unique challenges and opportunities compared to traditional software systems. There are more deployment patterns available than ever before, each with distinct trade-offs.
graph TD
A["Can your data leave<br/>your infrastructure?"] -->|"No: regulated,<br/>classified, or air-gapped"| B{"Does a small model<br/>meet the requirement?"}
A -->|"Yes, within a<br/>trusted boundary"| C{"Sustained<br/>high volume?"}
B -->|"Yes"| D["Edge / On-Device"]
B -->|"No: needs a<br/>larger model"| E["Self-Hosted<br/>(on-prem or private cloud)"]
C -->|"Yes: API spend<br/>exceeds GPU cost"| F["Self-Hosted<br/>or Hybrid"]
C -->|"No: variable or<br/>low volume"| G{"Need data residency<br/>or a single vendor bill?"}
G -->|"Yes"| H["Serverless Inference<br/>(Bedrock / Foundry / Vertex)"]
G -->|"No"| I["Cloud API<br/>(pay-per-token)"]
style A fill:#1565c0,color:#fff
style B fill:#1565c0,color:#fff
style C fill:#1565c0,color:#fff
style G fill:#1565c0,color:#fff
style D fill:#2d5016,color:#fff
style E fill:#2d5016,color:#fff
style F fill:#a85800,color:#fff
style H fill:#2d5016,color:#fff
style I fill:#2d5016,color:#fff
Note what the first question is. Not budget, not model quality – can the data leave? That constraint eliminates more options than any other, and it is the one you cannot engineer around later.
The five patterns at a glance
Read this table before opening the detail below. The patterns only make sense against each other.
| Pattern | Data boundary | Cost shape | Capability ceiling | Operational burden | Primary security concern |
|---|---|---|---|---|---|
| Cloud API | Leaves your network entirely; provider-controlled | Per token, no floor | Highest – newest frontier models | Near zero | Data exposure to a third party; credential theft |
| Serverless inference | Stays inside your cloud tenancy | Per token, often at a premium | High – broad catalogue, slight lag on newest | Low | Misconfigured tenancy and IAM; over-broad access |
| Self-hosted | Never leaves your infrastructure | Fixed: hardware plus power, idle or not | Limited to open-weight models | High – you run everything | Malicious model files; you own the entire stack |
| Edge / on-device | Never leaves the device | Fixed per device, no per-request cost | Lowest – small models only | Medium – fleet updates | Physical access to weights; weakened guardrails |
| Hybrid | Split by design; depends on routing | Mixed, optimized by routing | Highest tier available on demand | Highest – two or more stacks | Routing logic itself becomes a trust decision |
Where the Trust Boundary Sits
The five patterns differ on cost and capability, but the difference that matters for the rest of this course is who holds which control. Deployment is the decision that sets your trust boundary, and everything Chapters 2 and 3 discuss happens relative to that line.
Three questions locate it:
- Where does the data go? Out to a vendor, into your cloud tenancy, onto your own hardware, or nowhere at all. This determines your exposure if the other party is breached – a risk you cannot patch.
- Who can see the traffic? Provider-side logging and retention, your own audit trail, or neither. You cannot investigate an incident in a system that logs nothing.
- Which controls do you own, and which are you inheriting silently? This is the one teams get wrong. Moving from an API to self-hosting does not just transfer the hardware bill – it transfers rate limiting, abuse detection, input and output moderation, and logging, all of which were included and none of which arrive with the weights.
| What you control | Cloud API | Serverless | Self-hosted | Edge |
|---|---|---|---|---|
| Where data is processed | Provider | You (region) | You | The user’s device |
| Model weights and behaviour | Provider | Provider | You | You, and whoever holds the device |
| Guardrails and moderation | Provider | Shared | You | You |
| Audit logging | Provider (partial) | You | You | Largely unavailable |
| Patching the serving stack | Provider | Provider | You | Via app/firmware release |
The pattern to carry forward
Control you did not build is control you can lose by changing deployment. Every column you take over is a column you now have to staff. This is why Chapter 3’s Blueprint is organized by layer – data, models, infrastructure, users, access – rather than by product: the layers stay constant, but which ones are yours is set by the choice you make here.
Each pattern also opens a different attack surface, covered in Chapter 2:
- Cloud API and serverless – stolen API credentials used at your expense, and infrastructure-level attacks: Chapter 2, Section 4
- Self-hosted – malicious model files, and the full serving stack as attack surface: Chapter 2, Section 4
- Edge / on-device – weakened guardrails, weight extraction, and physical access: Chapter 2, Section 7
- Every pattern – prompt injection, which does not care where the model runs: Chapter 2, Section 2
Model examples on this page were verified in August 2026. The AI landscape moves fast. Model names below were verified at the date shown; the concepts they illustrate outlast any particular release. Always check a vendor's current documentation before making a deployment decision.
Cost Considerations
Token-based pricing is the standard for cloud API and serverless deployments, but costs vary dramatically across providers and models. Understanding the cost structure helps inform deployment decisions.
Rather than memorize prices that change every few months, learn the shape of the pricing. Every major provider offers roughly three tiers, and the ratios between them are remarkably stable:
| Tier | Typical relative cost | Good for |
|---|---|---|
| Small / fast | 1x (baseline) | Classification, routing, extraction, simple summarization |
| Mid / balanced | ~10-20x | General Q&A, drafting, most application workloads |
| Frontier / flagship | ~50-100x | Complex analysis, hard coding problems, agentic work |
| High reasoning effort | Adds to any tier | Multi-step problems – you pay for the reasoning tokens too |
Two further rules of thumb: output tokens cost several times more than input tokens, and the whole curve trends downward – equivalent capability has become dramatically cheaper year over year.
Self-hosting inverts the shape entirely. Instead of a per-request cost with no floor, you get a fixed cost with no ceiling on usage: the GPUs cost the same whether you serve one request or a million. That makes the decision a crossover calculation rather than a comparison – below some daily volume the API is cheaper, above it the hardware is, and the crossover point moves every time either side changes price.
Always Check Current Pricing
Deliberately, no dollar figures appear above. Provider pricing changes frequently – sometimes by a large factor in a single announcement – and any number printed in a course is wrong within months. Get current pricing from the provider’s own pricing page before making a deployment decision. What lasts is the ratio: a small model is one to two orders of magnitude cheaper than a frontier one, which is why routing by task complexity is the biggest cost lever you have.
Model Serialization and Security
When deploying models yourself, serialization stops being an implementation detail and becomes a security decision.
Concept: Serialization
Serialization is the process of converting a model into a format that can be saved to disk and later loaded into memory for inference. Every self-hosted or edge deployment does it. The choice of format determines whether “loading a model” is reading data or running code.
Serialization Formats
Two formats dominate LLM deployment, and the difference between them is not performance:
-
Pickle: Python’s native serialization format, long the default for PyTorch checkpoints. It is flexible because it can serialize arbitrary Python objects – and insecure for exactly the same reason. Deserializing a pickle file can execute arbitrary code by design, not as a bug. A tampered checkpoint runs whatever the attacker put in it, the moment you load it.
-
Safetensors: A format that stores only raw tensor data and a JSON header, with no mechanism for executing code on load. It has become the default across the Hugging Face Hub and joined the PyTorch Foundation in 2026, so the “immature ecosystem” objection that once counted against it no longer holds.
| Format | Advantages | Limitations |
|---|---|---|
| Pickle | Flexible; can serialize arbitrary Python objects | Arbitrary code execution on load – by design |
| Safetensors | No code execution on load; fast zero-copy loading; now the ecosystem default | Stores tensors only, so anything that relied on embedding Python objects needs reworking |
Prefer Safetensors. Treat a repository that ships only Pickle checkpoints as a finding to investigate, not a routine download.
Key Security Considerations
-
Verify the publisher, not the platform. “Download from Hugging Face” is not a trust decision – the Hub is an open upload platform where anyone can publish, and a typo-squatted repository looks much like the real one. What counts is the specific publishing organization, whether it is verified, and whether the file you got matches the hash the publisher states. The Hub does run automated scanning for unsafe pickle payloads, which is useful but is a filter, not a guarantee.
-
Isolate the load. Load model files in a sandboxed or containerized environment with no credentials and no outbound network access, so that if deserialization does execute something, it executes somewhere harmless. Where your framework supports restricting deserialization to tensor data only, turn it on.
-
Treat models as supply chain. Model repositories carry the same risks as package registries such as npm and PyPI, with less mature tooling. Pin versions, verify checksums and signatures, and keep a record of which artifact is running in production.
Security Preview
This is a live attack vector, not a theoretical one, and it is where AI security diverges sharply from prompt-level concerns: a tampered model file executes code the moment it is loaded, before a single prompt is sent. Chapter 2, Section 4 covers serialization exploits in depth, mapped to OWASP LLM04: Supply Chain.
Safety Guardrails
“The model refuses harmful requests” describes one layer of a stack, and the least controllable one. It helps to see the whole stack, because deployment choice determines which layers you actually have.
graph LR
U["User input"] --> A["1. Input moderation<br/>external classifier"]
A --> B["2. Refusal behaviour<br/>trained into the weights"]
B --> C["3. Output moderation<br/>external classifier"]
C --> D["4. Application policy<br/>your code"]
D --> O["Response"]
style U fill:#4a4a4a,color:#fff
style A fill:#a85800,color:#fff
style B fill:#2d5016,color:#fff
style C fill:#a85800,color:#fff
style D fill:#1565c0,color:#fff
style O fill:#4a4a4a,color:#fff
Only layer 2 travels with the weights. Layers 1 and 3 are services you either buy from a provider or build yourself – which is why a team that moves from a cloud API to self-hosting can silently lose two thirds of its guardrail stack while believing the model is “the same”. Layer 4 is always yours and is always required, because none of the layers below it know your business rules.
Refusal Pathways
Refusal pathways are the mechanisms that lead a model to decline harmful, unethical, or otherwise undesirable requests – the behaviour behind responses like “I’m sorry, but I can’t assist with that.” They are instilled during post-training alignment, chiefly through RLHF, rather than being a filter bolted on afterwards.
How Refusal Pathways Work
Refusal is not a rule the model looks up; it is a learned behaviour distributed across the weights. Interpretability research has found it to be strikingly concentrated: Arditi et al. (2024) showed that across 13 open-weight chat models up to 72B parameters, refusal is mediated by a single direction in the model’s residual stream. Erase that one direction from the activations and the model stops refusing harmful instructions; add it and the model refuses harmless ones.
That is a remarkable scientific result and an uncomfortable security one, and it explains a defence gap this section keeps returning to:
Why this matters for deployment
If a behaviour is mediated by one direction, and you hold the weights, you can remove it – cheaply, without retraining, and without much skill. This is the basis of “abliteration”, and it is why “the model will refuse” is not a control you can rely on for any open-weight deployment. It is a property of the artifact you shipped, and the artifact is now in someone else’s hands.
The finding is established on open-weight models, since it requires access to internal activations. Whether proprietary frontier models are organized the same way is not publicly verifiable – which is itself a reason not to treat refusal as a security boundary.
Challenges of Refusal Mechanisms
Refusal behaviour is necessary and also unreliable in three distinct ways:
-
Over-refusal. Models become overly conservative and decline legitimate work that superficially resembles harmful requests – security research, medical and pharmacological questions, chemistry education. For a security team, this is not a minor annoyance: it is the failure mode you will hit most often in your actual job.
-
Inconsistency. Refusal is highly sensitive to phrasing, and varies enormously between models. A benchmark of dual-use scientific requests found refusal rates on the same prompt set ranging from near-total refusal by one frontier model to no refusals at all by another. More tellingly, when the same request was rephrased five ways, a model’s own consistency fell from roughly 85% to 65%. A safety mechanism that answers differently depending on wording is not a boundary; it is a probability.
-
Removability on open weights. As above – weights you can download are weights whose refusal behaviour can be stripped. This is a direct, structural consequence of the deployment choice, not a defect in any particular model.
Bypassing refusal
Chapter 2, Section 2 covers jailbreaking and prompt injection – the techniques used to get past these mechanisms without touching the weights at all. For now the useful takeaway is the shape of the limitation: refusal behaviour is probabilistic, phrasing-sensitive, and, on open weights, removable.
Ethical and Policy Implications
Refusal tuning is where a vendor’s values get compiled into a product. Set it too loose and the model assists harm; too tight and it obstructs legitimate research, education, and defensive security work. There is no setting that satisfies everyone, which is part of why organizations with specialized needs end up self-hosting – and then owning the consequences.
Moderation Endpoints
Where refusal behaviour is embedded in the model, moderation endpoints sit outside it: separate classifiers that score content against categories of harm, applied to input before the model sees it and to output before the user does. They are layers 1 and 3 of the stack above.
How Moderation Endpoints Work
A moderation classifier evaluates content against predefined harm categories – hate speech, violence, sexual content, self-harm, and others. On a positive result, your application can:
- Block the response entirely
- Flag the content for human review
- Return a safe alternative response
The major cloud providers offer these as managed services, and open-weight safety classifiers exist for self-hosted and air-gapped deployments – which matters, because a self-hosted stack has no provider moderation to inherit and must supply this layer itself.
Advantages of Moderation Endpoints
- Independent of the model. Because moderation is a separate component, it still applies when the model is swapped, fine-tuned, or has had its refusal behaviour weakened. This is the layer that survives the failure modes above.
- Configurable thresholds. You choose the sensitivity per category, rather than accepting a vendor’s single global judgement.
- Applies to both directions. Screening output catches harm the model produced despite refusing nothing – including cases where the model was manipulated.
- Auditable. Moderation decisions are events you can log, count, and alert on. Refusals inside the model are not.
Challenges
- False positives and negatives. Automated classification of intent remains imperfect, and the thresholds need continuous tuning against real traffic.
- Latency and cost. Each check is another inference call in the request path, on both directions.
- Provider dependence. Managed moderation ties your policy to the provider’s category definitions and their changes to them.
- Privacy. Sending content to an external moderation service is another boundary crossing – one that is easy to overlook precisely because it is a security feature.
Neither layer is a boundary on its own
Refusal is removable and phrasing-sensitive; moderation is imperfect and can be routed around. Together they raise the cost of misuse rather than preventing it, which is exactly why Chapter 3 builds defence in layers instead of relying on any single control.
Key Takeaways
- Five deployment patterns – cloud API, serverless inference, self-hosted, edge/on-device, and hybrid – each with a distinct security and compliance profile. The first question is not budget or model quality, but whether the data can leave
- Deployment sets your trust boundary. Taking on a pattern means taking on every control the previous one supplied for free: rate limiting, abuse monitoring, moderation, and logging do not arrive with the weights
- Loading a model file is a code-execution decision. Pickle executes arbitrary code on load by design; Safetensors cannot, and is now the ecosystem default. Verify the publisher, not the platform
- Guardrails are a stack, not a switch. Only refusal behaviour travels with the model, and on open weights it can be removed – research shows it is mediated by a single direction in the activations
- Cost differs in shape, not just amount: per-token pricing has no floor, self-hosting has no ceiling on usage. Routing by task complexity is the largest lever, worth orders of magnitude more than any other optimization
- Data sovereignty and compliance requirements dictate deployment architecture more often than performance does
Test Your Knowledge
Ready to test your understanding of deployment considerations? Head to the quiz to check your knowledge.
Up next
You now know where a model can run and what that choice costs you in control. Next we open up the model itself – tokenization, embeddings, attention, and context windows – so that the attacks in Chapter 2 land on a mechanism you understand rather than a black box.