3. Deployment Considerations

Introduction

The AI landscape can be confusing when it comes to deployment choices, particularly because similar names often mask very different security and operational implications. For instance, when someone mentions “using GPT,” they might be referring to a vendor’s consumer chat interface, its developer API, or a cloud provider’s enterprise deployment of the same model – each with a vastly different security profile.

This distinction becomes especially important when evaluating AI solutions for enterprise use. Consider the controversy around DeepSeek: while some organizations banned its use over data privacy concerns, they often failed to distinguish between DeepSeek’s web platform (where processing happens on the vendor’s servers) and its open-weight models, which can be downloaded and run locally with full control over data flows. Those are two different risk decisions wearing the same name.

What will I get out of this?

By the end of this section, you will be able to:

  1. Differentiate the five deployment patterns – cloud API, serverless inference, self-hosted, edge/on-device, and hybrid – and describe the security, scalability, and compliance profile of each.
  2. Locate the trust boundary in a given deployment: where your data crosses out of your control, and which security controls you own versus inherit from a provider.
  3. Evaluate the trade-offs between capability, cost, control, and operational burden, and justify a deployment choice for a given scenario.
  4. Explain why loading a model file is a code-execution decision, comparing the Pickle and Safetensors serialization formats.
  5. Describe how safety guardrails are layered – model-internal refusal behaviour and external moderation – and where each one fails.

Why Deployment Choices Matter

The way a model is deployed fundamentally shapes its security, scalability, and compliance profile. The same model, reached three different ways, gives you three different systems to defend.

A web-based consumer platform offers ease of access but requires sending all input data to vendor-controlled servers. This raises questions about data residency, retention policies, and geopolitical risk depending on where those servers sit – and, critically, it is usually adopted by employees without a security review at all.

API services provide a middle ground. They let organizations integrate AI capabilities into their own systems while retaining control over what gets sent and what gets logged. Even so, APIs require careful scrutiny of terms of service – some providers retain data temporarily for abuse monitoring unless the account is explicitly configured otherwise.

Self-hosted open-weight models offer unparalleled control over data flows and compliance, but come with significant operational overhead: GPU infrastructure, technical expertise for setup and tuning, and ongoing monitoring. They also transfer every security responsibility onto you – including the ones a provider had been quietly handling.

Navigating Misconceptions

These distinctions also clear up two common misconceptions:

  • Naming a model does not describe a deployment. “We use ChatGPT” conflates a consumer web platform with the API that serves the same underlying model. Same technology, different trust boundary, different controls.
  • Banning a vendor is not the same as banning a model. Blocking a provider’s hosted service says nothing about whether its open weights can be run safely in an air-gapped environment – which is often exactly where they belong.
The question to actually ask

Not “which model is best?” but “where does my data go, and who holds each control?” Every pattern below is a different answer to that question. Keep it in mind as you read – it is also the question Chapters 2 and 3 are built on.


Model Deployment Options

Deploying large language models presents unique challenges and opportunities compared to traditional software systems. There are more deployment patterns available than ever before, each with distinct trade-offs.

graph TD
    A["Can your data leave<br/>your infrastructure?"] -->|"No: regulated,<br/>classified, or air-gapped"| B{"Does a small model<br/>meet the requirement?"}
    A -->|"Yes, within a<br/>trusted boundary"| C{"Sustained<br/>high volume?"}
    B -->|"Yes"| D["Edge / On-Device"]
    B -->|"No: needs a<br/>larger model"| E["Self-Hosted<br/>(on-prem or private cloud)"]
    C -->|"Yes: API spend<br/>exceeds GPU cost"| F["Self-Hosted<br/>or Hybrid"]
    C -->|"No: variable or<br/>low volume"| G{"Need data residency<br/>or a single vendor bill?"}
    G -->|"Yes"| H["Serverless Inference<br/>(Bedrock / Foundry / Vertex)"]
    G -->|"No"| I["Cloud API<br/>(pay-per-token)"]

    style A fill:#1565c0,color:#fff
    style B fill:#1565c0,color:#fff
    style C fill:#1565c0,color:#fff
    style G fill:#1565c0,color:#fff
    style D fill:#2d5016,color:#fff
    style E fill:#2d5016,color:#fff
    style F fill:#a85800,color:#fff
    style H fill:#2d5016,color:#fff
    style I fill:#2d5016,color:#fff

Note what the first question is. Not budget, not model quality – can the data leave? That constraint eliminates more options than any other, and it is the one you cannot engineer around later.

The five patterns at a glance

Read this table before opening the detail below. The patterns only make sense against each other.

Pattern Data boundary Cost shape Capability ceiling Operational burden Primary security concern
Cloud API Leaves your network entirely; provider-controlled Per token, no floor Highest – newest frontier models Near zero Data exposure to a third party; credential theft
Serverless inference Stays inside your cloud tenancy Per token, often at a premium High – broad catalogue, slight lag on newest Low Misconfigured tenancy and IAM; over-broad access
Self-hosted Never leaves your infrastructure Fixed: hardware plus power, idle or not Limited to open-weight models High – you run everything Malicious model files; you own the entire stack
Edge / on-device Never leaves the device Fixed per device, no per-request cost Lowest – small models only Medium – fleet updates Physical access to weights; weakened guardrails
Hybrid Split by design; depends on routing Mixed, optimized by routing Highest tier available on demand Highest – two or more stacks Routing logic itself becomes a trust decision
Cloud API (Hosted Inference)

How it works: You send requests to a provider’s API endpoint. The provider manages all infrastructure, scaling, and model updates.

When to use:

  • Rapid prototyping and getting started quickly
  • Applications that need the latest models without infrastructure investment
  • Variable or unpredictable workloads

Advantages:

  • Zero infrastructure management
  • Immediate access to the newest model versions
  • Pay-per-use pricing, so no idle costs
  • Safety features and moderation are supplied by the provider

Considerations:

  • Data leaves your network for processing
  • Token-based billing can become expensive at scale
  • Vendor lock-in risk
  • Rate limits may constrain high-volume applications
Provider-Managed APIs

Each provider has different data retention policies, rate limits, and pricing. Retention is the one to check first: several providers hold inputs for a period for abuse monitoring by default, and turning that off may require a specific enterprise agreement rather than a settings toggle. Always review the data processing agreement before sending anything sensitive.

Serverless Inference (Cloud-Managed)

How it works: Your existing cloud provider hosts models on your behalf. You select from a catalogue, and the platform handles deployment, scaling, and access management. Requests stay within your own tenancy on that cloud, under the IAM, logging, and networking controls you already use.

When to use:

  • Enterprise deployments requiring data residency controls
  • Organizations already invested in a cloud provider ecosystem
  • Need for multiple model providers through a single interface and a single bill

Current platforms:

Platform Provider Key characteristic
Amazon Bedrock AWS Unified API across many third-party model families, with VPC integration, guardrails, and IAM-native access control
Microsoft Foundry Microsoft A single catalogue spanning OpenAI’s models alongside open-weight and third-party options, with content filtering, RBAC, and regional data residency. Formerly Azure AI Foundry, and before that Azure OpenAI Service
Google Vertex AI Google Gemini models plus a Model Garden of open-weight options, integrated with Google Cloud IAM

Advantages:

  • Enterprise security controls and compliance certifications you may already be audited against
  • Data stays within your cloud provider’s boundary and chosen region
  • No GPU management or model deployment overhead
  • Unified billing and identity through an existing cloud account

Considerations:

  • Higher per-token cost than direct API access in some cases
  • Vendor lock-in to a specific cloud provider
  • Model availability varies by platform and region, and newest releases often land on the vendor’s own API first
  • The security boundary is now your cloud IAM configuration – a misconfigured role is the whole control
Self-Hosted (On-Premises or Private Cloud)

How it works: You download model weights and run inference on your own hardware. Common serving stacks include vLLM, Ollama, llama.cpp, and TGI (Text Generation Inference).

When to use:

  • Strict data sovereignty requirements (healthcare, defense, finance)
  • High-volume inference where API costs would be prohibitive
  • Need for full control over model behaviour, including custom fine-tuning
  • Air-gapped environments with no internet connectivity

Advantages:

  • Complete control over data – nothing leaves your infrastructure
  • No per-token costs, just hardware and power
  • Full customization: fine-tune, adjust serving parameters, change safety settings
  • Runs air-gapped

Considerations:

  • Significant upfront hardware investment; GPUs are expensive and often supply-constrained
  • Requires ML engineering expertise for deployment and optimization
  • You are responsible for security patches, model updates, monitoring, and abuse detection
  • Limited to open-weight models – currently Alibaba's Qwen 3.x (Apache 2.0), DeepSeek-V4 (MIT), Mistral's open models (Apache 2.0), and Microsoft's Phi-4 family (MIT)
  • You inherit the provider’s security work. Rate limiting, abuse monitoring, input and output moderation, and audit logging were all being done for you. None of them come with the weights
Hardware Requirements

A rough guide, scaling by parameter count rather than model name. Quantization – storing weights at reduced numerical precision – shifts these down substantially, at some cost to quality:

  • ~7B parameters: Single consumer GPU (e.g. 24GB VRAM)
  • ~13-34B: Two enterprise GPUs, or one high-VRAM consumer card
  • ~70B: Multiple enterprise GPUs (A100/H100 class), or quantized onto fewer
  • Frontier-scale (hundreds of billions): A GPU cluster with specialized interconnect. Mixture-of-Experts designs reduce the compute per request, but you still need memory for all the weights
Edge and On-Device AI

How it works: Small Language Models run directly on end-user devices – smartphones, laptops, IoT devices, or embedded systems. Models are compressed through quantization and distillation to fit hardware constraints.

When to use:

  • Privacy-critical applications where data must never leave the device
  • Offline functionality, with no network dependency
  • Latency-sensitive interactions, where a network round trip is the bottleneck
  • Cost-sensitive deployments at massive scale, with no per-request API costs

Current reality:

  • Apple Intelligence runs on-device models for summarization, image generation, and smart replies across iPhone and Mac, escalating harder requests to a cloud tier
  • Google ships an on-device Nano tier through Android’s AICore, powering local features in Pixel devices and Chrome
  • Within the SLM tier from Section 2, the upper end runs on laptops and tablets while the lower end runs on phones – the constraint is device memory, not the model’s name

Advantages:

  • Zero data transmission for anything handled locally
  • No internet required
  • No network round trip, so consistently low interactive latency
  • Scales to millions of devices without server-side inference cost

Considerations:

  • Limited to small models, so a lower capability ceiling than any cloud option
  • Hardware fragmentation across the device fleet
  • Model updates ship as app or firmware updates, so patching is slow
  • Physical access is part of the threat model. The weights are on a device you do not control, which makes extraction and guardrail removal far easier than in any server-side deployment
Smaller is not safer

It is tempting to treat a small model on a phone as a small risk. The opposite is often true: SLMs tend to carry weaker safety training than frontier models, aggressive quantization can degrade that training further, and an attacker holding the device holds the weights. Chapter 2, Section 7 is dedicated to this.

Hybrid Approach

How it works: Combine multiple patterns based on task requirements. Route simple or sensitive queries to local models and complex ones to cloud APIs.

When to use:

  • Applications with widely varying task complexity
  • Organizations balancing cost, performance, and privacy
  • Systems needing both offline capability and frontier model access

Example architecture:

  1. Edge model (a small model on the device): auto-complete, simple classification, basic summarization
  2. Self-hosted model (a mid-size open-weight model on an internal GPU cluster): sensitive documents and medium-complexity tasks
  3. Cloud API (a frontier model, direct or via a serverless platform): complex analysis, reasoning, and work needing the latest capabilities

Advantages:

  • Routes each request to the cheapest capable option
  • Keeps sensitive operations inside your boundary while retaining frontier access for the rest
  • Provides fallback if one tier is unavailable
  • Balances latency requirements across use cases

Considerations:

  • Complex routing logic required
  • Testing multiplies across models and deployment targets
  • Orchestration overhead
  • The router becomes a security control. It decides which data crosses which boundary, so a routing bug is a data leak, and a router that can be influenced by the input it is classifying is an attack surface

Where the Trust Boundary Sits

The five patterns differ on cost and capability, but the difference that matters for the rest of this course is who holds which control. Deployment is the decision that sets your trust boundary, and everything Chapters 2 and 3 discuss happens relative to that line.

Three questions locate it:

  1. Where does the data go? Out to a vendor, into your cloud tenancy, onto your own hardware, or nowhere at all. This determines your exposure if the other party is breached – a risk you cannot patch.
  2. Who can see the traffic? Provider-side logging and retention, your own audit trail, or neither. You cannot investigate an incident in a system that logs nothing.
  3. Which controls do you own, and which are you inheriting silently? This is the one teams get wrong. Moving from an API to self-hosting does not just transfer the hardware bill – it transfers rate limiting, abuse detection, input and output moderation, and logging, all of which were included and none of which arrive with the weights.
What you control Cloud API Serverless Self-hosted Edge
Where data is processed Provider You (region) You The user’s device
Model weights and behaviour Provider Provider You You, and whoever holds the device
Guardrails and moderation Provider Shared You You
Audit logging Provider (partial) You You Largely unavailable
Patching the serving stack Provider Provider You Via app/firmware release
The pattern to carry forward

Control you did not build is control you can lose by changing deployment. Every column you take over is a column you now have to staff. This is why Chapter 3’s Blueprint is organized by layer – data, models, infrastructure, users, access – rather than by product: the layers stay constant, but which ones are yours is set by the choice you make here.

Each pattern also opens a different attack surface, covered in Chapter 2:

  • Cloud API and serverless – stolen API credentials used at your expense, and infrastructure-level attacks: Chapter 2, Section 4
  • Self-hosted – malicious model files, and the full serving stack as attack surface: Chapter 2, Section 4
  • Edge / on-device – weakened guardrails, weight extraction, and physical access: Chapter 2, Section 7
  • Every pattern – prompt injection, which does not care where the model runs: Chapter 2, Section 2

Model examples on this page were verified in August 2026. The AI landscape moves fast. Model names below were verified at the date shown; the concepts they illustrate outlast any particular release. Always check a vendor's current documentation before making a deployment decision.


Cost Considerations

Token-based pricing is the standard for cloud API and serverless deployments, but costs vary dramatically across providers and models. Understanding the cost structure helps inform deployment decisions.

Rather than memorize prices that change every few months, learn the shape of the pricing. Every major provider offers roughly three tiers, and the ratios between them are remarkably stable:

Tier Typical relative cost Good for
Small / fast 1x (baseline) Classification, routing, extraction, simple summarization
Mid / balanced ~10-20x General Q&A, drafting, most application workloads
Frontier / flagship ~50-100x Complex analysis, hard coding problems, agentic work
High reasoning effort Adds to any tier Multi-step problems – you pay for the reasoning tokens too

Two further rules of thumb: output tokens cost several times more than input tokens, and the whole curve trends downward – equivalent capability has become dramatically cheaper year over year.

Self-hosting inverts the shape entirely. Instead of a per-request cost with no floor, you get a fixed cost with no ceiling on usage: the GPUs cost the same whether you serve one request or a million. That makes the decision a crossover calculation rather than a comparison – below some daily volume the API is cheaper, above it the hardware is, and the crossover point moves every time either side changes price.

Always Check Current Pricing

Deliberately, no dollar figures appear above. Provider pricing changes frequently – sometimes by a large factor in a single announcement – and any number printed in a course is wrong within months. Get current pricing from the provider’s own pricing page before making a deployment decision. What lasts is the ratio: a small model is one to two orders of magnitude cheaper than a frontier one, which is why routing by task complexity is the biggest cost lever you have.

Cost Optimization Strategies

Strategies to reduce AI costs:

  1. Model selection by task: Use small-tier models for simple tasks. Reserve frontier models for complex analysis. This is the largest lever by a wide margin – it changes cost by orders of magnitude, while everything below changes it by fractions.
  2. Prompt caching: Most providers charge a reduced rate for repeated prompt prefixes, which is substantial for applications with a large fixed system prompt.
  3. Batching: Group non-urgent requests for bulk pricing, typically in exchange for a slower turnaround.
  4. Output length control: Set an appropriate max_tokens – output tokens are the expensive ones.
  5. Self-hosting for volume: Above the crossover point, fixed hardware cost beats per-token pricing. Budget for the engineering time too, not just the GPUs.
  6. Hybrid routing: Automate the routing decision rather than choosing one tier for everything.
Cost control is also a security control

Per-token billing means an attacker who can make your application generate tokens can generate a bill. Unbounded consumption and credential abuse are covered in Chapter 2, Section 4 – the max_tokens limit above is a defensive measure as much as a budgetary one.


Model Serialization and Security

When deploying models yourself, serialization stops being an implementation detail and becomes a security decision.

Concept: Serialization

Serialization is the process of converting a model into a format that can be saved to disk and later loaded into memory for inference. Every self-hosted or edge deployment does it. The choice of format determines whether “loading a model” is reading data or running code.

Serialization Formats

Two formats dominate LLM deployment, and the difference between them is not performance:

  • Pickle: Python’s native serialization format, long the default for PyTorch checkpoints. It is flexible because it can serialize arbitrary Python objects – and insecure for exactly the same reason. Deserializing a pickle file can execute arbitrary code by design, not as a bug. A tampered checkpoint runs whatever the attacker put in it, the moment you load it.

  • Safetensors: A format that stores only raw tensor data and a JSON header, with no mechanism for executing code on load. It has become the default across the Hugging Face Hub and joined the PyTorch Foundation in 2026, so the “immature ecosystem” objection that once counted against it no longer holds.

Format Advantages Limitations
Pickle Flexible; can serialize arbitrary Python objects Arbitrary code execution on load – by design
Safetensors No code execution on load; fast zero-copy loading; now the ecosystem default Stores tensors only, so anything that relied on embedding Python objects needs reworking

Prefer Safetensors. Treat a repository that ships only Pickle checkpoints as a finding to investigate, not a routine download.

Key Security Considerations

  1. Verify the publisher, not the platform. “Download from Hugging Face” is not a trust decision – the Hub is an open upload platform where anyone can publish, and a typo-squatted repository looks much like the real one. What counts is the specific publishing organization, whether it is verified, and whether the file you got matches the hash the publisher states. The Hub does run automated scanning for unsafe pickle payloads, which is useful but is a filter, not a guarantee.

  2. Isolate the load. Load model files in a sandboxed or containerized environment with no credentials and no outbound network access, so that if deserialization does execute something, it executes somewhere harmless. Where your framework supports restricting deserialization to tensor data only, turn it on.

  3. Treat models as supply chain. Model repositories carry the same risks as package registries such as npm and PyPI, with less mature tooling. Pin versions, verify checksums and signatures, and keep a record of which artifact is running in production.

Security Preview

This is a live attack vector, not a theoretical one, and it is where AI security diverges sharply from prompt-level concerns: a tampered model file executes code the moment it is loaded, before a single prompt is sent. Chapter 2, Section 4 covers serialization exploits in depth, mapped to OWASP LLM04: Supply Chain.


Safety Guardrails

“The model refuses harmful requests” describes one layer of a stack, and the least controllable one. It helps to see the whole stack, because deployment choice determines which layers you actually have.

graph LR
    U["User input"] --> A["1. Input moderation<br/>external classifier"]
    A --> B["2. Refusal behaviour<br/>trained into the weights"]
    B --> C["3. Output moderation<br/>external classifier"]
    C --> D["4. Application policy<br/>your code"]
    D --> O["Response"]

    style U fill:#4a4a4a,color:#fff
    style A fill:#a85800,color:#fff
    style B fill:#2d5016,color:#fff
    style C fill:#a85800,color:#fff
    style D fill:#1565c0,color:#fff
    style O fill:#4a4a4a,color:#fff

Only layer 2 travels with the weights. Layers 1 and 3 are services you either buy from a provider or build yourself – which is why a team that moves from a cloud API to self-hosting can silently lose two thirds of its guardrail stack while believing the model is “the same”. Layer 4 is always yours and is always required, because none of the layers below it know your business rules.

Refusal Pathways

Refusal pathways are the mechanisms that lead a model to decline harmful, unethical, or otherwise undesirable requests – the behaviour behind responses like “I’m sorry, but I can’t assist with that.” They are instilled during post-training alignment, chiefly through RLHF, rather than being a filter bolted on afterwards.

How Refusal Pathways Work

Refusal is not a rule the model looks up; it is a learned behaviour distributed across the weights. Interpretability research has found it to be strikingly concentrated: Arditi et al. (2024) showed that across 13 open-weight chat models up to 72B parameters, refusal is mediated by a single direction in the model’s residual stream. Erase that one direction from the activations and the model stops refusing harmful instructions; add it and the model refuses harmless ones.

That is a remarkable scientific result and an uncomfortable security one, and it explains a defence gap this section keeps returning to:

Why this matters for deployment

If a behaviour is mediated by one direction, and you hold the weights, you can remove it – cheaply, without retraining, and without much skill. This is the basis of “abliteration”, and it is why “the model will refuse” is not a control you can rely on for any open-weight deployment. It is a property of the artifact you shipped, and the artifact is now in someone else’s hands.

The finding is established on open-weight models, since it requires access to internal activations. Whether proprietary frontier models are organized the same way is not publicly verifiable – which is itself a reason not to treat refusal as a security boundary.

Challenges of Refusal Mechanisms

Refusal behaviour is necessary and also unreliable in three distinct ways:

  1. Over-refusal. Models become overly conservative and decline legitimate work that superficially resembles harmful requests – security research, medical and pharmacological questions, chemistry education. For a security team, this is not a minor annoyance: it is the failure mode you will hit most often in your actual job.

  2. Inconsistency. Refusal is highly sensitive to phrasing, and varies enormously between models. A benchmark of dual-use scientific requests found refusal rates on the same prompt set ranging from near-total refusal by one frontier model to no refusals at all by another. More tellingly, when the same request was rephrased five ways, a model’s own consistency fell from roughly 85% to 65%. A safety mechanism that answers differently depending on wording is not a boundary; it is a probability.

  3. Removability on open weights. As above – weights you can download are weights whose refusal behaviour can be stripped. This is a direct, structural consequence of the deployment choice, not a defect in any particular model.

Bypassing refusal

Chapter 2, Section 2 covers jailbreaking and prompt injection – the techniques used to get past these mechanisms without touching the weights at all. For now the useful takeaway is the shape of the limitation: refusal behaviour is probabilistic, phrasing-sensitive, and, on open weights, removable.

Ethical and Policy Implications

Refusal tuning is where a vendor’s values get compiled into a product. Set it too loose and the model assists harm; too tight and it obstructs legitimate research, education, and defensive security work. There is no setting that satisfies everyone, which is part of why organizations with specialized needs end up self-hosting – and then owning the consequences.

Moderation Endpoints

Where refusal behaviour is embedded in the model, moderation endpoints sit outside it: separate classifiers that score content against categories of harm, applied to input before the model sees it and to output before the user does. They are layers 1 and 3 of the stack above.

How Moderation Endpoints Work

A moderation classifier evaluates content against predefined harm categories – hate speech, violence, sexual content, self-harm, and others. On a positive result, your application can:

  • Block the response entirely
  • Flag the content for human review
  • Return a safe alternative response

The major cloud providers offer these as managed services, and open-weight safety classifiers exist for self-hosted and air-gapped deployments – which matters, because a self-hosted stack has no provider moderation to inherit and must supply this layer itself.

Advantages of Moderation Endpoints

  1. Independent of the model. Because moderation is a separate component, it still applies when the model is swapped, fine-tuned, or has had its refusal behaviour weakened. This is the layer that survives the failure modes above.
  2. Configurable thresholds. You choose the sensitivity per category, rather than accepting a vendor’s single global judgement.
  3. Applies to both directions. Screening output catches harm the model produced despite refusing nothing – including cases where the model was manipulated.
  4. Auditable. Moderation decisions are events you can log, count, and alert on. Refusals inside the model are not.

Challenges

  1. False positives and negatives. Automated classification of intent remains imperfect, and the thresholds need continuous tuning against real traffic.
  2. Latency and cost. Each check is another inference call in the request path, on both directions.
  3. Provider dependence. Managed moderation ties your policy to the provider’s category definitions and their changes to them.
  4. Privacy. Sending content to an external moderation service is another boundary crossing – one that is easy to overlook precisely because it is a security feature.
Neither layer is a boundary on its own

Refusal is removable and phrasing-sensitive; moderation is imperfect and can be routed around. Together they raise the cost of misuse rather than preventing it, which is exactly why Chapter 3 builds defence in layers instead of relying on any single control.

Key Takeaways
  • Five deployment patterns – cloud API, serverless inference, self-hosted, edge/on-device, and hybrid – each with a distinct security and compliance profile. The first question is not budget or model quality, but whether the data can leave
  • Deployment sets your trust boundary. Taking on a pattern means taking on every control the previous one supplied for free: rate limiting, abuse monitoring, moderation, and logging do not arrive with the weights
  • Loading a model file is a code-execution decision. Pickle executes arbitrary code on load by design; Safetensors cannot, and is now the ecosystem default. Verify the publisher, not the platform
  • Guardrails are a stack, not a switch. Only refusal behaviour travels with the model, and on open weights it can be removed – research shows it is mediated by a single direction in the activations
  • Cost differs in shape, not just amount: per-token pricing has no floor, self-hosting has no ceiling on usage. Routing by task complexity is the largest lever, worth orders of magnitude more than any other optimization
  • Data sovereignty and compliance requirements dictate deployment architecture more often than performance does

Test Your Knowledge

Ready to test your understanding of deployment considerations? Head to the quiz to check your knowledge.


Up next

You now know where a model can run and what that choice costs you in control. Next we open up the model itself – tokenization, embeddings, attention, and context windows – so that the attacks in Chapter 2 land on a mechanism you understand rather than a black box.