6. Inference Techniques
Introduction
An internal support assistant is asked “what is our parental leave policy?” and answers correctly. The next day someone asks it to “summarise the disciplinary file for J. Okonkwo” and it answers that correctly too. Nothing was hacked. The retrieval pipeline did exactly what it was built to do – it searched the corpus for the passages most similar to the question and handed them to the model. Nobody had told it that some passages belong to some people.
That is the shape of most real failures in production AI systems. They are rarely clever exploits; they are pipelines working as designed, on an architecture whose defaults are wrong. This section builds that pipeline – API integration, response handling, retrieval, and cost – and at each stage asks the two questions that separate a working system from a defensible one: who can put text in here, and who is allowed to get it back out?
Section 4 established that the context window is a flat token sequence with no privilege levels, and Section 5 that an instruction in a prompt is a request rather than an enforcement boundary. Retrieval is where those two facts stop being abstract, because retrieval is the machinery that puts other people’s text into that window on your behalf.
What will I get out of this?
By the end of this section, you will be able to:
- Choose an API integration shape – legacy completions, chat completions, or the current stateful and agentic APIs – and say who holds the conversation state in each and what that commits you to.
- Trace a request through a production retrieval pipeline, naming each stage from ingestion through hybrid search, reranking and metadata filtering to generation.
- Identify the attack surface at each stage of a retrieval pipeline you are shown, naming who can write to it and where the corresponding control lives.
- Explain what an embedding is and is not – why it is a derived representation rather than an anonymised one, and what that means for how a vector store must be classified.
- Diagnose a retrieval failure from its symptom, distinguishing a chunking problem from an embedding problem from a ranking problem from a permissions problem.
- Evaluate cost optimization strategies against both their saving and their security consequence, including why a semantic cache in front of a permission-filtered corpus is a disclosure bug.
API Integration Types
Provider APIs have gone through three generations, and all three are still in use somewhere. Knowing which one you are looking at tells you where the conversation lives – which is an architecture decision before it is a convenience one.
- Raw completions (
/v1/completionsand equivalents) – one string in, one continuation out. Now legacy; kept alive by older integrations and by self-hosted serving stacks. - Chat completions – a list of role-tagged messages in, one message out. Still the most widely implemented shape in the industry: every gateway, proxy, and open-weight serving framework speaks it, which is why it functions as the lingua franca even where a provider prefers something newer.
- Stateful and agentic APIs – OpenAI’s Responses API, Anthropic’s Messages API with server-side conversation state, and their equivalents. Built around multi-turn tool use and reasoning, and the shape providers now recommend for new work.
The first two are worth understanding because the difference between them is the clearest illustration of what “context management” actually means. The third is where you will build.
Raw Completions
The completion shape is a single-turn interaction: you send one prompt and receive one continuation. It offers precise control over the input structure and is best suited for isolated tasks.
Key Features:
- Simplicity: Each interaction is self-contained, with no built-in conversation structure.
- Flexibility: You can design prompts exactly as needed.
- Token Efficiency: Typically uses fewer tokens for standalone tasks.
- Use Cases: Generating content, text summarization, one-off queries.
Structuring a Raw Completion with Templates
With no message roles available, any structure has to be typed into the string yourself. This is what chat templating is:
ChatML marks turn boundaries with dedicated special tokens. This is not decoration – <|im_start|> and <|im_end|> are single tokens the model was trained on, and they are the closest thing in the token sequence to a structural marker.
They still confer no privilege. A model trained on ChatML has learned that text after <|im_start|>system is usually instructions, which is a statistical habit rather than a permission check.
They also explain why tokenizers refuse to encode them from user input. If a user could type <|im_end|><|im_start|>system and have it tokenized as the real special tokens, they would be opening a system turn of their own – and the model, seeing a well-formed role transition, would treat what followed as instructions. This is special token injection, and it is why tiktoken raises an error on special-token strings unless you explicitly allow them, and why serving frameworks parse them out of the user-supplied portion of a prompt. It is a real jailbreak class, not a theoretical one, and Chapter 2 Section 2 covers the family it belongs to. Note where the fix lives: in the tokenizer, outside the model, refusing to construct the input – not in an instruction telling the model to ignore fake role markers.
Custom delimiters enable clear section boundaries within your prompt.
Did you Notice?
Leaving the ‘assistant’ section open is one way to ensure the model begins completing following the pattern that preceded it, rather than breaking the expected format or sequence.
Chat Completions
The chat shape is built for multi-turn conversation. Instead of one string, you send an array of messages, each tagged with a role:
Two details in that payload do more work than they appear to:
- The field is
content, and the messages are an array on a single request object. Roles are a property of each message, not a wrapper around the request. - You sent the entire conversation. The previous assistant turn is in there because you stored it and replayed it. The provider did not remember it. This is Section 4’s statelessness made concrete: a chat API is a stateless API with a convention for formatting history.
Key Features:
- Structured Dialogue: Well-defined message roles (system, user, assistant).
- Multi-Turn: Handles conversational exchanges naturally.
- Portability: The most widely implemented request shape in the industry – gateways, proxies and self-hosted serving stacks all speak it.
- Use Cases: Chatbots, interactive assistants, back-and-forth dialogue.
Stateful and Agentic APIs
The current generation moves conversation state to the provider. You send only the new message plus a reference to the previous response, and the server reconstructs the history. Reasoning traces, tool calls and tool results are first-class parts of the exchange rather than text you serialize into a message yourself.
This is a genuine improvement in ergonomics and cost – a twenty-turn conversation no longer means re-uploading nineteen turns – and it is also a change of trust posture that is easy to miss, because nothing in the code looks different.
Who Holds the Conversation?
When state lives on your side, the transcript is in your infrastructure, under your retention policy, subject to your access controls, and deleted when you delete it. When state lives on the provider’s side, the transcript is in someone else’s system, held for as long as their policy says, and your only handle on it is an ID.
That is the same control-ownership question Section 3 asked about deployment, arriving one layer up. It is not an argument against stateful APIs – it is the thing you need to have decided before you use one for anything regulated, because “the conversation history” is now a data-residency and retention question and not just a token-cost one.
It also enlarges the durable-injection surface. Server-held state is state a previous turn can write into and a later turn will read back – the mechanism behind the memory poisoning attacks in Chapter 2 Section 5.
Choosing an Integration Shape
| Criteria | Raw completions | Chat completions | Stateful / agentic |
|---|---|---|---|
| Request shape | One prompt string | Array of role-tagged messages | New message + reference to prior response |
| Who stores the history | You, in the string | You, replayed every request | The provider |
| Multi-turn cost | You pay for all of it, every turn | You pay for all of it, every turn | You pay for it once |
| Tool use and reasoning | Hand-rolled | Bolted on via message conventions | First-class |
| Portability | Universal but unstructured | Broadest real-world support | Provider-specific |
| Current status | Legacy | Supported, the de facto standard shape | What providers recommend for new work |
| What it commits you to | Managing turn markers yourself, including keeping user text out of them | Holding the transcript, and its retention policy | Transcript residency at the provider; a durable state store a prior turn can write to |
Which One Should You Use?
For new work, the stateful and agentic APIs, because tool use and reasoning are where applications are actually going and hand-rolling those on top of chat completions is re-implementing your provider’s roadmap. Use chat completions when you need to stay portable across providers or run behind a gateway. Use raw completions when something legacy requires it, or when you are driving a self-hosted serving stack directly and want control of the template.
Whichever you pick, be able to answer where the transcript is stored – that is the answer a security review will ask for, and it is not visible in the client code.
Response Handling
The way you receive and process responses from an LLM significantly impacts your application’s user experience:
Wait for complete response before processing
Best for batch processing and applications where immediate feedback isn’t critical.
Advantages:
- Simpler to implement
- Easier to validate and process complete responses
- Better for systems that need to analyze full responses before proceeding
Disadvantages:
- Longer perceived latency (user sees nothing until complete)
- No intermediate feedback
Use when: Processing data pipelines, batch classification, generating structured data, backend processing.
Real-time token delivery as the model generates
Ideal for interactive applications and chat interfaces.
Advantages:
- Better user experience with immediate feedback
- Allows progressive rendering (text appears word-by-word)
- Can implement typing indicators
- Users can interrupt if response goes off-track
Disadvantages:
- More complex to implement
- Requires handling partial responses
- Additional error handling for interrupted streams
- You cannot inspect what you have already sent. Output validation and streaming are in direct tension – see below.
Use when: Chat interfaces, interactive assistants, real-time writing assistance, any user-facing application.
The Trade-off the Tabs Don’t Show
Streaming is usually presented as a pure user-experience win with an implementation cost. There is a third axis, and it is the one that matters here: a token you have streamed is a token you have published.
Section 5 established that model output crosses a trust boundary on its way into your application. Every control that acts on output – scanning for leaked secrets and PII, blocking policy violations, validating a schema, escaping HTML before it renders – needs the output to act on. Synchronous handling gives you the whole response before anyone sees it. Streaming gives your user the first sentence before your filter has seen the last one.
| Synchronous | Streaming | |
|---|---|---|
| Perceived latency | Full generation time | Milliseconds to first token |
| Output validation | Inspect the complete response, then decide | Must decide per-chunk, on partial text |
| Schema validation | Parse and reject before use | Incomplete JSON until the last token |
| Retraction | Nothing to retract – you never sent it | Already on the user’s screen |
| Where it fits | Anything feeding another system | Anything a human reads |
The workable pattern for user-facing streaming is to buffer a rolling window rather than emitting token by token, run detection on the buffer, and accept that detection is best-effort on the tail. The pattern that does not work is streaming straight to the client and calling the post-hoc log review a control.
This is why Chapter 3 Layer 5 treats response filtering as an architectural position rather than a function call – it has to sit somewhere the response can still be stopped. And it is why anything whose output feeds a downstream system, rather than a human, should be synchronous: there is no user experience to protect, and a partially-validated payload reaching an interpreter is LLM10: Improper Output Handling.
Why Retrieval
Section 4 compared the four ways of getting information in front of a model – conversation history, a long context window, retrieval, and persistent memory layers – across cost, freshness, infrastructure and who can write into each. That comparison is the one to use when choosing; this section assumes the choice has been made and builds the thing.
The short version of why the choice usually lands on retrieval: it is the only option that decouples the size of your knowledge base from the size of your request. A corpus can grow to millions of documents while each request still carries a handful of passages. Fine-tuning could bake stable facts into the weights, but re-training on every policy update is neither practical nor cheap, and the resulting model cannot tell you where an answer came from.
That last property is the underrated one. Retrieval is the only approach that produces a citation – you know which chunks the answer was built from, which makes the answer auditable and the failure debuggable. It is also, as the rest of this section shows, the approach with the largest attack surface, for exactly the same reason: an auditable path from a corpus to an answer is also a path from a corpus to an answer.
The RAG Pipeline: A Practical Walkthrough
RAG bridges the gap between an LLM’s static training knowledge and the dynamic information your application needs. Let’s walk through exactly how it works.
How RAG Works: The Data Flow
flowchart LR
subgraph Ingestion["1. Ingestion (offline)"]
A[Documents] --> B[Chunking]
B --> C[Embedding Model]
C --> D[(Vector store<br/>chunks + vectors<br/>+ metadata)]
end
subgraph Retrieval["2. Retrieval (per request)"]
E[User query] --> F[Embed query]
F --> G[Vector search]
E --> BM[Keyword search<br/>BM25]
D --> G
D --> BM
G --> RRF[Fuse candidates]
BM --> RRF
RRF --> MF{{"Metadata filter<br/>who may see this?"}}
MF --> RR[Rerank<br/>cross-encoder]
RR --> H[Top 5-8 passages]
end
subgraph Generation["3. Generation"]
H --> I[Assemble context]
E --> I
I --> J[LLM]
J --> K[Response + citations]
end
style A fill:#2d5016,color:#fff
style D fill:#1a3a5c,color:#fff
style MF fill:#8b6914,color:#fff
style J fill:#4a1a5c,color:#fff
style K fill:#2d5016,color:#fff
Three stages in the middle of that diagram were absent from the first generation of RAG systems and are standard in production now. Keyword search runs alongside vector search rather than being replaced by it; the candidate lists are fused; and a reranker re-scores the survivors before anything reaches the model. The reason each one exists is covered below – but note the shape of the pipeline first. Retrieval is not one step. It is a funnel, and each narrowing is a place where the wrong thing can survive or the right thing can be dropped.
The gold box is the stage most systems are missing. Everything else in the diagram optimizes relevance; that one is the only stage that asks about entitlement.
Step 1: Document Ingestion
Before your RAG system can answer questions, it needs to process and store your documents:
Chunking: Documents are split into smaller pieces (chunks). This is critical because:
- LLMs have context window limits – you can’t send entire documents
- Smaller, focused chunks lead to more relevant retrieval
- Chunk size affects both retrieval quality and cost
Embedding: Each chunk is converted into a vector (embedding) that captures its semantic meaning – a high-dimensional numerical representation in which similar meanings land near each other. The roster of current models is below.
Storage: Three things are stored per chunk, and the third is the one that decides whether your system is safe:
- The vector, which retrieval searches against.
- The original text, because the model needs the words, not the numbers.
- The metadata – source document, section, date, author, and whatever your application uses to decide who may read it.
Vector stores range from dedicated services (Pinecone, Weaviate, Qdrant, Milvus, Chroma) to extensions of databases you already run (pgvector in PostgreSQL, and vector types in Elasticsearch, Redis and MongoDB). For most corpora under a few million chunks, the extension in your existing database is the stronger default – vectors, text and metadata sit in one system you can already query, back up, and permission. That last point is not incidental. A bolt-on vector store is frequently deployed with no authentication at all, which is precisely the weakness catalogued as lack of access controls in Chapter 2 Section 3.
Ingestion Is a Write Path From Outside Your Trust Boundary
Ask of any corpus: who can cause a document to be indexed? If the answer includes a shared drive anyone can drop a file into, a ticketing system customers can write to, a mailbox, a wiki, or a crawler pointed at the public web, then untrusted parties can put text into your context window. They do not need access to your model, your prompt, or your application.
Ingestion is asynchronous, so this is also a stored attack: the document is indexed today and retrieved next month, into someone else’s session. Chapter 2 Section 3 shows the exploitation, including the finding that a handful of documents crafted for embedding similarity can dominate retrieval in a corpus of millions. The corresponding controls are in Chapter 3 Layer 1, and they are the ordinary ones – vet the source, validate at the ingestion gate, record provenance per chunk, and be able to un-index.
Step 2: Retrieval
When a user asks a question, a production pipeline runs a funnel rather than a lookup.
1 · Embed the query. The same embedding model converts the question into a vector. It must be the same model – vectors from different models are not comparable, which makes changing your embedding model a full re-index rather than a config change.
2 · Search twice. Vector search finds passages that are semantically close. Keyword search (BM25) finds passages that contain the actual words. Neither is sufficient alone: vector search retrieves a passage about “annual leave” for a query about “holiday allowance” and keyword search does not; keyword search finds the error code E-4471 or the product name XR-500 and vector search often does not, because rare exact strings are exactly what embeddings blur. Running both and fusing the candidate lists – usually with Reciprocal Rank Fusion, which scores by rank position rather than by incomparable raw scores – is the production standard, and measurably beats semantic search alone.
3 · Filter on metadata. Restrict candidates to what this user is entitled to see, and to whatever else scopes the query – tenant, date range, document class. This is covered in its own subsection below, because getting it wrong is the most common serious defect in deployed RAG.
4 · Rerank. Retrieval so far has compared the query vector to chunk vectors that were computed before the query existed. A cross-encoder reranker instead scores the query and each candidate passage together, which is far more accurate and far too slow to run over a whole corpus. So the funnel does both: cheap search to get from millions of chunks to perhaps fifty candidates, expensive reranking to get from fifty to the handful you actually send.
5 · Keep few. Five to eight passages is a typical final context. Sending more is not safer: Section 4’s “lost in the middle” effect means a model attends unevenly across a long context, and a marginal passage that displaces attention from a good one makes the answer worse while costing more.
Step 3: Generation
The surviving passages are assembled with the user’s question into a prompt:
The model then generates a response grounded in the retrieved passages. This substantially reduces confabulation – the model inventing a return policy it never saw – and it is why RAG is the default architecture for factual applications.
It does not reduce being wrong. It changes the failure mode, and in one respect makes it harder to catch:
Grounded Is Not the Same as Correct
A confabulating model produces an answer with no source, which a careful reader can challenge. A RAG system produces an answer with a citation – and will do so just as fluently when the retrieved passage is outdated, when it is the wrong policy for this user’s region, when a better passage was ranked ninth and dropped, or when an attacker wrote it. The citation raises the reader’s confidence without raising the answer’s accuracy.
This is the mechanism behind the over-trust and hallucination-weaponization material in Chapter 2 Section 6. It also has a design consequence worth taking from this section: a RAG answer’s citations are only worth what the corpus is worth, so “the system cites its sources” is an auditability property, not a correctness guarantee.
Note also which sentence in that prompt is doing the security work: “If the answer is not in the context, say so.” Per Section 5, that is steering. It reduces off-corpus answers on ordinary traffic and enforces nothing. The retrieved passages sit in the same flat token sequence as your instruction, and if one of them says “ignore previous instructions,” your sentence does not outrank it.
The Default Is a Shared Index
This is the most important paragraph in the section.
A vector index has no concept of a user. Similarity search compares one vector to every vector in scope and returns the closest, and “in scope” means the whole index unless you narrowed it. Every chunk in the corpus is retrievable by every question, from every user, by default. Nothing about the architecture leans towards restriction; you have to build it.
That is how the opening scenario happens. The HR corpus was indexed as one collection because it was one folder. The assistant answered the parental-leave question from it and the disciplinary-file question from it, using identical code paths, because from the pipeline’s point of view those are the same operation performed twice.
Permission filtering must happen inside the retrieval query, not after it. The distinction is not stylistic:
| Approach | What happens | Verdict |
|---|---|---|
| Filter in the vector query | The store returns only chunks this user may see; ranking happens within that set | Correct. The user’s top 5 is the top 5 of their own corpus |
| Retrieve top 50, then drop unauthorized ones in application code | Ranking happened over everything; you may discard 48 and answer from two weak passages – or from none, on a question the corpus can answer | Fragile. Quality silently depends on how much the user cannot see |
| Ask the model not to reveal restricted content | The restricted text is already in the context window | Not a control – see Section 5 |
| Index each tenant separately | Physical separation; no filter to forget | Strongest, and the right answer for multi-tenant systems |
Two failure modes follow from this that are worth being able to name, because both look like bugs rather than breaches:
- The metadata is the access control, so tampering with metadata is privilege escalation. If
department: hr-restrictedis what keeps a chunk out of the wrong results, then anyone who can write that field can change who reads the document – without touching the document. Chapter 2 Section 3 catalogues this as metadata exploitation. - Permissions drift after indexing. Access was correct on the day the document was indexed. Then someone changed the folder’s permissions, or left the company, or the document was withdrawn – and the index still holds the copy with the old label. A vector store is a cache of your permissions model, and caches go stale. Re-indexing on permission change, or resolving entitlement at query time against the live source, is the fix; Chapter 3 Layer 1 covers the operational side.
Every Stage of This Pipeline Is an Attack Surface
Read the pipeline the way Section 4 taught you to read the context window – one row at a time, asking who controls it.
| Stage | Who can influence it | What goes wrong | Attacked in | Defended in |
|---|---|---|---|---|
| Ingestion | Anyone who can get a document into the corpus | Poisoned documents crafted to be retrieved; stored injection that fires in a later session | Ch2 s3 – RAG poisoning (LLM09) | Ch3 L1 – corpus vetting, provenance |
| Chunking | You | Splitting a caveat away from its claim, so a passage retrieves as unconditional advice | – | Design-time: chunk on structure, keep qualifiers with claims |
| Embedding / vector store | Whoever can write to the store directly | Unauthenticated store; embeddings inverted to recover source text; vectors are a second copy of your data | Ch2 s3 – LLM09 | Ch3 L1 – store hardening, encryption |
| Metadata | Whoever can write metadata | Relabelling a chunk to change who retrieves it – privilege escalation with no code execution | Ch2 s3 – metadata exploitation | Ch3 L1 + L5 |
| Query | The user | Query phrased to steer retrieval towards material they should not see | Ch2 s2 – direct injection (LLM01) | Ch3 L5 – gateway filtering |
| Retrieval / ranking | Whoever wrote the retrieved documents | Missing or post-hoc permission filter; adversarial documents outranking legitimate ones | Ch2 s3 | Filter in-query; per-tenant indexes |
| Context assembly | Everyone above, jointly | Retrieved text sits unlabelled beside your instructions in one flat sequence | Ch2 s2 – indirect injection | Ch3 L5 |
| Generation / output | The model, on everyone’s input | Fluent, cited, wrong; sensitive retrieved content passed straight through to the user | Ch2 s6 – LLM02, LLM10 | Ch3 L5 – output validation |
The pattern in the last two columns is the same one Section 5 found for prompts: every real control either removes access or inspects traffic. None of them is an instruction to the model. The pipeline you have just built is the one Chapter 2 Section 3 attacks stage by stage – it is worth reading that section with this table beside it.
Technical Components
Vector Operations and Embeddings
To understand how LLMs and RAG systems process text, we need to grasp two fundamental concepts: vectors and embeddings.
What is a Vector?
A vector is simply a list of numbers that can represent a point in space:
- A 2D vector [3, 4] represents a point 3 units east and 4 units north
- A 3D vector [1, 2, 3] represents a point in three-dimensional space
In LLMs, we use vectors with hundreds or thousands of dimensions. But at their core, vectors are just containers for numbers – they don’t inherently mean anything.
Think of it This Way…
Think of vectors like barcodes – they’re just sequences of numbers that don’t mean anything on their own. Embeddings are like barcodes that have been configured to represent specific products. When you scan a barcode, the numbers suddenly have meaning because they’ve been trained to represent that product.
All embeddings are vectors, but not all vectors are embeddings!
What is an Embedding?
An embedding is a specific use of vectors for representing meaning. What makes embeddings special is that:
- They are learned through training – the model learns what numbers to put in each vector
- They have meaningful relationships – similar concepts get similar numbers
Key Takeaway
The power of embeddings lies in how they capture relationships between concepts. Just like how we naturally understand that “kitten” is related to “cat” and “puppy” is related to “dog,” embeddings allow AI models to understand these connections through carefully chosen numbers.
Embeddings: Beyond RAG
Embeddings are not used exclusively for RAG or vector databases! They are fundamental to how all modern LLMs work internally – as Section 4 covered, the model works in embedding space whether or not you attach a retrieval system. And while we’ve focused on text, the same techniques apply to images, audio, and multimodal systems.
An Embedding Is Derived, Not Anonymised
It is tempting to treat a vector store as containing “just numbers” – a mathematical index rather than a copy of the documents. That intuition drives real decisions: vector stores get provisioned outside the controls applied to the source data, excluded from classification exercises, and replicated to regions the source data may not leave.
The intuition is wrong twice over. First, most vector stores keep the original chunk text alongside the vector, because the model needs words – so it is a literal copy. Second, even where they do not, embeddings retain enough of their input to reconstruct meaningful content from the vectors alone. Embedding inversion is an active research area and a catalogued risk; Chapter 3 Layer 1 covers it directly.
The rule to carry forward: a vector store inherits the classification of its most sensitive source document. If the corpus contains regulated data, so does the index – with the same residency, retention, encryption and access requirements. Embedding is a transformation, not a de-identification step.
Cost Optimization Techniques
One of the most practical aspects of inference is managing costs. Token-based pricing means every interaction has a direct cost, and these costs can escalate quickly at scale.
The Cost Optimization Decision Tree
flowchart TD
A[Incoming request] --> P{"Answer may vary<br/>by user?"}
P -->|Yes| Q[Cache per identity<br/>or not at all]
P -->|No| C{Cached?}
C -->|Yes| E[Return cached response]
C -->|No| B
Q --> B
B{Simple or complex?}
B -->|Simple| F[Small tier<br/>fast + cheap]
B -->|Complex| D{Needs data the<br/>model lacks?}
D -->|Yes| G[Retrieval + mid tier]
D -->|No| H{Needs multi-step<br/>reasoning?}
H -->|Yes| I[Raise reasoning effort<br/>on a capable model]
H -->|No| M[Mid tier<br/>direct answer]
F --> J[Cache response<br/>scoped to entitlement]
G --> J
I --> J
M --> J
style E fill:#2d5016,color:#fff
style F fill:#1a3a5c,color:#fff
style G fill:#4a1a5c,color:#fff
style I fill:#8b0000,color:#fff
style P fill:#8b6914,color:#fff
style Q fill:#8b6914,color:#fff
The gold branch at the top is not a cost optimization – it is the guard that makes the rest of the tree safe, and the reason for it is below.
Key Optimization Strategies
Use the right model for each task
Not every request needs a frontier model. Model selection by task complexity is the single biggest cost lever:
| Task Type | Recommended Tier | Rough input-price ratio |
|---|---|---|
| Classification, routing, extraction | Small tier | 1x |
| Summarization, Q&A, most RAG | Mid tier | ~10-30x |
| Complex analysis, agentic work | Frontier tier | ~100x and up |
Treat those as orders of magnitude, not figures. The spread between the cheapest usable models and the top of the frontier line is currently around two orders of magnitude on input tokens and has been widening, not converging, as reasoning tiers pull away from commodity ones.
Reasoning Cost Is a Volume Problem, Not a Price Problem
Raising reasoning effort is a bigger cost change than moving up a tier, and the reason is easy to miss on a price list. Reasoning tokens are billed as output – the expensive direction, typically 4-8x the input rate – and a high-effort request can emit many times the tokens a direct answer would.
Those multiply. A tier change moves the unit price; an effort change moves the unit price and the quantity. This is the line item that surprises teams on their first bill after enabling reasoning, and it is why effort belongs under per-request routing rather than being set globally and forgotten. Section 5 covers the dial itself.
Rule of thumb: Start with the smallest model that produces acceptable results, then scale up only where quality demands it.
Don’t generate the same response twice
Three different mechanisms travel under the name “caching”, and they have very different properties:
- Prompt caching – the provider caches the input prefix, so a repeated system prompt or a stable retrieved context is not re-processed. Now standard across major providers and largely automatic: on OpenAI models it applies with no code change to any prompt whose prefix exceeds roughly 1,024 tokens, and Anthropic exposes it as explicit cache breakpoints. Cached input reads at a large discount – typically 75-90% off the input rate.
- Exact-match response caching – your own cache, keyed on the request. Same question, same answer, no model call.
- Semantic response caching – your own cache, keyed on embedding similarity. A similar question returns a previous answer deemed close enough.
Prompt caching rewards a stable prefix, which is the opposite of the usual advice. The cache is a prefix match, so it survives only as long as everything before your variable content is byte-identical. Put the system prompt and few-shot examples first, the retrieved passages and user turn last, and never interpolate a timestamp or a request ID near the front – one changed token at position 20 invalidates the entire prefix. On current models a cache write also costs slightly more than an uncached request, so a prefix that misses every time is worse than no caching at all.
Best for: FAQ-style applications, customer support, common queries where answers don’t change frequently.
A Semantic Cache in Front of a Permission-Filtered Corpus Is a Disclosure Bug
The two mechanisms above compose badly, and the failure is silent.
Retrieval was filtered by entitlement, so the answer a user receives is a function of who they are. A response cache keyed only on the question throws that away. User A, who may read the restricted corpus, asks a question; the answer is cached. User B asks the same question – or with a semantic cache, merely a similar one – and is served A’s answer without any retrieval happening at all. Every control you built into the pipeline was bypassed, by a cache hit.
Semantic caching makes it worse in two ways: the match is fuzzy, so B never has to guess A’s phrasing, and the similarity threshold is a tuning parameter, which means someone can widen your disclosure boundary while trying to improve your hit rate.
The fix is that the cache key must include everything the answer depends on: identity or entitlement set, tenant, and corpus version. Which mostly removes the benefit for personalized answers – and that is the correct outcome. Cache aggressively where answers are public and identically true for everyone; cache per-identity or not at all where they are not.
Process efficiently at scale
-
Batching: Submit many requests as one job and collect results asynchronously. OpenAI and Anthropic both price batch at a flat 50% discount, with completion guaranteed within 24 hours rather than in real time. Genuinely free money for anything not user-facing – back-fills, bulk classification, nightly enrichment.
-
Request throttling: Cap the request rate so a bug or a burst cannot run up an unbounded bill. Note this is a security control as much as a cost one: unbounded consumption is LLM06, and an attacker who can trigger expensive requests can turn your API budget into a denial-of-wallet attack. Chapter 3 Layer 5 treats rate limiting as an abuse control for this reason.
-
Bounding output length: Set
max_tokensto what the task actually needs – a classification does not need room for 500 words. Be precise about the mechanism, though: as Section 5 covered,max_tokensis a hard cut, not a hint to be brief. The model does not wrap up as it approaches the limit; it stops mid-sentence, and structured output becomes unparseable. Check the finish reason. To make responses genuinely shorter, ask for brevity in the prompt and setmax_tokensas the backstop.
Minimize token usage without sacrificing quality
- Concise system prompts: Every token in your system prompt is paid for on every request
- Selective context: In retrieval, send the 5-8 passages that survived reranking, not the top 20 – this improves answer quality as well as cost
- Output length limits: Set
max_tokensas a backstop, and ask for brevity in the prompt - Compression: Summarize conversation history instead of replaying full transcripts
Quick math: If your system prompt is 500 tokens and you make 100,000 requests/day:
- 500 x 100,000 = 50M input tokens/day
- At a representative mid-tier input price of $2.50/1M: $125/day just for the system prompt
- Cutting that prompt to 200 tokens saves $75/day ($27,000/year)
Why That Example Works – and Where It Inverts
That saving is real because the prompt is 500 tokens. Prompt caching needs a prefix of roughly 1,024 tokens before it engages, so a 500-token system prompt is below the threshold and is paid for at full rate on all 100,000 requests. Shortening it is the only lever available.
Above the threshold the advice reverses. A 4,000-token prompt that stays byte-identical is cached at 75-90% off, which beats trimming it to 3,000 tokens of unstable text. Below the cache threshold, shorten. Above it, stabilize. Trimming a long prompt in a way that changes its prefix each request can cost more than leaving it alone – the two optimizations are in tension, and knowing which regime you are in is the whole skill.
Comparing the Strategies
Each tab above describes a technique in isolation. Applied to a real workload they differ in effort, in how much they save, and – the column usually missing – in what they change about your security posture.
| Strategy | Typical saving | Effort | What it costs you | Security consequence |
|---|---|---|---|---|
| Model routing by complexity | Largest single lever; 70-90% on the routed share | Medium – needs a classifier and quality thresholds per route | Quality regressions if routing is wrong; two models to evaluate | The router reads user input, so it is an attack surface: a crafted request can steer itself to a weaker model with weaker guardrails (Ch2 s7) |
| Prompt caching | 75-90% of input cost on the cached prefix | Low – often automatic | Prefix rigidity; a cache write costs slightly more than a miss | Benign. Caches an input you already own; no cross-user path |
| Exact-match response cache | 100% on every hit | Low | Staleness | Safe only if the key includes identity and corpus version |
| Semantic response cache | 100% on a much larger share of requests | Medium | Wrong answers to near-miss questions | Highest-risk item here. Serves one user’s answer to another; the threshold that tunes hit rate also tunes your disclosure boundary |
| Batching | Flat 50% | Low | Up to 24 hours’ latency | Benign, and it removes a real-time DoS path |
| Request throttling | Prevents runaway spend rather than reducing unit cost | Low | Rejected requests at peak | Improves posture – it is the control for unbounded consumption (LLM06) |
| Prompt and context trimming | 10-40% of input cost | Low, ongoing | Quality loss if you cut load-bearing instructions | Benign, with one trap: trimming the system prompt is often trimming safety instructions |
Read the last column against the second. The two cheapest wins – batching and throttling – are also the two that improve your security posture, and the technique with the broadest saving is the one that can quietly disclose data across users. Cost optimizations are architecture changes. They alter what is stored, who can reach it, and which model sees which request, and they are usually shipped as performance work without review.
Key Takeaways
- Provider APIs come in three shapes – legacy completions, chat completions (the de facto standard), and the current stateful/agentic APIs. The difference that matters is who holds the conversation state: moving it to the provider is a data-residency and retention decision, not just a cost saving
- Streaming publishes as it generates. A token you have streamed cannot be filtered, so output validation and streaming are in tension – anything feeding a downstream system rather than a human should be synchronous
- Production retrieval is a funnel, not a lookup: hybrid search (vector + keyword, fused) → metadata filter → cross-encoder rerank → 5-8 passages. Each narrowing is a place where the wrong passage survives or the right one is dropped
- A vector index has no concept of a user. Every chunk is retrievable by every question until you filter, and the filter must be inside the retrieval query. Metadata is the access control, which makes metadata writes privilege escalation
- A vector store inherits the classification of its most sensitive source document. Embedding is a transformation, not de-identification – the chunk text is usually stored alongside, and embeddings can be inverted
- Grounded is not correct. RAG reduces confabulation and replaces it with fluent, cited answers from outdated, mis-scoped or poisoned passages – a failure that is harder to spot because the citation raises confidence
- Cost optimizations are architecture changes. Model tier is the biggest lever; caching is the sharpest edge – a response cache that ignores identity serves one user’s permission-filtered answer to another
Test Your Knowledge
Ready to test your understanding of inference techniques? Head to the quiz to check your knowledge.
Up next
Retrieval gave the model a channel to read from systems you do not fully control. The next section gives it the ability to write – agentic AI, where the model calls tools, takes actions with real consequences, and decides for itself which to call. Every question this section asked about the corpus (“who can put text in here?”) returns immediately, with the stakes raised: an agent’s tool output arrives in the same flat context window, and its actions are not reversible by declining to display them.