Section 6 Quiz
Test Your Knowledge: Inference Techniques
Let’s see how much you’ve learned!
This quiz tests your understanding of API integration shapes, the retrieval pipeline and its attack surface, embeddings, streaming trade-offs, and cost optimization.
---
shuffle_answers: true
shuffle_questions: false
---
## Your team migrates from a chat completions API to a stateful API that keeps conversation history server-side. What changes about your security posture?
> Hint: Think about where the transcript physically lives afterwards.
- [ ] Nothing changes, because the provider already stored conversation history on your behalf before the migration
> Chat completions is stateless. You stored the transcript and replayed it on every request; the provider kept nothing between calls.
- [x] The transcript moves to the provider, and prior turns can write state later turns read back
> Correct! Two things change. Residency and retention move to the provider, so "where is the conversation history" becomes a data-governance question you answer with someone else's policy. And server-held state is durable state a previous turn can write into, which is the mechanism behind the memory-poisoning attacks in Chapter 2 Section 5.
- [ ] Prompt injection becomes impossible, because message roles are now enforced on the provider's server
> Roles are a formatting convention, not a privilege boundary, wherever the state is stored. Section 4 established that the context window has no privilege levels.
- [ ] Your token costs rise, because the provider replays the entire history on every turn
> Costs generally fall -- you send only the new message rather than re-uploading the whole transcript each turn.
## In a retrieval pipeline, what is the purpose of the "chunking" step during document ingestion?
> Hint: Think about why you can't just send entire documents to the model.
- [ ] To encrypt the documents so they can be stored safely in the vector database
> Chunking splits documents. It applies no cryptography and provides no confidentiality.
- [ ] To remove redundant words and shrink the documents before they are indexed
> Chunking preserves the original text within each chunk -- it splits without removing content.
- [x] To split documents into focused pieces, so only the relevant portions occupy the context window
> Correct! Chunking serves two purposes: context windows are finite so whole documents often can't be sent, and smaller focused chunks retrieve more precisely. It also carries a correctness risk -- splitting a claim away from the sentence that qualifies it creates a chunk that retrieves as unconditional and is wrong.
- [ ] To translate documents into a single language for consistent embedding
> Chunking involves splitting, not translating. Multilingual retrieval is a property of the embedding model.
## Your system prompt is 3,000 stable tokens. To cut costs, a developer proposes trimming it to 2,200 tokens and appending the current timestamp at the top for logging. What is the effect?
> Hint: Prompt caching matches on a prefix.
- [ ] Costs drop about 27%, proportional to the tokens removed
> That would hold if no caching were involved. The timestamp is the decisive detail here.
- [x] Costs rise, because the timestamp changes the prefix every request and defeats the cache
> Correct! At 3,000 stable tokens the prompt was above the roughly 1,024-token threshold and cached at 75-90% off. A timestamp at the top changes the prefix on every request, so every call becomes a cache miss -- and a cache write costs slightly more than an uncached request. The 800 tokens saved are dwarfed by losing the discount on the remaining 2,200. Below the threshold, shorten; above it, stabilize.
- [ ] Costs are unchanged, since cached tokens are billed at the standard input rate
> Cached input reads at a large discount -- typically 75-90% off -- which is exactly what makes the prefix worth protecting.
- [ ] Costs drop sharply, because shorter prompts are cached more aggressively
> Caching engages *above* a minimum prefix length. Shortening a prompt toward that threshold reduces what can be cached.
## A retrieval system returns chunks about "Python programming" when a user asks about "python snakes." Which stage failed, and what fixes it?
> Hint: Think about which component decides what counts as "similar," and what the other kind of search is good at.
- [ ] The chunking strategy split the snake material across too many small pieces
> If the material is in the corpus, chunk size wouldn't cause programming content to outrank it. Chunking affects what a chunk contains, not what the ranker prefers.
- [x] Semantic search ranked on embedding proximity alone; keyword search and reranking would disambiguate
> Correct! Both senses of "python" occupy overlapping regions of embedding space, so vector similarity cannot separate them. This is why production pipelines run keyword search alongside vector search and fuse the results, then rerank with a cross-encoder that scores the query and passage together rather than comparing pre-computed vectors.
- [ ] The language model hallucinated programming content in place of snake information
> The model answers from what it was given. If programming chunks were retrieved, this is a retrieval failure, not a generation failure.
- [ ] The vector database corrupted the stored embeddings during indexing
> Corruption would degrade retrieval broadly and unpredictably, not produce a coherent wrong topic.
## An internal assistant correctly answers an HR policy question, then correctly summarises a restricted personnel file for a user with no HR access. Which stage failed?
> Hint: Nothing malfunctioned. Ask what the pipeline was never told.
- [ ] Generation -- the model should have refused a request for restricted content
> The restricted text was already in the context window by then. Asking the model to decline is steering, not enforcement, and the disclosure happened at retrieval.
- [x] Retrieval -- the index was searched with no metadata filter scoping it to this user
> Correct! A vector index has no concept of a user. Similarity search compares the query against everything in scope, and scope is the whole index unless you narrowed it. Both answers came from identical code paths because to the pipeline they were the same operation. The filter has to be inside the retrieval query -- filtering after ranking means the user's top results were chosen from documents they cannot see.
- [ ] Ingestion -- the personnel file should never have been indexed at all
> Sometimes true, but it isn't the defect. The corpus legitimately contains restricted documents; the failure is retrieving them for the wrong person.
- [ ] Embedding -- the model placed the file too close to routine policy queries
> Embedding quality determines relevance, never entitlement. A perfect embedding model produces exactly this outcome.
## Why should a response that feeds a downstream system, rather than a human reader, be handled synchronously rather than streamed?
> Hint: Think about what you can still do with output you have not yet sent.
- [ ] Streaming produces measurably lower-quality output than synchronous generation does
> Identical tokens are generated either way. Only delivery differs.
- [ ] Streaming consumes more tokens overall, so each request costs noticeably more
> Token usage is unaffected by delivery method.
- [x] A streamed token is already published, so validation has nothing left to block
> Correct! Output validation needs the output before anyone sees it. Streaming hands over the first sentence before your filter has seen the last, and there is no retraction. With no human waiting there is no latency benefit to trade for it -- and an unvalidated payload reaching an interpreter is LLM10: Improper Output Handling.
- [ ] Downstream systems cannot parse server-sent event streams without additional client libraries
> A tooling inconvenience at most, and not the reason. The issue is that validation must precede delivery.
## A 500-token system prompt is sent on every request. The application serves 100,000 requests per day on a model priced at $2.50 per 1M input tokens. What does the system prompt alone cost annually?
> Hint: Find the daily token volume first, then the daily cost, then scale up.
- [ ] About $1,250 -- and prompt length is a minor cost factor at this volume
> Check the math. This underestimates the daily token count by an order of magnitude.
- [ ] About $4,562 -- roughly one month of the true annual figure
> Not the right calculation. Work through tokens, then daily cost, then the year.
- [x] About $45,625 -- 50M tokens per day, $125 per day, multiplied across the year
> Correct! 500 x 100,000 = 50M input tokens daily; at $2.50/1M that is $125/day, or $45,625/year. Trimming to 200 tokens saves about $27,375. Note *why* trimming is the right lever here: at 500 tokens the prompt sits below the roughly 1,024-token caching threshold, so it is billed at full rate every time.
- [ ] About $125 -- the same as the daily figure, since caching absorbs the rest
> That is the daily cost. Caching does not apply: the prefix is under the minimum length.
## For an application where 70% of queries are simple FAQ lookups and 30% need complex analysis, which optimization has the largest impact -- and what does it introduce?
> Hint: One change affects the majority of requests. Ask what reads the request to make that decision.
- [ ] Capping max_tokens at 50 for every response, which bounds spend predictably
> This truncates the 30% mid-sentence. max_tokens is a hard cut, not an instruction to be concise.
- [x] Routing by complexity, which introduces a router that reads untrusted user input
> Correct! Moving 70% of traffic to a model costing 10-30x less is the biggest available lever. The cost is a new component: the router reads the user's request to classify it, so a crafted request can steer itself toward a weaker model with weaker guardrails -- Chapter 2 Section 7 covers reduced-guardrail exposure on smaller models.
- [ ] Consolidating onto one self-hosted model, which removes per-token billing entirely
> It substitutes fixed infrastructure cost for per-token cost and risks quality on the complex 30%.
- [ ] Prompt caching on the system prompt, which requires no application changes
> Worth doing and nearly free, but it discounts a prefix rather than the whole request, so the saving is much smaller than re-routing 70% of traffic.
## A support assistant retrieves from a corpus filtered by each user's entitlements. To cut costs, a team adds a semantic cache keyed on question similarity. Assess this design.
> Hint: Ask what the cached answer depended on, and what the key records.
- [ ] Sound -- semantic caching only ever returns answers to genuinely equivalent questions
> The match is deliberately fuzzy. That is the feature, and here it is the problem.
- [x] Unsafe -- the answer depends on identity but the cache key does not
> Correct! Retrieval was entitlement-filtered, so the answer is a function of who asked. A key built only from the question discards that: a privileged user's answer is cached, then served to anyone asking a *similar* question, with no retrieval running at all. Every pipeline control is bypassed by a cache hit. Worse, the similarity threshold that tunes hit rate also widens the disclosure boundary, so a performance tweak can enlarge a leak.
- [ ] Sound -- the retrieval filter still runs and re-checks entitlement on every cache hit
> A cache hit is a cache hit precisely because retrieval did not run.
- [ ] Unsafe, but only for exact-match caching; semantic matching avoids the problem
> Reversed. Exact-match caching has the same flaw, and semantic matching widens it.
## Your vector store holds embeddings of documents classified as confidential. What classification does the store itself require, and why?
> Hint: Ask what is physically in the store, and what can be recovered from what remains.
- [ ] Lower -- embeddings are numerical vectors, so the sensitive text is not present
> This is the common and consequential mistake. Most stores keep the chunk text beside the vector, because the model needs words.
- [x] The same as its most sensitive source document -- chunk text is stored alongside the vectors
> Correct! Embedding is a transformation, not de-identification. The original text is usually stored with each vector, and even where it isn't, embedding inversion can reconstruct meaningful content from the vectors alone. So the store inherits the source's residency, retention, encryption and access requirements -- which is exactly what gets missed when a vector store is provisioned as "just an index."
- [ ] Lower, provided the store is encrypted at rest and reachable only inside the VPC
> Good controls, but they don't change what the data *is*. Classification determines which controls apply, not the reverse.
- [ ] Unclassified -- vectors are derived data and fall outside the scope of data governance
> Derived data carrying recoverable sensitive content is in scope. Being derived is not being anonymised.