Section 4 Quiz

Test Your Knowledge: Technical Foundations

Let’s see how much you’ve learned!

This quiz tests your understanding of tokenization, the context window as a trust boundary, Transformer architecture and attention, the four memory approaches, and how training data shapes what a model knows and what it can be made to do.

--- shuffle_answers: true shuffle_questions: false --- ## A 4,000-word legal contract tokenizes to roughly 6,500 tokens. A 4,000-word blog post tokenizes to roughly 5,300. What explains the gap? > Hint: Tokenizers are built by measuring which character sequences are common. - [x] Rare specialized terms are split into several sub-word tokens each > Correct! At the usual ratio of about 0.75 words per token, 4,000 words of ordinary prose lands near 5,300 tokens -- which is what the blog post does. Legal jargon is uncommon in the training corpus the tokenizer was built from, so terms like "indemnification" get broken into pieces rather than encoded as one token. The rule of thumb holds for ordinary text and quietly stops holding for specialized text, code, names, and non-English content. - [ ] Legal documents contain more punctuation and formatting > Punctuation does add tokens, but not on this scale, and formatting is stripped before tokenization in most pipelines. - [ ] The tokenizer duplicates tokens to improve accuracy > Tokenizers never duplicate tokens. Each token occupies one position in the sequence. - [ ] Longer documents always tokenize less efficiently > Both documents are the same length. Efficiency depends on vocabulary, not document size. ## A team blocks a banned phrase with a keyword filter. An attacker gets the phrase through by inserting zero-width Unicode characters into it, and the model still acts on the request. Why did the filter fail? > Hint: What unit does each component operate on? - [x] The filter matches characters, the model reads tokens -- they resolve differently > Correct! This is token smuggling. The filter sees a string that does not match its blocklist because of the injected characters; the model's tokenizer resolves the same bytes into a token sequence that still carries the original meaning. A tokenizer belongs to a model, not to the language -- so any control that does not tokenize the way the model tokenizes is inspecting a different input than the one the model will read. Chapter 2, Section 2 covers this family of evasions. - [ ] The zero-width characters raise the model's temperature setting > Temperature is a sampling parameter set by the caller. Input content cannot change it. - [ ] The model decodes the hidden characters before it processes the prompt > There is no decoding step. The characters are simply tokenized along with everything else. - [ ] The filter runs after generation rather than ahead of it > An input filter runs before generation. The ordering is not the failure here. ## What does the "lost in the middle" finding describe? > Hint: Think about how attention is distributed across a long input. - [x] Uneven attention across a long input, favouring its beginning and its end > Correct! Liu et al. (TACL 2024) found that model performance on retrieval-style tasks degrades sharply when the relevant passage sits in the middle of a long input -- in some settings falling below what the same model achieved with no retrieved documents at all. Every token remains in the context; the model simply does not weigh them uniformly. A large window is not the same thing as even attention across it. - [ ] Deletion of middle-position tokens to conserve context space > Nothing is deleted. All tokens stay in the window; the issue is how they are attended to. - [ ] Models physically reading only the first and last paragraphs > Models process every token in the window. The effect is a matter of degree, not access. - [ ] Garbled output appearing midway through a generated response > This concerns how input is attended to, not the quality of generated output. ## How does a Mixture-of-Experts model reduce cost relative to a dense model of the same total size? > Hint: Consider what happens to parameters during a single request. - [x] It routes each token to a small subset of specialized expert sub-networks > Correct! A gating mechanism selects which experts handle each token, so a model holding hundreds of billions of total parameters may activate only a small fraction per request. This is why "how many parameters?" is no longer a useful question on its own: total parameters set your memory requirement, active parameters set your compute bill. - [ ] It compresses every parameter into a smaller memory footprint > MoE does not compress anything. All expert weights are held in full and must be resident. - [ ] It discards unnecessary parameters once training completes > No parameters are removed. Every expert stays available for the tokens routed to it. - [ ] It trains each expert separately, which shortens training time > Experts do specialize, but MoE's headline advantage is inference efficiency, not training speed. ## FlashAttention is described as computing *exact* attention. Compared with a naive implementation, what changes? > Hint: The section separates the arithmetic from the plumbing. - [x] Memory use drops to linear; the FLOP count stays quadratic > Correct! FlashAttention never writes the full token-by-token attention matrix to GPU memory, which cuts memory from quadratic to linear and delivers large real-world speedups -- because attention is bottlenecked by memory traffic rather than by arithmetic. The maths is untouched, which is why the result is exact. Approximate schemes such as sliding-window attention are the ones that genuinely reduce asymptotic compute, and they pay for it in what the model can see in one hop. - [ ] Memory use stays quadratic; the FLOP count drops to linear > This reverses the actual result, and it would not be exact attention if the FLOP count changed. - [ ] Both memory use and the FLOP count drop to linear > Reducing FLOPs below quadratic requires approximating attention. FlashAttention does not approximate. - [ ] Neither changes; only the model's output quality improves > Output is identical by construction. The entire benefit is in memory and wall-clock time. ## A chatbot appears to "forget" something a user shared 30 messages earlier. What is the most likely explanation? > Hint: Recall what an LLM does and does not retain between requests. - [x] The conversation outgrew the context window and older turns were dropped > Correct! LLMs are stateless -- each request is independent, and what feels like memory is the conversation history being re-sent in every prompt. Once that history exceeds the window, older turns must be truncated or summarized away. Note the failure mode: a truncated prompt does not announce itself, so the application looks like it is ignoring the user rather than like it has run out of room. - [ ] The model's weights degraded over the course of the conversation > Weights are fixed at inference. Nothing about a conversation modifies them. - [ ] The model chose to disregard information it judged unimportant > LLMs have no intent. This is a capacity limit, not a decision. - [ ] The user's earlier messages were corrupted in the application database > Possible in principle, but far less likely than the ordinary context-window explanation. ## An engineer proposes auditing a model for output bias by inspecting its bias parameters. Why will this not work? > Hint: The section warns that two different things share the name. - [x] Output bias is spread across the whole parameter set, not held in the bias terms > Correct! A bias term is arithmetic: a learned baseline attached to each unit that shifts its output before any input is considered. Bias in a model's *outputs* -- associating certain jobs with certain genders, for instance -- is learned from patterns across the training data and is distributed throughout the weights. The two share a name and nothing else, which is why output bias is addressed at the data and post-training layers rather than by inspecting parameters. - [ ] Bias parameters are encrypted and cannot be read directly > Bias parameters are ordinary numbers, readable in any open-weight model. - [ ] Bias terms are discarded after training and no longer exist > Bias terms are part of the trained model and are used at every inference. - [ ] Output bias appears only in fine-tuned models, not base models > Base models learn output bias from pretraining data. Fine-tuning can add or reduce it. ## A company chooses RAG over simply loading everything into a long-context model. Which situation best justifies that decision? > Hint: Weigh cost, freshness, and scale against architectural simplicity. - [x] A 100,000-document knowledge base that changes every day > Correct! Sending 100,000 documents on every request would be prohibitively expensive and slow, and re-sending them is the only way a long-context approach reflects an update. RAG retrieves a handful of relevant chunks per query and stays current through live updates to the index. The trade is real infrastructure -- a vector store and an embedding pipeline -- plus a new input path into the context window that anyone with write access to the corpus can use. - [ ] A requirement to avoid embeddings anywhere in the stack > RAG is built on embeddings for similarity search. This argues against RAG, not for it. - [ ] A need for the simplest architecture with no additional services > RAG adds a vector database and an embedding pipeline. Long context is the simpler option. - [ ] A need to reason across relationships spanning every document at once > Long context suits this better, since RAG only ever surfaces the chunks it retrieved. ## An architect proposes hardening an assistant by "putting the security rules in the system prompt so they take priority over user input." What is wrong with the reasoning? > Hint: What does the model actually receive? - [x] The model receives one flat token sequence with no enforced priority by source > Correct! The system prompt, the user message, retrieved documents, tool output and stored memory all arrive as undifferentiated tokens. Those categories exist in your architecture diagram, not in the model's input, and attention weighs a sentence from a poisoned wiki page by the same mechanism it weighs your instructions. This is why prompt injection has no clean fix at the model layer, and why Chapter 3's defenses sit around the model -- validating what enters the window and what leaves it -- rather than inside it. - [ ] The system prompt is capped at 4,096 tokens, too few for real rules > There is no such cap. System prompts can be long; length is not the problem. - [ ] The system prompt is processed after the user message, not before > Ordering is not what grants priority, and a longer system prompt would not fix it either. - [ ] Security rules must be expressed in a structured format, not prose > Format does not confer enforcement. Structured rules in the window are still just tokens. ## Two base models score identically on capability benchmarks. One ships a documented, versioned, signed training corpus; the other's training data is undisclosed. From a security standpoint, what does that difference buy you? > Hint: The section separates properties that make a model good from properties that make it trustworthy. - [x] Evidence about what could have been planted before the model reached you > Correct! Diversity and accuracy determine whether a model is good; provenance and integrity determine whether it can be trusted, and only the second pair is something an attacker acts on. A documented, signed corpus lets you reason about who could have contributed content and verify that the data used was the data intended. Without it, a dormant backdoor triggered by a specific phrase is not something you can rule out by testing -- which is the substance of OWASP LLM05, covered in Chapter 2, Section 3. - [ ] A guarantee that the model will not exhibit bias in its outputs > Disclosure is not review. A fully documented corpus can still be skewed. - [ ] Lower inference cost, since documented corpora encode more efficiently > Training-data documentation has no bearing on inference cost. - [ ] Assurance that the model refuses harmful requests more consistently > Refusal is a post-training property, and Section 3 showed it is probabilistic regardless.