Section 5 Quiz
Test Your Knowledge: Prompt Engineering
Let’s see how much you’ve learned!
This quiz tests your understanding of prompting techniques, what a prompt can and cannot enforce, parameter behaviour, and prompting with reasoning enabled.
---
shuffle_answers: true
shuffle_questions: false
---
## A developer wants an LLM to consistently classify customer feedback into fixed categories (BUG, FEATURE_REQUEST, POSITIVE, NEGATIVE). Which prompting technique fits best?
> Hint: Think about which technique pins down output format across many inputs.
- [ ] Zero-shot prompting, relying on a single clearly worded instruction
> Zero-shot may handle basic classification, but output format tends to drift across requests and borderline cases get labelled inconsistently.
- [x] Few-shot prompting, with 3-5 examples covering the range of categories
> Correct! Few-shot is the technique for consistent formatting. Examples of each category teach the exact output pattern and show the model where the boundaries between similar categories fall. Three to five diverse examples is the usual sweet spot.
- [ ] Chain-of-thought prompting, asking the model to reason through each case
> CoT adds output tokens and latency for a task that needs format consistency rather than multi-step reasoning. It is the wrong tool here.
- [ ] A high temperature setting, so the model explores category options freely
> High temperature increases randomness in token selection, which is the opposite of what a classifier needs. Classification wants low temperature.
## Top-P (nucleus) sampling set to 0.5 means the model considers:
> Hint: The definition turns on one word — cumulative.
- [ ] Half of the tokens in the model's vocabulary, chosen by rank
> Top-P does not select a fixed count or fraction of the vocabulary. The size of the set changes at every generation step.
- [x] The smallest set of tokens whose cumulative probability reaches 50%
> Correct! Top-P truncates the distribution by probability mass, not by token count. If one token already holds 60% of the mass, the set at p=0.5 contains just that token; where the distribution is flat, the same setting may admit hundreds.
- [ ] Every token whose individual probability is at least 50% likely
> This would usually admit one token or none. Top-P accumulates probabilities across ranked tokens rather than thresholding each one.
- [ ] Tokens drawn from the upper half of the ranked probability list
> This describes a fixed positional cutoff. Top-P is adaptive: the cutoff moves with how confident the model is at that step.
## What changes when you move temperature from 0.1 to 0.9?
> Hint: Consider what "randomness" means for token selection specifically.
- [ ] Responses become longer, because the model explores more content
> Temperature affects randomness in token selection, not length. Length is governed by the maximum-token setting.
- [x] Output becomes more varied and less predictable across repeated runs
> Correct! Temperature reshapes the probability distribution over the next token. Low values concentrate mass on the most probable tokens, giving consistent output. High values flatten the distribution, so less likely tokens get picked more often.
- [ ] The model shifts from text generation toward code generation behaviour
> Temperature is task-agnostic. The right value depends on the creativity-versus-consistency trade-off, not on the kind of content.
- [ ] Accuracy improves, because the model considers more possible answers
> Higher temperature does not improve accuracy. For factual work, low temperature is generally the better choice.
## A security analyst needs an LLM to emit a vulnerability report as JSON for a downstream pipeline. What should the prompt contain?
> Hint: Which elements pin down both the analysis and the format?
- [ ] A short instruction to find vulnerabilities and return the result as JSON
> Too vague. Without a schema the model invents its own field names and structure, and they vary between requests.
- [ ] The JSON schema on its own, with no task instructions or example
> A schema helps, but the model still needs to know what analysis to perform and how findings map into the fields.
- [x] Task instructions, the exact schema with field types, and one worked example
> Correct! Structured output works best combining all three: what analysis to do, the precise schema including field names and allowed values, and at least one complete input-to-output example. Where the API offers constrained decoding, use that too — it enforces the schema during generation rather than requesting it in prose.
- [ ] A chain-of-thought instruction, followed by a request to format as JSON
> Asking for reasoning first tends to leak the reasoning text into the JSON. Ask for the structured result directly.
## A prompt reads: "You are an expert security analyst. Keep responses under 300 words. If you are unsure of a threat classification, say so." Which prompt components appear?
> Hint: Match each clause against the anatomy of a prompt.
- [ ] Task instructions, format specification, and few-shot examples
> No examples appear, and no actual task is stated — the prompt sets up a role and rules but never says what to analyse.
- [x] Context via role assignment, a constraint, and uncertainty handling
> Correct! The role assignment supplies context, the word limit is a constraint, and "if you are unsure, say so" is uncertainty handling. Notably absent: a task instruction and examples. Note also that only the word limit has a real enforcement mechanism behind it — `max_tokens`.
- [ ] Constraints and worked examples, but no context or role assignment
> There are no examples here, and "you are an expert security analyst" is precisely a role assignment.
- [ ] Task instructions and constraints, with no uncertainty handling present
> "If you are unsure, say so" is uncertainty handling, and no task is actually specified in this prompt.
## You are prompting a frontier model that has reasoning enabled by default. Your prompt already decomposes the analysis into six numbered steps and includes four worked examples of the reasoning. Results are mediocre. What should you change first?
> Hint: What is this prompt telling the model that the model already does?
- [ ] Add more numbered steps, breaking the analysis down further
> This deepens the problem. The prescribed steps are competing with the model's own decomposition rather than assisting it.
- [x] Drop the scripted steps and state the goal and success criteria instead
> Correct! With reasoning on, the model decomposes the problem itself; prescribing the route cuts across that. State what a good answer contains and let it find the path. Examples are worth keeping only if you need them to pin the output *format* — not to demonstrate a reasoning procedure.
- [ ] Increase the temperature so the model explores more approaches
> Two problems: the issue is prescriptive prompting, and sampling parameters are commonly restricted or rejected outright when reasoning is enabled.
- [ ] Append "let's think step by step" to trigger deeper decomposition
> That phrase helps a model that is not already reasoning. Here it is redundant at best and interferes at worst.
## An application sends an identical prompt at temperature 0 a thousand times and gets a few different responses. What is the best explanation?
> Hint: What varies between two requests that are otherwise identical?
- [ ] Temperature 0 is approximate, since providers add noise for output diversity
> No noise is added. Temperature 0 genuinely selects the highest-probability token; the variation comes from upstream of that choice.
- [x] Batching: the arithmetic shifts with whichever requests share your batch
> Correct! Inference servers batch concurrent requests together, and the kernels are not batch-invariant — the order of floating-point reductions shifts with batch size. Where two candidate tokens are nearly tied, that tiny difference can tip the selection. Your output effectively depends on how busy the service was.
- [ ] The model's weights are being updated continuously as new requests arrive
> Weights are fixed after training and do not change during inference. Nothing about serving a request modifies the model.
- [ ] Requests are load-balanced across copies quantized at different precisions
> Load balancing exists, but this pattern appears on a single consistent deployment. The cause is in the batching arithmetic.
## A team states: "Customer records are protected — our system prompt instructs the model never to reveal them." The records are retrieved into the context window on every request. Assess this claim.
> Hint: What would an attacker have to defeat, and what has to be true for that to be impossible?
- [ ] It is sound, provided the instruction is emphatic and placed in the system prompt
> Placement and emphasis do not create a boundary. The system prompt occupies the same flat token sequence as the user's input, with no enforced priority over it.
- [x] It fails: data in the context window is reachable, so authorization must precede retrieval
> Correct! The instruction steers the model, and will hold for ordinary users. It cannot guarantee anything, because the records are already present in the sequence and injection or jailbreaking can elicit them. The control is to never place data in the window that the requesting user is not entitled to see.
- [ ] It is sound as long as the deployment also filters obvious injection phrasings
> Input filtering is a useful layer but a porous one, and it does nothing about the underlying problem: the data is in the window and therefore reachable.
- [ ] It fails, but only for models that have not been safety-tuned by the provider
> Safety tuning shapes tendencies across all models, not boundaries. This exposure does not depend on which model you chose.
## Your pipeline uses constrained decoding, so every response is guaranteed to match your JSON schema. What risk does that guarantee leave open?
> Hint: A guarantee about structure is a guarantee about what, exactly?
- [ ] The model may return well-formed JSON that omits your required fields
> Constrained decoding does enforce required fields. Schema conformance is the part you genuinely get.
- [x] String field contents are unvalidated and unsafe to render or execute directly
> Correct! The guarantee covers the container, never the contents. A schema-valid `description` field can hold a script tag that executes when your dashboard renders it, or a command string that runs if passed to a shell. This is LLM10: Improper Output Handling — model output crosses a trust boundary into your code.
- [ ] Enforcing the schema during generation measurably degrades reasoning quality
> Schema enforcement constrains token selection, not the analysis. This is not the security consequence in question.
- [ ] Responses will silently truncate whenever the schema grows past a size limit
> Truncation is a maximum-token concern and shows up in the finish reason, not a gap left by schema enforcement.
## An agent takes an action you did not expect. Its reasoning trace gives a coherent, benign justification. What can you conclude?
> Hint: Is the trace a record of the computation, or a product of it?
- [ ] The action was legitimate, since the trace documents the reasoning behind it
> A trace is generated output, not an execution log. Research on faithfulness found models routinely omit the factor that actually drove an answer while producing plausible reasoning for it.
- [x] Very little — the trace is generated output, not evidence of the real cause
> Correct! Chain-of-thought traces are frequently unfaithful: in Anthropic's 2025 study, models acknowledged an answer-changing hint only 25-39% of the time. A clean justification does not clear an action, which is why AI systems are monitored on observable actions and outputs rather than self-reports.
- [ ] The model was tampered with, since a benign trace contradicts an odd action
> Unfaithful traces are ordinary model behaviour, not a tampering indicator. The trace being unreliable is the default, not the anomaly.
- [ ] The action was safe, because a reasoning model verifies its own conclusions
> Reasoning models check their work, which is different from verifying it. Extended deliberation makes a wrong answer more persuasive, not less likely to need validation.