Section 6 Quiz
Test Your Knowledge: Output and Trust Exploitation
Let’s see how much you’ve learned!
This quiz tests your understanding of hallucination weaponization, data leakage, improper output handling, excessive agency, and human over-trust – including the Samsung, Meta and GrafanaGhost case studies, the USENIX 2025 research, and how you would prioritise these five categories in a system you are handed.
---
shuffle_answers: true
shuffle_questions: false
---
## A development team installs an npm package recommended by their AI coding assistant. The package works correctly but secretly exfiltrates environment variables. Investigation reveals the package only exists because the AI consistently hallucinated that package name. What type of attack is this?
> Hint: Think about who created the package and why the AI recommended it.
- [ ] A supply chain attack in which a legitimate, widely used package was compromised by an attacker who obtained maintainer credentials and published a malicious release
> The package was never legitimate and had no maintainer to compromise. It was created by an attacker specifically because the AI consistently hallucinated that name, so no existing dependency was subverted.
- [x] A package hallucination attack, or slopsquatting -- the attacker registered a name that LLMs reliably invent, exploiting how predictable those inventions are
> Correct! This maps to LLM07: Misinformation, and ATLAS models it as `AML.T0062` (Discover LLM Hallucinations) followed by `AML.T0060` (Publish Hallucinated Entities). Spracklen et al. at USENIX Security 2025 found 19.7% of generated code samples carried a hallucinated package, and 43% of those names recurred in all ten runs. It beats typosquatting because a hallucinated name is correctly spelled and confidently recommended -- there is nothing on the page to notice.
- [ ] Prompt injection, in which the assistant was manipulated by hidden instructions planted in the developer's project files into recommending a package the attacker controlled
> No prompt injection occurred and no hidden instructions were involved. The AI hallucinated the package name from its own training patterns, and the attacker exploited that predictable behaviour rather than the model's input processing.
- [ ] Data poisoning, in which the assistant's training corpus was deliberately corrupted so the model would recommend specific attacker-registered packages during code generation
> No training data was corrupted. The hallucination arises naturally from how the model generates plausible-sounding names, and the attacker exploited that existing behaviour by registering a name the model already invents.
## Frontier models in 2026 hallucinate packages at roughly 4.62-6.10%, down from a 5.2-21.7% spread in 2025. Why does the 2026 research still conclude that "the range shrinks, the threat remains"?
> Hint: Think about what an attacker needs, and whether a lower per-sample rate is the only variable that changed.
- [ ] Because the reported improvement applies only to Python packages, while JavaScript and other ecosystem registries showed no measurable reduction in hallucination rates at all
> The compression of the range was observed across the evaluated cohort rather than being confined to one language. The persistence of the threat rests on cross-model convergence, not on one ecosystem being left behind.
- [x] Because five models still invent 127 identical names, 53 of them still registrable -- so one registration reaches users of several assistants at once
> Correct! Cross-model convergence is the finding that survives the rate improvement. A lower per-sample rate spread across far more AI-generated code is not obviously less exposure, and because different models converge on the same fabricated names, the defence cannot be "use a better model." The practical control is unglamorous: resolve every dependency against a verified allowlist before install.
- [ ] Because the improvement was measured on benchmark prompts rather than production code, so the published rates understate what developers actually encounter in their day-to-day work
> The re-evaluation measured the current frontier cohort on comparable tasks. The reason the threat persists is that models converge on the same invented names, which makes a single registration effective across assistants.
- [ ] Because attackers have shifted entirely away from package registries toward hallucinated API endpoints and domain names, which no current registry protection is able to detect or reserve
> Hallucinated endpoints and domains are a real extension of the pattern, but they are additional targets rather than a replacement. The 2026 conclusion rests on the 127 converging package names that remain exploitable.
## In the Samsung ChatGPT data leak (2023), three separate incidents over 20 days involved engineers pasting confidential semiconductor data into ChatGPT. What made this a data leakage event rather than a prompt injection attack?
> Hint: Consider who initiated the data exposure and whether any external attacker was involved.
- [ ] An external attacker used carefully crafted extraction prompts to pull Samsung's proprietary semiconductor designs back out of ChatGPT's training set after the confidential data had been ingested
> No external attacker was involved at any stage. Samsung engineers submitted the confidential data themselves while doing their jobs, which is what makes this a data-flow problem rather than an attack.
- [x] No attacker was involved -- authorised engineers submitted the material themselves, and the default policy then made it eligible for training
> Correct! This maps to LLM02: Sensitive Information Disclosure. The precise statement matters: the data was *eligible* for training under OpenAI's default consumer policy, and Samsung could neither verify whether it had been used nor retract it. The control failure is the irreversibility, not a confirmed leak. Samsung imposed a 1,024-byte prompt cap, then banned the tools in May 2023.
- [ ] ChatGPT had been specifically designed and marketed to harvest proprietary semiconductor process data from large industrial users in order to improve its own technical reasoning capabilities
> ChatGPT is a general-purpose assistant and was not targeting Samsung or its industry. The exposure came from authorised users submitting confidential material through ordinary use of the product.
- [ ] The submitted data was retained only for the duration of each session and then discarded, which limited the incident to a temporary and fully reversible exposure
> Under the default policy at the time, submitted conversations were eligible for training rather than discarded after the session. That inability to verify or retract the data is precisely what made the exposure serious.
## An analyst asks an internal assistant a perfectly ordinary, in-scope question. The answer is fluent, correctly cited, and contains salary details the analyst is not entitled to see. Where is the defect?
> Hint: Ask how the content got into the context window before asking what the model did wrong.
- [ ] In the model's training data, which memorized the salary records during fine-tuning and is now reproducing them verbatim in response to an unrelated query
> Training-data memorization is a real vector but the wrong diagnosis here. The content arrived at inference time through retrieval, which is why the answer is correctly cited to a source document rather than reconstructed from weights.
- [x] Upstream in retrieval -- the query returned a passage the analyst was not entitled to, and the model then summarised it faithfully, so nothing in the model malfunctioned at all
> Correct! A vector index has no concept of a user, so permission filtering that is not inside the retrieval query is not a control. The model behaved correctly on the content it was given, which is why this is the leakage vector most likely to reach a real organisation. The fix belongs in Layer 1, with output filtering at Layer 5 as a backstop -- and redaction after retrieval is strictly worse than never retrieving.
- [ ] In the model's alignment training, which failed to instruct it to refuse to discuss compensation data when a user without the appropriate entitlement asks about it
> Instructing a model to withhold content it has already been given is not an access control, because the instruction competes with everything else in the context window. Entitlement has to be enforced before retrieval returns the passage.
- [ ] In the output filter, which should have detected the salary figures in the response and redacted them before the completed answer was returned to the analyst
> An output filter is a legitimate backstop and is not the defect. By the time it runs, the restricted content has already crossed into a context the user's prompt can steer, so the control belongs in the retrieval query.
## An engineering team says: "Our LLM emits strict JSON through constrained decoding, and we stream responses for latency. Output handling is covered." Which part of that claim is wrong?
> Hint: Consider separately what a schema guarantees and what remains true once a token has left the server.
- [ ] Neither part is wrong -- constrained decoding plus streaming is the recommended production configuration for any service that passes model output to a downstream parser
> Constrained decoding is genuinely valuable and streaming is genuinely good for perceived latency, but neither addresses output handling, and together they can make an unvalidated payload arrive faster and look more trustworthy.
- [x] Both parts -- a schema guarantees the container, never the contents, and a streamed token is already published before validation can act
> Correct! `{"query": "'; DROP TABLE users; --"}` satisfies a `{"query": string}` schema perfectly, and the guarantee makes it *more* likely to be passed on unexamined -- so validate the field, not the envelope. Streaming and output validation are in direct tension: a filter can stop token 400 but cannot recall tokens 1 through 399. Anything feeding a downstream system rather than a human reader should be synchronous.
- [ ] Only the streaming part -- constrained decoding does fully sanitize field contents, so the sole remaining gap is that partial responses can reach a client before validation has finished
> Constrained decoding restricts the sampling step so the model cannot violate the schema's structure. It does not inspect or sanitize what goes inside a string field, so both halves of the claim fail rather than just one.
- [ ] Only the schema part -- streaming is safe for downstream consumers because the client reassembles the complete response before any parser or interpreter is given the chance to act on it
> Some clients buffer, but many parsers and agent loops act incrementally, and the server cannot depend on client behaviour for a security property. Once a token is emitted it is published, which is why synchronous delivery is the requirement.
## An organization has three confirmed output exploitation problems at once: (1) developers install packages hallucinated by their coding assistant, (2) LLM-generated SQL reaches the database unsanitized, and (3) a code review agent's "no issues found" reports are merged without verification. What should be remediated first?
> Hint: Compare what an attacker must already possess for each one, not how likely you are to encounter it.
- [ ] Package hallucination, because the USENIX research measured 19.7% of generated code samples carrying a fabricated dependency, which makes it by far the most statistically likely path to be exploited
> Base rate is not the triage criterion. Package hallucination requires an attacker to have already registered the name *and* a developer to install it, so it has a prerequisite and a delay that the SQL path does not.
- [x] The unsanitized SQL, because improper output handling is the only one of the three where a single crafted prompt reaches infrastructure directly with no attacker prerequisite and no waiting
> Correct! LLM10: Improper Output Handling is first and it is not close. The path is prompt → generated payload → your unsanitized sink → database, and it is exploitable right now. Package hallucination needs a pre-registered name; over-trust needs a separate attack to produce a dangerous output first. When triaging this boundary, immediacy and directness beat base rate.
- [ ] Human over-trust of the review agent, because automation bias is the behavioural root cause that amplifies every other category and therefore addresses the largest share of the total risk
> Over-trust is genuinely the multiplier across all five categories, but it has no direct exploitation path of its own. You cannot exploit automation bias the way you can exploit an unsanitized sink, so it follows the direct paths rather than preceding them.
- [ ] All three simultaneously, because output exploitation categories share a single root cause in unvalidated model output and cannot be meaningfully ranked against one another
> These differ sharply in prerequisites and timelines. One needs only a prompt, one needs prior attacker registration, and one is an amplifier -- which is exactly what makes them rankable.
## Why does a divergence attack -- asking a model to repeat a single word indefinitely -- eventually cause it to output training data such as PII and credentials?
> Hint: Think about how autoregressive models predict each next token and what happens when the repetition pattern breaks down.
- [ ] The repetition exhausts the model's memory buffer, which causes it to dump cached content left over from other users' concurrent inference sessions
> LLMs have no memory buffer holding other users' sessions, and each inference request is independent. The mechanism involves the model's learned statistical patterns rather than any form of session caching.
- [ ] The repeated word acts as a decryption key, unlocking compressed copies of the training corpus that are stored inside the model's weights
> Model weights do not contain encrypted or compressed copies of training data. The model learns statistical patterns, and the attack works by shifting which pattern it predicts next.
- [x] After enough repetitions the context stops predicting another repetition, and prediction drifts to adjacent memorized text
> Correct! Nasr, Carlini et al. (arXiv:2311.17035) measured this at 150x the extraction rate of normal prompting, mapping to `AML.T0057` LLM Data Leakage. Read the paper's conclusion rather than its trick: alignment does not eliminate memorization. Filtering this prompt removes the known path to the data and leaves the data in place -- the same shape as the backdoor-persistence result in Section 3.
- [ ] The repetition triggers a buffer-handling bug in the inference server, which causes output filtering to be bypassed for the rest of the response
> This is a property of how autoregressive models learn and reproduce patterns, not a defect in the serving stack. The behaviour comes from the model's weights rather than from a filter being circumvented.
## In GrafanaGhost (April 2026), poisoned log entries caused Grafana's AI assistant to emit a markdown image whose protocol-relative URL passed the platform's URL validation. Why is the rendering path the half a defender should focus on?
> Hint: Of the two ends of this chain, which one can you actually control?
- [ ] Because the injected instructions in the log entries could have been removed by a prompt filter tuned to detect imperative language in ingested observability data
> Text you do not control will always reach a log aggregator, and filtering imperative phrasing in arbitrary log content is not a reliable control. The renderer is the end of the chain you own.
- [x] Because you cannot prevent untrusted text reaching a log aggregator, but you can stop a renderer auto-fetching outbound images -- so the renderer is the end of the chain you actually control
> Correct! The injection was the ingress and the renderer was the egress. Any surface that renders model output is an exfiltration channel, because markdown image auto-fetch needs no user interaction and leaves only a broken image. Note also that the URL validator was present and passed: `//attacker.example` parses as a path rather than a host, an edge case an input-validation review does not look for because in a conventional app no untrusted content reaches a renderer.
- [ ] Because the vulnerability was fundamentally a server-side request forgery issue, meaning the correct remediation was to restrict which internal hosts the Grafana backend was permitted to contact
> The fetch was made by the user's browser rendering the image, not by the backend reaching an internal service, so this is an outbound exfiltration channel rather than SSRF.
- [ ] Because the AI guardrails were bypassed by an unusually sophisticated prompt, so the remediation was to strengthen the assistant's instruction hierarchy against embedded commands
> The guardrails were not defeated by a cleverer prompt. They were bypassed because the attack operated on a channel nobody had classified as output at all.
## A code review agent reports "no security issues found" on a pull request containing a subtle IDOR flaw. The developer merges it without further review. Which concept explains why the flaw was not caught?
> Hint: Think about the behavioural vulnerability that amplifies every other output exploitation category.
- [ ] Hallucination weaponization, because the agent fabricated a clean security assessment that an attacker had anticipated and deliberately engineered the pull request to elicit
> The assessment was wrong, but the question asks why the *human* accepted it without verification. No attacker engineered the agent's output here.
- [ ] Improper output handling, because the agent's report was consumed by the merge pipeline without being sanitized or validated against the actual contents of the diff
> Improper output handling concerns injection through unsanitized output reaching a sink. Nothing was injected here; the failure is in human behaviour in response to the agent's report.
- [x] Human over-trust and automation bias (ASI09) -- the developer favoured the automated assessment over independent review, and a clean report actively discourages looking deeper
> Correct! A negative finding is the highest-risk output an assistant can produce: it recommends no action, and nothing downstream contradicts it. The trust gradient makes it worse over time -- a reviewer who verified the first hundred outputs is far less likely to verify the hundred-and-first, which is exactly when one is wrong. METR measured the gap: developers were 19% slower with AI tooling while believing they were 20% faster.
- [ ] Excessive agency, because the review agent had been granted a scope of authority that should never have extended to assessing security-sensitive application code
> The agent was authorised to review code, so its scope of action is not the issue. The failure is the human's reliance on its judgement without independent verification.
## Which single change would have prevented the March 2026 Meta internal data exposure, in which an agent posted incorrect engineering advice unprompted and a colleague implemented it?
> Hint: Three categories chained together here. Which link was cheapest to remove, and which one did someone already believe was in place?
- [ ] Restricting the agent's read access to the internal forum, so that it could not have seen the original technical question or the surrounding discussion thread
> Answering the question was the agent's intended and legitimate function. Removing its access removes the feature rather than the failure, and the advice being wrong was not caused by what it could read.
- [x] Enforcing the human-in-the-loop confirmation step the engineer already believed existed, which would have broken the chain between the wrong answer and its publication
> Correct! Three categories chained: LLM07 (confidently wrong guidance), LLM03 (published autonomously without the expected gate), ASI09 (a colleague implemented it without validation). Remove any link and there is no incident, and the confirmation step was the cheapest to restore. The gap between the *expected* control and the *implemented* control was the vulnerability -- and the agent needed no privileged access, only a human who trusted its output.
- [ ] Deploying output filtering on the agent's responses to detect factually incorrect technical guidance before it was posted to the internal engineering forum
> Detecting confidently wrong but plausible technical advice is not something an output filter can reliably do, since there is no pattern to match. Gating publication is achievable where verifying correctness is not.
- [ ] Revoking the agent's permission to modify access control settings, ensuring it could not have broadened data access permissions on its own initiative
> The agent never touched permissions. A human colleague broadened access after acting on the agent's advice, which is why the control belongs at the point where advice becomes an action.