Section 3 Quiz

Test Your Knowledge: Layer 1 - Secure Your Data

Let’s see how much you’ve learned!

This quiz tests what Layer 1 owns and what it hands forward, the inheritance rule for derived stores, which controls exist for your deployment pattern, and what DSPM and encryption can and cannot do.

--- shuffle_answers: true shuffle_questions: false --- ## An organization's vector database has no authentication -- any application on the network can read from and write to it. Which Layer 1 control is missing, and what does it enable? > Hint: Think about what an attacker gains from unrestricted write access to a store the model reads from. - [ ] Encryption at rest is missing, which allows the stored corpus to be exfiltrated > Encryption protects data from someone who obtains the storage without the credentials. Here no credentials are needed at all, so the failure is access control. Encryption would also be transparent to the write. - [x] Authentication and RBAC are missing, which allows writes that bypass ingestion > Correct! Without authentication on the store, an attacker writes straight into it and every validation the ingestion pipeline performs is skipped -- source allowlisting, provenance, integrity hashing. RBAC should split ingestion (write), the application (query-only), and administration. - [ ] Data lineage tracking is missing, which allows stale documents to accumulate > Lineage records provenance so you can find and un-index affected documents. It is genuinely important and it is a detective control -- it makes a write reconstructable afterwards rather than preventing it. - [ ] DSPM discovery is missing, which leaves the database invisible to security teams > Discovery finds unknown assets and would surface this store's exposure as a finding. But the store is known here, and DSPM reports rather than enforces, so it is not the missing control. ## A RAG corpus is assembled from three sources: a public product wiki, a legal folder, and an HR share. How should the resulting vector store be classified? > Hint: The store is built from the sources, not merely pointed at them. - [ ] MEDIUM, because the store holds numerical vectors rather than documents > This is the reasoning the inheritance rule exists to block. Classifying a derived store by what it *is* rather than what it was *made from* is how a store over confidential documents ends up rated MEDIUM and deployed with a connection string. - [ ] By source, with each chunk carrying the tier of the document it came from > Per-chunk labelling is useful for entitlement filtering, but the store is one system with one access model. Anyone who can read it reads everything in it, so its protection must match its most sensitive contents. - [x] At the tier of its most sensitive source -- governed as HR-and-legal data > Correct! A derived store inherits the classification of its most sensitive source document. The store also holds the original chunk text in plaintext, because the pipeline needs the passage for the prompt, and even the vectors alone are invertible. It is a copy of your data, not an index over it. - [ ] At the tier of the public wiki, since that content is already disclosed > This inverts the rule -- it takes the *least* sensitive source as the ceiling. The presence of public content in a corpus does nothing to reduce the sensitivity of the HR and legal content sitting beside it. ## A team's DSPM tool reports that their fine-tuning dataset is writable by eleven service principals. What has the tool accomplished, and what still has to happen? > Hint: Consider which verbs appear in the definition of DSPM. - [ ] It has enforced least privilege, and the finding is a record of the change it made > DSPM does not modify access. Describing a detective control as though it enforces produces an architecture that looks defended and intercepts nothing -- the same error Section 2 found in the Blueprint's own defense-in-depth diagram. - [x] It has produced a finding; the IAM policy and gate you build in response are the control > Correct! DSPM discovers, classifies, and monitors -- it provides visibility into where sensitive data is and who can reach it. Every output is a finding. The enforcement is the IAM policy and ingestion gate you implement in response. DSPM is how you learn the gate is missing; it is not the gate. - [ ] It has blocked ten of the eleven principals, leaving the pipeline identity in place > No such action occurs. An organization with excellent DSPM coverage and no ingestion gate has perfect visibility into a dataset that anyone on that list can poison. - [ ] It has satisfied the control, since continuous monitoring is itself a preventive measure > Continuous monitoring is valuable and it is detective. It shortens the time to notice a poisoned record; it does not stop the record being written. ## A security team hardens their RAG ingestion pipeline: source allowlisting, provenance metadata, integrity hashes at ingestion. The vector store itself still accepts writes from any service on the network. How much of the poisoning risk have they closed? > Hint: Trace both attacker edges on the RAG supply chain diagram. - [ ] Nearly all of it, because a poisoned document has to pass ingestion to be embedded > This is the assumption the direct-write path breaks. The attacker does not need the ingestion pipeline to produce a vector; they can compute an embedding themselves and insert the record. - [x] Roughly half -- they closed the front door and left the direct-write path open > Correct! Every control they added evaluates at ingestion. A direct write to the store is caught by none of them, because the validations already ran at an earlier step. Hardening ingestion while leaving the store writable moves the attacker one hop and costs them nothing. Authentication on the store is the control that closes it. - [ ] All of it, since integrity hashes will detect any record that arrives out of band > Hashing detects a modification to a chunk it already knows about, on the schedule you re-verify. A wholly new record inserted directly has no prior hash to contradict, and detection is not prevention regardless. - [ ] None of it, because ingestion controls only address staleness rather than poisoning > The ingestion gate is the load-bearing control against poisoned documents arriving through legitimate feeds, and it does real work. The problem is that it is not the only path into the store. ## Your company uses a vendor SaaS assistant with a persistent memory feature. A researcher demonstrates that processed documents can write to that memory. Which Layer 1 control should you implement? > Hint: Check the ownership table before reaching for a control. - [ ] A write-path allowlist on the memory store, excluding document-derived content > This is exactly the right control and it is not yours to implement. The memory store is vendor-owned; you have no access to its write path. Recommending it in an internal review produces a finding nobody can action. - [ ] Encryption at rest on the memory store, with per-entry provenance metadata > Encryption is transparent to the injection, which arrives through the supported write path. Provenance is the right instrumentation and, like the allowlist, it sits entirely inside the vendor's product. - [x] None -- you own none here; it becomes vendor assessment and Layer 4 governance > Correct! When you consume a finished assistant, the memory store, corpus and retention policy are product behaviour. Your surface is contractual and configurational, not architectural. In the SpAIware case every control was OpenAI's, and the vendor chose to fix the exfiltration channel rather than the write primitive. The Layer 1 lesson applies to teams *building* agents with memory. - [ ] Classify the memory store as CRITICAL so that write validation becomes mandatory > Classification is what makes write validation required rather than optional, and it presumes you can then require it. You can classify a vendor feature in your own risk register; you cannot make the vendor validate writes. ## Which statement correctly describes the PoisonedRAG result, and what does it imply for defense? > Hint: The unit of the measurement is the part most often repeated wrongly. - [ ] Five documents compromise a corpus of millions, so detection at scale is hopeless > This is the common misstatement. It overstates the attack by implying general corpus control, and the fatalistic conclusion does not follow -- it points away from the control that works. - [x] Five documents per attacker-chosen question, so enumerate your high-value questions > Correct! PoisonedRAG injected five malicious texts *per targeted question* and reached roughly 90% attack success. It is narrower than the popular version, since the five documents buy the answer to one pre-chosen question. It is also worse, because attackers want specific answers -- the refund policy, the wire instructions -- and each costs only five. That makes the defense tractable: audit provenance on the chunks answering the questions that matter. - [ ] Poisoning requires a percentage of the corpus, so large corpora dilute the attack > Dilution is not a defense in retrieval. Similarity search does not care how many documents it did not return, which makes corpus size irrelevant to the attack. - [ ] Five documents per corpus, which is why ingestion filters must inspect every chunk > The unit is per question, not per corpus. Content inspection also struggles here, because a fact-bearing poisoned document contains no instructions to detect -- it reads as an ordinary plausible passage. ## An auditor notes that a RAG corpus is encrypted at rest with AES-256 and asks whether the poisoning risk is addressed. What is the correct answer? > Hint: Ask who the attacker is in each poisoning scenario from this section. - [ ] Yes, because an attacker cannot craft a poisoned chunk without reading the corpus > Poisoned documents are written, not derived from existing content. PoisonedRAG optimises texts against the embedding model, which requires no access to the corpus at all. - [ ] Partly -- encryption raises the cost of the write without eliminating it > There is no added cost. Transparent encryption is invisible to an authorized writer; the record is encrypted on the way in exactly as a legitimate one would be. - [x] No -- these attacks use legitimate write access, which encryption is transparent to > Correct! Corpus poisoning, direct store writes, metadata tampering, feedback-loop poisoning and memory poisoning are all carried out by a principal with legitimate write access through a supported path. Encryption defends a narrower threat: someone who obtains the storage without the credentials -- a stolen snapshot, a mis-scoped bucket, a decommissioned disk. Keep it, and do not count it against poisoning. - [ ] Yes, provided keys are rotated and key access is logged separately from data access > Rotation and key-access logging are sound practice against key compromise. Neither changes what an authorized writer can put into the corpus. ## A user asks an internal assistant a legitimate question and receives an accurate answer built from a salary document they were never entitled to read. Which layer owns the failure, and which one owns the backstop? > Hint: Locate the moment the failure occurs relative to the moment the request exists. - [ ] Layer 5 owns it, because the response filter should have redacted the salary figures > Response filtering is the backstop, not the owner, and redaction after retrieval is strictly worse than never retrieving -- the content has already entered a context the user's prompt can steer. - [ ] Layer 1 owns it entirely, since classifying the corpus would have prevented the disclosure > Classification is how you know the content is in the corpus, and it is necessary. It does not by itself decide whether *this* reader may see *this* passage, which is the decision that failed. - [x] Layer 1 for having the content classified and indexed; Layer 5 for entitlement at query time > Correct! This is LLM02's most common route -- retrieved content passed to a reader with no entitlement, no attacker involved. Layer 1 owns knowing the content is there and how sensitive it is; the entitlement decision happens when a request exists, so it belongs to Layer 5 and must be enforced *inside* the retrieval query rather than applied to its results. - [ ] Layer 4 owns it, because the user acted on AI output without verifying the source > The user did nothing wrong and the answer was accurate. Layer 4 addresses over-trust in AI output, which is a different failure from a correct answer drawn from an unentitled source. ## A team is choosing between a dedicated vector service and pgvector in the PostgreSQL instance they already operate. From a Layer 1 standpoint, what is the argument they should weigh? > Hint: Consider which controls already point at each candidate store. - [ ] The dedicated service is stronger, since purpose-built stores ship better security features > The major dedicated services do ship RBAC, encryption and SOC 2 attestations, so this is not a feature gap. The 2026 problem is the distance between features a product supports and what a given deployment turns on. - [x] The extension keeps vectors, chunk text and permissions in a system your controls already cover > Correct! Audit tooling, backup policy, IAM integration and DLP coverage already point at the database you run. A dedicated service is a new data store with its own access model, and every one of those controls has to be re-established around it. Choose the dedicated service for scale or features, then budget the governance work -- rather than discovering it during an audit. - [ ] They are equivalent at Layer 1, because both stores hold identical data and identical vectors > The contents are comparable; the governance around them is not. Layer 1 is about who can reach a store and what watches it, which differs sharply between the two options. - [ ] The extension is weaker, because relational databases cannot enforce per-chunk entitlement > Metadata filtering works in both, and a relational store is if anything better placed for it, since entitlement data can live beside the chunks and be joined at query time.