9. The AI Application Security Continuous Loop

Introduction

A guardrail rule set is the only control in the Blueprint with a shelf life.

Every other control you have built holds its value while you leave it alone. An encrypted volume stays encrypted. A signed model artifact stays signed. An egress allowlist keeps denying the same destinations next quarter. But a filter is a description of attacks that were known when it was written, and Chapter 2 Section 2 spent a whole section on why that description is never finished: encoding, Unicode and token smuggling, language switching and cross-modal payloads each defeat a filter that was correct the day it shipped. Section 7 put a number on the residual – the best-instrumented production guardrail on record still leaves a low-single-digit bypass rate against dedicated attackers, and that figure came from a frontier lab spending 1,700 hours of red-teaming on it.

So the operational question is not “is the filter good?” It is “what process keeps it aimed at the attacks that actually arrive, and how would I know if it stopped being aimed?” That process is what this section covers.

Two TrendAI Vision One components instrument it. AI Scanner probes an AI application before and during deployment and reports what it was able to make it do. AI Guard sits in the live request path and inspects prompts and responses as they flow. Neither is interesting on its own – Scanner produces a report nobody actions, Guard produces rules nobody updates. They are interesting as a loop: scan, protect, validate, improve, with each phase producing the input to the next.

The loop is not a product feature. MITRE ATLAS names the same cycle as a mitigation in its own right – AML.M0035 AI Red Team – and specifies that red-teaming is “a continuous process” whose findings must be converted into “regression tests, evaluation datasets, detection logic, monitoring requirements, or deployment criteria.” Read this section as that practice, with two products supplying the instrumentation. The practice is what transfers if you buy neither.

What will I get out of this?

By the end of this section, you will be able to:

  1. State what AI Scanner can and cannot establish, and read a scan result as evidence of a specific strength rather than as a pass or fail.
  2. Place AI Guard correctly in the request path, including the availability decision – fail-open or fail-closed – that an inline inspection call forces on you.
  3. Run the scan-protect-validate-improve loop as a practice, naming the artifact each phase must produce for the next phase to have anything to work with.
  4. Assign each phase of the loop to an owner, distinguishing the work a product performs from the work no product arrives configured to do.
  5. Make the deployment and data-flow decision for an inline inspection service, on the axes that actually decide it.

What the Loop Owns – and What It Cannot Stop

The loop is presented in most vendor material as closing a gap. It is worth being precise about which gap, because two of the three things people expect it to cover are covered elsewhere in the Blueprint, and one of them is not covered at all.

What the loop genuinely owns is decay. A control that was correct at deployment drifts out of alignment with the traffic in front of it: the model gets updated, the application gains a new retrieval source, a technique that did not exist last quarter starts arriving. Nothing else in the Blueprint notices that. Layer 5’s filter does not know its own bypass rate. Layer 6’s baseline does not know it was built during an attack. The loop’s product is evidence about your own controls – and evidence is the thing an AI security program is usually shortest of.

What it cannot do is find a technique nobody has described. This is a property of the design, not a maturity gap. Scanner assesses by running a library of known attacks; a technique absent from the library returns clean, and a clean result means “none of these worked,” never “nothing works.” That is precisely the boundary Section 8 is built on: Layer 6 exists because signature-shaped controls need the attack to be described first, and the honest reading of a Scanner report is that it is a signature-shaped control with the same limit.

And it cannot bound consequence. This is the one that gets missed, and the case study below is an entire incident built on it. Scanner and Guard both operate on text – prompts in, responses out. An attack whose damage is done by an action the agent was permitted to take passes both of them cleanly, because at no point is there a hostile-looking string to catch. Bounding consequence is Layer 4’s human gate and Layer 3’s permission scoping, and no amount of loop maturity substitutes for either.

The characteristic failure is treating a clean scan as an assurance artifact

A Scanner report showing no successful injections is very frequently forwarded to a risk committee as evidence that the application is safe to ship. It is not that, and the asymmetry is worth stating in the same terms Section 4 uses for model integrity checks: a finding is strong evidence, a clean run is weak evidence, and both are inputs to a deployment decision rather than the decision.

A finding is strong because it is a demonstration – someone made the system do the thing, and the transcript is in the report. A clean run is weak because its scope is the library, its adversary is automated, and its budget is minutes. Compare that with the 1,700 hours behind the figures in Section 7 and the calibration is obvious.

The residual, then, is not a defect of the tooling. It is the operating assumption. Design the system so that a bypass costs you an unhelpful answer rather than an unrecoverable action, and the loop’s job becomes what it can actually deliver: keeping the bypass rate low, and telling you when it starts climbing.


Which Loop Responsibilities Are Yours

Chapter 1 Section 3 promised that “the layers stay constant, but which ones are yours is set by the choice you make here.” For the loop the split does not run along deployment patterns – it runs between what a product performs and what only you can supply, and the second column is the larger one.

Loop responsibility What the product does What only you can do
Deciding what to assess Runs against an endpoint or application you point it at Enumerate the AI surface. A scanner never discovers the assistant a team stood up on a personal key – that is Layer 4’s shadow-AI problem
Attack library coverage Maintains and updates a library of known techniques Judge whether the library covers your threat model – your modalities, your languages, your retrieval sources
Interpreting a finding Reports which probes succeeded, with severity Decide whether the finding matters given what the application can do. A jailbreak in a read-only summariser and in a tool-calling agent are not the same finding
Configuring the filter Supplies detectors, classifiers and policy primitives Define the content policy. What counts as sensitive, prohibited, or out-of-scope is a business decision no vendor default encodes
Placing the inspection point Exposes an API and integration hooks Decide where in your request path it is called, on which inputs, and what happens when it is unavailable
False-positive tuning Reports rates and provides tuning controls Own the threshold. The cost of a blocked legitimate request is yours, and so is the decision about what rate is tolerable
Re-scan cadence and triggers Supports scheduled and pipeline-invoked scans Define the triggers – model version change, new retrieval source, new tool, prompt template edit
Converting findings into tests Records what succeeded Turn a confirmed failure into a regression test that runs forever (AML.M0035)
Acting on a validated gap Alerts Change the architecture. If the gap is consequence rather than detection, the fix is a permission or a gate, not a rule

Three readings, and they are why this table is here rather than a feature list:

1 · The purchasable half is the middle of the loop. Detectors, libraries and dashboards are products. Scoping at the front and remediation at the back are not, and a program that buys the middle and skips both ends produces a great deal of telemetry and no change in risk.

2 · The bottom row is the one that decides outcomes. Every other row makes detection better. That row is the only one that can move an incident from “detected late” to “not possible,” and it is the row a tool cannot even suggest, because it requires knowing what the application is for.

3 · The failure is silent. A loop where nobody owns the last two rows still runs. Scans complete, Guard blocks, dashboards populate. Nothing signals that findings are accumulating without being converted into anything. Assign those rows to a named owner before you tune a single detector.


AI Scanner: Proactive Assessment

AI Scanner assesses AI applications – not only models – for exploitable weaknesses, and ranks what it finds. Point it at an application or model endpoint and it probes with known attack techniques, then reports what succeeded. Its documented scope includes prompt injection susceptibility, data leakage, and authentication bypass, which is worth noting because it is the one class that is not a model property at all: it is a finding about the application wrapped around the model, and it is the reason the tool is scoped to applications rather than weights.

What AI Scanner Does

  • Attack technique testing. Known injection variants, jailbreak patterns and adversarial inputs, run against the target rather than reasoned about. The deliverable is a transcript of what worked.
  • Vulnerability identification and ranking. Which categories the target is susceptible to, mapped to OWASP, with severity – so that a susceptibility reachable by a first-time user is separated from one that took a chained multi-turn setup.
  • Data leakage probing. Whether the model has memorised and will reproduce sensitive training data, which is the extraction route of LLM02: Sensitive Information Disclosure. Layer 1 owns the fix: by the time this fires the data is already in the weights, and the remediation is upstream in corpus curation.
  • Harmful content generation. Whether safety alignment holds under adversarial pressure, which is a Layer 4 design input – it tells you how much verification the humans downstream will need.
  • Configuration and access findings. Authentication and authorisation weaknesses on the serving endpoint itself, which belong to Layer 3 and are where a scan most often produces something immediately fixable.

This is ATLAS AML.M0016 Vulnerability Scanning applied to a running application rather than to an artifact – the artifact half sits at Layer 2, and the two answer different questions about the same model.

When a Scan Runs

Scanning is worthless as a one-off and cheap as a trigger, so define the triggers explicitly:

  • Pre-deployment, as a pipeline gate. This is the Validate stage of the AI-SDLC from Section 1, and the gate should fail the build on a finding above an agreed severity – otherwise it is a report, not a gate.
  • On model version change, including a provider’s silent update to a hosted model. A model you did not change can behave differently next month, and this is the trigger teams most often lack.
  • On surface change – a new retrieval source, a new tool binding, an edited system prompt. Each of these changes what an injection can reach, which changes the meaning of the last clean scan.
  • On a schedule, to catch library growth. A quarterly re-scan against an unchanged application still finds things, because the library moved even if you did not.
Defense Connection

Scanner is how the prompt-level attacks of Chapter 2 get tested against your deployment rather than read about. The difference matters for a specific reason: susceptibility is model- and prompt-specific, so the published jailbreak that fails against your system and the obscure one that works are indistinguishable until someone runs both. What Scanner cannot tell you is what a successful injection would then be able to do – that is the action scope you set at Layer 5, and it is the number that decides whether a finding is an inconvenience or an incident.


AI Guard: Runtime Protection

If Scanner asks “what can this system be made to do,” Guard asks “is someone doing it right now.” It reached general availability on 1 December 2025 as an API endpoint within TrendAI Vision One that inspects AI system inputs and outputs in real time.

What AI Guard Does

  • Inbound guardrails. Injection detection on prompts and on the retrieved content, tool output and documents that enter the context alongside them – which, as Section 7 establishes, is where the payload actually arrives in every case study in this course.
  • Sensitive information leakage prevention. Outbound scanning for PII, credentials and system prompt fragments, with blocking or redaction.
  • Harmful content filtering. Responses that are violent, illegal, abusive or policy-violating, refused before delivery.
  • Policy and compliance enforcement. Organisational content policy applied consistently across every application that routes through it, which is the property that makes a central inspection point worth having at all.

In ATLAS terms this is AML.M0020 Generative AI Guardrails: “safety controls placed between users, tools, and generative AI models to evaluate prompts, retrieved context, model outputs, and agent actions before they are accepted, executed, or shown to a user.” Note the four objects in that sentence. A deployment that inspects only the user’s prompt is implementing one of them.

Where Guard Sits in the Request Path

Guard is consumed as an API endpoint – Trend-hosted at a regional address (api.<region>.xdr.trendmicro.com), or self-hosted in your own environment. Either way it does not intercept traffic by itself: something in your path has to call it, and that something is your Layer 5 gateway or LLM proxy. TrendAI documents the pattern with LiteLLM: the proxy sends each incoming request to AI Guard for scanning before forwarding it to the target LLM provider, and sends the provider’s response to AI Guard again before returning it to the caller.

Hold on to that sentence, because it is the part people get wrong when they compare deployment options. Self-hosting Guard changes where the inspector runs; it does not make the inspector a network device. The gateway is still yours, the call is still explicit, and coverage is still a property of your routing.

That architecture has a consequence the product literature does not dwell on, and it is the most important operational decision in this section.

An inline inspection call has an availability mode, and you must choose it deliberately

Guard is a network call on the critical path of every request. Sooner or later it will time out. Your gateway does one of two things, and the choice is yours to make in advance:

Fail-open – on error or timeout, forward the request unfiltered. Availability is preserved; the guardrail silently disappears exactly when something is stressing the system, which is also when an attack is most likely to be under way.

Fail-closed – on error or timeout, refuse the request. The guarantee holds; a Guard outage becomes an outage of your AI application, and an attacker who can degrade the inspection service now has a denial-of-service primitive.

Neither answer is universally right, and the useful move is to stop asking it as one question. Decide per action class, the same way Layer 4 decides what to gate: fail-open on a read-only summarisation path where an unfiltered answer is an acceptable worst case, and fail-closed on anything that writes, spends, or reaches a third party. Then alert on the fail-open path itself – a filter that is bypassed and silent is worse than no filter, because the architecture around it was sized assuming it was there.

A second consequence: because inspection is a call your gateway makes, coverage is a property of your routing, not of the product. Section 7’s finding applies unchanged here – “a correctly configured gateway that 40% of traffic never traverses” protects 60% of your traffic, and a Guard subscription does nothing about the other 40%.

What Guard Does Not Change

Guard is a guardrail, and Chapter 2’s verdict on guardrails stands: a cost-raiser, not a boundary. The four evasion families are not addressed by procuring a better one, the measured residual in Section 7 is low single digits under sustained attack, and the OWASP LLM01 entry states outright that there is no fool-proof prevention. Deploy it – raising attack cost is a real outcome – and size the blast radius as though it will occasionally fail, because it will.


Deployment and Data Flow

Objective 5 is a decision, so here is the decision rather than a menu. Both AI Scanner and AI Guard ship in two deployment modes – Trend-hosted, using TrendAI’s cloud infrastructure, and self-hosted, deployed and run in your own environment, on-premises or in a private cloud. TrendAI’s own documentation states the deciding axis in one line: “If your data needs to remain local, choose self-hosted instead of hosted by TrendAI.”

So the question people ask – “cloud or self-hosted?” – has a documented answer, and the question that produces it is: what is the maximum sensitivity of the content that will cross the inspection boundary, and does your jurisdiction, contract or classification permit it to cross? Everything else follows.

Note what makes inspection different from every other AI service you evaluate: to inspect a prompt, the inspector must receive the prompt. A guardrail service therefore sees, by construction, one hundred percent of your AI traffic – including the traffic your DLP policy exists to keep in-house. It is one of the highest-sensitivity data flows in the architecture, and it is routinely approved as a security tool without the review a data flow of that sensitivity would otherwise get.

Axis What to establish Why it decides the outcome
Content sensitivity The most sensitive class of prompt or response that will traverse the path Sets whether an external inspection endpoint is permissible at all. This is the gating axis; the rest only matter if it passes. Where it fails, it selects self-hosted rather than ending the conversation
Jurisdiction and residency Which region the endpoint terminates in, and whether that satisfies your obligations Trend-hosted exposes regional endpoints (us, eu, jp, au, in, sg, mea), which is the documented residency mechanism. A requirement no region satisfies is the self-hosted case
Connectivity Whether the environment can reach an external endpoint at all A disconnected or air-gapped environment cannot call a Trend-hosted endpoint, and no configuration of the hosted mode fixes that. It is a constraint on the mode, not on the control
Operational ownership Who patches the inspector, updates its detection content, and carries its uptime The axis self-hosting turns around. A hosted inspector is updated for you; a self-hosted one inherits this section’s own thesis – a rule set decays, and now the decay is yours to manage
Latency budget Added round-trip on every request, in and out Two extra calls per interaction, in either mode. Measure against your p99 target before you commit the architecture
Availability coupling Your fail-open / fail-closed choice, per action class Determines whether the inspector’s uptime becomes your application’s uptime
Coverage The share of AI traffic that traverses a path you control An inspection point protects what routes through it and nothing else

Where that leaves the hard cases. If content sensitivity and connectivity both pass, Trend-hosted at the region that satisfies your obligations is the straightforward answer, and the remaining work is placement and tuning. If either fails – classified material, a residency rule no region matches, a disconnected environment – the control does not disappear with the hosted mode. Self-hosted is the documented answer to exactly that case, and the price of it is the Operational ownership row: you now run, patch and update the inspector, and the detection content that made it worth having is on your maintenance schedule.

There is a third case the product cannot answer, and it is the one to plan for. Where neither mode is available to you – no budget for the component, an environment nothing may be deployed into, a platform decision made elsewhere – the requirement still does not disappear. It has to be met another way: a locally-run guardrail model in your own serving stack, deterministic policy checks at the gateway, and correspondingly more weight on the controls that do not require inspection at all – action scoping, least-privilege tool credentials and human gates on the irreversible subset.

That is the transferable lesson, and it is why the axes matter more than the modes. When the inspecting control is unavailable to you, the answer is not a weaker inspector – it is an architecture that needed the inspector less.

Check the current deployment matrix before you design around it

AI Scanner and AI Guard are actively developed components and their supported deployment topologies change between releases – which is not a hypothetical caution. The facts stated here were verified against the TrendAI Online Help Center on 17 August 2026: Guard’s 1 December 2025 GA as an API endpoint, the Trend-hosted regional addresses, the documented self-hosted mode for both components, and the gateway-mediated integration pattern. An earlier revision of this section described the hosted endpoint as the only mode, which was wrong on the day it was written.

Confirm the current matrix against the vendor’s documentation before committing an architecture to it, and treat the axes in the table above as the durable part. Deployment modes are a fact with an expiry date; “the inspector must receive what it inspects” is not.


The Continuous Loop

Each phase exists to produce a specific artifact that the next phase consumes. Name the artifact and the loop is auditable; leave it unnamed and the loop degrades into four activities that happen near each other.

The Scan-Protect-Validate-Improve Cycle

graph LR
    S["<b>SCAN</b><br/><small>Probe the application<br/>with known techniques.<br/><i>Artifact: ranked findings<br/>with transcripts</i></small>"]
    P["<b>PROTECT</b><br/><small>Configure rules and<br/>bound consequence.<br/><i>Artifact: rules, plus a<br/>scope or gate change</i></small>"]
    V["<b>VALIDATE</b><br/><small>Measure the controls,<br/>not the threat.<br/><i>Artifact: detection and<br/>false-positive rates</i></small>"]
    I["<b>IMPROVE</b><br/><small>Convert failures into<br/>permanent tests.<br/><i>Artifact: regression suite,<br/>new detection logic</i></small>"]

    S -->|"Ranked<br/>findings"| P
    P -->|"Runtime logs,<br/>blocked patterns"| V
    V -->|"Measured<br/>gaps"| I
    I -->|"Regression tests,<br/>new techniques"| S

    style S fill:#1565c0,color:#fff
    style P fill:#2d5016,color:#fff
    style V fill:#7a6a00,color:#fff
    style I fill:#6a1b9a,color:#fff

1 · SCAN – Produce Ranked Findings

Scanner runs its library against the target and reports which probes succeeded, how hard they were, and what the successful prompts were. Read the output against two questions rather than one:

  • What worked? – the finding, with its transcript. This is the strong evidence.
  • What could it reach? – not in the report, and not knowable by the scanner. A susceptibility in an application that can only return text is a different risk from the same susceptibility in an application holding a write-capable tool.

The second question is why a Scanner report should never be triaged by severity alone. Severity describes how easily the model was manipulated; impact is set by the permissions on the other side of it, and those live in your architecture.

2 · PROTECT – Configure Rules and Bound Consequence

Guard rules are configured against what Scanner actually found, which is the property that makes them better than defaults: if the finding was a role-play jailbreak, the detection targets that pattern; if the finding was training-data reproduction, output rules target the specific data classes that came back.

But a phase that only writes rules has already conceded the point made at the top of this section. Every finding gets two responses: a detection change and a consequence change. The detection change lowers the odds. The consequence change decides what happens on the occasion the detection loses – narrower tool credentials, a human gate on the irreversible subset, a scope reduction on what the compromised path can reach. A loop whose PROTECT phase only ever produces rules is a loop that will keep detecting the same class of incident more accurately and never stop one.

3 · VALIDATE – Measure the Controls, Not the Threat

This is the phase that is most often skipped, because it produces no artifact anyone is asking for until an incident. Four measurements, and each has a decision attached:

Measurement The decision it drives
Detection rate against a known-attack set replayed on a schedule Whether the rule set still covers what it covered at deployment
False-positive rate on legitimate traffic The tuning threshold – and this one is a security metric, not a UX metric. A guardrail with an intolerable FP rate gets switched off by the team it obstructs, which is the most common way a deployed control stops existing
Coverage – share of AI traffic traversing an inspected path Whether the number above is being measured on 60% of reality
Novel-pattern volume from Guard’s blocked logs Whether attackers have moved to techniques the last scan did not test

Two things make this phase real rather than nominal.

Replay, don’t assume. A detection rate you have not measured recently is a detection rate from the day you deployed. Maintain a corpus of attacks that previously succeeded and replay it against the current configuration on a schedule. This is where AML.M0035 is explicit: convert confirmed failures into “regression tests, evaluation datasets, detection logic, monitoring requirements, or deployment criteria.” A failure that becomes a permanent test can only regress once.

Automated assessment is not red-teaming, and the gap is the interesting part. ATLAS scopes an AI red team across “models and data, agents (including memory and tools), data flows, decision processes, application logic, retrieval systems, identities and permissions, software dependencies, non-AI system components, infrastructure, user interfaces, and human workflows.” A scanner covers the first item and part of the second. The rest – the chained, multi-step, human-in-the-workflow attacks that produced every case study in Chapter 2 – needs people. Section 11 covers how that team is scoped and run; the loop’s job is to make sure their findings do not stay in a report.

4 · IMPROVE – Make the Failure Permanent

The improve phase has one test of success: something changed that cannot silently revert.

  • A regression test exists for every confirmed failure, running in the pipeline.
  • Guard rules updated from novel patterns in the blocked logs.
  • Scanner’s library extended, where the technique is new rather than merely new to you.
  • Thresholds retuned from false-positive analysis, with the new rate recorded so the next cycle has a comparison.
  • Findings routed to the layer that owns them. A volume of exfiltration attempts is a Layer 1 data-classification signal. A cluster of tool-call anomalies is a Layer 6 baseline input. A finding that keeps recurring after two cycles of rule tuning is telling you the fix is architectural.

That last item is the loop’s exit condition, and it deserves stating plainly: if the same class of finding survives two full cycles, stop tuning. The loop has done its job – it has produced evidence that detection is the wrong control for this problem. What follows is a permission change, a gate, or a design change, and it belongs to a different layer.


Scanner and Guard Across the Blueprint Layers

Neither tool belongs to a single layer. What each contributes per layer is worth stating precisely, because the contributions are uneven and the uneven ones are where the misconceptions live.

Blueprint Layer AI Scanner contribution AI Guard contribution
Layer 1: Data Tests whether training data was memorised and is reproducible – evidence about your corpus, not a model defect Blocks sensitive data in responses; enforces classification policy at the boundary
Layer 2: Models Assesses the model’s behavioural properties: adversarial robustness, alignment under pressure, injection susceptibility None. Layer 2 secures a static artifact – weights, provenance, serialisation, container. Guard never sees an artifact; it sees traffic, and traffic is Layer 5
Layer 3: Infrastructure Finds authentication and authorisation weaknesses on serving endpoints; feeds AI-SPM risk scoring Emits runtime telemetry into posture monitoring (AML.M0024)
Layer 4: Users Tests whether the model produces confidently wrong or harmful output – a design input telling you how much verification the humans downstream need Filtering itself sits in Layer 5; what Layer 4 adds is the gate and the calibration around what does reach a user
Layer 5: Access Tests susceptibility to injection, jailbreaking and system prompt extraction – the exact classes Layer 5 filters The enforcement engine. Guard is what the AI Gateway calls to inspect; the gateway routes, Guard decides
Layer 6: Zero-Day Supplies a description of normal that was not learned from production traffic, which is what a behavioural baseline is worst at getting on its own Contributes labelled events – the blocked-attack corpus a baseline needs to be calibrated against something other than noise

The Layer 2 row is the one that teaches the distinction, and it is worth reading twice. It is not that Guard is weaker at Layer 2; it is that Layer 2’s subject does not exist at runtime. A signature on a weights file is either valid or not before a single request arrives, and there is no prompt that can be inspected to establish it.


Defense Perspective: GitHub Copilot CVE-2025-53773

An incident the loop would not have prevented, and what that tells you

The attack (from Chapter 2 Section 2): prompt injection delivered through any content Copilot read as context – a source file, a README, an issue, a fetched page – caused the agent to write "chat.tools.autoApprove": true into .vscode/settings.json. That setting disables every confirmation prompt for shell commands. Having silently escalated its own privileges, the agent then executed arbitrary commands with the developer’s rights. Wormable, because the payload could be committed back to the repository. Reported 29 June 2025, patched in the August 2025 Patch Tuesday.

Note what did not happen: no malicious suggestion was ever accepted. The settings file was written by the agent without a suggestion and without an approval prompt, so code review never entered the path. This is the detail that makes the case worth its place here.

Where the loop would have helped:

  • SCAN would have found the susceptibility. Probing the assistant with instructions embedded in repository content establishes that it follows them, and that finding – the agent acts on instructions from files it reads – is exactly what a scan is good at producing. It was reported by a researcher doing manually what a scan does automatically.
  • VALIDATE and IMPROVE would have shortened the tail. Once the technique was public, a replayed regression test would show whether your configuration was still exposed after the August patch, and across every subsequent agent version.

Where the loop would have stopped, and why:

  • PROTECT could not have caught the payload as text. The injected instruction is an ordinary sentence about a configuration setting. There is no encoding, no jailbreak framing, nothing that distinguishes it from documentation. And Chapter 2’s EchoLeak case already established the stronger version of this point: a payload can be deliberately phrased to survive a production injection classifier.
  • Output filtering had nothing to inspect. The damaging output was not shown to anyone. It was a JSON key written to disk by a permitted file-write tool. A guardrail on prompts and responses does not sit on that path.

The lesson is the one Chapter 2 draws: injection was the entry, agency was the impact. The single control that would have prevented this is the one from the ownership table’s bottom row – denying the agent write access to the file governing its own approvals. No task Copilot performs requires it, and an agent that can rewrite its own approval settings has no meaningful approval control at all. That is a Layer 4 gate and a Layer 3 permission scope, reached through the loop but not implemented by it.

Which is precisely the value of running the loop honestly. A SCAN finding that says “this agent follows instructions from files it reads” is not a request for a better filter. It is a request to look at what the agent is permitted to do – and a loop whose PROTECT phase only ever writes rules would have filed it as a detection gap and moved on.


The Loop → OWASP Mapping

Scanner and Guard are named in the course-wide mapping against a specific set of categories. For each one, the same two-part honesty applies as in every layer section – what the loop does, and what it does not reach.

Category What the loop does What it cannot do Completed by
LLM01: Prompt Injection Scanner establishes susceptibility pre-deployment; Guard raises the cost at runtime; the loop keeps rules aimed at arriving techniques Close it. The residual is low single digits against a dedicated attacker, and the payload may be phrased to survive the classifier L5 action scope and L4 gates decide what a bypass costs
LLM02: Sensitive Information Disclosure Scanner probes for memorised training data; Guard redacts or blocks PII, credentials and prompt fragments on the way out Remove data from the weights. A positive scan is evidence about a pipeline that already ran L1 corpus curation and classification
LLM05: Data and Model Poisoning Scanner detects behavioural indicators – a model that responds anomalously to particular triggers Prove absence. A backdoor with an unguessed trigger is behaviourally invisible L1 lineage and L2 provenance and integrity
LLM08: Hidden Context Exposure Scanner tests extraction; Guard blocks instruction-shaped fragments in responses Make the system prompt secret. Anything in the window is reachable by anyone who can talk to the model L5 – keep credentials and authorisation out of the window entirely
LLM10: Improper Output Handling Guard inspects generated content before it is returned, which reduces what reaches an unhardened sink Encode for a sink it cannot see. Safe rendering is a property of the consumer, not of the producer Each sink encodes for itself – Chapter 2 Section 6; L6 virtual patching covers the unhardened ones
WarningASI01: Agent Goal Hijacking Scanner establishes that the agent follows injected instructions; Guard inspects the content entering context Stop the hijack. Every subsequent tool call is authorised and looks normal L4 gates the consequence; L3 scopes the credential

Read the third column down and the pattern is consistent: the loop’s limit is always the same limit. It acts on text, and in every row the thing it cannot reach is either a property fixed before runtime or an action taken after the text was accepted. That is not a gap in the tooling – it is the definition of where assessment and filtering sit in a defense-in-depth architecture, and it is why this section comes after the six layers rather than instead of them.

Note also which categories are absent. LLM03: Excessive Agency and LLM04: Supply Chain have no row, because there is no prompt or response whose inspection establishes anything about either. The Copilot case above is exactly what an LLM03 incident looks like from inside a loop that has no LLM03 row.


TrendAI Vision One Integration

Scanner and Guard are GA components of TrendAI Vision One, and the integration argument is about correlation rather than features. Scanner runs from CI/CD as a pre-deployment gate; Guard is called inline by your gateway; both report into the Vision One console alongside AI-SPM posture findings and Layer 6 network detections.

The value of that convergence is the same one Section 2 works through: single-layer signals that are individually unalarming and jointly unmistakable. A Scanner finding that a model leaks its system prompt, a posture finding that the same endpoint is reachable without authentication, and a Guard log showing extraction attempts against it are three low-priority tickets in three consoles and one incident in one timeline.

Read that against the ownership table before sizing the investment. The platform correlates and detects; the scoping at the front of the loop and the remediation at the back are yours in every deployment pattern, and a console arrives configured for neither.

Key Takeaways
  • A guardrail rule set is the only Blueprint control with a shelf life. The loop exists because filters decay and scans age, and its product is evidence about your own controls – not a new class of protection.
  • A finding is strong evidence; a clean scan is weak evidence. A scanner runs a library of known techniques, so clean means “none of these worked,” never “nothing works.” Treat both as inputs to a deployment decision, not as the decision.
  • Every finding deserves two responses: a detection change and a consequence change. A PROTECT phase that only ever writes rules will detect the same class of incident more accurately forever and never prevent one.
  • The loop acts on text, so it cannot bound consequence. CVE-2025-53773 was an ordinary sentence about a config setting and a JSON key written to disk – no hostile string to catch, no output to inspect. Injection was the entry; agency was the impact.
  • An inline inspection call has an availability mode. Choose fail-open or fail-closed per action class, and alert on the fail-open path – a filter that is bypassed and silent is worse than none, because the architecture was sized assuming it was there.
  • Content sensitivity is the gating axis of the deployment decision, because inspection requires the inspector to receive one hundred percent of your prompts. Where a hosted endpoint is not permissible it selects the self-hosted mode, and where no mode is available the answer is an architecture that needs the inspector less – never a weaker inspector.
  • If a class of finding survives two cycles, stop tuning. The loop has produced its most valuable output: evidence that detection is the wrong control for that problem.

Test Your Knowledge

Ready to test your understanding of the AI security continuous loop? Head to the quiz to check your knowledge.


Up next

Scanner and Guard operationalize the Blueprint for security teams. But none of it answers the developer’s question: how do we know the application we wrote is secure? In Section 10 you’ll work through OWASP AISVS 1.0 – 191 testable requirements across 12 chapters – with the LEARN mnemonic as the recall index into it. Where this section keeps controls from decaying, that one defines what “correct” means in the first place.