6. Layer 4: Secure Your Users

Introduction

In January 2024, a finance employee at the Hong Kong office of the engineering firm Arup joined a video conference with the group CFO and several colleagues. Everyone on the call was synthetic. Over the course of a day the employee executed 15 transfers totalling HK$200 million (about US$25.6 million) to five Hong Kong bank accounts. The fraud surfaced only when they later contacted headquarters to discuss the “confidential transaction” — and none of the money has been recovered. MITRE ATLAS catalogues the case under AML.T0052.001 Deepfake-Assisted Phishing, and the source material the attackers used to build the faces and voices was Arup’s own conference footage and earnings calls.

Nothing in that incident touched a model, a vector store, or a serving container. Every control that failed was a control over a person and the device in front of them — and that is the whole of Layer 4.

The layer has two faces, and they point in opposite directions. Users are targets: cloned voices, composited faces, and AI-generated phishing make impersonation cheap and convincing. Users are also a source of risk: they adopt unsanctioned AI tools, paste proprietary data into consumer chatbots, and accept AI output that a track record has taught them not to check. Layer 4 protects users from AI-powered attacks and protects the organisation from users’ uncontrolled AI adoption — and, most consequentially, it is the layer that decides which AI actions a human has to approve before they take effect.

In Chapter 2 you saw the attacks this layer answers: output and trust exploitation turning a wrong answer into an action, human-agent trust exploitation exploiting the trust a reliable agent accumulates, and rogue agents that nobody registered and nobody owns.

What will I get out of this?

By the end of this section, you will be able to:

  1. State what Layer 4 owns and what it cannot stop, and explain why a layer that touches no traffic is still the one that bounds the damage.
  2. Decide which Layer 4 controls are yours given your device, identity, and contractor estate — rather than your model deployment pattern.
  3. Design a synthetic-media defense that does not depend on detection, using the provenance and out-of-band verification controls that survive a detector being wrong.
  4. Design a shadow AI discovery and governance program that finds local model runtimes as well as network-visible services, and that balances security with productivity.
  5. Place a human-in-the-loop gate on the irreversible subset of an AI system’s actions, and explain why gating everything produces a gate on nothing.
  6. Map Layer 4 to its four OWASP categories, and say which of this section’s threats has no OWASP LLM slot at all and why.

What Layer 4 Owns – and What It Cannot Stop

Section 2’s comparison table records Layer 4 as human-boundary, with a flat No under “stops a live request” — because “it acts on people and devices, not on traffic” — and coverage of 4 of 20 OWASP categories, the narrowest in the Blueprint. Every part of that is correct, and read carelessly it makes Layer 4 look like the least important layer in the course. Section 2 already supplies the correction: Layer 4 “bounds what a wrong or leaked answer can cause.”

That is a different job from the other five layers, and it is the only one that survives their failure. Layers 1 and 2 have finished before the request arrives. Layer 5 filters the request. Layers 3 and 6 watch and alert. If every one of them misses — and Chapter 2’s case studies are mostly cases where they did — the incident still requires a human to act, or a machine to act on a human’s behalf without being asked. Layer 4 is the last place that can be false.

What Layer 4 owns:

  • The approval boundary. Which AI-initiated actions take effect on their own, and which wait for a person. This is the layer’s highest-value control and the one most often absent.
  • The device. What AI tooling runs on it, what leaves it, what is installed that nobody approved.
  • The verification procedure. How a person confirms that a request arriving through a channel is what it appears to be.
  • The knowledge estate. Which AI services exist in the organisation, who owns each one, and what data is permitted to reach them.
  • The reviewer. Training, calibration, and the honest admission that a reviewer’s attention is a depleting resource.

What Layer 4 cannot stop:

  • Anything in the request path. It does not read a prompt or a response. A prompt injection is Layer 5’s to filter and Layer 6’s to notice; Layer 4 only limits what the injection can cause once it has succeeded.
  • The generation of synthetic media. Detection is a signal about content already delivered. Nothing here prevents a voice from being cloned.
  • Trust exploitation, technically. Chapter 2’s ASI09 row records “not technically detectable,” and the glossary is blunter: no technical signal exists for a trust gradient, so the control is procedural. Layer 4’s answer to ASI09 is a process, not a detector — see User Behavior Analytics for what monitoring can and cannot see here.
The characteristic Layer 4 failure is buying the detectors and skipping the gate

Deepfake detection and shadow AI discovery are products. You can procure them, point them at the estate, and show a dashboard. Human-in-the-loop enforcement is a design decision about your own systems, it slows people down, and nobody sells it to you.

So the layer gets deployed detector-first, which inverts its value. In the Meta case study under Human-in-the-Loop Enforcement below, no detector was relevant and the missing control cost nothing to build: a confirmation step the engineer already believed was there.


Which Layer 4 Controls Are Yours

Chapter 1 Section 3 promised that “the layers stay constant, but which ones are yours is set by the choice you make here.” Layer 4 is the one layer where that choice is not the deployment pattern. A cloud-API consumer and a self-hosting team own identical Layer 4 obligations, because the asset being protected is the person and the device — and those are the same in both.

The axis that does vary Layer 4 ownership is the device and identity estate. Read this table against how your users actually reach AI.

Layer 4 control Managed corporate device BYOD / unmanaged Contractor or partner identity Personal account, corporate work
Human-in-the-loop gates on AI actions You — it lives in your application, not on the device You — unchanged You, and the approver may not be your employee You, if the action touches your systems
Endpoint AI tool inventory You — full visibility Partial at best; treat as unknown Rarely yours; contract for attestation Not yours. Assume it exists
Local model runtime discovery You Partial No No
Endpoint DLP toward AI services You Managed browser or profile only Contractual No technical control exists
Verification procedure for high-value actions You You You — and this is where they are weakest You
Identity lifecycle (joiner/mover/leaver) You You You, and the leaver event is the one that gets missed Outside your directory entirely
AI-literacy and skepticism training You You Usually excluded — a real gap Excluded
Approved AI tool catalog You You — and it is your only lever You You — same
Provenance checking on inbound media You, at the mail/collab gateway Gateway only Gateway only No

Three consequences, and they are the point of the table:

1 · Ownership does not shrink with the deployment pattern — it shrinks with device control. Consuming a managed API removes almost all of Layer 3 from your plate (and Section 5 has the table). It removes none of Layer 4. If anything it adds: the fewer components you run, the larger the share of your remaining risk that is a person pasting something.

2 · The rightmost columns are where real incidents live, and they are the columns with the fewest controls. Arup’s loss ran through a video call and a wire authorisation, neither of which a device control touches. Read down the “Contractor or partner” column: the only entries that are reliably yours are the procedural ones. That is not a coincidence — it is why the procedure has to be the primary control rather than the fallback.

3 · One row is yours in every column. Human-in-the-loop gating lives in your application code, so no device posture, BYOD policy, or third-party contract can take it away from you. It is the only Layer 4 control with that property, which is a strong argument for building it first.


Synthetic Media, Impersonation, and Verification

AI-generated deepfakes — synthetic audio, video, and images that impersonate real people — attack the oldest trust assumption in an organisation: that recognising a face or a voice is evidence of identity.

The attacker’s chain

ATLAS models this as a sequence rather than a single act, which is useful because the defensible link is not the first one:

graph LR
    OSINT["Harvest source material<br/><small>earnings calls, conference<br/>recordings, podcasts,<br/>social video</small>"]
    CAP["<b>AML.T0016.002</b><br/>Obtain Generative AI<br/><small>off-the-shelf or<br/>purpose-built tooling</small>"]
    GEN["<b>AML.T0088</b><br/>Generate Deepfakes<br/><small>cloned voice, composited<br/>face, real-time swap</small>"]
    PH["<b>AML.T0052.001</b><br/>Deepfake-Assisted Phishing<br/><small>vishing call, video<br/>conference, layered with<br/>email and SMS</small>"]
    ACT["Target acts<br/><small>transfers funds, resets<br/>a credential, grants access</small>"]

    OSINT --> CAP --> GEN --> PH --> ACT

    style OSINT fill:#a85800,color:#fff
    style CAP fill:#8b0000,color:#fff
    style GEN fill:#8b0000,color:#fff
    style PH fill:#8b0000,color:#fff
    style ACT fill:#8b0000,color:#fff

You cannot defend the first two links. The source material is public by design — an executive who never appears on a recording is not an executive — and ATLAS notes that a voice can be cloned from a few seconds of it. The tooling is commodity, including projects built specifically for fraud. Every control you own acts on the last two links, and the last one is the only link where a policy, rather than a model, decides the outcome.

Detection, and why it degrades outside the lab

ATLAS carries deepfake detection as a named mitigation, AML.M0034 — apply detection to untrusted or user-provided media, using model-based classification, inconsistency analysis, and biometric cues such as blinking and microexpressions. It is a real control. It is not a reliable one, and the reason is structural rather than a matter of product quality.

Media type What detection looks at Where it breaks
Audio Spectral artifacts, voice biometric comparison, frequency-domain irregularities Cloning from a few seconds of public audio; real-time analysis during a live call is expensive and adds latency to the call it is meant to protect
Video Facial landmark consistency, temporal coherence, lighting and shadow physics, compression artifacts Pre-recorded fakes get unlimited post-processing passes; live face-swap quality has moved faster than detector retraining
Image Generator fingerprints, metadata analysis, pixel-level artifacts Diffusion models leave far weaker fingerprints than the GANs most detectors were trained on; recompression and resizing strip what remains
Biometric liveness Challenge-response, depth and motion cues Injection attacks feed synthetic frames to the camera pipeline directly, bypassing the capture step the check assumes

The generalisation problem underneath all four rows is documented and consistent: detectors that exceed 99% AUROC on a benchmark such as FaceForensics++ lose a large fraction of that — reported at tens of percentage points — when tested against a generator they were not trained on, or a newer version of the same generator, and degrade further under ordinary recompression and resizing. A detector’s benchmark score describes the synthesis pipelines it has seen. An attacker’s choice of tool is not drawn from that set.

Read the consequence, not the accuracy figure

Whatever number a detector reports, an attacker with a current tool gets more attempts than the detector gets retrainings. So detection is a triage signal, not an authorisation decision. A flagged call warrants a callback; an unflagged call does not warrant a wire transfer.

Provenance: Content Credentials and the absence problem

Because detection asks an unanswerable question — was this generated? — the industry moved to an answerable one: where did this come from? The C2PA standard (Coalition for Content Provenance and Authenticity, whose steering committee includes Adobe, Amazon, BBC, Google, Meta, Microsoft, OpenAI, Sony, TikTok, and Truepic) embeds a cryptographically signed Content Credentials manifest recording the capture device or generating tool, whether AI was involved, and each subsequent edit. Specification 2.4 is current as of mid-2026, and 2.1 added redactable assertions and zero-knowledge identity proofs for cases where provenance would otherwise leak the author.

Provenance is a stronger primitive than detection because it is verifiable rather than probabilistic. It also has one property you must teach explicitly, because getting it wrong is worse than not deploying it:

A valid credential proves origin. A missing credential proves nothing.

Most media in circulation carries no manifest — legacy files, screenshots, re-encodes, anything through a platform that strips metadata, and every deliberate fake. If a control treats “no credential” as “suspicious,” it fires on the majority of legitimate content and gets switched off within a week.

Use provenance asymmetrically: a valid manifest from a trusted signer raises confidence and can shorten a verification path. Its absence returns you to the procedure below — which is where you were anyway.

The control that does not depend on being right

The defense that would have stopped Arup is not on the detection table. It is procedural, cheap, and unaffected by which generator the attacker used:

  • Out-of-band confirmation on high-value actions. Wire transfers, credential and MFA resets, privileged access grants, and payment-detail changes require confirmation on a separate channel the requester did not choose — a callback to a number from the directory, not one supplied in the request. Note where this sits relative to the fake: the call can be perfect and the control still holds.
  • Pre-arranged verification phrases for high-stakes verbal instructions, agreed in advance and never transmitted over the channel they authenticate.
  • Authority to pause without penalty. Every one of these frauds runs on urgency and confidentiality. If declining to act immediately on an executive’s request is a career risk, the control does not exist regardless of what the policy says. This is the part that belongs to leadership rather than to security, and Section 11 is where it is developed.
  • Skepticism training (ATLAS AML.M0018 User Training) that demonstrates current cloning quality rather than describing it. Training built on the assumption that fakes look fake teaches a tell that no longer exists.
Defense Connection

Deepfake-enabled impersonation has no OWASP LLM Top 10 category, and that is worth stopping on. It is not an attack on an LLM application — the target is a person, and the model is the attacker’s tool. Its identifier lives in MITRE ATLAS: AML.T0052.001, with AML.M0034 and AML.M0018 as the named mitigations. This is the framework-selection judgement Chapter 2 Section 1 set up and Section 4 applied to adversarial perturbation: when OWASP has no slot for a technique, you are holding the wrong reference work, not looking at an unimportant technique. Filing this under LLM07: Misinformation — which OWASP defines as the model’s own incorrect or misleading output — would send you to review grounding and output validation, neither of which has anything to do with a cloned voice on a phone call.


Endpoint Security for AI-Era Threats

Traditional endpoint security matches malware signatures, blocks known-bad URLs, and watches the filesystem. AI-era threats push it in two directions: defending against AI-enhanced attacks, and seeing what AI tooling the endpoint is running.

Defending against AI-enhanced social engineering

AI-generated phishing lacks the signals a generation of awareness training was built on — broken grammar, awkward register, generic greetings. Worse, it lacks them at scale:

  • Personalised spear-phishing referencing the target’s real projects, colleagues, and last week’s activity, generated per-target rather than per-campaign
  • Voice cloning for vishing, the ATLAS AML.T0052.001 path, using a few seconds of public audio
  • Real-time video impersonation in conference calls, as at Arup
  • Multi-channel sequencing — ATLAS notes deepfake campaigns typically layer email, SMS, and messaging around the call, so each channel corroborates the others

The evolution required is from content-based to intent-based detection: classify the message by what it asks the recipient to do — click an unfamiliar link, open an unexpected attachment, move money through a non-standard path, bypass a normal approval — rather than by how it is written. Fluency is no longer evidence of anything.

Monitoring AI tool usage on endpoints

The other direction is visibility into how the device is used with AI:

  • Browser and extension telemetry — which AI interfaces are reached, and which extensions inject themselves into pages
  • Clipboard and upload monitoring — large code or data blocks moving toward AI service domains
  • Application inventory — desktop apps, browser extensions, and IDE plugins with AI capability
  • Local model runtime detection — this is the one that gets skipped, and Chapter 2 Section 7 exists partly to explain why it cannot be. An SLM running under a local runtime is a full inference stack on the endpoint, and its inference never crosses a network boundary, so every network-side signal in the discovery list below is blind to it. Detect the runtime, not the traffic: the process, the listening port, the weights on disk, the GPU or NPU utilisation with no attributable application.
  • On-device interaction logging — for the same reason, a local deployment has no central log unless you build one. Chapter 2 Section 7 specifies the shape: an on-device audit log that syncs opportunistically, with a defined retention and an alert on gaps in the sync record. A log that only exists while the device is online is a log an attacker can end by going offline.

Shadow AI Discovery and Governance

Shadow AI — AI tools and services used without security review — is shadow IT with a shorter fuse. Shadow SaaS leaks data over months of accumulated use; shadow AI can move a codebase out of your control in a single paste. And the motive is almost never malice: people use these tools because they work.

Discovery, and what each technique misses

Discovery is usually taught as a list of signals. The list matters less than knowing what each one is blind to, because shadow AI’s fastest-growing form is invisible to most of it.

Technique What it finds Blind to
Network traffic analysis Connections to hosted AI service endpoints from the corporate network Anything off-network, on a personal device, or served locally
DNS monitoring Resolution of AI service domains Encrypted DNS (DoH/DoT), which most browsers now default to; and local models, which resolve nothing
CASB / SaaS discovery Tools reached through sanctioned identity or with an org subscription Free tiers on personal accounts — the dominant case
Expense and procurement records Paid subscriptions and API spend Free tiers, personal cards, and anything under an approval threshold
Endpoint agent inventory Installed AI applications, extensions, and IDE plugins Unmanaged and BYOD devices; portable binaries
Local runtime detection Model runtimes, loaded weights, unattributed accelerator load Nothing on a managed device — which is why it is the technique that completes the set

Read the “blind to” column and the design rule falls out: no single signal is a discovery program. A network-only program reports a clean estate for a team that has standardised on local models, and reports it confidently.

graph LR
    DISC["Discovery<br/><small>network, DNS, CASB,<br/>expense, endpoint,<br/>local runtimes</small>"]
    CAT["Catalog &<br/>Assign Owner<br/><small>inventory each service,<br/>classify risk, name a<br/>human accountable</small>"]
    POL["Policy &<br/>Vendor Review<br/><small>acceptable use, data<br/>handling, retention and<br/>training terms</small>"]
    ENF["Enforcement<br/><small>approved catalog,<br/>endpoint DLP, block<br/>the unreviewed</small>"]
    MON["Continuous<br/>Monitoring<br/><small>usage analytics,<br/>drift, compliance</small>"]

    DISC --> CAT --> POL --> ENF --> MON
    MON -.->|"new services,<br/>new runtimes"| DISC

    style DISC fill:#a85800,color:#fff
    style CAT fill:#2d5016,color:#fff
    style POL fill:#2d5016,color:#fff
    style ENF fill:#1565c0,color:#fff
    style MON fill:#2d5016,color:#fff

Governance

Discovery without governance is just surveillance. The program has four parts, and the second is the one most often missing.

Approved AI tool catalog. A vetted list of sanctioned tools with per-tool guidance on what data may reach them. Make the approved path the easy path — if the sanctioned option is slower or less capable than the unsanctioned one, people keep using the unsanctioned one, and the catalog becomes a document rather than a control.

A named owner per AI service and per agent. This is the Layer 4 half of ASI10: Rogue Agents. Chapter 2 Section 5 makes the point that ASI10 needs no attacker at all, and Section 5’s mapping hands Layer 4 “governance and ownership” specifically. An agent with credentials, tools, and no accountable human is a rogue agent that has not gone wrong yet. Record, per agent: who owns it, what it may do, what it may reach, when its authorisation is next reviewed, and who can turn it off.

Acceptable use policy stating what may not be sent to an AI tool — PII, credentials and tokens, proprietary source code, unreleased financials, legal communications, customer data — expressed per approved tool rather than as one global rule, because the answer legitimately differs between a contained enterprise deployment and a consumer chatbot. ATLAS carries this as AML.M0021 Generative AI Guidelines.

Vendor assessment, which is where Layer 1 explicitly hands you work. When you consume a finished assistant, the corpus, the memory store, and the retention policy are product behaviour — not architecture you can change. The controls become contractual and configurational, and they are Layer 4’s because the decision being made is which tool people are allowed to use. Four questions carry most of the weight: does the vendor train on your submissions by default, what is the retention period and can you set it, what is the tenancy boundary for any memory or cache feature, and what happens to your data when you stop paying.

Endpoint DLP toward AI services. Content inspection on data leaving the device for an AI endpoint — API keys and tokens, PII patterns, proprietary code signatures, document classification markings.

Where the Layer 4 / Layer 5 seam sits

Section 2 records Layer 4 as acting “on people and devices, not on traffic,” and DLP plainly inspects traffic — so be precise about which traffic. Layer 4’s DLP acts at the device, on data leaving for a service you do not run. Layer 5 acts in the request path of AI services you do run. Different position, different control owner, and the distinction decides where a finding goes: data heading to a consumer chatbot is a Layer 4 problem, and data coming back out of your own assistant is a Layer 5 problem.

Samsung, 2023 – and why it is not a Layer 4 win

Between 11 and 30 March 2023 — within 20 days of Samsung authorising ChatGPT use in its Device Solutions division — engineers submitted semiconductor measurement source code, defect-detection code, and a confidential meeting transcript to ChatGPT. Under OpenAI’s default consumer policy at the time, submissions were eligible for use in training; the opt-out arrived the following month. State the exposure precisely, as Chapter 2 Section 6 does: there is no public evidence the data ever surfaced in another user’s output, and Samsung could neither verify nor retract it. The control failure is the irreversibility, not a confirmed leak.

Chapter 2 assigns the two controls that would have caught it to Layer 1 (classification) and Layer 5 (egress inspection), because both act on traffic rather than on intent. That assignment stands. What Layer 4 changes is different and worth being honest about: it lowers the probability the paste happens at all — an approved contained alternative, a per-tool data rule, and a device-side inspection point — and it makes the exposure known rather than discovered by a newspaper.

Samsung’s own first response is the instructive part. Before the May 2023 ban, they imposed a 1,024-byte cap on prompt length — a control that limits the size of a leak without touching its nature. A kilobyte is ample for a credential, a schema, or the one architectural decision that mattered.


Human-in-the-Loop Enforcement

This is the control that makes Layer 4 the layer that bounds consequence, and it is the one the rest of the course keeps pointing here for. Section 2’s card lists it among Layer 4’s key controls; Section 5 hands Layer 4 “human gate on irreversible actions” for LLM03: Excessive Agency; Chapter 2 Section 6 hands Layer 4 “human-in-the-loop enforcement” for automation bias; Section 11 maps it to the EU AI Act’s human-oversight requirement. ATLAS names it AML.M0029 Human In-the-Loop for AI Agent Actions.

Gate the irreversible subset, not everything

The design rule is a single sentence, and the glossary states the failure mode it protects against: a gate a reviewer clicks through fifty times a day stops functioning as a gate. Approval fatigue is not a training problem to be solved with reminders; it is the predictable result of asking for attention more often than attention exists. A gate on everything is a gate on nothing.

So the work is classification, done once per action an AI system can take:

Ask If yes Where the gate goes
Can the effect be undone by us, within minutes, without asking anyone? Let it run. Log it. No gate
Does it move money, change access, or contact a third party? Gate it, always Before execution, with the full action shown
Does it write to a system of record, or delete? Gate it, or make it staged and reversible Before commit, or convert to a draft
Is it externally visible and unretractable — a post, an email, a filing? Gate it Before send
Does it grant the system new capability — a credential, a tool, a scope? Gate it, and review it separately Before grant, and at review

Chapter 1 Section 7 supplies the reason this belongs to whoever designs the agent rather than to whoever operates it: excessive autonomy is one of the three separate causes of excessive agency, and it is fixed at design time. Chapter 2 Section 5 states the same rule as the primary control for rogue agents: identify the irreversible subset of its actions and put a gate there.

Three ways a gate that exists still fails

  • It was never implemented, only expected. The Meta case below is exactly this, and it is the most common form.
  • It presents a decision the reviewer cannot make. “Approve this action?” with no diff, no target, and no stated effect is a button, not a gate. Show what will change, on what, and what happens if it is wrong.
  • The reviewer is the same kind of thing as the actor. Chapter 2 Section 5’s cascade makes the point: a reviewing agent built on the same model, reading the actor’s output as authoritative, is one control wearing two hats. A gate adds assurance only when it can fail differently.

Confidence calibration

The other half of Layer 4’s answer to over-trust is presentational, and it is a control on your own product rather than on the user. AI output arrives fluent, formatted, and uniformly confident whether or not it is grounded — which strips the reader of the signal they would use on a human colleague. Surface, in the interface: whether the answer is grounded in retrieved sources or generated unsupported, the sources themselves and not merely their existence, and where the model’s own uncertainty was high. Chapter 2 Section 6 records the finding that makes this necessary rather than cosmetic: a citation raises reader confidence without raising accuracy. A citation the reader cannot open is decoration with a cost.

And label AI-generated content as AI-generated inside your own systems. A summary that circulates without provenance becomes a fact three forwards later.

Defense Perspective: The Meta Agent Post (March 2026)

The incident (from Chapter 2 Section 6): An engineer posted a technical question to an internal forum. A second engineer routed it to an internal agentic AI system. The agent posted its own reply to the thread — without review, though the engineer had expected a human-in-the-loop confirmation step. The advice was wrong. A colleague implemented it, broadening access permissions to sensitive company and user data for roughly two hours. Meta classified it SEV1, its second-highest internal severity, and confirmed no external party accessed the data.

Chapter 2 calls this the most important case study in the section, because it chains three OWASP categories with no attacker anywhere in it: LLM07 (confidently wrong guidance), LLM03 (published autonomously), ASI09 (implemented on trust).

What Layer 4 owned, and in which order:

  1. The gate that did not exist. Publishing to a shared forum is externally visible and unretractable — row four of the table above. The engineer’s expectation was correct about what the control should have been; it simply was not built. Remove this link and there is no incident, at a cost of one confirmation dialog.
  2. Confidence calibration. The reply arrived with the formatting and assurance of a verified answer, carrying nothing to mark it as generated or ungrounded. The colleague who acted on it had no signal to weigh.
  3. AI content labeling. An agent’s post in a human thread, indistinguishable from a human post, inherits the thread’s credibility without having earned it.
  4. Verification requirements for consequential change. Broadening access permissions is a Layer 4-gated action in its own right, independently of where the advice came from — the second place the chain could have been broken.

Note what is not on this list: no detector was relevant, no deepfake, no shadow AI, no anomalous usage pattern for analytics to catch. The gap was between the expected control and the implemented control — which is why “does the gate exist, or does everyone merely believe it does?” is a question worth asking about your own systems this week.


Identity and the Human Side of Privilege Abuse

ASI03: Identity and Privilege Abuse is one of Layer 4’s four OWASP categories, and Section 5 says why it is shared: Layer 3 owns the technical identity — scoping, rotation, per-function service accounts — and cannot prevent abuse of a permission that was correctly granted. What remains is a human process:

  • Joiner / mover / leaver, applied to AI access. The mover event is where AI entitlements accumulate: a role change adds new access and rarely removes the old, and an agent acting on a user’s behalf inherits whatever that user has accreted. The leaver event is where they persist — a departed employee’s personal API key funding an agent that is still running is a Chapter 2 rogue-agent scenario with a payroll record attached.
  • Periodic human review of agent authorisation. Not “does this agent work,” but: is this scope still the minimum, is the owner still here, and is the gate still on the actions that need one. Pair it with Layer 3’s credential inventory — the inventory tells you what exists, the review decides whether it should.
  • User-derived credentials over service identities where the choice exists. Section 5’s confused-deputy section covers the mechanism. The Layer 4 consequence is auditability: an action attributable to a person can be reviewed by a person.

User Behavior Analytics for AI Interactions

Monitoring how users interact with AI systems surfaces two things well and one thing not at all. Be precise about which, because the third is the one people expect it to deliver.

What to monitor

Signal Normal pattern Anomalous pattern Likely indicator
Query volume Consistent daily usage in a stable band Sudden spike, especially outside working hours Automated extraction; compromised account
Data sensitivity General and role-appropriate questions Queries targeting named customers, financials, credentials Exfiltration attempt; insider risk
Copy/paste and upload volume Occasional small snippets Large code or data blocks, repeatedly, toward AI domains Bulk exfiltration through a chatbot
Output destinations Screen, or authorised storage Output forwarded to personal storage, external mail, unmanaged services Leakage through an AI intermediary
Prompt patterns Task-related, consistent with the role Injection probes, system-prompt extraction attempts Vulnerability testing, or a compromised account
Approval behaviour Gates are read, and sometimes rejected Approval latency collapsing; rejection rate near zero across hundreds of gates Approval fatigue — the gate has stopped being a control

The last row is the one worth adding to a program that does not have it. It monitors your own defense rather than your users, and it is the only listed signal that detects a Layer 4 control silently failing.

What UBA cannot see

Analytics does not detect trust exploitation, and Chapter 2 says so explicitly

It is tempting to file ASI09: Human-Agent Trust Exploitation under behavioural monitoring — anomalous interaction patterns, detected. Chapter 2 Section 5’s comparison table records ASI09’s runtime signal as “not technically detectable,” and the glossary’s trust-gradient entry is unambiguous: no technical signal exists for it, so the control is procedural.

The mechanism is why. In a trust-exploitation incident nothing anomalous happens. The agent behaves as it always has, the volumes are normal, the queries are role-appropriate, and the human approves as they have approved a hundred times before. The change is in the reviewer’s attention, and attention emits no telemetry. Chapter 1 Section 7 carries the measurement: developers using AI tooling were about 19% slower while believing they were 20% faster — a 39-point error in the direction of trusting the tool, which means practitioners are poor judges of how much verification an output needs.

Layer 4’s real answer to ASI09 is the three controls above it in this section — a gate on the irreversible subset, confidence calibration so the reviewer has something to weigh, and a rotation or sampling scheme so verification does not depend on one person’s sustained vigilance. UBA’s contribution is the approval-fatigue row: it cannot see the trust gradient, but it can see the gate going quiet.


Layer 4 → OWASP Mapping

Layer 4 is named as a primary defense in four of the twenty categories across the LLM Top 10 and the Agentic AI Top 10 — the narrowest coverage in the Blueprint, and, as Section 2 insists in the other direction, coverage is not stopping power.

Category The route Layer 4 acts on What Layer 4 does What it cannot do Completed by
LLM07: Misinformation A person or a build acting on confidently wrong output Verification requirements before consequential action, dependency resolution against a verified allowlist, confidence calibration, AI content labeling Prevent the model from generating it, or judge a specific answer’s truth L5 — grounding checks and output validation
WarningASI03: Identity and Privilege Abuse The human lifecycle behind every AI credential Joiner/mover/leaver applied to AI access, periodic human review of agent authorisation, attributable user-derived identity Scope, rotate, or technically constrain a credential L3 — the technical identity and its scope
WarningASI09: Human-Agent Trust Exploitation The reviewer’s degrading attention over a track record Gate the irreversible subset, calibrate confidence, rotate or sample verification, monitor approval behaviour Detect it. Nothing anomalous happens at the moment of exploitation Nothing else — this category has no other layer, which is why the process is the control
WarningASI10: Rogue Agents Agents nobody registered, and agents nobody owns Shadow AI discovery including local runtimes, a named owner and review date per agent, the gate that bounds an unintended action Judge whether a discovered agent’s behaviour is malicious L3 — credential and endpoint inventory; L6 — behavioural baselines

Two things to read off the table. First, the fourth column: three of the four rows say Layer 4 cannot detect its own category — which is the correct reading of a human-boundary layer and the reason its controls are procedural. Second, the ASI09 row has no “completed by” entry. It is the only category in the course with a single owning layer, so a Layer 4 that has been reduced to a detector leaves it uncovered entirely.

And the section’s largest threat is absent from the table by design: deepfake-enabled impersonation has no OWASP LLM slot at all, and lives in ATLAS as AML.T0052.001.


AI Scanner Cross-Reference

AI Scanner contributes to Layer 4 indirectly, by testing the properties that make human over-trust exploitable: whether a model produces confidently wrong output, how it behaves when pushed outside its knowledge, and whether it can be manipulated into generating content that would be presented to a user as authoritative. Those are pre-deployment findings about a model, and they tell you how much verification the humans downstream will need — which is a Layer 4 design input rather than a Layer 4 control. Section 9 covers the full scan-protect-validate-improve loop.

TrendAI Vision One’s Endpoint Security component includes deepfake detection for inbound communications, and its Email Security layer classifies AI-generated phishing on behavioural indicators rather than content patterns — both of which map to the triage-signal role above, not to an authorisation decision. Vision One’s discovery capability inventories unsanctioned AI services across the network and supports DLP policy toward them. Read that against the discovery blind-spot table: a network- and identity-side console is strong on hosted services and weak on a model running locally, so pair it with endpoint runtime detection rather than treating the console as the whole program. And note what no console supplies at all — the approval gate, the verification procedure, and the owner named against each agent. Those are decisions, and buying the platform does not make them.


Layer 4 User Protection Checklist

Run the ownership table before this checklist — a finding written against a device you do not manage is not a finding.

The gate (start here — it is yours in every column)

  • Classify every AI-initiated action as reversible or not, and gate the irreversible subset — money, access changes, external contact, writes to systems of record, new capability grants
  • Verify the gate exists in code, not only in expectation. Ask the team that built it to show you the branch that blocks
  • Make each gate decidable — show the action, its target, and its effect, not just an approve button
  • Ensure at least one gate in any agent chain is a different kind of check than another agent
  • Monitor approval behaviour — collapsing latency and a near-zero rejection rate mean the gate has stopped working

Synthetic media and impersonation

  • Require out-of-band confirmation for high-value actions, on a channel the requester did not supply
  • Establish verification phrases for high-stakes verbal instruction, never sent over the channel they authenticate
  • State explicitly that pausing to verify is never penalised, including for executive requests
  • Deploy detection as triage on high-risk inbound media, with a defined action for a flag and no assumption that a clean result authorises anything
  • Check Content Credentials asymmetrically — a valid manifest raises confidence; a missing one changes nothing
  • Train on current cloning quality, demonstrated rather than described

Endpoint and shadow AI

  • Use more than one discovery signal, and include local model runtime detection — process, port, weights on disk, unattributed accelerator load
  • Build on-device interaction logging for local deployments, with retention and a sync-gap alert
  • Maintain an approved AI tool catalog that is genuinely easier to use than the alternatives, with per-tool data rules
  • Name an owner and a review date for every AI service and every agent
  • Complete a vendor assessment per approved tool — default training use, retention and whether you can set it, tenancy of memory and cache features, exit terms
  • Deploy endpoint DLP toward AI services for credentials, PII, code signatures, and classification markings
  • Inventory AI applications, extensions, and IDE plugins across the managed fleet

Identity and the reviewer

  • Apply joiner/mover/leaver to AI access, with the mover event explicitly removing what the old role needed
  • Review agent authorisation periodically — scope, owner, and gates, not just whether it still runs
  • Calibrate confidence in your own interfaces — grounded versus generated, openable sources, flagged uncertainty
  • Label AI-generated content as AI-generated inside your own systems
Key Takeaways
  • Layer 4 touches no traffic and stops no request. What it does is bound what a wrong or leaked answer can cause — which makes it the only layer that still works after the other five have missed
  • The gate is the layer’s highest-value control and the one that is usually absent, because detectors are products and approval boundaries are design decisions nobody sells you. Gate the irreversible subset: a gate on everything is a gate on nothing
  • Detection is a triage signal, not an authorisation decision. A detector’s benchmark score describes the generators it has seen; the attacker’s tool is not drawn from that set. Provenance is the stronger primitive, used asymmetrically — a valid Content Credential raises confidence, a missing one proves nothing
  • The control that would have stopped Arup is procedural and free: out-of-band confirmation on a channel the requester did not choose, and permission to pause without penalty
  • No single signal is a shadow AI program. Every network-side technique is blind to a local model runtime, whose inference never crosses a network boundary — so detect the runtime, not the traffic
  • Analytics cannot see trust exploitation. In an ASI09 incident nothing anomalous happens; the change is in the reviewer’s attention, which emits no telemetry. ASI09 is the only category in the course with a single owning layer, so a Layer 4 reduced to dashboards leaves it uncovered
  • Deepfake impersonation has no OWASP LLM category at all — its identifier is MITRE ATLAS AML.T0052.001. Filing it under LLM07 would send you to fix grounding, which has nothing to do with a cloned voice

Test Your Knowledge

Ready to test your understanding of AI user security? Head to the quiz to check your knowledge.


Up next

Layer 4 bounds what a user or an agent can do with an answer. The next layer sits in the path the answer travels. In Section 7 you’ll meet Layer 5: Secure Access to AI Services – AI Gateway architecture, Zero Trust Secure Access (ZTSA), prompt and response filtering, and rate limiting. It is the only layer in the Blueprint that can stop a live request, and the only one that sees every prompt and every response.