1. Integrating Security into AI Architectures
Introduction
Every attack you studied in Chapter 2 shares a common thread: they exploit gaps where security wasn’t designed in from the start. Prompt injection succeeds when input validation is an afterthought. Data poisoning persists when training pipelines lack integrity checks. Model theft goes undetected when monitoring is bolted on rather than built in. The attacks are varied, but the root cause is the same – security was treated as a feature to add later, not a principle to build around.
This section bridges the gap between knowing the attacks and building the defenses. You’ll learn how traditional DevSecOps extends to AI systems, how to threat-model an LLM deployment, and how a shared responsibility model divides defense duties between AI providers and enterprise security teams. By the end, you’ll have the architectural mindset needed for the Blueprint framework introduced in Section 2 and the six layers that follow in Sections 3-8.
What will I get out of this?
By the end of this section, you will be able to:
- Extend DevSecOps to AI systems by identifying where AI-specific security checkpoints fit into the development lifecycle.
- Apply threat modeling to an LLM deployment using STRIDE adapted for AI – and choose a different framework when STRIDE stops fitting, which it does as soon as the system becomes agentic.
- Locate the responsibility boundary for a given deployment and predict how it moves between hosted, self-hosted and on-device – including the responsibilities no provider ever held.
- Map Chapter 2’s attack domains to Blueprint layers, and recognise the two cases the mapping hides: attacks that need several layers, and attacks no layer defends.
DevSecOps for AI
Traditional DevSecOps integrates security into every phase of the software development lifecycle – planning, coding, building, testing, deploying, and monitoring. AI systems follow the same principle but add stages that traditional software doesn’t have: data curation, model training, fine-tuning, and inference monitoring. Each of these stages requires its own security checkpoints.
The AI-Adapted DevSecOps Pipeline
graph LR
P["Plan<br/><small>Threat Model</small>"]
D["Develop<br/><small>Secure Coding</small>"]
T["Train<br/><small>Data Integrity</small>"]
V["Validate<br/><small>Red Team Testing</small>"]
De["Deploy<br/><small>Container Scanning</small>"]
M["Monitor<br/><small>Runtime Protection</small>"]
F["Feedback<br/><small>Incident Response</small>"]
P --> D --> T --> V --> De --> M --> F
F -.->|"Continuous<br/>Improvement"| P
SP["Security Checkpoint:<br/>AI Threat Model Review"]
SD["Security Checkpoint:<br/>Prompt Hardening, Dependency Audit"]
ST["Security Checkpoint:<br/>Data Lineage, Poisoning Detection"]
SV["Security Checkpoint:<br/>AI Scanner Assessment,<br/>Adversarial Testing"]
SDe["Security Checkpoint:<br/>Model Integrity Verification,<br/>AI-SPM Posture Check"]
SM["Security Checkpoint:<br/>AI Guard Runtime Filters,<br/>Anomaly Detection"]
SP -.-> P
SD -.-> D
ST -.-> T
SV -.-> V
SDe -.-> De
SM -.-> M
style P fill:#2d5016,color:#fff
style D fill:#2d5016,color:#fff
style T fill:#2d5016,color:#fff
style V fill:#2d5016,color:#fff
style De fill:#2d5016,color:#fff
style M fill:#2d5016,color:#fff
style F fill:#2d5016,color:#fff
style SP fill:#1565c0,color:#fff
style SD fill:#1565c0,color:#fff
style ST fill:#1565c0,color:#fff
style SV fill:#1565c0,color:#fff
style SDe fill:#1565c0,color:#fff
style SM fill:#1565c0,color:#fff
The key difference from traditional DevSecOps: the Train and Validate stages are entirely new, and the Monitor stage gains AI-specific responsibilities (prompt/response filtering, behavioral anomaly detection, cost monitoring).
What Changes for AI?
| DevSecOps Stage | Traditional Focus | AI-Specific Additions |
|---|---|---|
| Plan | Threat modeling, architecture review | AI-specific threat categories (OWASP LLM Top 10), data supply chain risk assessment |
| Develop | Secure coding, dependency scanning | Prompt hardening, system prompt design, guardrail configuration |
| Train | (not applicable) | Data lineage tracking, poisoning detection, bias evaluation, training data access controls |
| Validate | Penetration testing, code review | AI red-teaming, adversarial input testing, prompt injection testing (AI Scanner) |
| Deploy | Container scanning, config management | Model integrity verification, AI-SPM posture checks, model artifact signing |
| Monitor | SIEM, intrusion detection | Runtime prompt/response filtering (AI Guard), cost anomaly detection, drift monitoring |
| Feedback | Incident response, patch management | AI incident playbooks, model retraining triggers, guard rule updates |
Defense Connection
The DevSecOps testing phase directly addresses the prompt injection and jailbreaking techniques covered in Chapter 2 by integrating AI-specific security testing before deployment. Rather than discovering prompt injection vulnerabilities in production, organizations catch them during the Validate stage through AI Scanner assessment and adversarial testing.
AI Threat Modeling
Threat modeling for AI systems builds on established methodologies but adapts them for the unique attack surfaces of LLM deployments. The goal is the same as traditional threat modeling: systematically identify what can go wrong, assess the risk, and design mitigations before deployment.
STRIDE Adapted for AI
STRIDE (Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege) is one of the most widely used threat modeling frameworks. Here’s how each category applies to AI systems:
| STRIDE Category | Traditional Application | AI-Specific Application |
|---|---|---|
| Spoofing | Impersonating users or services | Agent identity spoofing, deepfake-generated credentials, impersonating legitimate MCP servers |
| Tampering | Modifying data in transit or at rest | Training data poisoning, RAG corpus manipulation, model weight tampering, memory/context poisoning |
| Repudiation | Denying actions were performed | AI-generated content attribution, agent action audit gaps, missing tool call logs |
| Information Disclosure | Unauthorized data access | System prompt leaking, training data extraction, PII in model responses, embedding inversion attacks |
| Denial of Service | Resource exhaustion | Unbounded consumption, context window stuffing, GPU exhaustion, API cost amplification |
| Elevation of Privilege | Gaining unauthorized access | Prompt injection escalating to tool use, agent privilege abuse, excessive agency exploitation |
Where STRIDE Runs Out, and What to Use Instead
STRIDE earns its place because it is already in your organization’s process and it forces coverage — six categories you cannot skip. But it was designed for systems with predictable logic and fixed trust boundaries, and Chapter 2 Section 5 described a system that has neither. An agent is not one entity in a data flow diagram. It is a caller, a callee and a data store at once, its output becomes its own next input, and the boundary you drew at design time moves whenever it acquires a tool.
The Cloud Security Alliance published MAESTRO in February 2025 for exactly this gap. It replaces STRIDE’s threat categories with seven layers to reason across – foundation model, data operations, agent frameworks, deployment infrastructure, evaluation, orchestration, and the wider ecosystem – so multi-agent interaction and cross-layer effects are in scope by construction rather than by the modeller remembering to look.
Which one to reach for:
| Your system | Framework | Why |
|---|---|---|
| LLM behind an API, no tool use | STRIDE adapted for AI (table above) | Trust boundaries are static and the six categories cover them |
| Adds RAG, memory, or fine-tuning | STRIDE, with the data layer modelled explicitly | Still tractable, but retrieval and memory are inputs, not storage |
| Agents, tool use, or multiple agents | MAESTRO, with STRIDE inside a layer | Boundaries change at runtime; per-entity analysis stops working |
| Any of the above, for coverage assurance | MITRE ATLAS alongside | Neither framework enumerates techniques; ATLAS does, and it is what a detection team will ask for |
The pairing matters more than the choice. A threat model tells you what could go wrong; ATLAS tells you what has actually been done. Running one without the other produces either a tidy diagram with no detections behind it, or a technique list nobody can trace to a design decision.
Threat Model for an LLM Deployment
The diagram below shows a threat model for a typical enterprise LLM deployment. Each numbered threat maps to an OWASP category, and each defense control maps to a Blueprint layer you’ll learn in subsequent sections.
graph TB
subgraph "External"
U["End Users"]
A["Attackers"]
end
subgraph "Access Layer"
GW["API Gateway /<br/>AI Gateway"]
PF["Prompt Filter"]
RF["Response Filter"]
end
subgraph "Application Layer"
LLM["LLM Engine"]
RAG["RAG Pipeline"]
TOOLS["Tool / Agent<br/>Executor"]
end
subgraph "Data Layer"
VS["Vector Store"]
TD["Training Data"]
MEM["Conversation<br/>Memory"]
end
U -->|"Prompts"| GW
A -->|"1. Prompt Injection<br/>(LLM01)"| GW
GW --> PF --> LLM
LLM --> RF --> GW
LLM -->|"Retrieval queries"| RAG
RAG --> VS
LLM -->|"Tool calls"| TOOLS
A -.->|"2. Data Poisoning<br/>(LLM05)"| TD
A -.->|"3. RAG Poisoning<br/>(LLM09)"| VS
A -.->|"4. Memory Poisoning<br/>(ASI06)"| MEM
TD -.-> LLM
MEM --> LLM
style A fill:#8b0000,color:#fff
style U fill:#2d5016,color:#fff
style GW fill:#1565c0,color:#fff
style PF fill:#1565c0,color:#fff
style RF fill:#1565c0,color:#fff
style LLM fill:#2d5016,color:#fff
style RAG fill:#2d5016,color:#fff
style TOOLS fill:#a85800,color:#fff
style VS fill:#a85800,color:#fff
style TD fill:#a85800,color:#fff
style MEM fill:#a85800,color:#fff
The numbered threats in this diagram correspond to the attack categories you studied in Chapter 2. The blue-shaded defense controls (API Gateway, Prompt Filter, Response Filter) correspond to Blueprint Layer 5 (Secure Access to AI Services), which you’ll study in detail in Section 7. The orange-shaded components (Tool Executor, Vector Store, Training Data, Memory) are the at-risk assets that multiple Blueprint layers protect.
Defense Connection
This threat model directly maps to the AI lifecycle attack surface you studied in Chapter 2 Section 1. Where that diagram showed attack vectors at each lifecycle stage, this diagram shows where defense controls intercept those vectors. The two diagrams are complementary – attack surface mapping drives threat model design.
The Shared Responsibility Model for AI Security
Cloud computing established the shared responsibility model: the provider secures the infrastructure, the customer secures their data and applications. AI systems follow a similar pattern, but with additional layers that reflect the unique components of AI deployments.
Who Secures What?
| Responsibility Area | AI Provider (e.g., OpenAI, Anthropic) | Enterprise Security Team |
|---|---|---|
| Base model safety | Model alignment, safety training, content policies | Validate model safety for your use case, monitor for drift |
| API infrastructure | API availability, DDoS protection, encryption in transit | API key management, rate limiting, cost controls |
| Model weights | Protect proprietary weights from extraction | Protect fine-tuned weights, adapter security, model artifact integrity |
| Training data | Curate pre-training data, remove PII at scale | Curate fine-tuning data, RAG corpora integrity, data classification |
| Prompt/response filtering | Basic content filtering in hosted APIs | Custom guardrails, domain-specific filtering, AI Guard deployment |
| Tool and agent security | (typically not provided) | Tool allowlisting, permission scoping, execution monitoring, agent guardrails |
| User authentication | API authentication mechanisms | User identity management, RBAC for AI features, session controls |
| Compliance and auditing | SOC 2, data processing agreements | Industry-specific compliance (HIPAA, PCI-DSS), audit logs, AI governance |
The critical gap: most of the security responsibility for production AI deployments falls on the enterprise, not the AI provider. Providers offer safe base models and basic API infrastructure, but everything from prompt hardening to agent security to data classification is the enterprise’s responsibility.
The Boundary Moves With the Deployment Model
The table above describes a hosted API. That boundary is not fixed – it slides, and the direction is always the same. Knowing where it sits for a given system is the single most useful thing this section gives you, because it determines which of the six Blueprint layers you are obliged to staff.
| Hosted API | Self-hosted | On-device | |
|---|---|---|---|
| Base model safety | Provider aligns it | You validate it | You validate it, and it can be fine-tuned away by whoever holds the device |
| Model weights | Provider protects them | You protect them | Published. Nobody protects them |
| Rate limiting / spend control | Shared – provider meters, you cap | Yours, at the serving layer | No equivalent exists |
| Prompt and response filtering | Provider gives a floor, you add depth | Entirely yours | Whatever fits on the device |
| Interaction logging | Provider logs the API, you log the app | Entirely yours | Only if you build it and it syncs |
| Tool and agent security | Never the provider’s | Yours | Yours |
Two rows deserve a second look, because they are where the model breaks rather than merely shifts.
“Base model safety” is not a responsibility you can hold once. Chapter 2 Section 7 showed that alignment does not survive the journey from a provider’s release to a deployed artifact: distillation, community fine-tuning and quantization each change safety behaviour without changing capability benchmarks. A provider’s safety claim covers the artifact they shipped. If you distilled, fine-tuned, quantized or downloaded a variant, the claim does not transfer – and neither does the responsibility.
“Tool and agent security” was never anyone’s but yours, at every deployment model. It is the only row with no provider column at all, and it is the row that Chapter 2 Section 5 demonstrated does the most damage. Read the two facts together: the highest-consequence attack surface in the course is the one no vendor has ever offered to cover.
The self-hosted column is where the infrastructure attacks from Chapter 2 Section 4 land – serialization exploits, adversarial inputs and denial of service all target components a hosting provider was previously absorbing.
The question to ask, in this order
1. Where does the model execute, and who owns that hardware? 2. What was done to the model after the provider released it? 3. Can it call tools? The first answer tells you which layers you must staff, the second tells you which vendor assurances still apply, and the third tells you the blast radius when something gets through.
From Attack Taxonomy to Defense Strategy
Chapter 2 gave you the attack vocabulary – the OWASP LLM Top 10, the Agentic AI Top 10, and the MITRE ATLAS framework. This chapter translates that vocabulary into defense architecture.
The mapping follows a simple principle: every attack category implies a defense requirement, and every defense requirement maps to one or more Blueprint layers.
| Attack Domain (Chapter 2) | Defense Approach (Chapter 3) | Primary Blueprint Layer(s) |
|---|---|---|
| Prompt-level attacks (S2) | Input/output filtering, prompt hardening | L5 Access, L6 Zero-Day |
| Data and training attacks (S3) | Data integrity, DSPM, supply chain scanning | L1 Data, L2 Models |
| Model and infrastructure attacks (S4) | Container security, posture management | L2 Models, L3 Infrastructure |
| Agentic attack vectors (S5) | Tool controls, execution monitoring, ZTSA | L5 Access, L3 Infrastructure, L2 Models |
| Output and trust exploitation (S6) | Response filtering, human-in-the-loop | L4 Users, L5 Access |
| SLM threats (S7) | Endpoint governance, model provenance, edge posture | L3 Infrastructure, L4 Users, L2 Models |
Read the mapping as “where to start”, not “where it ends”
Two habits this table should not encourage.
A domain rarely lands on one layer, and the layer that owns the control is often not the one that owns the asset. Agentic attacks are the clearest case: Section 5 established that supply-chain compromise of an MCP server is defended at Layer 2 (what you install) and Layer 3 (what it can reach), while the tool call itself is constrained at Layer 5. Fixing only the layer where the attack was observed is how organizations end up monitoring tool calls without ever reviewing what tools got installed.
Some attacks have no defending layer, and the table will not tell you. Section 7 is the example: once model weights ship to a device you do not control, no layer recovers them. That is a constraint on the architecture rather than a control to select — and a mapping table is exactly the artifact that hides it, because every row has an entry.
The next section introduces the Blueprint framework in full detail. You’ll see how these six layers stack together into a defense-in-depth architecture, with TrendAI Vision One providing unified visibility across all layers.
Defense Connection
The attack-to-defense mapping above connects every section from Chapter 2 to a specific defense strategy. As you progress through the Blueprint layers in Sections 3-8, you’ll see these mappings become concrete controls with real product capabilities behind them.
Defense Perspective: Microsoft v. Storm-2139
The attack (from Chapter 2 Section 1): a cybercrime network harvested exposed Azure OpenAI API keys from public sources, fronted them behind reverse-proxy software, and resold access to generate content the provider’s policies prohibited. Victims found out when the charges appeared on their bills. Note what Section 1 established about the numbers: the widely quoted “$100,000 per day” is a modelled ceiling from Sysdig’s research on the technique, not a damages figure for this case – use it to reason about exposure, not to quantify an incident.
What integrated security architecture would have caught:
An organization with AI-adapted DevSecOps would have had multiple detection points for this attack:
- Plan phase: Threat modeling would have identified API credential theft as a high-risk scenario, leading to credential rotation policies and monitoring requirements.
- Deploy phase: AI-SPM posture management would have flagged API keys stored in public repositories or without rotation policies (a common credential harvesting vector).
- Monitor phase: Cost anomaly detection would have triggered alerts when consumption patterns deviated from baseline – catching the unauthorized usage within hours, not days.
- Feedback phase: Incident response playbooks with automated credential revocation would have contained the damage immediately.
The key insight: no single control stops this attack. It’s the integration of security across the full lifecycle – posture management (Layer 3), access controls (Layer 5), and anomaly detection (Layer 6) working together – that turns a catastrophic breach into a contained incident.
Key Takeaways
- AI security requires integrating security checkpoints into every phase of the AI-adapted DevSecOps lifecycle, including the new Train and Validate stages
- STRIDE fits an LLM behind an API and stops fitting once the system is agentic, because an agent is caller, callee and data store at once and its boundaries move at runtime. Use MAESTRO’s seven layers for agentic systems, and pair either framework with MITRE ATLAS – a threat model says what could go wrong, ATLAS says what has actually been done
- The responsibility boundary slides with the deployment model, always in the same direction. Establish where the model executes, what was done to it after release, and whether it can call tools – in that order
- A provider’s safety claim covers the artifact they shipped. Distillation, fine-tuning and quantization each void it without changing benchmark scores, and tool security was never the provider’s at any deployment model
- Attack domains map to Blueprint layers, but the mapping hides two cases: attacks needing several layers at once, and attacks no layer defends – which are architecture constraints, not controls to select
Test Your Knowledge
Ready to test your understanding of AI security architecture? Head to the quiz to check your knowledge.
Up next
Now that you understand how security integrates into AI architectures, it’s time to see the defense framework that organizes it all. In the next section, you’ll explore the Security for AI Blueprint – a 6-layer defense architecture that maps every control to a specific protection domain.