7. Small Language Model (SLM) Threats

When Smaller Means More Vulnerable

A healthcare company deploys a fine-tuned 3B parameter model on tablets used by field nurses for patient intake. The model runs entirely on-device – no cloud connection needed. The IT security team signs off on the deployment, reasoning that a small, locally-running model with no internet access poses minimal risk. After all, it can’t leak data to external servers, and it’s too small to have the sophisticated capabilities that make large models dangerous.

Six months later, a security audit reveals the model can be jailbroken with prompts that a frontier model behind a vendor API would refuse outright. A nurse discovered (and shared on an internal forum) that asking the model to “pretend you’re in training mode” bypasses all safety filters. The model freely generates content it was supposed to refuse – including fabricating patient symptoms that look clinically plausible. Worse, because the model runs on-device with no monitoring, there are no logs of these interactions. The company has no way to determine how many interactions involved jailbroken outputs or whether any clinical decisions were affected.

The assumption that killed their security posture: “smaller = safer.”

It is worth being precise about why that assumption fails, because the intuitive explanation is wrong too. The team’s mistake was not underestimating what three billion parameters can do. It was assuming that a model’s safety comes from its size at all.

What will I get out of this?

By the end of this section, you will be able to:

  1. Explain why SLM safety tracks training investment, not parameter count – and why that means you cannot infer a model’s robustness from its size
  2. Refute the “smaller = safer” assumption with current evidence, including the finding that a third of tested SLMs comply with harmful requests that carry no jailbreak at all
  3. Identify which security controls disappear when a model moves from a vendor API to a device you do not control, and name the compensating control for each
  4. Assess an SLM supply chain – base model, distillation, quantization and community fine-tunes each change safety behaviour independently of capability, so each needs its own test
  5. Decide whether a proposed edge-SLM deployment is acceptable, and place the residual risk on the Blueprint layers that carry it

Why SLMs Deserve Dedicated Coverage

What Are SLMs?

Chapter 1 Section 2 placed SLMs on the size spectrum at roughly 1B-15B parameters – the tier that runs on phones, laptops and embedded hardware rather than on someone else’s GPU cluster. That is the definition this course uses throughout. Be aware that it is a convention, not a standard: the security literature often draws the line lower, at 5B or below, so always check what a given paper counted before comparing its numbers to another’s.

Named Model Examples

Microsoft's Phi-4-mini (3.8B), Google's Gemini 3.5 Flash-Lite, Anthropic's Claude Haiku 4.5, and DeepSeek-V4-Flash.

These change every few months. The deployment property that makes them a distinct security problem does not.

The number is not what matters here. What matters is the deployment model the number enables. Chapter 1 Section 3 covered edge and on-device deployment as an architectural choice; this section covers what that choice does to your threat model. A model small enough to run on the device is a model whose weights, prompts, outputs and logs are all on hardware you do not control – and that single fact moves almost every control you have relied on for the last six sections.

The “Smaller = Safer” Misconception

Security attention follows capability, so it concentrates on frontier models. That leaves a blind spot around the tier that is quietly shipping in the largest numbers, sustained by several assumptions that do not survive contact with the evidence:

Assumption Reality
“SLMs are too small to generate harmful content” Harmful output needs far less capability than useful output. SLMs clear the bar, often with weaker refusal behaviour than large models
“SLMs can’t be used for sophisticated attacks” Phishing text, social-engineering scripts and malware templates are not capability-limited tasks
“Edge deployment means no data exfiltration risk” Edge devices sync, connect to networks, and can be picked up
“SLMs don’t have enough knowledge to be dangerous” A fine-tuned SLM can hold deep domain expertise – including in dangerous domains
“Nobody would bother attacking a small model” The weights are in the attacker’s hands, which makes the model cheaper to attack, not less interesting

The evidence. The reference study is Wang et al., Can Small Language Models Reliably Resist Jailbreak Attacks? (Zhejiang University; ACM CCS ‘26), the first systematic evaluation of the question. Across 59 SLMs from 15 model families against 12 jailbreak methods:

  • 61.0% averaged an attack success rate above 40% across the attack suite
  • 37.3% failed on more than half of direct harmful queries – no jailbreak technique at all, just the request stated plainly

The second number is the one to carry. Most of this chapter has been about defeating a model’s safety training; a third of these models had no safety training worth defeating.

The intuitive explanation is wrong, and it matters

It is tempting to reason that a 3B model has “fewer neurons to spare” for safety, so safety degrades with size. The study’s correlation analysis says otherwise: SLM vulnerability tracks training details – dataset, alignment pipeline, distillation source – rather than model size scaling.

The evidence is that robustness splits by family, not by parameter count. Llama-3.2, Phi-3, Qwen, Gemma and MiniCPM hold up well at sizes where other families collapse; StableLM-2-1.6B complies with 75.7% of direct harmful queries while a well-aligned model of the same size refuses most of them.

This changes what you do about it. If size caused the problem, the only remedy would be a bigger model. Because training causes it, the remedy is testing the specific model build you are shipping – and vendor size tier tells you nothing about the answer.

Size does still matter within a family, where the training recipe is held roughly constant: the 12B StableLM refuses direct harmful queries about four times more reliably than its 1.6B sibling. That is the correct shape of the claim – compare within a family, never across.


Reduced Guardrails

The security disadvantage of SLMs is not that they are small. It is that safety alignment is expensive, invisible in benchmarks, and therefore the first thing cut when the goal is to fit a useful model onto a phone.

Why SLMs Have Weaker Safety

Safety is a budget line, and it loses. Alignment (RLHF, constitutional AI, red-teaming) costs compute and specialist time, and unlike capability it does not show up in the benchmark scores a release is judged on. Teams shipping an SLM are usually optimising a capability-per-byte target, and safety coverage is what gives. The Wang study’s central correlation is exactly this: robustness tracks the training dataset and alignment pipeline, not the parameter count.

Alignment does not survive the compression step. Most SLMs are not trained from scratch – they are derived from larger models by distillation, and often quantized afterwards. Each derivation step is a chance for safety behaviour to be left behind while capability is carried across, because the data used to distil is chosen for capability. The DeepSeek-R1-Distill case below is what that looks like in practice.

Nuanced rules degrade first. Safety instructions are usually the conditional part of a system prompt – “generate code but never exploit code”, “discuss medical topics but never give diagnoses”. Weaker instruction-following means the model reliably obeys the first clause and drops the second. This is a capability limitation with a safety consequence, and it is the one place where “smaller” genuinely is the cause.

Jailbreak Susceptibility in Practice

The evaluation produced a counter-intuitive result that is worth understanding, because it changes which attacks you should test for:

  • Direct harmful queries are the dominant risk. For 37.3% of tested models, more than half of plainly-stated harmful requests were answered. The sophisticated techniques from Section 2 are largely wasted effort against a model with no refusal behaviour to defeat.
  • Sophisticated attacks generalise worse to SLMs, not better. Multi-turn escalation such as Crescendo – highly effective against frontier models – largely fails against SLMs. So do several gradient-based methods. The reason is unflattering: the attack requires the model to track a long conversation and follow the attacker’s logical chain, and the weakest models cannot do it. They refuse or answer something irrelevant.
  • Capability is a double-edged sword. Model capability (measured by MMLU) correlates negatively with success on direct attacks and positively with success on multi-turn ones. A longer context window makes it worse: Phi-3-mini with a 128k window is about 15% more susceptible to long-context jailbreaks than the same model at 4k, because it can now process the whole adversarial prompt instead of truncating it.
What to test, and in which order

The practical consequence is that an SLM red-team is not a scaled-down LLM red-team. Against a hosted frontier model you start with sophisticated multi-turn and encoded attacks, because the simple ones are all blocked. Against an SLM you start by asking plainly for the thing it should refuse – for a third of models that is the whole attack, and it is the test most likely to be skipped as too obvious.

The Fine-Tuning Removal Problem

Whatever safety training an SLM does have can be removed cheaply. Qi et al. (2023) stripped the guardrails from a commercial aligned model by fine-tuning on 10 adversarially designed examples, for under $0.20 through the vendor’s own API – and showed that fine-tuning on entirely benign data degrades safety too, without anyone intending harm. On open weights the ceiling is even lower: one consumer GPU and an afternoon. This is why “uncensored” variants of open-weight SLMs are trivially easy to produce and impossible to recall, and why fine-tuning access is a privilege to be governed, not a feature to be exposed.


Edge Deployment Vulnerabilities

SLMs are specifically designed for edge deployment – running on devices outside the traditional data center security perimeter. This deployment model introduces a unique set of security challenges that don’t apply to cloud-hosted large models.

Physical Access to Model Weights

When a model runs on an edge device, anyone with physical access to that device can potentially extract the model weights. Unlike cloud APIs where the model is a black box, edge-deployed SLMs are white-box targets.

What physical access enables:

  • Weight extraction: Copying the model file from the device’s storage for analysis, replication, or malicious fine-tuning
  • Adversarial analysis: Studying the model’s weights to craft precise adversarial inputs that exploit specific weaknesses
  • Model modification: Replacing the legitimate model with a backdoored version that functions normally except when triggered
graph TB
    subgraph "Cloud LLM Security Model"
        CUser["User"] -->|"API Request"| CProxy["API Gateway<br/>(Authentication,<br/>Rate Limiting)"]
        CProxy -->|"Validated Request"| CModel["Model<br/>(Isolated Server)"]
        CModel -->|"Filtered Response"| CProxy
        CProxy -->|"Monitored Output"| CUser
        CLogs["Centralized<br/>Logging & Monitoring"]
        CProxy -.-> CLogs
        CModel -.-> CLogs
    end

    subgraph "Edge SLM Security Model"
        EUser["User"] -->|"Direct Access"| EDevice["Edge Device"]
        EDevice -->|"Local Inference"| EModel["Model<br/>(On-Device)"]
        EModel -->|"Unfiltered Output"| EDevice
        EDevice -->|"No Monitoring"| EUser
        EPhys["Physical Access<br/>to Device"]
        EPhys -.->|"Extract Weights<br/>Modify Model<br/>Bypass Controls"| EDevice
        ENoLog["No Centralized<br/>Logging"]
    end

    style CProxy fill:#2e7d32,stroke:#1b5e20,color:#fff
    style CLogs fill:#2e7d32,stroke:#1b5e20,color:#fff
    style EPhys fill:#b71c1c,stroke:#7f0000,color:#fff
    style ENoLog fill:#b71c1c,stroke:#7f0000,color:#fff

Limited Monitoring and Observability

Cloud-deployed models benefit from centralized logging, real-time monitoring, usage analytics, and anomaly detection. Edge-deployed SLMs often have none of these:

  • No interaction logging: Conversations with on-device models may not be recorded, making it impossible to detect misuse or jailbreaking after the fact
  • No real-time alerts: There’s no security operations center watching edge device AI interactions for anomalous patterns
  • Delayed updates: Security patches and model updates for edge devices depend on device connectivity and user compliance – some devices may run vulnerable model versions for months or years
  • No rate limiting: On-device models can be queried at unlimited speed, making brute-force attacks against safety training practical

Resource-Constrained Security

Edge devices have limited compute, memory, and storage. This creates a direct tension between security features and model performance:

  • No room for guardrail models: Large cloud deployments often use separate classifier models to filter inputs and outputs. Edge devices typically can’t run both a task model and a safety model
  • Simplified input/output filtering: Resource constraints force simpler, more easily bypassed filtering rules
  • No encryption at rest (sometimes): Some edge devices sacrifice model encryption for performance, leaving weights accessible on disk

Deciding Whether an Edge Deployment Is Acceptable

Listing what edge deployment costs you is only half the job. The question you will actually be asked is “can we ship this?” – and answering it means knowing which of the controls you are giving up can be bought back, and where.

Read the table as a checklist against a proposed deployment. A control you lose and cannot compensate is not a finding to note; it is a constraint on what the deployment is allowed to do.

Control you had in the cloud Why edge takes it away What buys it back Blueprint layer
Guardrail model on input and output No compute headroom for a second model Narrow the task until the model’s own refusals are sufficient; move anything higher-risk to a cloud tier Layer 2 · Models
Centralized interaction logging Inference never crosses a network boundary On-device audit log that syncs opportunistically, with a defined retention and a gap-detection alert Layer 4 · Users – and the discovery blind-spot table for why network signals miss this entirely
Rate limiting No gateway between user and model Nothing equivalent exists. Assume unlimited query volume and design so that unlimited querying is not itself a compromise Layer 3 · Infrastructure
Model confidentiality Weights ship to the device Nothing durable. Treat the weights as public and remove anything proprietary from them Layer 2 · Models
Prompt-side attack filtering No proxy to filter at Constrain the input surface – fixed prompt templates, no free-text passthrough to the model Layer 5 · Access
Fast patching Updates need device connectivity and user consent Signed model updates with a maximum-staleness policy that disables the feature rather than running an old build Layer 3 · Infrastructure
Knowing the deployment exists An SLM on a laptop needs no procurement Endpoint discovery for local model runtimes – this is shadow AI with weights Layer 4 · Users

Two rows have no compensating control at all, and those are the ones that decide the answer. If unlimited free querying of the model is a problem for you, or if the weights themselves are the asset, edge deployment is the wrong architecture – not a risky one to be accepted with mitigations. Everything else on the list is a cost you can choose to pay.


Model Theft Amplified

LLM06: Unbounded Consumption

Section 4 covered model extraction against a hosted API, where the attacker has to reconstruct a model they can only query. Edge deployment removes that whole problem: the weights are already on the device. Extraction stops being an attack and becomes a file copy, and every subsequent step is easier too:

Factor Large Model Small Model
File size 50-400+ GB 1-15 GB
Exfiltration time Hours to days Minutes to hours
Hardware to run Enterprise GPU cluster Consumer laptop or phone
Fine-tuning cost Thousands to millions of dollars Tens to hundreds of dollars
Distribution Requires specialized infrastructure Can be shared via file hosting

The amplification effect: When a proprietary SLM is stolen, the attacker doesn’t just get a copy – they get a starting point for malicious fine-tuning. Removing safety guardrails from a stolen 3B model can be done on a single consumer GPU in hours. The resulting “uncensored” model can then be distributed anonymously through file sharing platforms, model hubs, or peer-to-peer networks.

graph LR
    subgraph "SLM Theft-to-Weaponization Pipeline"
        A["Proprietary SLM<br/>on edge device"]
        B["Attacker extracts<br/>model weights<br/>(minutes, not hours)"]
        C["Malicious fine-tuning<br/>on consumer GPU<br/>(remove safety, add backdoors)"]
        D["Distribution via<br/>model hubs, torrents,<br/>file sharing"]
        E["Uncensored model<br/>used for phishing,<br/>malware, fraud"]

        A -->|"Physical access<br/>or exfiltration"| B
        B -->|"~10-100 examples<br/>to remove safety"| C
        C -->|"Anonymous<br/>upload"| D
        D -->|"Unrestricted<br/>generation"| E
    end

    style A fill:#2e7d32,stroke:#1b5e20,color:#fff
    style B fill:#7a6a00,stroke:#4d4300,color:#fff
    style C fill:#a85800,stroke:#6e3900,color:#fff
    style D fill:#b34700,stroke:#7a3000,color:#fff
    style E fill:#7f0000,stroke:#4a0000,color:#fff

Real-world context: The proliferation of “uncensored” and “unfiltered” model variants on model hubs demonstrates this risk. Many are fine-tuned versions of legitimate models with safety training deliberately removed. Some providers prohibit it in their license terms; enforcement is essentially impossible once weights are distributed.

Where the defence lives: nothing in this pipeline can be stopped after step B, so the controls that matter are the ones applied before shipping – model integrity and provenance in Layer 2, and treating on-device weights as published. Chapter 3 has no control that recovers a distributed model, and you should be suspicious of any vendor claiming otherwise.


Supply Chain Risks for SLMs

LLM04: Supply Chain

Section 3 covered supply chain risk for models in general. The SLM ecosystem adds a specific problem: an SLM is rarely one artifact with one provenance. By the time a model reaches a device it has typically been distilled from a larger model, fine-tuned by someone else, and quantized by a third party – and each of those steps can change safety behaviour while leaving capability intact, so none of them is visible in the benchmark scores on the model card.

The four transformations, and what each one can break:

Step What it is for What it can silently change
Distillation Producing a small model from a large one Safety behaviour is only carried across if the distillation data covers it. Capability transfers; refusals frequently do not
Community fine-tuning Adapting the model to a domain Alignment degrades even on benign data (Qi et al.). A single base model spawns hundreds of variants, none systematically reviewed
Quantization Fitting the model into device memory Contested – see below. Either way, the model you tested is not the model you shipped
Packaging Distribution through a model hub Pickle-based serialization formats execute code on load, before any inference. Section 4 covers the mechanism
Quantization and safety: what the evidence actually says

It is widely repeated that aggressive quantization degrades safety training disproportionately. The evidence is genuinely mixed, and the largest SLM-specific study finds the opposite. Wang et al. tested AWQ, GPTQ-Int4 and GPTQ-Int8 across the Qwen2.5 series and found quantization slightly improved jailbreak robustness, with AWQ reducing attack success rate by 15.9% – attributed to activation compression suppressing the noisy features gradient-based attacks exploit. Other 2025-2026 work reports the reverse for post-training quantization on larger models, including safety failures that appear only after 4-bit conversion.

Treat the direction as unsettled and method-dependent. The operational rule survives either way, and it is the part worth remembering: safety is a property of the specific build, so re-run the safety evaluation on the quantized artifact you are actually shipping. Do not inherit a safety result from the full-precision model, in either direction.

Tooling is the fourth supply chain. As SLMs gain tool use through MCP and function calling, they inherit everything Section 5 covers under ASI04 – including the postmark-mcp rug pull, where fifteen clean versions preceded the backdoor. The SLM-specific twist is on the receiving end: a tool description is untrusted text placed in the context window, and a model that complies with 50% of plainly stated harmful requests is not going to resist one hidden in a tool’s prose.


OWASP LLM Top 10: How Each Category Manifests in SLMs

Every vulnerability in the OWASP LLM Top 10 applies to SLMs, but many manifest differently – and often more severely. The table below maps each category to its SLM-specific characteristics. Read the severity column as a comparison against the same category on a hosted frontier model, not as an absolute rating:

OWASP Category SLM-Specific Manifestation Severity vs. Large Models
LLM01: Prompt Injection Weaker refusal behaviour means the injected instruction rarely needs disguising; sophisticated multi-turn injection generalises worse, but direct injection succeeds outright Higher for simple attacks, lower for elaborate ones – the reverse of the frontier-model pattern
LLM02: Sensitive Information Disclosure Edge deployment means leaked data may not be logged; physical access enables direct weight analysis for memorized data extraction Different – less training data memorized, but harder to detect leaks
LLM04: Supply Chain Four transformation steps between base model and device, each able to change safety without changing benchmarks; hundreds of unreviewed variants per base model Higher – far more unvetted variants, and provenance is a chain rather than a source
LLM05: Data and Model Poisoning Easier to poison a smaller model with fewer training examples; fine-tuning to insert backdoors requires minimal resources Higher – lower poisoning threshold
LLM10: Improper Output Handling Same risk as large models, but edge deployment may skip output filtering due to resource constraints Higher – less room for output safety layers
LLM03: Excessive Agency SLMs gaining tool use through MCP and function calling; weaker judgment about when to refuse actions Higher – less capability to reason about appropriate action boundaries
LLM08: Hidden Context Exposure SLMs are worse at maintaining system prompt confidentiality; simpler extraction techniques succeed Higher – weaker prompt boundary enforcement
LLM09: Vector and Embedding Weaknesses Smaller embedding spaces are more susceptible to adversarial perturbations and collision attacks Higher – less dimensional space for robust embeddings
LLM07: Misinformation Higher hallucination rates due to less training data and knowledge; no real-time fact-checking on edge devices Higher – more hallucinations, less detection
LLM06: Unbounded Consumption (model extraction) Small file size enables rapid exfiltration; consumer hardware sufficient to run stolen model; trivial to fine-tune for malicious purposes Much Higher – dramatically easier to steal and weaponize

Case Studies

Case Study 1: The First Systematic SLM Jailbreak Evaluation

LLM01: Prompt Injection

Researchers: Wang, Zhang, Xu, He, Zhu and Ren – State Key Laboratory of Blockchain and Data Security, Zhejiang University Published: arXiv, March 2025; revised August 2026; ACM CCS ‘26 Scope: 59 SLMs from 15 model families, against 12 jailbreak methods, with two LLMs as baselines

This is the study every claim in this section rests on, and it is worth knowing its shape rather than just its headline number – because two of its results contradict what “everybody knows” about small models.

  • 61.0% of evaluated SLMs averaged an attack success rate above 40% across the twelve methods
  • 37.3% failed on more than half of direct harmful queries – the request stated plainly, with no jailbreak technique applied
  • Vulnerability correlates with training details, not with model size. Robustness splits by family: Llama-3.2, Phi-3, Qwen, Gemma and MiniCPM hold up where other families of identical size do not
  • Multi-turn attacks such as Crescendo largely fail against SLMs – not because of any defence, but because the models cannot follow the attacker’s logical chain. Model capability correlates negatively with simple-attack success and positively with multi-turn success
  • Quantization slightly improved robustness in their tests (AWQ reduced ASR by 15.9%), against the common assumption
  • Practical harm is capped by capability. The jailbroken output of a weak SLM is frequently incoherent or non-actionable – a real mitigating factor, and the one place where “smaller” genuinely helps

Why this changes practice: the field’s defensive playbook was built against frontier models, where the simple attacks are already blocked and the sophisticated ones are the threat. On SLMs that ordering inverts. The authors also evaluated five defences and found prompt-level ones inconsistent across models and attacks, while model-level adversarial training generalised poorly to multi-turn attacks – which is why their conclusion is a call for security-by-design in SLM development rather than a filter to bolt on afterwards.

Case Study 2: Distillation Strips Safety – The DeepSeek-R1-Distill Family

LLM04: Supply Chain LLM05: Data and Model Poisoning

Models: DeepSeek-R1-Distill-Qwen (1.5B and 7B), distilled from Qwen2.5-Math Date: Released January 2025; safety evaluated through 2025-2026 Scope: Among the most widely downloaded open-weight small reasoning models

Most SLMs are not trained from scratch. They are distilled from a larger model – the big model generates outputs, the small model is fine-tuned to imitate them. The DeepSeek-R1 distills are the best-documented example of what that process does to safety.

The distillation data was reasoning traces generated by DeepSeek-R1, selected to transfer reasoning capability. It transferred. Safety behaviour was not part of the objective, and it did not come along:

  • The distilled 1.5B and 7B models reach direct-harmful-query attack success rates of 0.271 and 0.514 – the 7B model higher than every Qwen-family SLM tested but one, and substantially worse than the base models they were derived from
  • The failure is visible in the reasoning traces themselves. Nearly all of them open with Okay, so I need to figure out [the harmful request]. Hmm, where do I even start? – the model treats a request it should refuse as a problem to be solved. There is no refusal step for the jailbreak to defeat
  • Independent evaluations found the same pattern in the family: Cisco’s red-team reported that DeepSeek-R1 itself failed to block a single prompt from a 50-prompt HarmBench sample, and FAR.AI showed the remaining guardrails could be removed by fine-tuning without degrading response quality

Outcome and why it generalises: these models were not attacked – they shipped this way, from a reputable lab, and were downloaded at scale. The lesson is not about DeepSeek. It is that distillation is how most SLMs are made, capability is what distillation optimises for, and safety is therefore the property most likely to be lost in transit. A model card reporting strong benchmark scores is evidence about the thing distillation was designed to preserve, and no evidence at all about the thing it was not.

This is why the supply chain table above lists four transformation steps rather than one source. The provenance question for an SLM is not “who made this?” but “what was done to it after they made it, and what was each step optimising for?”


Chapter 2 Summary: The Complete Attack Taxonomy

You’ve now completed a comprehensive tour of the AI attack landscape. Let’s step back and see the full picture of what you’ve learned across all seven sections:

Section Focus Primary Framework Key Takeaway
S1: The AI Attack Surface Attack taxonomy and OWASP overview OWASP LLM Top 10 (2026) Every AI capability has a corresponding attack surface
S2: Prompt-Level Attacks Injection, jailbreaking, prompt leaking LLM01, LLM08 The model’s input is the primary attack vector
S3: Data and Training Attacks Poisoning, backdoors, supply chain LLM04, LLM05 Attacks can happen before the model ever sees a user
S4: Model and Infrastructure Serialization, adversarial inputs, extraction, DoS LLM09, LLM06 The model itself and its infrastructure are targets
S5: Agentic Attack Vectors Tool exploitation, cascading failures, rogue agents OWASP Agentic AI Top 10 (2026) Agency turns text vulnerabilities into real-world actions
S6: Output and Trust Exploitation Hallucinations, data leakage, improper output handling LLM02, LLM10, LLM03, LLM07 AI outputs are attack vectors against downstream systems
S7: SLM Threats Weak alignment, edge risks, extraction as file copy Full LLM Top 10, SLM lens Safety is a property of a build, not of a size

The thread connecting everything: From prompt injection to agentic cascading failures, from data poisoning to SLM jailbreaking, the fundamental challenge is the same – AI systems blur the boundary between data and instructions, between trusted and untrusted, between capability and vulnerability. Every feature is a potential attack surface. Every integration point is a trust boundary.

And this section adds the boundary that decides which defences are even available to you. Sections 1-6 assumed a model behind an API you control: a place to filter input, inspect output, count requests and write logs. That assumption is a deployment property, not a property of AI. Remove it and most of the chapter’s defences have nowhere to run. When you assess a system, establish where the model executes before you reason about anything else – it determines which half of your toolkit exists.

In Chapter 3, you’ll learn how to defend against everything you’ve seen here – from foundational security architectures to specific tools and frameworks for protecting AI systems at every layer.

Key Takeaways
  • Safety tracks training, not size. 61.0% of 59 tested SLMs averaged an attack success rate above 40%, but robustness splits by family rather than by parameter count – so a model’s size tier tells you nothing about its robustness, and only testing the specific build does.
  • Start the red-team with the obvious attack. 37.3% of tested SLMs complied with more than half of plainly-stated harmful requests. Sophisticated multi-turn attacks generalise worse to SLMs, because the model cannot follow the chain.
  • Edge deployment removes controls, and two of them cannot be bought back: there is no equivalent of rate limiting, and no way to keep weights confidential on a device you do not control. Those two decide whether the architecture is viable at all.
  • Extraction stops being an attack and becomes a file copy. Small weights on hardware the attacker holds means minutes-to-hours exfiltration, consumer hardware to run it, and roughly 100 fine-tuning examples to strip whatever alignment remains.
  • Provenance is a chain, not a source. Distillation, community fine-tuning, quantization and packaging each change safety without changing benchmark scores – re-test the artifact you actually ship.

Test Your Knowledge

Ready to test your understanding? The quiz covers SLM threats, then closes with two questions that draw across all seven sections of the chapter.


Chapter 2 Complete!

You’ve explored the full AI attack landscape – from prompt injection through agentic cascading failures to SLM-specific threats. In Chapter 3, you’ll learn how to defend against these attacks using TrendAI’s 6-layer Security for AI Blueprint.