The model is only one component. The attack surface is the entire inference pipeline.
Traditional threat modeling identifies assets, entry points, trust boundaries, and potential attack paths. AI inference systems require the same discipline, but introduce a complication: untrusted data can influence not only what a system returns, but also what it decides to do.
A production inference pipeline may look like this:
User
↓
API Gateway
↓
Application
↓
Context Builder ← RAG / External Data
↓
Model Endpoint
↓
Output Validation
↓
Tool Executor → APIs / Database
↓
Response
Each component introduces different assets, privileges, and failure modes. Threat modeling must examine the transitions between them, not just the model endpoint.
Identify the Assets and Entry Points
Before identifying threats, establish what needs protection.
In an inference architecture, critical assets may include system instructions, retrieved enterprise documents, model credentials, conversation history, tool permissions, and access to internal services.
Entry points extend beyond the user prompt. Retrieved documents, API responses, uploaded files, and tool outputs can all introduce attacker-controlled content.
Consider a RAG application that retrieves documents from an internal knowledge base. An attacker modifies a document to include instructions directing the model to disclose sensitive information. The retrieval service functions correctly. The model receives valid context. Yet the application may still produce an unauthorized response. The failure is not necessarily in retrieval or inference. It is in how the architecture handles untrusted content.
Model the Attack Paths
A useful threat model connects attacker control to a security impact.
For example:
Attacker-Controlled Document
↓
RAG Retrieval
↓
Model Context
↓
Malicious Tool Call
↓
Privileged Tool Executor
↓
Unauthorized Data Access
This path is possible when several conditions align: the attacker can influence retrieved content, the model follows malicious instructions, and the tool executor lacks independent authorization checks.
The model’s behavior alone does not determine whether the attack succeeds.
The surrounding architecture does.
Map Threats to Controls
| Threat | Architectural control |
|---|---|
| Indirect prompt injection | Preserve instruction/data separation; restrict downstream authority |
| Sensitive data exposure | Retrieval authorization and data minimization |
| Unauthorized tool execution | User-scoped authorization and policy enforcement |
| Excessive tool privileges | Least-privilege identities and scoped credentials |
| Malicious model output | Schema validation and context-specific output handling |
These controls should operate independently of the model’s ability to recognize malicious instructions.
For example, even if a model generates a valid database query, the execution layer must verify that the requesting identity is authorized to access the requested records.
Threat Model the Failure, Not Just the Component
A component-by-component review might confirm that the API is authenticated, the vector database is private, and the model endpoint is isolated.
All three observations can be true while the system remains vulnerable to an attack crossing those components.
Effective threat modeling therefore asks:
What can an attacker control, which boundaries can that influence cross, what authority becomes reachable, and what prevents the final action?
The goal is not to prove that an LLM will always behave correctly.
It is to design an inference architecture where incorrect or manipulated model behavior cannot automatically become a security compromise.
A secure AI system is not one where the model never makes a mistake. It is one where the architecture limits what that mistake can do.
Leave a Reply