Threat Modeling the AI Inference Pipeline

The model is only one component. The attack surface is the entire inference pipeline.

Traditional threat modeling identifies assets, entry points, trust boundaries, and potential attack paths. AI inference systems require the same discipline, but introduce a complication: untrusted data can influence not only what a system returns, but also what it decides to do.

A production inference pipeline may look like this:

User
  ↓
API Gateway
  ↓
Application
  ↓
Context Builder ← RAG / External Data
  ↓
Model Endpoint
  ↓
Output Validation
  ↓
Tool Executor → APIs / Database
  ↓
Response

Each component introduces different assets, privileges, and failure modes. Threat modeling must examine the transitions between them, not just the model endpoint.

Identify the Assets and Entry Points

Before identifying threats, establish what needs protection.

In an inference architecture, critical assets may include system instructions, retrieved enterprise documents, model credentials, conversation history, tool permissions, and access to internal services.

Entry points extend beyond the user prompt. Retrieved documents, API responses, uploaded files, and tool outputs can all introduce attacker-controlled content.

Consider a RAG application that retrieves documents from an internal knowledge base. An attacker modifies a document to include instructions directing the model to disclose sensitive information. The retrieval service functions correctly. The model receives valid context. Yet the application may still produce an unauthorized response. The failure is not necessarily in retrieval or inference. It is in how the architecture handles untrusted content.

Model the Attack Paths

A useful threat model connects attacker control to a security impact.

For example:

Attacker-Controlled Document
          ↓
      RAG Retrieval
          ↓
    Model Context
          ↓
   Malicious Tool Call
          ↓
   Privileged Tool Executor
          ↓
   Unauthorized Data Access

This path is possible when several conditions align: the attacker can influence retrieved content, the model follows malicious instructions, and the tool executor lacks independent authorization checks.

The model’s behavior alone does not determine whether the attack succeeds.

The surrounding architecture does.

Map Threats to Controls

ThreatArchitectural control
Indirect prompt injectionPreserve instruction/data separation; restrict downstream authority
Sensitive data exposureRetrieval authorization and data minimization
Unauthorized tool executionUser-scoped authorization and policy enforcement
Excessive tool privilegesLeast-privilege identities and scoped credentials
Malicious model outputSchema validation and context-specific output handling

These controls should operate independently of the model’s ability to recognize malicious instructions.

For example, even if a model generates a valid database query, the execution layer must verify that the requesting identity is authorized to access the requested records.

Threat Model the Failure, Not Just the Component

A component-by-component review might confirm that the API is authenticated, the vector database is private, and the model endpoint is isolated.

All three observations can be true while the system remains vulnerable to an attack crossing those components.

Effective threat modeling therefore asks:

What can an attacker control, which boundaries can that influence cross, what authority becomes reachable, and what prevents the final action?

The goal is not to prove that an LLM will always behave correctly.

It is to design an inference architecture where incorrect or manipulated model behavior cannot automatically become a security compromise.

A secure AI system is not one where the model never makes a mistake. It is one where the architecture limits what that mistake can do.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *