Traditional threat modeling starts with assets, trust boundaries, identities, data flows, and attacker goals.
AI agents still need all of that.
What changes is the execution model.
An autonomous agent can consume untrusted content, reason over it, select tools, execute commands, invoke APIs, and alter external systems—all during one workflow.
A useful agent threat model therefore needs to ask:
What happens if the agent uses its legitimate capabilities in an unintended way?
Step 1: Define the agent as a system, not a model
Do not threat-model only the LLM.
Map the whole system:
User
|
v
Agent Runtime
+-- Model Provider
+-- Filesystem
+-- Shell
+-- APIs
+-- MCP Servers
+-- Databases
+-- Cloud
`-- Memory / State
The attack surface exists across all of those components.
Step 2: Inventory the agent's capabilities
Create a capability table.
| Domain | Capability |
|---|---|
| Filesystem | Read/write workspace |
| Process | Execute shell |
| Network | HTTPS outbound |
| GitHub | Read/write repo |
| Jira | Read/write issues |
| MCP | Three servers |
| Cloud | Object-store read |
| Inference | External model |
Threat severity depends heavily on what the agent can actually do.
Prompt injection against a read-only research assistant is different from prompt injection against a production deployment agent.
Step 3: Identify assets
List what you are trying to protect:
- source code;
- customer records;
- PHI/PII;
- intellectual property;
- API credentials;
- cloud roles;
- CI/CD;
- production systems;
- financial transactions;
- model context;
- prompts;
- logs.
Then classify them.
Step 4: Draw trust boundaries
Example:
Employee
|
| Trust Boundary 1
v
Agent Sandbox
|
| Trust Boundary 2
+------> External Model
|
| Trust Boundary 3
+------> MCP Server
|
| Trust Boundary 4
`------> Production API
For each boundary ask:
- Which identity crosses it?
- Which credential is used?
- Which data crosses it?
- What authorization is enforced?
- Is the destination trusted?
- Is traffic logged?
Step 5: Map data flows
Agents continuously move data.
Jira ticket
|
v
Agent context
+--> Model provider
+--> GitHub
`--> MCP tool
Now ask whether the ticket contains sensitive data, whether that data leaves the enterprise, whether MCP output can influence GitHub writes, and whether any credential enters the model context.
Step 6: Model the major agent-specific threats
Prompt injection
Untrusted content attempts to influence the agent's behavior. Sources include websites, repository files, tickets, emails, documents, database records, and tool responses.
Excessive agency
The agent can perform more actions than the workflow requires.
Credential theft
The agent, a child process, or malicious content causes credentials to be exposed.
Data exfiltration
Authorized data is transmitted to an unauthorized destination.
Tool abuse
The agent invokes a legitimate tool for an unintended purpose.
MCP compromise
An MCP server or tool returns malicious instructions or receives overly broad authorization.
Privilege escalation
The workload expands beyond its intended process or system boundary.
Supply-chain compromise
The agent installs a malicious dependency, executes untrusted code, or modifies CI/CD.
Cross-workspace access
One agent reaches another agent's data, secrets, or environment.
Model-provider exposure
Sensitive data is sent to an unapproved inference provider.
Unauthorized persistence
The agent creates long-lived credentials, jobs, branches, infrastructure, or other durable changes beyond its intended workflow.
Audit failure
Security teams cannot determine what the agent did after an incident.
Step 7: Build abuse cases
A threat model should contain concrete stories.
Abuse case: malicious README
Precondition: Coding agent reads repository content.
Attack: A README tells the agent to read cloud credentials and upload them.
Potential impact: Credential disclosure and cloud compromise.
Controls: filesystem deny for credential paths, default-deny egress, credential mediation, endpoint policy, and security logging.
Validation test: place a simulated malicious instruction in repository content and confirm the agent cannot access or transmit protected credentials.
That is much more actionable than simply writing "Risk: prompt injection."
Step 8: Separate agent failure from control failure
Assume the model may eventually misunderstand, hallucinate, follow malicious context, or choose the wrong tool.
Then ask two questions.
Agent behavior
Did the agent attempt an unsafe action?
Security control
Did the infrastructure prevent the action?
A secure system can tolerate some model-level failures because surrounding controls constrain impact.
Step 9: Score risk using autonomy and sensitivity
Traditional impact and likelihood remain useful.
For agents, add two contextual factors:
- Autonomy: how independently can the agent act?
- Resource sensitivity: what assets can it reach?
Use the result for prioritization rather than pretending the score is mathematically exact.
Step 10: Map every threat to an enforceable control
| Threat | Control |
|---|---|
| Data exfiltration | Egress allowlist |
| Credential theft | Brokered credentials |
| File disclosure | Filesystem policy |
| MCP abuse | Tool-level authorization |
| Privilege escalation | Restricted process identity/seccomp |
| Unapproved model | Provider/network policy |
| Missing audit | SIEM export |
The threat model should lead directly to architecture.
Step 11: Test the control
For each high-risk threat, define a negative test:
Attempt unauthorized domain
Attempt restricted file read
Attempt credential print
Attempt prohibited MCP tool
Attempt production API action
Attempt privilege escalation
Document expected outcome, actual outcome, evidence, and severity if it fails.
Step 12: Revisit the model when capabilities change
A seemingly small change such as "add the AWS MCP server" may materially change the threat model.
Re-evaluate when new tools are added, credentials change, production access is granted, model providers change, filesystem scope expands, internet access changes, or autonomy increases.
Where OpenShell fits
OpenShell can provide enforcement primitives for several common findings: filesystem restrictions, process restrictions, outbound network policy, provider credential mediation, controlled inference, workspace separation, and security logging.
It does not replace the threat model.
The threat model tells you what policy should exist. OpenShell is one technology that can enforce parts of it.
Threat-model the system you actually have
The biggest mistake is using a generic AI threat list without mapping it to real enterprise access.
The meaningful question is not:
Could prompt injection happen?
It is:
If prompt injection succeeds against this specific agent, what can the attacker cause the agent to access or change?
That answer determines the security priority.
How Anpu Labs approaches agent threat modeling
Our Secure Agent Infrastructure Assessment starts by building a capability and dependency map. We then model realistic abuse paths through identity, files, credentials, models, MCP, APIs, cloud infrastructure, and network egress.
The result becomes the foundation for the target OpenShell architecture and implementation backlog.
References
- NVIDIA OpenShell — Overview: https://docs.nvidia.com/openshell/latest/about/overview
- NVIDIA OpenShell — How OpenShell Works: https://docs.nvidia.com/openshell/about/how-it-works
- MCP TypeScript SDK v2: https://ts.sdk.modelcontextprotocol.io/v2/




