P59 · AI & Agent Application Security
prompt injection, tool boundaries, agent auth & sandboxing
Securing LLM and agent applications: quarantining untrusted content, keeping the instruction channel trusted, redacting what travels, gating an agent's tool calls with human-in-the-loop, sandboxing, and threat-modelling the whole system.
Every other pillar in this discipline defends a system that executes only the instructions you wrote. An LLM application does not. It reads text from places you do not control and acts on it, and the component doing the reading has no reliable way to distinguish here is some information from here is what to do next. Attach tools to that component and the gap stops being theoretical: a line hidden in a retrieved document becomes a call to your API, made with your agent's credentials and your agent's permissions.
This pillar builds defences that survive a fully persuaded model, because defences made of prompt wording do not. The instruction channel comes first — keeping your instructions structurally separate from arriving content, quarantining what retrieval returns so it is unambiguously data, and separating privilege inside prompt assembly — alongside the outbound question of what leaves: PII in a log, a system prompt handed over on request, one tenant's documents surfacing in another tenant's answer. Then the load-bearing principle of the whole pillar: the model proposes, your code decides. A toolset scoped to least privilege, a human in the loop before anything irreversible, and a sandbox as the isolation floor for code you did not write. Finally the system as a whole, with STRIDE adapted to LLM and agent architectures — assets, trust boundaries, abuse cases, mitigations — which you produce yourself in the capstone threat model. Attack techniques appear here only as the thing a named defence stops, and the exercises grade the defence.