Agent Boundaries
A track of P59 · AI & Agent Application Security.
The model proposes, your code decides: least-privilege toolsets, human-in-the-loop for irreversible actions, and the sandbox isolation floor untrusted code must run inside.
You give an agent a delete_files tool so it can tidy a working directory, and
a task that says "clean up the old build artefacts". It reads a README in the
repository to understand the layout. Buried in that README is a line placed
there by whoever contributed it last, addressed to any assistant that reads it.
The agent has a plausible instruction, a capable tool, and your credentials.
Nothing in this chain requires the model to be defective — it requires only that
the model was persuaded, which is a thing models do.
This track is built on one sentence: the model proposes, your code decides. A tool call emitted by a model is a request, not an action, and the boundary where that request is evaluated is the only place a defence can live that survives a fully convinced model. Everything else follows from taking that seriously. The toolset is scoped to least privilege — the agent that answers questions about invoices does not hold a tool that can issue refunds, because a tool it does not have cannot be talked into being used. Arguments are validated at that boundary against what the caller's identity is entitled to, exactly as in P56, rather than trusted because a model produced them.
Then irreversibility, which is the property worth organising around. Reading is recoverable; sending an email, deleting a record, moving money and merging a branch are not. Actions in the second category get a human in the loop by design, with enough context presented for the approval to be meaningful rather than a reflex click — an approval prompt nobody reads is worse than none, because it manufactures the appearance of oversight. The track closes on the floor beneath all of this: when an agent runs code, that code is untrusted, whoever wrote it, and it belongs in a sandbox with restricted filesystem, network and time, so that the worst outcome is a failed task rather than a compromised host.