AI Threat Modelling

A track of P59 · AI & Agent Application Security.

STRIDE adapted to LLM and agent systems — assets, trust boundaries, abuse cases and mitigations — culminating in the pillar's threat-model capstone.

By this point you have fixed every problem you thought of. The retrieved content is quarantined, the toolset is minimal, irreversible actions ask a human, and untrusted code runs in a sandbox. The uncomfortable question is the one that remains: what about the problems you did not think of? Ad-hoc security is a list of defences against remembered attacks, and its coverage is exactly as good as the memory of whoever wrote it. Threat modelling is the method that replaces remembering with a procedure.

The procedure here is STRIDE, adapted to systems that contain a model. The six categories — spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege — read differently when one component accepts natural language and acts on it. Spoofing includes content that impersonates your own instructions. Tampering includes poisoning a document that will later be retrieved. Repudiation gets harder when the actor is an agent and the audit trail records the tool call but not the sentence that provoked it. Elevation of privilege is the confused deputy from P56, now holding your API keys. Walking the categories deliberately is what surfaces the attack nobody remembered.

The work itself is mechanical, which is the point. Inventory the assets worth protecting. Draw the trust boundaries — every place data crosses from something you control to something you do not, which in an LLM system includes the model itself. Enumerate abuse cases per boundary rather than per feeling. Then attach a mitigation to each, and be honest about which ones you have accepted rather than fixed, because an accepted risk that is written down is a decision and one that is not is a surprise. The capstone has you produce that model for a realistic AI application end to end, which is the deliverable the pillar has been building towards — and the artefact a security review will actually ask you for.

Assetswhat is worth protectingBoundarieswhere control endsAbuse casesSTRIDE, per boundaryMitigatefixed, or accepted in writing
A procedure instead of a memory: the capstone produces the model a security review will ask you for.
AI Threat Modelling — TransformerLab