Prompt injection & untrusted content
P59.llm-threats.01 · Audience: guest, it-ml, language-pro · Prerequisites: Injection & the untrusted-input lens
Welcome to AI & Agent Application Security — the surface LLM and agent systems add on top of everything in P52–P57, and the discipline's distinctive pillar. Its defining problem is the instruction/data boundary: a model reads text, and text that looks like an instruction can be treated as one. When that text arrived from a web page, a document or a tool result, the attacker wrote it — indirect prompt injection.
This track builds the two defences that carry most of the weight: quarantining retrieved content so it enters as clearly-marked data, and privilege separation in prompt assembly so untrusted content can never reach the trusted instruction channel.
The honest framing this pillar keeps throughout: these mitigations reduce risk, they do not eliminate it. That is precisely why the agent track puts irreversible actions behind a gate. Every kata here is defence — hostile-looking strings appear only as inputs your code must handle safely.
The code runs entirely in your browser (Pyodide) — no model is called, no agent is run.
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Rung 1 — quarantine untrusted content
Loading exercise…
Rung 2 — keep the instruction channel trusted
Loading exercise…
My notes on this module
Loading your notes...
Where next?
Later in LLM Threats
This module unlocks