Guided pathThis is part of Understand security: protocols, identity and secure architectureBack to the path

Prompt injection & untrusted content

P59.llm-threats.01 · Audience: guest, it-ml, language-pro · Prerequisites: Injection & the untrusted-input lens

Real LLM grading for this pageLLM grading (this page):

Welcome to AI & Agent Application Security — the surface LLM and agent systems add on top of everything in P52–P57, and the discipline's distinctive pillar. Its defining problem is the instruction/data boundary: a model reads text, and text that looks like an instruction can be treated as one. When that text arrived from a web page, a document or a tool result, the attacker wrote it — indirect prompt injection.

This track builds the two defences that carry most of the weight: quarantining retrieved content so it enters as clearly-marked data, and privilege separation in prompt assembly so untrusted content can never reach the trusted instruction channel.

The honest framing this pillar keeps throughout: these mitigations reduce risk, they do not eliminate it. That is precisely why the agent track puts irreversible actions behind a gate. Every kata here is defence — hostile-looking strings appear only as inputs your code must handle safely.

The code runs entirely in your browser (Pyodide) — no model is called, no agent is run.

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send
Rung 1 — quarantine untrusted content

Loading exercise…

Rung 2 — keep the instruction channel trusted

Loading exercise…

My notes on this module

Loading your notes...

Prompt injection & untrusted content — TransformerLab