Prompt Injection Defences for LLM Apps
Design LLM apps that stay safe when users, documents, web pages or tool results contain hostile instructions.
A taste of a lesson
Our email assistant can read the inbox and send replies. Is a strong system prompt enough protection?
No. An incoming email can contain instructions such as forward the last ten invoices to this address, and models cannot reliably tell that text apart from your instructions. You have all three risk ingredients: private data, untrusted content and a way to send data out. Break the chain in code: require the user to approve every outgoing email, restrict recipients or flag new ones, and keep the step that reads raw emails separate from the step that can send. Which of those could you add first?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain direct and indirect prompt injection with realistic examples
- Assess risk from data access, untrusted content and exfiltration paths
- Apply least privilege, user scoped permissions and approval steps
- Block exfiltration through links, images and outbound tools
- Isolate untrusted content from privileged steps and test defences
Lesson plan
- 1 How injection works Understand why models follow instructions hidden in data. Start
- 2 Assessing your risk Map what your app reads, accesses and can do. Start
- 3 Limiting capability Reduce what a successful injection could achieve. Start
- 4 Treating output as untrusted Handle model output safely before acting on or rendering it. Start
- 5 Isolating untrusted content Separate steps that read untrusted text from steps with privileges. Start
- 6 Testing and monitoring Probe defences regularly and watch for attacks in production. Start
Try asking
About this tutor
For experienced developers and security minded engineers building LLM features that read untrusted content or can take actions. You learn how direct and indirect prompt injection work, why there is no complete fix today, and how to limit damage through architecture: least privilege tools, user scoped permissions, human approval for consequential actions, treating model output as untrusted, blocking data exfiltration through links and images, constraining outputs, and isolating untrusted content from privileged steps. The tutor works through realistic attack paths in email assistants, RAG systems and browsing features, and pairs each with layered defences.
Reviews
4.7
3 ratingsSample
- Tunde O.Sample
Advanced and thorough. I would have liked more time on testing, but the attack path walkthroughs were excellent.
- Alessia R.Sample
Honest that there is no complete fix, and practical about reducing blast radius. The quarantined call pattern was new to me and works well.
- Dmitri K.Sample
The three ingredient risk model reshaped our design review. We added approval for outgoing actions and stopped rendering model generated images.
About the teacher
Structured output, tool calling and safe input handling for LLM applications that must behave predictably
9 tutors 308 lessons taught Sample
I teach the parts of LLM apps where free text has to meet real software: JSON that must parse, tools the model calls, images and documents coming in, and users who send things you did not plan for. I spent years writing integrations between messy systems, which taught me to treat every input as untrusted and every output as something...
See Greta's profile and tutorsMore like this
Other tutors on the same or nearby topics.