Technique / LLM01:2026 / AML.T0051
What is prompt injection?
Prompt injection is an attack in which text that a large language model reads, typed by a user or hidden in a document, email or web page, overrides the instructions its developer gave it. OWASP ranks it first among the risks to LLM applications in 2026, as it did in 2025.
Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction
Water damage from a burst pipe is covered up to the policy limit.
A claim is assessed within ten working days of the report.
A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.
Questions this page does not answer go to the claims desk.
Prompt injection, defined
Prompt injection is an attack on an application built on a large language model (LLM): text the model processes changes what the model does, against the intent of the people who built the application. The text can come from the person in the chat, or from anything the model is given to read.
The name dates from September 2022, when Simon Willison proposed it and drew the obvious parallel with SQL injection: in both cases, data that was meant to be read ends up being obeyed.1 The parallel is useful for the first five minutes and misleading after that, for reasons this piece comes back to.
OWASP defines a prompt injection vulnerability as one where prompts alter the model’s behaviour or output in unintended ways, and places it first in its Top 10 for LLM Applications.2 It held that position in the 2025 edition and keeps it, as LLM01:2026, in the edition published on 3 August 2026.3
MITRE ATLAS catalogues the same attack as technique AML.T0051, LLM Prompt Injection, and describes it as a possible initial access vector: a foothold from which an attacker carries out the next steps of an operation.4
What is the difference between direct and indirect prompt injection?
In a direct injection the attacker is the user. They type into the chat box and try to talk the application out of its instructions. In an indirect injection the attacker never talks to the system at all: they place text where the system will read it later, in a web page, a shared document, an email, a support ticket or the output of a tool.5
The indirect form was set out in 2023 by Greshake and colleagues, who showed it against real LLM-integrated applications and argued that such applications blur the line between data and instructions.6 ATLAS has since added a third sub-technique, Triggered, for injections set off by a user action or an event in the victim’s own environment.7
| Direct | Indirect | |
|---|---|---|
| Who writes the text | The user of the system | Anyone who can write where the system reads |
| Channel | The chat or API input | Retrieved documents, web pages, email, tickets, files, tool results, memory |
| Who is harmed | Mostly the operator: policy, cost, reputation | Often another user, whose request triggers it with their permissions |
| What limits it | What the user may already see and do | What the system may see and do on the victim’s behalf |
That last row is the reason indirect injection matters most for agents. A direct attacker can only abuse their own session. An indirect attacker borrows someone else’s: the agent reads the planted text while working for a legitimate user, and acts with that user’s access. Fig. 1 follows one such line through a fictional insurer’s support agent, from its knowledge base into the stack beneath it.
MODEL LAYER
STACK LAYER
- KNOWLEDGE BASE
- RETRIEVAL
- AGENT
- TOOLS
EMAIL / CRM / REFUNDS - MCP SERVER
- WEB APP
- API GATEWAY
- IDENTITY
- TENANT DATA
- OBJECT STORAGE
- CLOUD ACCOUNT
- 01 UNREVIEWED TEXT
STORED AS POLICY LLM01:2026 / PROMPT
INJECTION - 02 RETRIEVED AS TRUSTED
CONTEXT ASI06 / MEMORY AND
CONTEXT POISONING - 03 INSTRUCTION FOLLOWED LLM03:2026 / EXCESSIVE
AGENCY - 04 THIRD-PARTY TOOL,
NEVER REVIEWED ASI04 / AGENTIC SUPPLY
CHAIN - 05 TOKEN SCOPED WIDER
THAN THE TASK API1:2023 / BROKEN
OBJECT LEVEL
AUTHORIZATION GET /claims/48213 / 200 - 06 ANOTHER TENANT’S
RECORD WSTG-ATHZ / TENANT
ISOLATION
Why does prompt injection work at all?
A model receives one stream of tokens. The developer’s instructions, the user’s message, a retrieved document and a tool result all arrive in that same stream, and nothing in it marks which parts are commands and which are material to work on (Fig. 2). Models are trained to give the developer’s instructions more weight, but that is a tendency learned from examples, not a rule the system enforces.
Developer instructions
User message
Retrieved document
Tool result
One context window / no type on any row
The highlighted row came from a document. Nothing marks it as data.
The UK National Cyber Security Centre made the point bluntly in December 2025: current LLMs do not enforce a security boundary between instructions and data, prompt injection may never be totally mitigated, and the realistic aim is to reduce the likelihood and the impact of attacks.8 Its author calls the model an inherently confusable deputy. That is where the SQL comparison breaks: SQL injection has a structural fix, the parameterised query, which keeps data out of the command channel. Language models have no equivalent channel to keep it out of.
OWASP draws the same conclusion from another angle: retrieval-augmented generation and fine-tuning make outputs more relevant, but they do not fully mitigate prompt injection.9 So the useful question is not how to make the model unpersuadable. It is what a persuaded model is able to do.
What can an attacker do with prompt injection?
The impact of a prompt injection is the reach of the system it lands in. The model’s words are rarely the damage; what the application does with them is. The consequences fall into four groups.
- Disclosure. The model reveals what it was given but should not pass on: hidden instructions, configuration, another user’s data, documents from the wrong tenant. OWASP is explicit that a system prompt should not be treated as a secret or used as a security control.10
- Actions. An agent calls the tools it holds on the attacker’s behalf: it sends the email, files the refund, updates the record. ATLAS describes how tool invocation can give an attacker the privileges of the integrated service, and how a tool that writes can carry data out inside a legitimate-looking action.11
- Output. The model’s answer is rendered or executed downstream without being treated as untrusted input, so the injection becomes a classic web or code execution flaw. OWASP files this under LLM10:2026, Improper Output Handling.3
- Persistence. The instruction is written into memory or a shared store, and keeps steering the system long after the original content is gone.
The OWASP Top 10 for Agentic Applications names the agent-level version of this risk as ASI01, Agent Goal Hijack: the attacker does not need the model to say anything wrong, only to pursue a different goal with the tools it already has.12
How do we test for prompt injection?
Testing for prompt injection is not a matter of trying a list of famous prompts. Those lists measure whether a filter has seen them before. We test the structure instead, in a scoped environment and under written rules of engagement.
- Map the channels. Every place text enters the context: user turns, uploads, retrieval sources, web access, email, tickets, tool results and memory, with who can write to each.
- Map the capabilities. Every tool, token, scope and renderer the model can reach, and whose permissions each one carries.
- Pair them. For each channel and capability, we try to make content from that channel steer that capability, with probes written in our own lab for this system.
- Repeat. Model output varies between runs, so each successful attempt is replayed and reported with its reproduction rate, model, version and settings.
- Follow it into the stack. When the model is steered, we follow the request it makes: does the token behind it reach objects, records or tenants it should not?
The report names the control that would have stopped each finding, maps it to LLM01:2026 and the ATLAS technique, and gives a severity that reflects the reach, not the cleverness of the wording.
How do you defend against prompt injection?
Assume the model can be talked into anything it is able to do, and make sure that what it is able to do is safe. The defences that hold sit outside the model, where behaviour is deterministic. The NCSC recommends exactly that emphasis, plus enough logging of inputs, outputs and tool calls to see abuse when it happens.13
| Defence | What it does | What it does not do |
|---|---|---|
| Least privilege per task | Scopes each tool and token to the task and the user it serves | Stop the injection; it limits what the injection reaches |
| Authorization outside the model | Checks every object at the API, whatever the agent asks for | Help if the token itself is broader than the user |
| Human approval for risky actions | Puts a person between the model and payments, deletions and outbound messages | Scale to every action; approval fatigue is real |
| Output treated as untrusted | Validates tool arguments against a schema; encodes before rendering | Prevent disclosure in plain text |
| Separating external content | Marks retrieved text as material, not instruction | Hold reliably; it raises the effort |
| Input and output filters | Catch known patterns and obvious abuse | Stop a novel phrasing; deny-lists age fast |
| Logging and monitoring | Records inputs, outputs and tool calls for detection and forensics | Prevent anything on its own |
| Adversarial testing | Finds the paths before an attacker does, and again after each model change | Prove absence; it measures one system at one time |
Most of these appear in OWASP’s own list of mitigations for LLM01, from constraining model behaviour to requiring human approval and running adversarial tests.14 The order above is ours: the first three cap the damage, the rest lower the odds.
The takeaway fits in one line: a prompt is not a vault, and a model is not a permission system. Put the rules that matter where text cannot argue with them.
Questions people ask
Is prompt injection the same as a jailbreak?
No, but they overlap. A jailbreak tries to make a model ignore its safety training; a prompt injection tries to make an application ignore its developer’s instructions. OWASP treats jailbreaking as a form of prompt injection, while MITRE ATLAS lists LLM Jailbreak as a technique of its own, AML.T0054.
Can prompt injection be fixed completely?
Not with today’s models. The UK NCSC wrote in December 2025 that current language models do not enforce a boundary between instructions and data. The realistic goal is a lower likelihood and a smaller impact: least privilege, checks outside the model, human approval for risky actions, and testing.
Does a system prompt that forbids it protect us?
No. The system prompt is text competing with other text, and OWASP advises against treating it as a secret or as a security control. Use it to shape behaviour, and enforce the rules that matter in code: at the API, in the permissions of each token, and in what the interface renders.
Which systems are most exposed to indirect prompt injection?
Any system where a model reads content that someone else can write and can then act on it: agents that browse, read email or tickets, retrieve from a shared knowledge base or call tools. A chatbot with no tools and no private data has little to lose; an agent holding a broad token has a lot.
Where we test this
AI Pentest / Chatbot and RAG
LLM and chatbot penetration testing
Prompt injection, data leakage and cross-tenant retrieval in chatbots and RAG applications.
AI Pentest / Agent and MCP
What an agent can reach: its tools, its memory, its tokens and the systems behind them.
AI Pentest / MCP
Model Context Protocol servers and the agents that call them.
Launch Clearance / Both layers
SaaS penetration testing, AI feature included
One scope across the model layer and the stack, with the paths between them.
Sources
Every factual sentence above carries a numbered note; these are the documents behind them.