Technique / LLM01:2026 / AML.T0051

What is prompt injection?

Prompt injection is an attack in which text that a large language model reads, typed by a user or hidden in a document, email or web page, overrides the instructions its developer gave it. OWASP ranks it first among the risks to LLM applications in 2026, as it did in 2025.

Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction

Water damage from a burst pipe is covered up to the policy limit.

A claim is assessed within ten working days of the report.

A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.

Questions this page does not answer go to the claims desk.

A page an assistant retrieves, as a reviewer sees it. Read it as the model does, and it has one more line.

Prompt injection, defined

Prompt injection is an attack on an application built on a large language model (LLM): text the model processes changes what the model does, against the intent of the people who built the application. The text can come from the person in the chat, or from anything the model is given to read.

The name dates from September 2022, when Simon Willison proposed it and drew the obvious parallel with SQL injection: in both cases, data that was meant to be read ends up being obeyed.1 The parallel is useful for the first five minutes and misleading after that, for reasons this piece comes back to.

OWASP defines a prompt injection vulnerability as one where prompts alter the model’s behaviour or output in unintended ways, and places it first in its Top 10 for LLM Applications.2 It held that position in the 2025 edition and keeps it, as LLM01:2026, in the edition published on 3 August 2026.3

MITRE ATLAS catalogues the same attack as technique AML.T0051, LLM Prompt Injection, and describes it as a possible initial access vector: a foothold from which an attacker carries out the next steps of an operation.4

What is the difference between direct and indirect prompt injection?

In a direct injection the attacker is the user. They type into the chat box and try to talk the application out of its instructions. In an indirect injection the attacker never talks to the system at all: they place text where the system will read it later, in a web page, a shared document, an email, a support ticket or the output of a tool.5

The indirect form was set out in 2023 by Greshake and colleagues, who showed it against real LLM-integrated applications and argued that such applications blur the line between data and instructions.6 ATLAS has since added a third sub-technique, Triggered, for injections set off by a user action or an event in the victim’s own environment.7

Direct and indirect prompt injection compared
DirectIndirect
Who writes the textThe user of the systemAnyone who can write where the system reads
ChannelThe chat or API inputRetrieved documents, web pages, email, tickets, files, tool results, memory
Who is harmedMostly the operator: policy, cost, reputationOften another user, whose request triggers it with their permissions
What limits itWhat the user may already see and doWhat the system may see and do on the victim’s behalf

That last row is the reason indirect injection matters most for agents. A direct attacker can only abuse their own session. An indirect attacker borrows someone else’s: the agent reads the planted text while working for a legitimate user, and acts with that user’s access. Fig. 1 follows one such line through a fictional insurer’s support agent, from its knowledge base into the stack beneath it.

MODEL LAYER

STACK LAYER

  • KNOWLEDGE BASE
  • RETRIEVAL
  • AGENT
  • TOOLS
    EMAIL / CRM / REFUNDS
  • MCP SERVER
  • WEB APP
  • API GATEWAY
  • IDENTITY
  • TENANT DATA
  • OBJECT STORAGE
  • CLOUD ACCOUNT
  1. 01 UNREVIEWED TEXT
    STORED AS POLICY LLM01:2026 / PROMPT
    INJECTION
  2. 02 RETRIEVED AS TRUSTED
    CONTEXT ASI06 / MEMORY AND
    CONTEXT POISONING
  3. 03 INSTRUCTION FOLLOWED LLM03:2026 / EXCESSIVE
    AGENCY
  4. 04 THIRD-PARTY TOOL,
    NEVER REVIEWED ASI04 / AGENTIC SUPPLY
    CHAIN
  5. 05 TOKEN SCOPED WIDER
    THAN THE TASK API1:2023 / BROKEN
    OBJECT LEVEL
    AUTHORIZATION
    GET /claims/48213 / 200
  6. 06 ANOTHER TENANT’S
    RECORD WSTG-ATHZ / TENANT
    ISOLATION
FIG. 1 / THE REACH. One line nobody reviewed, stored in a fictional insurer’s knowledge base, retrieved as trusted context, obeyed by the agent, passed through a third-party MCP server and into the API gateway, which hands over another tenant’s record.

Why does prompt injection work at all?

A model receives one stream of tokens. The developer’s instructions, the user’s message, a retrieved document and a tool result all arrive in that same stream, and nothing in it marks which parts are commands and which are material to work on (Fig. 2). Models are trained to give the developer’s instructions more weight, but that is a tendency learned from examples, not a rule the system enforces.

Developer instructions

User message

Retrieved document

Tool result

One context window / no type on any row

The highlighted row came from a document. Nothing marks it as data.

Fig. 2 / One context window Four sources, one stream. A sentence in the retrieved document is phrased as an instruction; to the model it is one more row of text.

The UK National Cyber Security Centre made the point bluntly in December 2025: current LLMs do not enforce a security boundary between instructions and data, prompt injection may never be totally mitigated, and the realistic aim is to reduce the likelihood and the impact of attacks.8 Its author calls the model an inherently confusable deputy. That is where the SQL comparison breaks: SQL injection has a structural fix, the parameterised query, which keeps data out of the command channel. Language models have no equivalent channel to keep it out of.

OWASP draws the same conclusion from another angle: retrieval-augmented generation and fine-tuning make outputs more relevant, but they do not fully mitigate prompt injection.9 So the useful question is not how to make the model unpersuadable. It is what a persuaded model is able to do.

What can an attacker do with prompt injection?

The impact of a prompt injection is the reach of the system it lands in. The model’s words are rarely the damage; what the application does with them is. The consequences fall into four groups.

  • Disclosure. The model reveals what it was given but should not pass on: hidden instructions, configuration, another user’s data, documents from the wrong tenant. OWASP is explicit that a system prompt should not be treated as a secret or used as a security control.10
  • Actions. An agent calls the tools it holds on the attacker’s behalf: it sends the email, files the refund, updates the record. ATLAS describes how tool invocation can give an attacker the privileges of the integrated service, and how a tool that writes can carry data out inside a legitimate-looking action.11
  • Output. The model’s answer is rendered or executed downstream without being treated as untrusted input, so the injection becomes a classic web or code execution flaw. OWASP files this under LLM10:2026, Improper Output Handling.3
  • Persistence. The instruction is written into memory or a shared store, and keeps steering the system long after the original content is gone.

The OWASP Top 10 for Agentic Applications names the agent-level version of this risk as ASI01, Agent Goal Hijack: the attacker does not need the model to say anything wrong, only to pursue a different goal with the tools it already has.12

How do we test for prompt injection?

Testing for prompt injection is not a matter of trying a list of famous prompts. Those lists measure whether a filter has seen them before. We test the structure instead, in a scoped environment and under written rules of engagement.

  1. Map the channels. Every place text enters the context: user turns, uploads, retrieval sources, web access, email, tickets, tool results and memory, with who can write to each.
  2. Map the capabilities. Every tool, token, scope and renderer the model can reach, and whose permissions each one carries.
  3. Pair them. For each channel and capability, we try to make content from that channel steer that capability, with probes written in our own lab for this system.
  4. Repeat. Model output varies between runs, so each successful attempt is replayed and reported with its reproduction rate, model, version and settings.
  5. Follow it into the stack. When the model is steered, we follow the request it makes: does the token behind it reach objects, records or tenants it should not?

The report names the control that would have stopped each finding, maps it to LLM01:2026 and the ATLAS technique, and gives a severity that reflects the reach, not the cleverness of the wording.

How do you defend against prompt injection?

Assume the model can be talked into anything it is able to do, and make sure that what it is able to do is safe. The defences that hold sit outside the model, where behaviour is deterministic. The NCSC recommends exactly that emphasis, plus enough logging of inputs, outputs and tool calls to see abuse when it happens.13

Defences against prompt injection
DefenceWhat it doesWhat it does not do
Least privilege per taskScopes each tool and token to the task and the user it servesStop the injection; it limits what the injection reaches
Authorization outside the modelChecks every object at the API, whatever the agent asks forHelp if the token itself is broader than the user
Human approval for risky actionsPuts a person between the model and payments, deletions and outbound messagesScale to every action; approval fatigue is real
Output treated as untrustedValidates tool arguments against a schema; encodes before renderingPrevent disclosure in plain text
Separating external contentMarks retrieved text as material, not instructionHold reliably; it raises the effort
Input and output filtersCatch known patterns and obvious abuseStop a novel phrasing; deny-lists age fast
Logging and monitoringRecords inputs, outputs and tool calls for detection and forensicsPrevent anything on its own
Adversarial testingFinds the paths before an attacker does, and again after each model changeProve absence; it measures one system at one time

Most of these appear in OWASP’s own list of mitigations for LLM01, from constraining model behaviour to requiring human approval and running adversarial tests.14 The order above is ours: the first three cap the damage, the rest lower the odds.

The takeaway fits in one line: a prompt is not a vault, and a model is not a permission system. Put the rules that matter where text cannot argue with them.

Questions people ask

Is prompt injection the same as a jailbreak?

No, but they overlap. A jailbreak tries to make a model ignore its safety training; a prompt injection tries to make an application ignore its developer’s instructions. OWASP treats jailbreaking as a form of prompt injection, while MITRE ATLAS lists LLM Jailbreak as a technique of its own, AML.T0054.

Can prompt injection be fixed completely?

Not with today’s models. The UK NCSC wrote in December 2025 that current language models do not enforce a boundary between instructions and data. The realistic goal is a lower likelihood and a smaller impact: least privilege, checks outside the model, human approval for risky actions, and testing.

Does a system prompt that forbids it protect us?

No. The system prompt is text competing with other text, and OWASP advises against treating it as a secret or as a security control. Use it to shape behaviour, and enforce the rules that matter in code: at the API, in the permissions of each token, and in what the interface renders.

Which systems are most exposed to indirect prompt injection?

Any system where a model reads content that someone else can write and can then act on it: agents that browse, read email or tickets, retrieve from a shared knowledge base or call tools. A chatbot with no tools and no private data has little to lose; an agent holding a broad token has a lot.

Where we test this

Sources

Every factual sentence above carries a numbered note; these are the documents behind them.

  1. 1

    Prompt injection attacks against GPT-3

    Simon Willison / 2022-09-12 / checked 2026-10-10

  2. 25914

    LLM01:2025 Prompt Injection

    OWASP GenAI Security Project / checked 2026-10-10

  3. 3

    OWASP Top 10 for LLM Applications 2026

    OWASP GenAI Security Project / 2026-08-03 / checked 2026-10-09

  4. 4711

    MITRE ATLAS data, release v2026.09 (ATLAS-2026.09.yaml)

    MITRE / 2026-09-15 / checked 2026-10-10

  5. 6

    Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

    Greshake, Abdelnabi, Mishra, Endres, Holz and Fritz (arXiv 2302.12173) / 2023-02-23 / checked 2026-10-10

  6. 813

    Prompt injection is not SQL injection (it may be worse)

    UK National Cyber Security Centre / 2025-12-08 / checked 2026-10-10

  7. 10

    LLM07:2025 System Prompt Leakage

    OWASP GenAI Security Project / checked 2026-10-10

  8. 12

    OWASP Top 10 for Agentic Applications for 2026

    OWASP GenAI Security Project / 2025-12-09 / checked 2026-10-09