Framework / OWASP GenAI / 2026 edition
The OWASP Top 10 for LLM Applications 2026, explained.
The OWASP Top 10 for LLM Applications 2026, the OWASP LLM Top 10 for short, is the OWASP GenAI Security Project’s ranked list of the most critical security risks in applications built on large language models, published on 3 August 2026. Prompt injection stays first, excessive agency rises to third, and only the first two entries keep their 2025 numbers.
Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction
Water damage from a burst pipe is covered up to the policy limit.
A claim is assessed within ten working days of the report.
A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.
Questions this page does not answer go to the claims desk.
What is the OWASP Top 10 for LLM Applications?
The OWASP Top 10 for LLM Applications is a ranked list of ten risk categories for software built on large language models: chatbots, retrieval-augmented generation (RAG) applications and agents. It names the risks, gives each an id, and lets builders, testers and buyers talk about the same thing.
It is published by the OWASP GenAI Security Project, which describes the 2026 edition as its community-driven guide to the most critical security risks facing applications powered by large language models.1 The edition is dated 3 August 2026 and maps its risks to NIST, MITRE ATLAS, CWE and the OWASP Top 10 for Agentic Applications.
What makes 2026 different is method. OWASP calls it the first edition to weigh the community’s ranking against a large body of real incident data, drawn from thousands of reported AI security incidents.2 Where expert judgement and the incident record agreed most clearly, the project says, was on excessive agency, now at number three.
Two things the list is not. It is not a test standard: it names categories of risk, not the steps to verify them. And it is not a certificate: an application cannot be compliant with it, only tested against it.
What changed between 2025 and 2026?
The ten categories are almost the same; their order is not. Only LLM01 and LLM02 keep their 2025 numbers (Fig. 1).3
Rank changes from 2025 to 2026: Prompt Injection 1 to 1, Sensitive Information Disclosure 2 to 2, Supply Chain 3 to 4, Data and Model Poisoning 4 to 5, Improper Output Handling 5 to 10, Excessive Agency 6 to 3, Vector and Embedding Weaknesses 8 to 9, Misinformation 9 to 7, Unbounded Consumption 10 to 6. System Prompt Leakage, 7 in 2025, has no entry of that name in 2026. Hidden Context Exposure, 8 in 2026, is new.
- 01Prompt Injection
- 02Sensitive Information Disclosure
- 03Supply Chain
- 04Data and Model Poisoning
- 05Improper Output Handling
- 06Excessive Agency
- 07System Prompt LeakageNot in 2026
- 08Vector and Embedding Weaknesses
- 09Misinformation
- 10Unbounded Consumption
- 01Prompt Injection
- 02Sensitive Information Disclosure
- 03Excessive Agency
- 04Supply Chain
- 05Data and Model Poisoning
- 06Unbounded Consumption
- 07Misinformation
- 08Hidden Context ExposureNew name
- 09Vector and Embedding Weaknesses
- 10Improper Output Handling
Our reading of the moves: the risks that climb are the ones that grow with what a model is allowed to do. Excessive agency and unbounded consumption rise as models gain tools and budgets, and send, buy, book and call other services. Improper output handling falls from fifth to tenth. It has not gone away; it is the part of LLM security that ordinary web testing already covers best.
The practical consequence is about ids. LLM06 meant Excessive Agency in a 2025 report and means Unbounded Consumption in a 2026 one. A control matrix or a security questionnaire mapped to bare numbers now points at the wrong risks. Always cite the edition: LLM03:2026, never LLM03 alone.
The ten risks, one by one
The names and the order below are OWASP’s. The explanations are ours: what each risk looks like in a chatbot, a RAG application or an agent, and how we test for it.
LLM01:2026 Prompt Injection
Text the model reads overrides the instructions its developer gave it, typed in the chat (direct) or planted in a page, file or email the model will process (indirect).4 We test every channel that feeds the context against every capability the model holds; our prompt injection explainer walks through it.
LLM02:2026 Sensitive Information Disclosure
The model reveals data it holds or can fetch: personal data, secrets left in prompts, documents belonging to another customer. In RAG the usual cause is retrieval that does not check the asking user’s permissions. We test as one user for what only another may see.
LLM03:2026 Excessive Agency
The application lets the model take damaging actions in response to unexpected, ambiguous or manipulated output. OWASP traces it to excessive functionality, excessive permissions or excessive autonomy.5 We list every tool and scope against the task it serves, then try to make each one act beyond the user’s intent.
LLM04:2026 Supply Chain
Anything the application inherits can carry a flaw: models, adapters, datasets, packages, plugins and, for agents, tool servers such as MCP. We inventory what is pulled in, from where and at which version, and review the tools an agent trusts.
LLM05:2026 Data and Model Poisoning
Manipulated training, fine-tuning or knowledge-base content changes how the model behaves. For most software teams the reachable part is the data they feed it, which ATLAS covers as RAG Poisoning and Training Data Poisoning.6 We trace who can write to each source and whether anything checks it before the model sees it.
LLM06:2026 Unbounded Consumption
Inference without limits: long inputs, runaway agent loops and expensive tool calls that exhaust capacity or budget. OWASP’s 2025 entry already listed denial of wallet among its examples.7 We test the caps at every layer; see our piece on denial of wallet.
LLM07:2026 Misinformation
Confident, wrong output that people or systems rely on: an invented policy, a citation that does not exist, a software package that was never published. ATLAS records attackers who look for such hallucinations on purpose.8 We test whether the application grounds high-stakes answers and shows how sure it is.
LLM08:2026 Hidden Context Exposure
The 2025 list had no entry by this name; its nearest relative was System Prompt Leakage, where OWASP warned that a system prompt should be neither a secret nor a security control.9 We read the 2026 name as it tests: everything in the model’s context that a user is not meant to see, from instructions and tool definitions to retrieved passages and other users’ memory.
LLM09:2026 Vector and Embedding Weaknesses
Weaknesses in how a RAG system stores and searches its vectors: indexes shared across tenants without access control, write paths that let outsiders add content, embeddings that leak what they encode. We test retrieval across tenants and every route into the index.
LLM10:2026 Improper Output Handling
Model output passed to a browser, a shell, a database or another service without being treated as untrusted input. The result is a familiar web flaw with a new source. We treat every output path as an input path, renderers and tool arguments included.
How should you use the list?
Use it as a coverage map, not a scorecard. Before a test, it shows which categories apply to your system: a chatbot without tools has little exposure to excessive agency, an agent with a payment tool has a lot. After a test, it gives every finding a name the next reader recognises.
Pair it with two others. The OWASP Top 10 for Agentic Applications, published in December 2025, covers what changes when models plan and act, from ASI01 Agent Goal Hijack onwards.10 MITRE ATLAS describes how attackers actually operate against AI systems, technique by technique; our MITRE ATLAS explainer shows how its ids become a test plan.
And keep one sentence in mind when you read any report mapped to it, ours included: a clean result against a list is evidence of what was tested, not proof that nothing else is there.
Questions people ask
What is new in the OWASP LLM Top 10 2026?
Three moves stand out. Excessive Agency rose from sixth to third, Unbounded Consumption from tenth to sixth, and Hidden Context Exposure appears at eighth, where System Prompt Leakage stood seventh in 2025. Only LLM01 Prompt Injection and LLM02 Sensitive Information Disclosure keep their 2025 numbers.
Is the OWASP LLM Top 10 a compliance standard?
No. It is a ranked list of risk categories, published as community guidance by the OWASP GenAI Security Project, and it certifies nothing. It works best as a shared vocabulary: findings mapped to its ids, with the year attached, are easier to compare, to prioritise and to act on.
How does it relate to the OWASP Top 10 for Agentic Applications?
They are companions. The LLM list covers risks in any application built on a language model. The Agentic list, published in December 2025, covers what changes when the model plans and acts with tools, memory and other agents, from ASI01 Agent Goal Hijack to ASI10 Rogue Agents.
Which ids should a pentest report use?
The full id with the edition, such as LLM03:2026, because the numbers moved: LLM06 meant Excessive Agency in 2025 and means Unbounded Consumption in 2026. We map each finding to its LLM id, to the Agentic id where an agent is involved, and to the MITRE ATLAS technique.
Where we test this
AI Pentest / Chatbot and RAG
LLM and chatbot penetration testing
Prompt injection, data leakage and cross-tenant retrieval in chatbots and RAG applications.
AI Pentest / Agent and MCP
What an agent can reach: its tools, its memory, its tokens and the systems behind them.
AI Pentest / Model layer
Adversarial testing of chatbots, RAG applications and agents, mapped to OWASP and MITRE ATLAS.
Launch Clearance / Both layers
SaaS penetration testing, AI feature included
One scope across the model layer and the stack, with the paths between them.
Sources
Every factual sentence above carries a numbered note; these are the documents behind them.