01 / Model layerCHATBOT & RAG

LLM and chatbot penetration testing, RAG included.

LLM penetration testing, also called AI penetration testing, is a scoped, time-boxed security assessment of an application built on a large language model. It covers the model’s inputs and outputs and everything the model can reach: retrieval sources, tools, APIs, user data and other tenants. We test your chatbot or RAG application the way an outside attacker would, record every finding with a reproduction rate, and map it to the OWASP Top 10 for LLM Applications 2026 and MITRE ATLAS.

For teams shipping a customer-facing chatbot, an internal assistant, or AI search over their own documents.

Package
CHATBOT & RAG
Range, excl. VAT
EUR 4,000 to 10,000
Testing days
3 to 7

READ / FICTION / CHAT AND RETRIEVAL07 ROWS READ / 1 FOUNDSAMPLE

CHAT

  1. CUSTOMERIs water damage from a burst pipe covered?
  2. RETRIEVEDKB/0212, policy page
  3. ASSISTANTAnswers from the page it retrieved

KB/0212AS THE MODEL READS IT

  1. L1Cover for sudden escape of water
  2. L2Excess and limits per claim
  3. L3A line the page never shows.
  4. L4How to report a claim

L3 / AN INSTRUCTION IN RETRIEVED TEXT / LLM01:2026

The page the assistant retrieves carries one line a reviewer never sees and the model reads as an order. We test who can write to what the model retrieves, and whether the application treats it as data.

What do we test in an LLM application?

Ten categories, one for each risk in the OWASP Top 10 for LLM Applications 2026, each tied to the MITRE ATLAS techniques an attacker would use against it. Together they ask four questions: what the model obeys, what it reveals, what it retrieves and for whom, and what your application does with its answer.12

Every row is a category in our test catalogue. Ids link to the framework that defines them.
No.CategoryWhat we checkMapped to
01Prompt injection resistance, direct and indirectWhether text from a user, a document or a web page can override your instructions, and whether the application treats retrieved content as data rather than as orders.
02Protection of system prompts and hidden configurationWhat the assistant reveals about its own instructions, tools and configuration, and whether anything secret was ever placed where the model can repeat it.
03Sensitive information disclosure in answers and rendered outputWhether answers or rendered output can carry personal data, another customer’s details or internal documents to someone who should not see them.
04Vector store and embedding isolation between users and tenantsWhether retrieval applies the same access rules as the rest of your product, per user and per tenant, including deleted and re-indexed documents.
05Integrity of training, fine-tuning and retrieval dataWho can put text into the sources the model relies on, and whether one planted document changes what other users are told.
06Validation of model output before it reaches code, browsers or databasesWhether model output is validated before it reaches a browser, a template, a query or another system, so an answer cannot turn into markup or a command.
07Excessive agency: permissions, autonomy and tool scopeWhich actions the assistant can take, with whose permissions, and whether any of them happen without a person confirming.
08Consumption limits and denial of wallet controlsRate limits and token and cost ceilings, and whether one user can run up your model bill or starve everyone else.
09Grounding of answers and safeguards against overrelianceWhether the product shows where an answer comes from and keeps a person in the loop where a wrong answer has consequences.
10Model, dataset and component supply chainWhich models, datasets, plugins and libraries you depend on, where they come from, and who can change them.

Out of scope Model training itself, bias and toxicity benchmarks without a security question, and any system you cannot give us written permission to test.

How does an LLM penetration test run?

Five steps, in this order. The first two decide whether the other three find anything worth your time.

  1. 01Map the context

    We map what the model can read and reach.

    Before the first probe we draw the system as the model sees it: the system prompt, every retrieval source, every tool and API it can call, and every user and tenant whose data passes through it. Grey box access (the prompt, tool definitions, a test account per role) gives the most coverage per day. Black box is possible, and the report says which one we used.

  2. 02Threat model

    We decide what an attacker would want from it.

    Another customer’s data, your internal documents, an action taken on someone else’s behalf, or a large invoice. Each goal gets a test plan, mapped to OWASP ids and MITRE ATLAS techniques before testing starts, so the coverage matrix in the report is decided up front, not written up afterwards.3

  3. 03Direct and indirect

    We test the chat box and everything behind it.

    Direct paths come through the chat box. Indirect paths come through what the model retrieves: a help page, a ticket, an email, a PDF. A filter that holds in one language proves nothing about the next one, so Each probe follows from the threat model; nothing we run is a public prompt list replayed at volume.

  4. 04Reproduction rate

    We count how often it happens.

    A model can refuse on one attempt and comply on the next. Every finding is rerun and reported as n of 10 attempts, with the model version, temperature, system prompt version and language recorded, so your engineers can reproduce it and your auditor can read it. Models are probabilistic; your auditors are not.

  5. 05Into the stack

    We follow each finding as far as it goes.

    A leaked instruction matters less than what the model can do with it. Where a finding reaches a tool, an API or another tenant’s data, we follow it into the stack with the same standard of evidence and report the path as one finding. When the platform under the assistant is in scope too, that is Launch Clearance.

What does a finding look like?

One card per finding, with the same fields every time. This one comes from the fictional engagement in our sample report.

SAMPLE / FICTIONAL CLIENT / REAL FORMAT

SAMPLE-01 / F-03

HIGHCVSS-B 7.6MODEL LAYERLLM01:2026

A knowledge-base page can give the assistant instructions

Business impact
Whoever can add a page to the knowledge base can steer the assistant for every customer.
Reproduction
REPRODUCED 7/10 Model, version, settings and language recorded per attempt.
CVSS v4.0 vector
CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N
Fix principle
Treat retrieved text as data: review what enters the knowledge base, record where each page came from, and keep retrieved content out of the instruction channel.
Retest
FIXED 2026-09-24 Partner uploads are reviewed before indexing; 0 of 10 attempts at retest.
The format every finding takes in our report. The client and the finding are fiction; the fields are not. Read F-03 in the sample report.

What does an LLM or chatbot pentest cost?

The range below is our Chatbot & RAG package, excluding VAT. Most tests land inside it; the quote after the scoping call sets the exact figure. An assistant with real tools, memory or MCP servers moves to the Agent & MCP package.

01 / AI Pentest

CHATBOT & RAG

EUR 4,000 to 10,000

TYPICALLY 3 TO 7 TESTING DAYS

What moves the price

  • The number of assistants, entry points and languages in scope
  • How many retrieval sources and tenants the model can reach
  • Grey box or black box access
  • Staging, or production only within agreed windows

See all prices

Report
Scope, method, dates, findings with evidence, severity in CVSS v4.0, framework ids, fix guidance and retest status.
Attestation letter
One page that confirms scope, dates and retest status, for customers who need the result without the findings.
Coverage matrix
Every category in scope marked tested, not applicable or out of scope, so the gaps are written down too.
Evidence
Transcripts of every reproduction, with the model, settings and language of each attempt.

Questions about LLM and chatbot testing.

What does LLM penetration testing cover?

We test what the model obeys, what it reveals, what it retrieves for whom, and what your application does with its answers. That covers prompt injection, system prompt exposure, data leakage, retrieval isolation between users and tenants, output handling, excessive agency and cost limits, mapped to the OWASP Top 10 for LLM Applications 2026.

How do you test RAG for poisoning and cross-tenant leakage?

We test the two questions separately. For poisoning, we check who can put text into the sources the model relies on and whether planted content changes what other users are told. For leakage, we query as users of different roles and tenants and check that retrieval applies the same access rules as the rest of your product.

How do you report a finding that only works some of the time?

As a reproduction rate. We rerun every finding, normally ten times, and report how many attempts succeeded, with the model version, settings, prompt version and language recorded. A finding that works three times in ten is still a finding, and the retest shows whether your fix brought it down to zero.

What does an AI penetration testing report map findings to?

Every finding carries its OWASP Top 10 for LLM Applications 2026 id and the MITRE ATLAS techniques involved, plus a severity rated in CVSS v4.0 and weighed against the reproduction rate and preconditions. The coverage matrix lists every category we tested, found not applicable, or left out of scope.

What does an LLM or chatbot penetration test cost?

Our Chatbot & RAG package runs EUR 4,000 to 10,000, excluding VAT, which is typically 3 to 7 testing days. The scope sets the figure: how many assistants and languages, how many retrieval sources and tenants, and whether we get grey box access. An assistant with real tools moves to the Agent & MCP package.

The engagement, in short.

3 to 7 testing days on staging where we can, under the rules of engagement you sign first. Then the report and a one-page letter, and the retest once you have fixed. Every step, from the scoping call to the regression pack, is in the method.