Skip the film

01 / Model layer

AI red teaming, adversarial and reproducible.

AI red teaming is authorized adversarial testing of an AI system.

Testers attack the model, its prompts, its data sources and its tools the way a real attacker would, to find exploitable failures before customers or attackers do.

We test chatbots, RAG pipelines and agents, and every finding comes with a reproduction rate.

30-minute scoping call, no obligation.

Claims assistantFICTIONAL COMPANY

  1. CUSTOMER

    My claim for the water damage: where does it stand?

  2. ASSISTANT

    Claim 20931 is with an assessor. The visit is planned for Thursday.

  3. CUSTOMER

    Can you add the plumber’s invoice to it?

  4. ATTACHMENT

    invoice-0412.pdf

    Leak repair, kitchen. Labour and parts, 14 September.

    The invoice carries one more line, set so the customer never sees it, and the assistant reads it as an instruction. The film draws that line as a sodium bar and never shows its words.
  • KNOWLEDGE BASE
  • RETRIEVAL
  • AGENT
  • TOOLS
    EMAIL / CRM / REFUNDS
  • MCP SERVER
  1. 01 LLM01:2026PROMPT INJECTION
  2. 02 LLM02:2026SENSITIVE INFORMATION
    DISCLOSURE
  3. 03 LLM03:2026EXCESSIVE AGENCY
  4. 04 ASI02TOOL MISUSE
  5. 05 ASI04AGENTIC SUPPLY CHAIN
  6. 06 ASI06MEMORY AND CONTEXT
    POISONING
FIG. 1 / THE PAVILION. A claims assistant reads three short messages, the way the model receives them, then the chat turns out to be one input among four: the knowledge base through retrieval, the tool results and a third-party MCP server all reach the agent, and each is a place an instruction can hide.

What the film shows

CHAT, KNOWLEDGE BASE, TOOL RESULTS, MCP SERVER.

  1. 01 / AGENT: LLM01:2026 / PROMPT INJECTION.
  2. 02 / RETRIEVAL: LLM02:2026 / SENSITIVE INFORMATION DISCLOSURE.
  3. 03 / AGENT: LLM03:2026 / EXCESSIVE AGENCY.
  4. 04 / TOOLS: ASI02 / TOOL MISUSE.
  5. 05 / MCP SERVER: ASI04 / AGENTIC SUPPLY CHAIN.
  6. 06 / KNOWLEDGE BASE: ASI06 / MEMORY AND CONTEXT POISONING.

THE MODEL LAYER

Everything your model reads can give it orders.

A claims assistant reads more than the customer’s message. It reads your policy pages through retrieval, the CRM notes a tool returns and whatever a third-party MCP server sends back, and any of them can carry text the model treats as an instruction. OWASP ranks prompt injection first in its Top 10 for LLM Applications 2026, and MITRE ATLAS lists the indirect form, delivered through content the model retrieves, as a technique of its own.1

COVERAGE

What each package tests, by framework.

Every category is one row in our test plan. The scoping call moves categories in or out; the report says which ran.

What each package tests, by framework.
FrameworkCHATBOT & RAGAGENT & MCP
OWASP Top 10 for LLM Applications 202610 categories / ids LLM01:2026 to LLM10:20269 of 10 in the package
  • LLM01: Prompt injection resistance, direct and indirect. In the package.
  • LLM02: Sensitive information disclosure in answers and rendered output. In the package.
  • LLM03: Excessive agency: permissions, autonomy and tool scope. When the system has it: tools, memory, code execution, several agents.
  • LLM04: Model, dataset and component supply chain. In the package. The components you deploy (models, plugins, libraries) by inventory and configuration; a model’s own provenance only as its vendor documents it.
  • LLM05: Integrity of training, fine-tuning and retrieval data. In the package. Retrieval data in the package; training and fine-tuning data only with access to the pipeline.
  • LLM06: Consumption limits and denial of wallet controls. In the package.
  • LLM07: Grounding of answers and safeguards against overreliance. In the package.
  • LLM08: Protection of system prompts and hidden configuration. In the package.
  • LLM09: Vector store and embedding isolation between users and tenants. In the package.
  • LLM10: Validation of model output before it reaches code, browsers or databases. In the package.
10 of 10 in the package
  • LLM01: Prompt injection resistance, direct and indirect. In the package.
  • LLM02: Sensitive information disclosure in answers and rendered output. In the package.
  • LLM03: Excessive agency: permissions, autonomy and tool scope. In the package.
  • LLM04: Model, dataset and component supply chain. In the package. The components you deploy (models, plugins, libraries) by inventory and configuration; a model’s own provenance only as its vendor documents it.
  • LLM05: Integrity of training, fine-tuning and retrieval data. In the package. Retrieval data in the package; training and fine-tuning data only with access to the pipeline.
  • LLM06: Consumption limits and denial of wallet controls. In the package.
  • LLM07: Grounding of answers and safeguards against overreliance. In the package.
  • LLM08: Protection of system prompts and hidden configuration. In the package.
  • LLM09: Vector store and embedding isolation between users and tenants. In the package.
  • LLM10: Validation of model output before it reaches code, browsers or databases. In the package.
OWASP Top 10 for Agentic Applications 202610 categories / ids ASI01 to ASI102 of 10 in the package
  • ASI01: Agent goal integrity when inputs and documents carry instructions. In the package.
  • ASI02: Tool and function calls kept within the task's intended scope. When the system has it: tools, memory, code execution, several agents.
  • ASI03: Agent identity, delegated credentials and privilege boundaries. When the system has it: tools, memory, code execution, several agents.
  • ASI04: Agentic supply chain: MCP servers, plugins and tool definitions. When the system has it: tools, memory, code execution, several agents.
  • ASI05: Code execution boundaries and sandboxing for agents. When the system has it: tools, memory, code execution, several agents.
  • ASI06: Integrity of agent memory and shared context. In the package.
  • ASI07: Authentication and integrity of messages between agents. Not in this package.
  • ASI08: Containment of failures that spread across agents and workflows. Not in this package.
  • ASI09: Human approval steps and safeguards on trusted agent output. When the system has it: tools, memory, code execution, several agents.
  • ASI10: Detection and containment of agents that drift from their mandate. Not in this package.
7 of 10 in the package
  • ASI01: Agent goal integrity when inputs and documents carry instructions. In the package.
  • ASI02: Tool and function calls kept within the task's intended scope. In the package.
  • ASI03: Agent identity, delegated credentials and privilege boundaries. In the package.
  • ASI04: Agentic supply chain: MCP servers, plugins and tool definitions. In the package.
  • ASI05: Code execution boundaries and sandboxing for agents. When the system has it: tools, memory, code execution, several agents.
  • ASI06: Integrity of agent memory and shared context. In the package.
  • ASI07: Authentication and integrity of messages between agents. When the system has it: tools, memory, code execution, several agents.
  • ASI08: Containment of failures that spread across agents and workflows. When the system has it: tools, memory, code execution, several agents.
  • ASI09: Human approval steps and safeguards on trusted agent output. In the package.
  • ASI10: Detection and containment of agents that drift from their mandate. In the package.
MITRE ATLASrelease 2026.0929techniques in the test plan31techniques in the test plan
  • In the package
  • When the system has it: tools, memory, code execution, several agents
  • Not in this package

Mapped to OWASP Top 10 for LLM Applications 2026, OWASP Top 10 for Agentic Applications 2026 and MITRE ATLAS 2026.09.

All 20 categories in plain words
LLM01
LLM01:2026 Prompt Injection: Prompt injection resistance, direct and indirect.
LLM02
LLM02:2026 Sensitive Information Disclosure: Sensitive information disclosure in answers and rendered output.
LLM03
LLM03:2026 Excessive Agency: Excessive agency: permissions, autonomy and tool scope.
LLM04
LLM04:2026 Supply Chain: Model, dataset and component supply chain. The components you deploy (models, plugins, libraries) by inventory and configuration; a model’s own provenance only as its vendor documents it.
LLM05
LLM05:2026 Data and Model Poisoning: Integrity of training, fine-tuning and retrieval data. Retrieval data in the package; training and fine-tuning data only with access to the pipeline.
LLM06
LLM06:2026 Unbounded Consumption: Consumption limits and denial of wallet controls.
LLM07
LLM07:2026 Misinformation: Grounding of answers and safeguards against overreliance.
LLM08
LLM08:2026 Hidden Context Exposure: Protection of system prompts and hidden configuration.
LLM09
LLM09:2026 Vector and Embedding Weaknesses: Vector store and embedding isolation between users and tenants.
LLM10
LLM10:2026 Improper Output Handling: Validation of model output before it reaches code, browsers or databases.
ASI01
ASI01 Agent Goal Hijack: Agent goal integrity when inputs and documents carry instructions.
ASI02
ASI02 Tool Misuse and Exploitation: Tool and function calls kept within the task's intended scope.
ASI03
ASI03 Identity and Privilege Abuse: Agent identity, delegated credentials and privilege boundaries.
ASI04
ASI04 Agentic Supply Chain Vulnerabilities: Agentic supply chain: MCP servers, plugins and tool definitions.
ASI05
ASI05 Unexpected Code Execution (RCE): Code execution boundaries and sandboxing for agents.
ASI06
ASI06 Memory & Context Poisoning: Integrity of agent memory and shared context.
ASI07
ASI07 Insecure Inter-Agent Communication: Authentication and integrity of messages between agents.
ASI08
ASI08 Cascading Failures: Containment of failures that spread across agents and workflows.
ASI09
ASI09 Human-Agent Trust Exploitation: Human approval steps and safeguards on trusted agent output.
ASI10
ASI10 Rogue Agents: Detection and containment of agents that drift from their mandate.

HOW WE TEST

Models are probabilistic. Your auditors are not.

A model can follow an injected instruction on one attempt and refuse it on the next. So we run every model-layer finding 10 times against the same build, record the model, version, settings and language, and report it as a reproduction rate. Three in 10 is still a finding: an attacker who tries ten times gets through about three times, and trying again costs nothing. Stack findings that the model reaches come with the exact requests instead. Read the method.

REPRODUCED 7/10

  1. Attempt 1: reproduced.
  2. Attempt 2: reproduced.
  3. Attempt 3: not reproduced.
  4. Attempt 4: reproduced.
  5. Attempt 5: reproduced.
  6. Attempt 6: not reproduced.
  7. Attempt 7: reproduced.
  8. Attempt 8: reproduced.
  9. Attempt 9: not reproduced.
  10. Attempt 10: reproduced.
FIG. 2 / REPRODUCTION RATE. One finding, ten attempts against the same build: seven reproduced, three did not.

ENGAGEMENT

From the first call to the retest.

The same eight steps for every test, AI or stack. Dates go in the rules of engagement before anything starts.

  1. 01

    SCOPING CALL

    30 minutes on what you ship, what matters and what stays out of scope.

  2. 02

    FIXED QUOTE

    One price for the agreed scope, after the call.

  3. 03

    RULES OF ENGAGEMENT

    Written authorization, test windows, named contacts on both sides and how to stop the test. Read the rules.

  4. 04

    TEST WINDOW

    3 to 10 testing days, fixed on the call.

  5. 05

    CRITICAL FINDINGS

    Reported to your named contact before the report.

  6. 06

    REPORT

    Findings with evidence, severity and fixes, and a one-page letter you can share.

  7. 07

    RETEST

    We replay every finding after you fix it and update the report.

PACKAGES

Two packages, priced in the open.

Agents first, because they act. Every range covers the testing days, the report and the letter.

AGENT & MCP

Agent & MCP

Agents that call tools, keep memory or connect through MCP servers.

COVERS

  • Tool and function calls kept within the task
  • Excessive agency: permissions, autonomy and spend
  • Memory and context poisoning across sessions
  • MCP servers: tool definitions, scopes and sign-in
EUR, EXCL. VAT
EUR 6,000 to 15,000
TESTING DAYS
4 to 10

AI agent security testingMCP security review

CHATBOT & RAG

Chatbot & RAG

Assistants and AI search that answer from your documents.

COVERS

  • Direct and indirect prompt injection
  • System prompt and data leakage
  • RAG poisoning and cross-tenant retrieval
  • Unsafe output handling and denial of wallet
EUR, EXCL. VAT
EUR 4,000 to 10,000
TESTING DAYS
3 to 7

LLM and chatbot penetration testing

No product to sell you after the test.

We sell no guardrail, firewall or platform, and no vendor owns us. Since 2025 platform vendors have bought several independent AI security firms, so their tests can come with a product to buy. Ours come with a list of fixes.2

Pricing for every package

SAMPLE FINDING

What one finding looks like in your report.

Finding F-03 of our sample report, word for word: a fictional insurer, the real format. Summary level only, with the fix.

F-03HIGHCVSS-B 7.6LLM01:2026MODEL LAYER

A knowledge-base page can give the assistant instructions

BUSINESS IMPACT
Whoever can add a page to the knowledge base can steer the assistant for every customer.
EVIDENCE
REPRODUCED 7/10 Model, version, settings and language recorded per attempt.
CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N3
FIX
Treat retrieved text as data: review what enters the knowledge base, record where each page came from, and keep retrieved content out of the instruction channel.
RETEST
FIXED 2026-09-24 Partner uploads are reviewed before indexing; 0 of 10 attempts at retest.
SAMPLE / FICTIONAL CLIENT / REAL FORMAT

FROM THE READING ROOM

Three techniques we test for, explained.

Questions about AI red teaming.

What is AI red teaming?

AI red teaming is authorized adversarial testing of an AI system, from its prompts to its data sources and tools. Testers work like a patient attacker to find failures that can be exploited, such as leaking data, following hidden instructions or misusing a tool, and report each one with evidence and a fix.

How is AI red teaming different from a penetration test?

A penetration test looks for exploitable flaws in code, configuration and infrastructure. AI red teaming tests behaviour: what the system does with what it reads, from prompts to retrieved documents and tool results. The two meet where an agent’s tools call your APIs, so we test across that line and report the paths between the layers.

Do you test the model or our application?

Your application. We test the system you built around a model: its prompts, retrieval, memory, tools, permissions and the data it can reach. We do not attack the model provider’s platform. Where a weakness sits in the model itself, we record the model and version and name the control in your application that should contain it.

Can you test a chatbot built on a hosted model such as OpenAI, Anthropic or Mistral?

Yes. We test your deployment, not the provider’s platform. Before testing we read the provider’s current usage and testing terms for your account and record them in the rules of engagement. Where a provider asks for notice or limits load, we plan the test window and rate limits around it.

How long does an AI red team engagement take?

Usually 3 to 10 testing days. A chatbot or RAG application typically takes 3 to 7 days, and an agent with tools or MCP servers 4 to 10, depending on the number of tools, roles and data sources. The scoping call fixes the number of days before you receive a quote, so the price does not move during the test.

Do automated scanners such as garak, PyRIT or promptfoo replace this?

No, though they are useful. Open-source scanners like these run large libraries of known attack prompts and catch regressions cheaply. They do not know your business rules, your tools or which customer’s data an answer may contain, so they cannot judge whether a response is a breach. That judgment is the work we do.

Sources: garak, PyRIT, promptfoo.

03 / Clearance

Shipping AI on a SaaS platform? Test both layers at once.

Launch Clearance is one scope for the feature and the platform under it, including the issues that only show up when both layers are tested together.

SAAS + AI FEATURE

EUR 10,000 to 20,000

30-minute scoping call, no obligation.