We pentest AI agents and the software they run on.

An AI security company and a penetration testing company in one: AI red teaming for chatbots, RAG and agents, human-led pentests for web, API and cloud. With written permission.

Scope a testSee a sample report

30-minute scoping call, no obligation.

Skip the reading

Slenkwater Verzekeringen, customer care knowledge base

Claims handling policy, v3.2. Applies to: support assistant, all channels. Owner: Customer Care. Reviewed 2026-09-02.

1. Scope. This policy tells our support assistant how to answer questions about open claims. It does not cover new claims or complaints.

2. Identity. Before sharing claim details, the assistant confirms the customer's policy number and date of birth. It never reads them back in full.

3. Refunds. Refunds above EUR 250 need a human agent's approval. The assistant may explain the status of a refund but cannot start one.

4. Escalation. If a customer is upset, offer a call back within one working day and log the request in the CRM with the claim number.

5. Language. Answer in the language the customer writes in. If unsure, use Dutch.

6. Records. Do not store chat transcripts outside the claims system. Transcripts are deleted after 90 days.

KB/0147 / Claims policy v3.2 / Fictional company
  1. This is a document your AI reads.
  2. This line was always there.
  3. It trusts every page equally.
  4. This is everything it can reach.
  5. We test both sides of the line.

MODEL LAYER

STACK LAYER

  • KNOWLEDGE BASE
  • RETRIEVAL
  • AGENT
  • TOOLS
    EMAIL / CRM / REFUNDS
  • MCP SERVER
  • WEB APP
  • API GATEWAY
  • IDENTITY
  • TENANT DATA
  • OBJECT STORAGE
  • CLOUD ACCOUNT
  1. 01 UNREVIEWED TEXT
    STORED AS POLICY LLM01:2026 / PROMPT
    INJECTION
  2. 02 RETRIEVED AS TRUSTED
    CONTEXT ASI06 / MEMORY AND
    CONTEXT POISONING
  3. 03 INSTRUCTION FOLLOWED LLM03:2026 / EXCESSIVE
    AGENCY
  4. 04 THIRD-PARTY TOOL,
    NEVER REVIEWED ASI04 / AGENTIC SUPPLY
    CHAIN
  5. 05 TOKEN SCOPED WIDER
    THAN THE TASK API1:2023 / BROKEN
    OBJECT LEVEL
    AUTHORIZATION
    GET /claims/48213 / 200
  6. 06 ANOTHER TENANT’S
    RECORD WSTG-ATHZ / TENANT
    ISOLATION
FIG. 0 / THE REACH. One line nobody reviewed, followed from the knowledge base through retrieval, the agent and a third-party MCP server into the API gateway and another tenant’s record.

  1. This is a document your AI reads.
  2. This line was always there.
  3. It trusts every page equally.
  4. This is everything it can reach.
  5. We test both sides of the line.

Not visible to a reviewer / read by the model

OWASP LLM01 / prompt injection, indirect

Lesson: retrieved text is data, not instructions

  1. 01 / unreviewed text stored as policy. LLM01:2026 / PROMPT INJECTION
  2. 02 / retrieved as trusted context. ASI06 / MEMORY AND CONTEXT POISONING
  3. 03 / instruction followed. LLM03:2026 / EXCESSIVE AGENCY
  4. 04 / third-party tool, never reviewed. ASI04 / AGENTIC SUPPLY CHAIN
  5. 05 / token scoped wider than the task. API1:2023 / BROKEN OBJECT LEVEL AUTHORIZATION
  6. 06 / another tenant’s record. WSTG-ATHZ / TENANT ISOLATION

Every system has a version of itself that only attackers see. We find it first, in the code and in the conversation, and turn it into a list your engineers can close.

Your product has two attack surfaces now: the code, and the conversation.

01 / Model layer

AI Pentest

What it says, shares and does when a stranger is typing.

  • Prompt injection, direct and indirect
  • System prompt and data leakage
  • RAG poisoning and cross-tenant retrieval
  • Tool misuse, excessive agency, MCP servers
  • Unsafe output handling
  • Denial of wallet

Mapped to the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications and MITRE ATLAS. Every finding carries a reproduction rate.

CHATBOT & RAG
EUR 4,000 to 10,000
AGENT & MCP
EUR 6,000 to 15,000

See AI Pentest

02 / Stack layer

Stack Pentest

What one user can do to another, and how far one weak spot reaches.

  • Web applications and business logic
  • Authorization and tenant isolation
  • REST and GraphQL APIs
  • Authentication and sessions
  • Cloud: AWS, Azure, GCP
  • External infrastructure

Mapped to the OWASP WSTG and the OWASP API Security Top 10 (2023), rated in CVSS v4.0.

WEB & API
EUR 5,000 to 15,000
CLOUD & INFRA
Quoted per scope

See Stack Pentest

30-minute scoping call.

Most tests start with a question someone just asked.

  1. A customer asked for our pentest report.

    A questionnaire, a SOC 2 window or an ISO 27001 audit is waiting on it. We scope in one call and write the report their security team expects.

  2. We’re shipping AI on our platform.

    A chatbot, RAG search or agent with real tools is about to go live on your SaaS. We test the feature and everything it can reach, before strangers do.

  3. We build agents for clients.

    Your clients want independent evidence for the agent and the app around it. We test before handover, report to you first, and never pitch your client.

Running security? Skip the doors and read the method.

Plate 01 / Fig. 1

Six places we look. Three in the model, three in the stack.

In the prompt

Who does it obey?

Direct and indirect prompt injection, jailbreaks and system prompt leakage.

Mapped to

  • LLM01:2026 / PROMPT INJECTION
  • LLM08:2026 / HIDDEN CONTEXT EXPOSURE

In the context

What does it remember, and for whom?

RAG poisoning, cross-tenant retrieval, sensitive data in embeddings and memory.

Mapped to

  • LLM02:2026 / SENSITIVE INFORMATION DISCLOSURE
  • ASI06 / MEMORY AND CONTEXT POISONING
  • LLM09:2026 / VECTOR AND EMBEDDING WEAKNESSES

In the tools

What can it do on someone else’s say-so?

Tool and function-call misuse, excessive agency, MCP server review, unsafe output handling and runaway spend.

Mapped to

  • ASI02 / TOOL MISUSE
  • ASI03 / IDENTITY AND PRIVILEGE ABUSE
  • ASI04 / AGENTIC SUPPLY CHAIN
  • LLM03:2026 / EXCESSIVE AGENCY
  • LLM06:2026 / UNBOUNDED CONSUMPTION

In the app

What can one user do to another?

Business logic, authorization, session handling and the multi-step abuse that scanners do not model.

Mapped to

  • WSTG-ATHZ / AUTHORIZATION
  • WSTG-BUSL / BUSINESS LOGIC
  • WSTG-SESS / SESSION MANAGEMENT

In the API

Which objects does it hand to whoever asks?

Object and function level authorization, property-level access, rate limits and the endpoints nobody documented.

Mapped to

  • API1:2023 / OBJECT LEVEL AUTHORIZATION
  • API3:2023 / PROPERTY LEVEL AUTHORIZATION
  • API4:2023 / RESOURCE CONSUMPTION
  • API5:2023 / FUNCTION LEVEL AUTHORIZATION
  • API9:2023 / INVENTORY MANAGEMENT

In the cloud

How far does one weak spot reach?

Identity and access, exposed storage, secrets in build pipelines, and the path from one service to the rest of the account.

Mapped to

  • IAM
  • STORAGE
  • CI/CD

MODEL LAYER

STACK LAYER

  • KNOWLEDGE BASE
  • RETRIEVAL
  • AGENT
  • TOOLS
    EMAIL / CRM / REFUNDS
  • MCP SERVER
  • WEB APP
  • API GATEWAY
  • IDENTITY
  • TENANT DATA
  • OBJECT STORAGE
  • CLOUD ACCOUNT
  • LLM01:2026
    LLM08:2026
  • LLM02:2026
    ASI06
    LLM09:2026
  • ASI02
    ASI03
    ASI04
    LLM03:2026
    LLM06:2026
  • WSTG-ATHZ
    WSTG-BUSL
    WSTG-SESS
  • API1:2023
    API3:2023
    API4:2023
    API5:2023
    API9:2023
  • IAM
    STORAGE
    CI/CD
FIG. 1 / WHERE WE LOOK. Three regions in the model layer and three in the stack layer. The selected region is hatched and keylined; its framework ids sit beside it.

Commitments

01 / 02

  1. Stack findings come with the exact requests. Model findings come with a reproduction rate.Note 1

  2. We sell no guardrail, firewall or platform. Our findings come without an upsell.Note 2

Plate 00 / Mariotte, 16681

Find your own blind spot first.

  1. Close your left eye.
  2. Look at the + with your right eye.
  3. Scroll slowly. Do not look at the dot.Move your head slowly towards the screen.
Skip
Fig. 00 / Seen from above

You did not see it go. Neither will your logs.

Somewhere along the way the dot disappeared, and the grid closed over the hole. Your brain filled the gap with what it expected to see. Every eye has a spot like this. Every system does too. That is the part we test.

SCOTOMA
/skuh-TOH-muh/3
n.
a blind spot in an otherwise normal field of view.

Every system has one. We find yours first.

We’re a new penetration testing company. Here’s how to check us anyway.

You can’t call our past clients yet. So we made everything else checkable: the method we test by, the rules we follow, a sample of the report you’d get, and this site itself.

Sample / fictional client / real format

Field Test / 24-2. 54 test categories laid out like a 24-2 visual field test. The fictional sample has 7 findings, 3 categories that do not apply and 8 scoped out, dashed, two inside the outline where the blind spot sits.
  • Critical
  • High
  • Medium
  • Low
  • Tested, nothing found
  • Not applicable
  • Out of scope

The only blind spot in our report is the one you scope out.

See a sample report

Method v1.0

  1. Scoping
  2. Threat model
  3. Testing
  4. Verification
  5. Severity
  6. Reporting
  7. Retest

v1.0 · 2026-10 · first public version

Read the method

Rules of engagement

  1. Written authorization first.

The full rules also cover test windows, how testing stops, how critical findings reach you and what happens to your data.

Rules of engagement

Standards map

  • OWASP Top 10 for LLM Applications 2026
  • OWASP Top 10 for Agentic Applications (2026)
  • MITRE ATLAS
  • OWASP WSTG
  • OWASP API Security Top 10 (2023)
  • CVSS v4.0
  • PTES

Mapped to the standards your auditors already use.

Counters

OWASP LLM categories we test
10/10
ATLAS techniques across both packages
32
Counted at build from MITRE ATLAS v2026.09.
Days since the Cyberbeveiligingswet applied
57
Counted from 15 August 2026, in your browser.

Scope builder

What would your test look like?

What are we testing?

Model layer

Stack layer

Pick ‘not sure’, scoping is our job.

How many apps or APIs?
Does the AI use tools or take actions?
Where does it run?

Used on this page only, for the percentage below. Never sent, never stored.

EUR 5,000 to 15,000See the quote
Indicative range

EUR 5,000 to 15,000

Stack Pentest / Web & API

    Testing days
    4 to 10
    Report + letter
    Included

    Indicative. The quote follows the scoping call.

    From the reading room.

    How we think about what we test, written up with the framework ids attached.

    Questions buyers ask before the call.

    AI pentest

    What is an AI pentest?

    An AI pentest is an authorized security test of a chatbot, RAG application or AI agent. We check whether it can be talked or tricked into leaking data, ignoring its instructions or misusing its tools, and we map every finding to the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications and MITRE ATLAS.

    How is it different from a regular pentest?

    A regular pentest examines code, configuration and infrastructure. An AI pentest also examines what the system does with what it reads: prompts, retrieved documents, tool results and memory. The worst issues cross into the stack, so we test both layers and report the paths between them.

    Do you need our system prompt?

    It helps and it shortens the test. Grey box access (system prompt, tool definitions, test accounts) gives the most coverage per day. We can also test black box, the way an outside attacker would, and report what we could infer on our own.

    Stack pentest

    Will our auditor and our customers accept the report?

    The report states scope, method, dates, tester qualifications, severity in CVSS v4.0 and retest status, which is what enterprise security teams and SOC 2 or ISO 27001 auditors look for. A one-page attestation letter lets you share the result without sharing the findings.

    Pricing

    What does a test cost?

    Our ranges, excluding VAT: AI chatbot or RAG EUR 4,000 to 10,000; AI agent and MCP EUR 6,000 to 15,000; web application and API EUR 5,000 to 15,000; a SaaS platform with its AI feature (Launch Clearance) EUR 10,000 to 20,000. A quote follows a 30-minute scoping call.

    Process

    Can you test a chatbot we bought from a vendor?

    Yes, with the vendor’s written permission for anything they host. We provide the authorization letter template and handle the paperwork with you.

    Regulation

    Do we need a pentest for SOC 2 or ISO 27001?

    SOC 2 does not strictly require one, but a point of focus under CC4.1 lists penetration testing among evaluation methods, and auditors and customers expect it. ISO 27001:2022 controls A.8.8 (technical vulnerabilities) and A.8.29 (security testing) are commonly evidenced with a pentest report.

    Does the Cyberbeveiligingswet (NIS2) affect us?

    If your organisation is a medium-sized or larger entity in one of its 18 sectors (some types regardless of size), it has applied since 15 August 2026. If your customers are, its supply chain duty means their security questionnaires now reach you, and a recent pentest report is the most direct answer.

    We are a bank or an insurer. Does DORA change the test?

    Yes. DORA, Regulation (EU) 2022/2554, has applied since 17 January 2025 and replaces NIS2 for financial entities on ICT risk and testing. Article 25 testing, the chatbot and AI layer included, is what we do. Threat-led testing under Article 26 is not ours yet: Article 27 sets tester requirements a firm as young as ours cannot meet.

    Does the EU AI Act require security testing of our chatbot?

    For high-risk AI systems, Article 15 requires accuracy, robustness and cybersecurity, from 2 December 2027 for Annex III systems under the AI Omnibus, Regulation (EU) 2026/1744. Most customer-service chatbots are not high-risk, but Article 50 transparency has applied since 2 August 2026. We produce test evidence; classification is for your counsel.

    Does the Cyber Resilience Act apply to SaaS?

    Mostly not. The CRA covers products with digital elements, such as installed software, apps and devices; SaaS is in scope only as the remote data processing of such a product. Whether NIS2 (the Cbw in the Netherlands) applies to you depends on your sector and size. The CRA's reporting duties started on 11 September 2026.

    Company

    Are you an independent AI security company?

    Yes. We sell no guardrail, firewall or platform, and no vendor owns us. Since 2025 platform vendors have bought several independent AI security firms, so their tests now come with a product to buy. Ours come with a list of fixes.

    Who does the testing?

    The founder leads every engagement and does the hands-on testing. Their credentials go on this site as soon as you can verify each one yourself, and not before.