01 / Model layerAGENT & MCP

AI agent security testing: tools, memory and reach.

AI agent security testing is an authorized attack on an AI system that plans and acts: it calls tools, keeps memory and works with delegated credentials. We test what the agent can be made to do on a stranger’s say-so, with whose permissions, and how far one manipulated step reaches into your APIs and data. Findings map to the OWASP Top 10 for Agentic Applications (2026), the OWASP Top 10 for LLM Applications 2026 and MITRE ATLAS.

For teams whose AI does more than answer: refunds, bookings, tickets, code, email, anything with a tool behind it.

Package
AGENT & MCP
Range, excl. VAT
EUR 6,000 to 15,000
Testing days
4 to 10

TRACE / FICTION / ONE TASK, THREE TOOL CALLS03 ROWS READ / 1 FOUNDSAMPLE

TASK: ADD THE INVOICE TO CLAIM 20931

  1. 01 lookup_claimCLAIM IN THIS SESSION / IN TASK
  2. 02 upload_documentCLAIM IN THIS SESSION / IN TASK
  3. 03 request_refundANOTHER CLAIM / OUTSIDE THE TASK

CALL 03 / TOOL USED OUTSIDE ITS TASK / ASI02

Two calls do what the customer asked; the third reaches a claim the conversation never named. We test which tools an agent can be talked into, with whose permissions, and whether the server binds each call to the task.

How does an AI agent security test run?

An agent is a chain of decisions. We test each link, then the chain.

  1. 01Inventory

    We list everything the agent is allowed to do.

    Every tool, its arguments, the credentials behind it and the data it returns; every memory store and who writes to it; every agent it hands work to. A permission the task never needed is a finding before the first test runs, so the inventory is part of the report, not a preamble to it.

  2. 02What it reads

    We test the indirect paths first.

    An agent is rarely attacked through its chat box. We place our own labelled test content where the agent will read it in normal work, such as a ticket, a document or a tool result, and observe whether it changes the agent’s goal. Every piece of test content is listed in the rules of engagement and removed at the end.12

  3. 03Tools, safely

    We exercise every tool against test data, never real customers.

    Tool calls run against test accounts and test records named in advance in the rules of engagement. Destructive actions are tested through dry runs or reversible test objects, and anything irreversible needs your written go-ahead first.

  4. 04Delegated credentials

    We follow the token.

    When the agent calls an API on a user’s behalf, we check that the token is scoped to that user and that task, then test what it opens beyond them. This is where an agent finding becomes a stack finding: the API must check object ownership against the end user, not against the agent.3

  5. 05Rate and containment

    We record how often it worked, and what would have stopped it.

    Each finding carries a reproduction rate, with the model, settings and tool versions recorded, and names the control that breaks it earliest: a narrower scope, a confirmation step, a sandbox or a limit. Breaking the chain early usually takes one change; patching every symptom takes many.

What does an agent finding look like?

One card per finding, with the same fields every time. This one comes from the fictional engagement in our sample report.

SAMPLE / FICTIONAL CLIENT / REAL FORMAT

SAMPLE-01 / F-02

HIGHCVSS-B 8.2MODEL LAYERASI02

The assistant can call its refund tool for a claim it is not handling

Business impact
A refund request can be raised on a claim the customer does not own. It stops at the human approval step, but it reaches the queue.
Reproduction
REPRODUCED 6/10 Model, version, settings and language recorded per attempt.
CVSS v4.0 vector
CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:N/VI:H/VA:N/SC:N/SI:N/SA:N
Fix principle
Give tools the least privilege the task needs: bind the tool to the claim in the session, not to an argument the model fills in.
Retest
FIXED 2026-09-24 0 of 10 attempts at retest.
The format every finding takes in our report. The client and the finding are fiction; the fields are not. Read F-02 in the sample report.

What do we test in an AI agent?

The ten risks in the OWASP Top 10 for Agentic Applications, plus the two LLM risks that turn dangerous once a model can act: excessive agency and unbounded consumption. Each row names what we check and the ids it maps to.4

Every row is a category in our test catalogue. Ids link to the framework that defines them.
No.CategoryWhat we checkMapped to
01Agent goal integrity when inputs and documents carry instructionsWhether a document, email, ticket or tool result can change what the agent is trying to achieve, and whether anything notices when it does.
02Tool and function calls kept within the task's intended scopeWhether each tool call stays inside the task the user asked for, with arguments the user could have chosen themselves.
03Agent identity, delegated credentials and privilege boundariesWhose credentials the agent acts with, whether they are scoped to the task, and whether one user’s session can borrow another’s.
04Integrity of agent memory and shared contextWho can write to the agent’s memory and shared context, and whether a poisoned entry follows other users into later sessions.
05Agentic supply chain: MCP servers, plugins and tool definitionsWhere tools, plugins and MCP servers come from, whether their definitions can change after you approved them, and who reviews the change.
06Code execution boundaries and sandboxing for agentsWhether code the agent writes or runs is sandboxed, with no route to secrets, the network or the host.
07Authentication and integrity of messages between agentsWhether agents authenticate each other, and whether a message between them can be forged, replayed or altered on the way.
08Containment of failures that spread across agents and workflowsWhether a failure in one step, tool or agent stays contained, or is repeated by every agent downstream of it.
09Human approval steps and safeguards on trusted agent outputWhich actions need a person’s approval, and whether the approval screen shows what will actually happen.
10Detection and containment of agents that drift from their mandateWhether you would notice an agent acting outside its mandate: the logs, the limits and a way to stop it.
11Excessive agency: permissions, autonomy and tool scopeWhich tools and permissions the agent holds that its task does not need.
12Consumption limits and denial of wallet controlsLoop limits, spend ceilings and timeouts, so a manipulated agent cannot run up cost or run forever.

Out of scope The agent framework’s own source code unless you maintain it, third-party services behind your tools without their owner’s written permission, and model training.

What does an AI agent security test cost?

The range below is our Agent & MCP package, excluding VAT, which also covers an MCP security review. The quote after the scoping call sets the exact figure.

01 / AI Pentest

AGENT & MCP

EUR 6,000 to 15,000

TYPICALLY 4 TO 10 TESTING DAYS

What moves the price

  • The number of tools, and how many of them write or spend
  • Memory stores and agent-to-agent hand-offs in scope
  • MCP servers you run, and third-party ones you depend on
  • Whether tool calls can run against test data in staging

See all prices

Report
Scope, method, dates, findings with evidence, severity in CVSS v4.0, framework ids, fix guidance and retest status.
Attestation letter
One page that confirms scope, dates and retest status, for customers who need the result without the findings.
Coverage matrix
Every category in scope marked tested, not applicable or out of scope, so the gaps are written down too.
Evidence
Tool-call traces for every reproduction, with the tool and model versions.

The engagement, in short.

4 to 10 testing days on staging where we can, under the rules of engagement you sign first. Then the report and a one-page letter, and the retest once you have fixed. Every step, from the scoping call to the regression pack, is in the method.

Questions about AI agent testing.

Which agent frameworks do you test?

We test the agent, whatever it is built with. Agents built on LangChain, LangGraph, LlamaIndex, Semantic Kernel, a model provider’s agent SDK or your own orchestration code reduce to the same questions: what it reads, which tools it calls with which credentials, what it remembers, and who approves its actions.

How do you test tool misuse without damaging production?

With test data and agreed limits. Tool calls run against test accounts and records named in the rules of engagement, destructive actions are tested through dry runs or reversible objects, and anything irreversible needs your written go-ahead first. We prefer staging; testing in production follows the windows and rate limits you set.

Do you test agent memory and multi-agent setups?

Yes. We check who can write to long-term memory and shared context, whether a poisoned entry carries into other users’ sessions, and whether agents authenticate the messages they pass to each other. These are ASI06 and ASI07 in the OWASP Top 10 for Agentic Applications, and each finding records the path it took.

What is excessive agency?

Excessive agency is the risk that an AI system can take damaging actions because it has more tools, permissions or autonomy than its task needs. OWASP ranks it LLM03 in the 2026 Top 10 for LLM Applications. We test it by comparing what the agent may do with what its task actually requires.

What does an AI agent security test cost?

Our Agent & MCP package runs EUR 6,000 to 15,000, excluding VAT, which is typically 4 to 10 testing days. The number of tools that write or spend, the memory stores and agent hand-offs in scope, and the MCP servers involved set the figure. An assistant without tools fits the smaller Chatbot & RAG package.