Guide / Model layer
MCP security: risks and best practices.
The Model Context Protocol (MCP) lets an AI agent discover and call tools on other servers. That puts every MCP server on your attack surface twice: its tool descriptions are text the model reads and may obey, and its tokens open the systems behind it. A review checks both, then what the agent may do without asking.
Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction
Water damage from a burst pipe is covered up to the policy limit.
A claim is assessed within ten working days of the report.
A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.
Questions this page does not answer go to the claims desk.
What MCP is, and where it sits
MCP is an open protocol that connects an AI application to outside tools and data. The host is the AI application, a client inside it holds each connection, and a server offers three kinds of capability: resources (context and data), prompts (templated messages) and tools (functions the model can call). The messages are JSON-RPC 2.0, and the current revision of the specification is dated 28 July 2026.1
The specification is frank about what that means. Tools represent arbitrary code execution, descriptions of tool behaviour should be treated as untrusted unless they come from a trusted server, and the protocol cannot enforce its own security principles: implementations have to.2 In practice an MCP server sits on the line between the two layers we test. On one side it talks to the model; on the other it holds credentials for the systems behind it.
MODEL LAYER
STACK LAYER
- KNOWLEDGE BASE
- RETRIEVAL
- AGENT
- TOOLS
EMAIL / CRM / REFUNDS - MCP SERVER
- WEB APP
- API GATEWAY
- IDENTITY
- TENANT DATA
- OBJECT STORAGE
- CLOUD ACCOUNT
- LLM01:2026
LLM08:2026 - LLM02:2026
ASI06
LLM09:2026 - ASI02
ASI03
ASI04
LLM03:2026
LLM06:2026 - WSTG-ATHZ
WSTG-BUSL
WSTG-SESS - API1:2023
API3:2023
API4:2023
API5:2023
API9:2023 - IAM
STORAGE
CI/CD
Risk 1: tool descriptions are instructions
When a client connects, it asks the server for its tools, and each tool arrives with a name, a description and an input schema. The client places those descriptions in the model’s context so the model can choose a tool. To the model they are instructions like any other text it reads.
In April 2025 Invariant Labs showed what follows from that. A tool description can carry instructions that the model sees and the user never does, telling the agent to read files or send data nobody asked for. A description on one server can change how the agent uses another server’s tools, which they called shadowing. And a server can change its descriptions after a user has approved them: a rug pull.3
The specification now requires clients to treat tool annotations as untrusted unless they come from a trusted server, and it allows a server’s tool list to change over time, announced by a notification.4 MITRE ATLAS catalogues both patterns: AI Agent Tool Poisoning, which names MCP servers explicitly, and AI Supply Chain Rug Pull.5
What to do about it
- Read every description the model will see, in full and as the model receives it, not the short label your client shows.
- Pin server versions and fingerprint the tool list. Keep a hash of every name, description and schema, and treat any change as a new release that needs review.
- Alert on list changes instead of accepting them silently.
- Keep servers apart. A server you trust should not share a context with one you do not, if a description in one can steer the tools of the other.
Risk 2: tool output is untrusted input
Whatever a tool returns goes back into the model’s context: a web page, a ticket, an email, a database row. If any of that text was written by someone else, it can carry instructions. That is indirect prompt injection, part of the first entry on the OWASP Top 10 for LLM Applications 2026,6 and MCP widens it: every connected tool is another inbox that strangers can write to.
The specification puts duties on both sides. Servers must sanitize what their tools return, and clients should validate tool results before passing them to the model.7 Neither makes the problem go away, so design as if some tool output will eventually be hostile: keep untrusted output away from tools that act, and read our guide to prompt injection for the defences that hold.
Risk 3: tokens, scopes and the confused deputy
Authorization in MCP is optional. Servers on an HTTP transport should follow the specification’s OAuth 2.1 based flow; servers on the local stdio transport take their credentials from the environment instead.8 Both roads end at the same place: a server that holds keys to your email, your CRM or your cloud account.
The specification closes the most common mistake explicitly. An MCP server must check that every access token was issued for it, and must not accept or pass on any other token.8 Forwarding a client’s token unchanged to a downstream API, known as token passthrough, is forbidden, because it bypasses the server’s own checks and breaks the audit trail.9
Two more patterns from the specification’s security guidance belong in every review. A server that proxies a third-party API under one shared client identity can become a confused deputy, letting an attacker obtain an authorization code without the user’s consent; such proxies must ask for consent per client. And scopes should start small and grow when a privileged operation needs them, rather than arriving as one wildcard grant at install time.9
Local servers have a quieter version of the same problem. Their keys usually sit in a configuration file next to the agent, and MITRE ATLAS records adversaries reading agent configuration to collect exactly those credentials.10 Keep them in a secret store, give each server its own narrowly scoped key, and rotate them like any other credential.
Risk 4: the server is software, and it runs somewhere
An MCP server is code, often installed with a single command. The specification warns that a local server runs with the same privileges as the client that starts it, requires clients that offer one-click setup to show the exact command before running it, and recommends a sandbox with minimal default access to files and the network.9
The client side deserves the same care. In July 2025 CVE-2025-6514 was published for mcp-remote, a bridge between local clients and remote servers: versions 0.0.5 to 0.1.15 could be made to run operating system commands when connecting to an untrusted server, through a crafted authorization URL. It was rated 9.6, critical, and later versions are fixed.11 The specification now tells clients never to open authorization URLs through a shell and to reject dangerous URL schemes.9
Two network checks complete the picture. During OAuth discovery a malicious server can point a client at internal addresses, such as a cloud metadata endpoint, so clients should block private address ranges. And because the protocol is now stateless, a server that keeps state across calls hands out a handle, and holding that handle must never count as proof of identity.9 That last rule is an old one in a new place: an object that answers to whoever names it is broken object level authorization.12
Risk 5: what the agent may do without asking
Every risk above grows with reach. OWASP ranks excessive agency third in its 2026 list for LLM applications.13 Put plainly: the agent can do more than its task needs. In MCP terms that is three things: which tools are exposed, which scopes they run with, and whether anyone confirms an action before it happens.
The specification’s answer is a person in the loop. There should always be a human able to deny a tool call; clients should ask for confirmation on sensitive operations and show the tool’s inputs before sending them; servers must enforce access control and rate limits.7 Rate limits also cap the bill, which matters when an agent loops: see denial of wallet.
Decide per tool, not per server. Reading a calendar can run unattended. Sending an email, issuing a refund, deleting a record or changing a permission should wait for a person, and the confirmation should show what will actually be sent, not the model’s summary of it.
How to review an MCP server
A review covers the server, the client that connects to it and the agent that decides when to call it. This is the checklist we work from, in the order we work through it.
| Area | What to check | Reference |
|---|---|---|
| Inventory | Every server in use, its version, who installed it, its transport, and where its credentials live | MCP specification, overview |
| Tool metadata | Descriptions and schemas read as the model sees them; versions pinned; a fingerprint of the tool list; changes reviewed like a release | AML.T0110, AML.T0109 |
| Tokens | Audience checked on every token; no passthrough; consent per client on proxies; minimal scopes, stepped up on demand | Authorization; Security Best Practices |
| Inputs and outputs | Inputs validated by the server; outputs sanitized; untrusted output kept away from tools that act | LLM01:2026; Tools, security considerations |
| Execution | Local servers sandboxed with minimal file and network access; one-click installs show the exact command; no URLs opened through a shell; private address ranges blocked | Security Best Practices; CVE-2025-6514 |
| State | Handles bound to the authenticated user and never treated as proof of identity | API1:2023 |
| Agency | Each tool’s worst case written down; confirmation for actions that send, pay, delete or grant; rate limits per tool | LLM03:2026, ASI02 |
| Evidence | Every tool call logged with its inputs, so an incident can be reconstructed | Tools, security considerations |
The output of a review is a short list, not a score: which tools can reach what, which of those paths an outsider can influence, and what closes each one.
How we test for this
Our MCP security review is part of AI Pentest. We read every tool description your agent receives and track it across versions, test each tool’s authorization with two users, and check whether content returned by one tool can steer the agent into another. AI findings come with a reproduction rate, because a model does not answer the same way twice.
See MCP security review and AI agent security testing for scope and method.
- AI Pentest, Agent & MCP
- EUR 6,000 to 15,000, 4 to 10 testing days
Questions people ask
What is MCP tool poisoning?
Tool poisoning is hiding instructions in an MCP tool’s description, schema or output, where the model reads them but the user usually does not. The model may follow them, for example by reading files or sending data it was never asked to. MITRE ATLAS lists it as AI Agent Tool Poisoning.
Are MCP servers safe to install?
Treat them like any other code you run with your own privileges. A local MCP server runs with the rights of the client that starts it, so check who publishes it, pin the version, read its tool descriptions, run it in a sandbox with minimal file and network access, and review every update before accepting it.
Does MCP have built-in authentication?
It has an optional authorization specification, not a mandatory one. Servers on HTTP should follow its OAuth 2.1 based flow, must check that every token was issued for them and must not pass tokens through to other services. Local servers using stdio take credentials from their environment, so protect those configuration files.
What is token passthrough, and why is it forbidden?
Token passthrough is an MCP server accepting a token issued for another service and forwarding it unchanged to a downstream API. The MCP specification forbids it because the server’s own checks, rate limits and logging are bypassed, and a stolen token can then be used through the server against every service that accepts it.
How do we secure an MCP server we built ourselves?
Expose the fewest tools with the narrowest scopes, validate every input, sanitize every output and check that each token was issued for your server. Bind state handles to the authenticated user, rate limit each tool, log every call with its inputs, and require a human confirmation for tools that send, pay, delete or grant access.
Where we test this
AI Pentest / MCP
Model Context Protocol servers and the agents that call them.
AI Pentest / Agent and MCP
What an agent can reach: its tools, its memory, its tokens and the systems behind them.
Sources
Every factual sentence above carries a numbered note; these are the documents behind them.