Guide / Model layer

MCP security: risks and best practices.

The Model Context Protocol (MCP) lets an AI agent discover and call tools on other servers. That puts every MCP server on your attack surface twice: its tool descriptions are text the model reads and may obey, and its tokens open the systems behind it. A review checks both, then what the agent may do without asking.

Retrieved / KB/0212 / Slenkwater Verzekeringen / fiction

Water damage from a burst pipe is covered up to the policy limit.

A claim is assessed within ten working days of the report.

A line a reviewer does not see and the model reads as an order. Its text is not shown on this site.

Questions this page does not answer go to the claims desk.

A page an assistant retrieves, as a reviewer sees it. Read it as the model does, and it has one more line.

What MCP is, and where it sits

MCP is an open protocol that connects an AI application to outside tools and data. The host is the AI application, a client inside it holds each connection, and a server offers three kinds of capability: resources (context and data), prompts (templated messages) and tools (functions the model can call). The messages are JSON-RPC 2.0, and the current revision of the specification is dated 28 July 2026.1

The specification is frank about what that means. Tools represent arbitrary code execution, descriptions of tool behaviour should be treated as untrusted unless they come from a trusted server, and the protocol cannot enforce its own security principles: implementations have to.2 In practice an MCP server sits on the line between the two layers we test. On one side it talks to the model; on the other it holds credentials for the systems behind it.

MODEL LAYER

STACK LAYER

  • KNOWLEDGE BASE
  • RETRIEVAL
  • AGENT
  • TOOLS
    EMAIL / CRM / REFUNDS
  • MCP SERVER
  • WEB APP
  • API GATEWAY
  • IDENTITY
  • TENANT DATA
  • OBJECT STORAGE
  • CLOUD ACCOUNT
  • LLM01:2026
    LLM08:2026
  • LLM02:2026
    ASI06
    LLM09:2026
  • ASI02
    ASI03
    ASI04
    LLM03:2026
    LLM06:2026
  • WSTG-ATHZ
    WSTG-BUSL
    WSTG-SESS
  • API1:2023
    API3:2023
    API4:2023
    API5:2023
    API9:2023
  • IAM
    STORAGE
    CI/CD
FIG. 1 / WHERE AN MCP SERVER SITS. Our system plate with the tools region selected: the MCP server stands on the line between the model layer and the stack layer, beside the agent’s own tools.

Risk 1: tool descriptions are instructions

When a client connects, it asks the server for its tools, and each tool arrives with a name, a description and an input schema. The client places those descriptions in the model’s context so the model can choose a tool. To the model they are instructions like any other text it reads.

In April 2025 Invariant Labs showed what follows from that. A tool description can carry instructions that the model sees and the user never does, telling the agent to read files or send data nobody asked for. A description on one server can change how the agent uses another server’s tools, which they called shadowing. And a server can change its descriptions after a user has approved them: a rug pull.3

The specification now requires clients to treat tool annotations as untrusted unless they come from a trusted server, and it allows a server’s tool list to change over time, announced by a notification.4 MITRE ATLAS catalogues both patterns: AI Agent Tool Poisoning, which names MCP servers explicitly, and AI Supply Chain Rug Pull.5

What to do about it

  • Read every description the model will see, in full and as the model receives it, not the short label your client shows.
  • Pin server versions and fingerprint the tool list. Keep a hash of every name, description and schema, and treat any change as a new release that needs review.
  • Alert on list changes instead of accepting them silently.
  • Keep servers apart. A server you trust should not share a context with one you do not, if a description in one can steer the tools of the other.

Risk 2: tool output is untrusted input

Whatever a tool returns goes back into the model’s context: a web page, a ticket, an email, a database row. If any of that text was written by someone else, it can carry instructions. That is indirect prompt injection, part of the first entry on the OWASP Top 10 for LLM Applications 2026,6 and MCP widens it: every connected tool is another inbox that strangers can write to.

The specification puts duties on both sides. Servers must sanitize what their tools return, and clients should validate tool results before passing them to the model.7 Neither makes the problem go away, so design as if some tool output will eventually be hostile: keep untrusted output away from tools that act, and read our guide to prompt injection for the defences that hold.

Risk 3: tokens, scopes and the confused deputy

Authorization in MCP is optional. Servers on an HTTP transport should follow the specification’s OAuth 2.1 based flow; servers on the local stdio transport take their credentials from the environment instead.8 Both roads end at the same place: a server that holds keys to your email, your CRM or your cloud account.

The specification closes the most common mistake explicitly. An MCP server must check that every access token was issued for it, and must not accept or pass on any other token.8 Forwarding a client’s token unchanged to a downstream API, known as token passthrough, is forbidden, because it bypasses the server’s own checks and breaks the audit trail.9

Two more patterns from the specification’s security guidance belong in every review. A server that proxies a third-party API under one shared client identity can become a confused deputy, letting an attacker obtain an authorization code without the user’s consent; such proxies must ask for consent per client. And scopes should start small and grow when a privileged operation needs them, rather than arriving as one wildcard grant at install time.9

Local servers have a quieter version of the same problem. Their keys usually sit in a configuration file next to the agent, and MITRE ATLAS records adversaries reading agent configuration to collect exactly those credentials.10 Keep them in a secret store, give each server its own narrowly scoped key, and rotate them like any other credential.

Risk 4: the server is software, and it runs somewhere

An MCP server is code, often installed with a single command. The specification warns that a local server runs with the same privileges as the client that starts it, requires clients that offer one-click setup to show the exact command before running it, and recommends a sandbox with minimal default access to files and the network.9

The client side deserves the same care. In July 2025 CVE-2025-6514 was published for mcp-remote, a bridge between local clients and remote servers: versions 0.0.5 to 0.1.15 could be made to run operating system commands when connecting to an untrusted server, through a crafted authorization URL. It was rated 9.6, critical, and later versions are fixed.11 The specification now tells clients never to open authorization URLs through a shell and to reject dangerous URL schemes.9

Two network checks complete the picture. During OAuth discovery a malicious server can point a client at internal addresses, such as a cloud metadata endpoint, so clients should block private address ranges. And because the protocol is now stateless, a server that keeps state across calls hands out a handle, and holding that handle must never count as proof of identity.9 That last rule is an old one in a new place: an object that answers to whoever names it is broken object level authorization.12

Risk 5: what the agent may do without asking

Every risk above grows with reach. OWASP ranks excessive agency third in its 2026 list for LLM applications.13 Put plainly: the agent can do more than its task needs. In MCP terms that is three things: which tools are exposed, which scopes they run with, and whether anyone confirms an action before it happens.

The specification’s answer is a person in the loop. There should always be a human able to deny a tool call; clients should ask for confirmation on sensitive operations and show the tool’s inputs before sending them; servers must enforce access control and rate limits.7 Rate limits also cap the bill, which matters when an agent loops: see denial of wallet.

Decide per tool, not per server. Reading a calendar can run unattended. Sending an email, issuing a refund, deleting a record or changing a permission should wait for a person, and the confirmation should show what will actually be sent, not the model’s summary of it.

How to review an MCP server

A review covers the server, the client that connects to it and the agent that decides when to call it. This is the checklist we work from, in the order we work through it.

An MCP review, area by area
AreaWhat to checkReference
InventoryEvery server in use, its version, who installed it, its transport, and where its credentials liveMCP specification, overview
Tool metadataDescriptions and schemas read as the model sees them; versions pinned; a fingerprint of the tool list; changes reviewed like a releaseAML.T0110, AML.T0109
TokensAudience checked on every token; no passthrough; consent per client on proxies; minimal scopes, stepped up on demandAuthorization; Security Best Practices
Inputs and outputsInputs validated by the server; outputs sanitized; untrusted output kept away from tools that actLLM01:2026; Tools, security considerations
ExecutionLocal servers sandboxed with minimal file and network access; one-click installs show the exact command; no URLs opened through a shell; private address ranges blockedSecurity Best Practices; CVE-2025-6514
StateHandles bound to the authenticated user and never treated as proof of identityAPI1:2023
AgencyEach tool’s worst case written down; confirmation for actions that send, pay, delete or grant; rate limits per toolLLM03:2026, ASI02
EvidenceEvery tool call logged with its inputs, so an incident can be reconstructedTools, security considerations

The output of a review is a short list, not a score: which tools can reach what, which of those paths an outsider can influence, and what closes each one.

01 / Model layer

How we test for this

Our MCP security review is part of AI Pentest. We read every tool description your agent receives and track it across versions, test each tool’s authorization with two users, and check whether content returned by one tool can steer the agent into another. AI findings come with a reproduction rate, because a model does not answer the same way twice.

See MCP security review and AI agent security testing for scope and method.

AI Pentest, Agent & MCP
EUR 6,000 to 15,000, 4 to 10 testing days

Questions people ask

What is MCP tool poisoning?

Tool poisoning is hiding instructions in an MCP tool’s description, schema or output, where the model reads them but the user usually does not. The model may follow them, for example by reading files or sending data it was never asked to. MITRE ATLAS lists it as AI Agent Tool Poisoning.

Are MCP servers safe to install?

Treat them like any other code you run with your own privileges. A local MCP server runs with the rights of the client that starts it, so check who publishes it, pin the version, read its tool descriptions, run it in a sandbox with minimal file and network access, and review every update before accepting it.

Does MCP have built-in authentication?

It has an optional authorization specification, not a mandatory one. Servers on HTTP should follow its OAuth 2.1 based flow, must check that every token was issued for them and must not pass tokens through to other services. Local servers using stdio take credentials from their environment, so protect those configuration files.

What is token passthrough, and why is it forbidden?

Token passthrough is an MCP server accepting a token issued for another service and forwarding it unchanged to a downstream API. The MCP specification forbids it because the server’s own checks, rate limits and logging are bypassed, and a stolen token can then be used through the server against every service that accepts it.

How do we secure an MCP server we built ourselves?

Expose the fewest tools with the narrowest scopes, validate every input, sanitize every output and check that each token was issued for your server. Bind state handles to the authenticated user, rate limit each tool, log every call with its inputs, and require a human confirmation for tools that send, pay, delete or grant access.

Where we test this

  • AI Pentest / MCP

    MCP security review

    Model Context Protocol servers and the agents that call them.

  • AI Pentest / Agent and MCP

    AI agent security testing

    What an agent can reach: its tools, its memory, its tokens and the systems behind them.

Sources

Every factual sentence above carries a numbered note; these are the documents behind them.

  1. 12

    Model Context Protocol specification (2026-07-28)

    Model Context Protocol project / checked 2026-10-10

  2. 3

    MCP Security Notification: Tool Poisoning Attacks

    Invariant Labs / checked 2026-10-10

  3. 47

    MCP specification 2026-07-28: Tools

    Model Context Protocol project / checked 2026-10-10

  4. 5

    MITRE ATLAS 2026.09: AML.T0110 AI Agent Tool Poisoning

    MITRE (ATLAS data release v2026.09) / checked 2026-10-10

  5. 613

    OWASP Top 10 for LLM Applications 2026

    OWASP GenAI Security Project / 2026-08-03 / checked 2026-10-09

  6. 8

    MCP specification 2026-07-28: Authorization

    Model Context Protocol project / checked 2026-10-10

  7. 9

    MCP specification 2026-07-28: Security Best Practices

    Model Context Protocol project / checked 2026-10-10

  8. 10

    MITRE ATLAS 2026.09: AML.T0083 Credentials from AI Agent Configuration

    MITRE (ATLAS data release v2026.09) / checked 2026-10-10

  9. 11

    CVE-2025-6514

    NIST National Vulnerability Database / checked 2026-10-10

  10. 12

    API1:2023 Broken Object Level Authorization

    OWASP API Security Project / checked 2026-10-10