MCP Security
MCP connects agents to tools and data. Malicious descriptions, untrusted results, excessive access and vulnerable server implementations create distinct risks. This guide separates demonstrated attacks, scanner findings and proposed controls, and explains what each control can actually enforce.
Builder Security · builders · security-teams What is MCP
The Model Context Protocol (MCP) is an open standard released by Anthropic in November 2024 for connecting AI agents to external tools and data sources. It defines how an agent discovers tools, calls them, and processes their responses. Think of it as the USB port for AI agents — a universal interface that lets any agent connect to any tool.
MCP servers expose tools (functions the agent can call), resources (data the agent can read), and prompts (templates the agent can use). The agent's MCP client discovers available tools, selects the right one based on the user's request, and calls it with parameters. The server executes and returns the result.
The Numbers
Six Attack Vectors
The examples below include research demonstrations, a maintainer-confirmed vulnerability and malicious packages. Scanner findings describe potential risk and inherent tool capabilities; they are not all confirmed exploitable vulnerabilities. Population and methodology differ between studies.
Malicious instructions in a tool description can influence selection or subsequent actions. Tool-result injection instead places instructions in returned data. Clients may display descriptions, but visibility and human review do not establish that the model will safely interpret them.
WhatsApp MCP Server (April 2025): Researchers demonstrated that hidden instructions in a tool description could exfiltrate complete WhatsApp chat histories through a legitimate MCP server.
GitHub MCP Server (May 2025): Researchers showed that malicious instructions embedded in GitHub Issues could hijack AI assistants using the official GitHub MCP server, leaking private repository source code and cryptographic keys into public pull requests.
A reviewed tool definition or implementation may change later. Clients differ in change detection and re-approval behavior. Test those properties in the actual client rather than assuming a connection approval pins future behavior.
A tool can change its description after approval, or change its remote implementation while retaining the same description. Whether re-approval occurs depends on the client. Description changes are metadata changes; implementation changes may be invisible to schema comparison.
A malicious MCP server exploits legitimate tools from OTHER trusted servers. The attacker controls one "trojan" server and uses it to poison tool descriptions that make the agent leak data through legitimate servers it already trusts.
AgentSeal reported 935 potential toxic-flow findings across 555 of 5,125 scanned servers in March 2026. Some findings reflect inherent capabilities, not implementation defects. The study reports classifier limitations and runtime probes separately. A shared client with access to sensitive data and an outbound tool may still enable cross-tool disclosure after injection.
MCP tools that make HTTP requests become SSRF proxies when destination URLs are attacker-controlled. The agent calls the tool with a URL parameter — the attacker controls where that request goes. Internal networks, cloud metadata endpoints, internal services on the same host become accessible.
HackMD MCP Server — CVE-2025-59155: the maintainer advisory describes SSRF through an attacker-supplied hackmdApiUrl. Affected versions are 1.4.0 up to, but excluding, 1.5.0; version 1.5.0 is patched. Impact depends on the server’s network reach and deployment. Maintainer advisory.
The November 2025 specification defines two standard transports: stdio and Streamable HTTP. Streamable HTTP can use SSE and replaces the older HTTP+SSE transport. Origin validation matters for HTTP-based local services, not only for SSE. Versioned specification.
Security depends on the server version and deployment. For remote connections use TLS plus authentication and resource authorization. Validate Origin as specified and bind local services to loopback. TLS does not determine which authenticated caller may invoke a tool. Local stdio requires trust in the launched process, executable path and host permissions.
Attackers publish "Trojanized" MCP servers to public registries with names similar to legitimate servers. Once installed, these servers can backdoor tool calls, exfiltrate data, or execute arbitrary code. This is the npm typosquatting problem applied to AI tool infrastructure.
The malicious postmark-mcp npm package was reported to copy outbound email content to an external recipient. Snyk identifies affected versions from 1.0.16. This is an observed supply-chain case, unlike the research demonstrations above. Kaspersky separately reported malicious PyPI packages disguised as MCP tooling. Snyk advisory.
Defending Your MCP Stack
These proposed controls address different trust boundaries. Their effectiveness needs verification in the client, server, identity system and network layer of the actual deployment.
Annotations such as readOnlyHint and destructiveHint describe expected behavior; they do not enforce it. Treat annotations from untrusted servers as untrusted, and base approval policy on independently established permissions. MCP tool specification.
Useful for interface hints and reviewed policy configuration; annotations alone do not prevent tool poisoning or SSRF.
Validate tool inputs server-side. For URLs, enforce the permitted scheme, host, resolved IP ranges and destination on every redirect; address DNS rebinding and block cloud metadata/internal ranges unless explicitly required. For files, resolve paths and symlinks within an enforced filesystem boundary.
Reduces SSRF and path traversal when the complete destination/path policy is enforced. Schema validity or a hostname allowlist alone does not establish that a destination is safe.
Pin approved descriptions and schemas by digest and require review when they change. This can detect definition changes. It cannot detect a remote implementation changing behind an identical schema; also verify software provenance and constrain runtime behavior.
Detects changes to the exact definition being hashed. Stable metadata does not prove stable implementation or safe tool behavior.
Use stdio for trusted local processes and authenticated TLS for remote HTTP. Validate Origin, bind local listeners to loopback, and enforce authorization per resource and operation. Review configuration defaults for the actual server version.
Protects transport and limits access when correctly configured; it does not make a malicious authenticated endpoint trustworthy.
Isolate server processes and credentials, and restrict each task’s tool set. Also control data flow in the client: an injected model can move data between separately isolated servers using its legitimate connections. Sensitive outputs require recipient and egress policy.
Isolation limits server compromise and credential sharing. Client-level authorization and data-flow controls are needed for cross-tool exfiltration through a shared orchestrator.
Verify MCP server packages before installation. Check publisher identity, package name spelling, download counts, and source code. Use lock files to pin exact versions. Scan for known malicious packages.
Publisher checks, pinned artifacts and review reduce supply-chain risk. Download counts, a familiar name or a clean scan do not establish trust.
MCP security is the tool-layer defense in the Agentic AI Kill Chain. These controls work alongside hook-based guardrails (agent-layer defense) to create defense in depth.
More builder security, red-team frameworks, and detection patterns coming.
This work represents the author's independent research and personal views. It is not related to or endorsed by the author's employer.