Issue 14 4 min read

Watch every MCP call, both ways

An MCP server can look harmless on the day you approve it and act differently after that. Most fleets approve the connection once and never look at what flows through it again.

Invariant Labs found a malicious MCP server that showed a harmless instruction when it was installed. After its second use, it switched to a malicious one. The agent was also connected to a trusted WhatsApp server, and the switch turned that trusted connection against its user. The NSA’s Artificial Intelligence Security Center retells the case in its May 20 guidance on MCP security. Elsewhere in that document it names the general problem: “a previously benign and approved AI service could later access sensitive resources on demand, without triggering any review.”

That sentence describes how most teams add MCP servers. Someone clicks approve, the server joins the config, and nobody looks again. We treat it like installing a plugin. It works more like opening a door out of your system.

MCP is now the main way agents reach tools and data. Palo Alto Networks put it that way on September 28. It counts more than 97 million monthly SDK downloads and more than 10,000 active public servers. Those are the vendor’s numbers, and it doesn’t say where they come from. The design point in the same post matters more: “the protocol, by design, doesn’t enforce security.” The agent sends arguments to a tool, and the tool sends back results. “In today’s implementations, neither side is typically inspected.”

NSA explains why that hurts. In a normal client-server setup, the client asks and the server answers. MCP, NSA writes, “often expects servers to query and sometimes execute actions.” Each server you connect can make requests into your system as well as answer them.

Both directions of that traffic go uninspected.

Outbound, the agent fills in tool arguments from whatever it just read. NSA cites research showing that tool parameters, sent unchecked in malformed messages, led open-source MCP agents to expose sensitive server data. Then there is the GitHub case. Teams often grant blanket read and write access across every repository, private and public. A poisoned tool can read private code and publish it to a public repo, and the user never asked for that. The permissions allowed it, so nothing stopped it.

Inbound, the response lands in the model’s context as if it were trusted. Palo Alto says most implementations treat responses that way by default. NSA warns that outputs passed down a chain of agents “may be misinterpreted as executable prompts rather than passive content.” One poisoned result can steer the next agent in the chain.

NSA also warns that this “should not be viewed as isolated problems that can be patched at the interface or endpoint level.” Even the tool for testing MCP servers had a hole. CVE-2025-49596 let crafted messages trigger remote code execution in MCP Inspector until version 0.14.1 fixed it.

The last two issues argued for keeping secrets out of the sandbox. That still matters. But a clean sandbox doesn’t help if every tool call leaves and returns without anyone reading it.

The market sees the gap. Palo Alto’s post sells a gateway that sits in the request path. On October 1, Unite.AI reported that DeepKeep’s new plug-in for coding agents inspects prompts, file reads, shell commands, and MCP tool calls “before and after they run.” Each call gets an allow, block, or audit decision on the developer’s own machine. It starts with Cursor and Claude Code. NSA adds a fair caution: security proxies built for MCP “remain limited and are still maturing.” Treat these products as signs of where controls are heading, not as the answer. The work starts with knowing what you run.

Here is what to change.

  • Keep one inventory that maps each agent to each MCP server and each tool. NSA asks for versions, patch history, and known issues. If a server isn’t on the list, it doesn’t get connected.
  • Scan your network for MCP servers nobody approved. NSA names open scanners, including MCP Scanner and Ramparts. Developers start local servers far more often than security teams hear about.
  • Check arguments before they leave. Validate every call against the tool’s schema and its expected values, as NSA recommends. Then treat every response as untrusted input to the next step.
  • Scope OAuth grants to the job. A GitHub token for one repo can’t leak a private repo into a public one. Drop blanket read and write.
  • Save each tool description when you approve it, and compare on every reconnect. If the text changes, the approval is void until someone reviews it again. This is how you catch a server that changes after install.
  • Log every call with its exact parameters and the identity that made it. NSA adds a hash of the output where you can. Rate-limit calls, and send the logs to your SIEM, your central security log.

Sources

Next issue in two days. See you then.

Subscribe