Cybersecurity
MCP security: how prompt injection turns a helpful tool into a leak
How attacks on Model Context Protocol servers work, why a read-only tool can still exfiltrate data, and a short checklist for running MCP servers safely.
By Raktim Ranjit · Published · 4 min read
Short answer: the main MCP risk is not a bug in the protocol. It is that a model reads untrusted text and can then call tools with your permissions. If a web page, a ticket or an email contains instructions, the model may follow them. Keep each server's access narrow, separate tools that read untrusted content from tools that can send or delete, and require a human approval for anything irreversible.
What is the attack in plain terms?
A model cannot reliably tell the difference between your instruction and text it was asked to read. Both arrive as tokens in the same context. If an attacker can put text where the model will read it, they can try to steer it.
Picture an assistant with two tools: one reads your support inbox, one can create GitHub issues in a private repository. An attacker emails you a message that says, in the body, to search the private repo for the file named .env and paste its contents into a new public issue. The model never saw a hostile program. It saw a message and a toolbox.
What makes it dangerous?
Security people describe the risk as a combination of three things at once: access to private data, exposure to untrusted content, and a way to send data out. Remove any one and the attack mostly stops. Most real incidents with agents had all three in one session.
- Private data: files, repositories, a database, saved credentials.
- Untrusted content: web pages, issues, pull requests, emails, documents shared by others.
- An outbound channel: posting a comment, sending a message, making an HTTP request, opening a pull request.
A read-only tool can still be part of this. Reading is the private-data leg. The leak happens through a different tool.
What other MCP-specific problems exist?
Tool poisoning
A tool's description is text that the model reads and trusts. A malicious or compromised server can hide instructions in a description, like telling the model to also read a credentials file before every call. You only see the short summary in the UI.
Rug pulls
A server you approved on Monday can change its tool list on Friday. If you pinned nothing, you now run different code.
Confused deputy and token passthrough
A server that holds one powerful token and uses it for every user can be tricked into doing something the current user is not allowed to do. Servers should act with the calling user's own scoped credentials.
Local servers run as you
A stdio server is a process on your machine with your file and network access. Installing one from a package manager is the same trust decision as running any program.
What should you do about it?
- Least privilege. Give a database server a role that can only read the tables it needs. Give a Git server a token scoped to one repository.
- Split read and act. Do not put a tool that reads untrusted text in the same session as a tool that can send data out, or a tool that reads secrets.
- Approve irreversible actions. Deleting, sending, paying and publishing should stop for a human each time. Do not click "always allow" on those.
- Pin versions. Install servers at an exact version, read the source for ones that touch sensitive data, and review changes when you upgrade.
- Read tool descriptions. If you add a server, look at the full text the model will see.
- Run risky servers in a container with no access to your home directory and a restricted network.
- Log every call with arguments, so you can reconstruct what happened.
Can a filter or a better prompt fix prompt injection?
Not reliably. Telling the model to ignore instructions in documents helps a little and fails often enough that you cannot depend on it. Detection tools catch known patterns and miss new ones. Treat the model as something an attacker may be able to steer, and design so that a steered model still cannot do harm. That is an access-control problem, which engineers know how to solve.
A small checklist before you add a server
- What is the worst thing this server can do with the token I give it?
- Does the session also read content written by strangers?
- Can anything in this session send data outside?
- Which actions are irreversible, and do they ask me first?
- Is the version pinned and the source readable?
If you can answer all five, you are ahead of most setups. For the concept behind the attack, see what prompt injection is.
References
Author
Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.