AI Agent Tool Descriptions Can Be Poisoned to Leak Corporate Data
Attackers can embed malicious instructions within Model Context Protocol (MCP) tool descriptions, effectively hijacking AI agents into exfiltrating sensitive company data while appearing to execute legitimate commands. The root cause is a lack of validation and integrity controls over the inputs that AI agents consume as trusted instructions, treating tool descriptions as implicitly safe. This matters because AI agents often operate with broad data access and elevated trust, meaning a single poisoned description can silently bypass traditional security controls. As enterprise AI adoption accelerates, this attack surface will grow rapidly if organizations do not establish governance around AI tool registries and inputs.
Tactical Insight
Immediate actions
- Audit all MCP tool descriptions and AI agent input sources for unexpected or unauthorized content.
- Restrict AI agent permissions to the minimum data access required to complete their defined tasks.
Long-term improvements
- Implement a signed and version-controlled registry for all MCP tool definitions to enforce integrity checks before agent execution.
- Establish a formal AI tool approval process requiring security review before any new tool description is deployed into production.
- Apply data loss prevention (DLP) controls on AI agent output channels to detect and block anomalous data exfiltration patterns.
Detection measures
- Enable detailed logging of all AI agent actions, tool invocations, and data access events for real-time behavioral analysis.
- Deploy anomaly detection rules that alert on AI agents accessing or transmitting data outside their expected operational scope.
- Conduct regular red-team exercises specifically targeting AI agent pipelines and MCP integrations to surface novel injection vectors.