Testing AI Agents Before Launch: Prompt Injection, Tool Permissions, and MCP
AI agents that read untrusted content and call tools open new attack paths. What to inventory, test, and govern before launch, mapped to OWASP and MCP.
- Author
- Ackatec
- Published
- Reading time
- 7 min read
- AI security
- Penetration testing
- Prompt injection
- MCP security
- AI governance
- Application security
Many AI features now take actions: they read inboxes, search documents, open tickets, and call internal APIs, often through Model Context Protocol (MCP) servers. The security question shifts from "could the model say something embarrassing?" to "what could someone make this system do?"
The short answer: before an AI agent reaches customers or sensitive data, map every input it reads and every tool it can call, then test whether untrusted content can steer those tools. Fix what you find in the architecture rather than relying on prompt wording.
How is testing an AI agent different from testing a web application?
A traditional application follows code paths a developer wrote. An AI agent is a system in which a language model decides at runtime which tool to call and with what arguments, often based on content it reads along the way: emails, PDFs, tickets, web pages, and other tools' output. A conventionally scoped web application test can miss that entirely.
What is indirect prompt injection?
Indirect prompt injection happens when instructions hidden in external content, such as a web page, file, or email, change how a model behaves once that content enters its context. OWASP ranks prompt injection first in its Top 10 for LLM Applications (opens in a new tab) (LLM01:2025). The OWASP Top 10 for Agentic Applications (opens in a new tab), published in December 2025, extends the concern to systems that act. Its first three entries are Agent Goal Hijack (ASI01), Tool Misuse and Exploitation (ASI02), and Identity and Privilege Abuse (ASI03).
What EchoLeak (CVE-2025-32711) showed
In June 2025, Aim Labs, the research team at Aim Security, disclosed EchoLeak (opens in a new tab), a zero-click flaw in Microsoft 365 Copilot tracked as CVE-2025-32711 (opens in a new tab) and rated critical by Microsoft. A crafted email, phrased as ordinary instructions to a human recipient, got past Microsoft's prompt injection classifier. When the user later asked Copilot a related question, the email was retrieved into context and steered Copilot to embed sensitive data in a reference-style Markdown image that the link filter missed. The client fetched that image automatically, and an allow-listed Microsoft Teams URL relayed the request past the content security policy. Microsoft fixed it server-side before public disclosure.
OWASP cites EchoLeak as an example of Agent Goal Hijack. It also illustrates what Simon Willison calls the lethal trifecta (opens in a new tab): access to private data, exposure to untrusted content, and a way to communicate externally. Many internal AI features combine all three.
Start with an inventory, not a scan
Before anyone sends a test payload, map each agent you plan to ship:
- Who can talk to it. Employees, authenticated customers, or anyone on the internet.
- What untrusted content it reads. Inbound email, uploads, shared drives, web results, and third-party tool output.
- What it can reach. Every tool, API, database, and MCP server, plus the identity and scopes it uses for each.
- What it can change. Anything it can write, send, delete, approve, or pay.
- Where its output lands. A chat window, a rendered web page, an email, a ticket, or another system's input.
This map becomes the test scope, and it is also governance work. If nobody can say who approved a tool connection or why the agent holds write access, that is a finding before testing begins. Setting that ownership, with human review points and an exception path, is the core of AI governance and readiness work, especially where AI arrived before the rules.
If the agent is a vendor product, testing the vendor's platform usually depends on its published testing rules. Your configuration is still yours to review: enabled connectors, searchable data sources, and granted permissions.
What should an AI agent security test cover?
1. Indirect prompt injection through every input channel
Plant instructions where the agent actually reads: documents and their metadata, email, ticket comments, file names, tool responses, and MCP tool descriptions, which OWASP flags as an agentic supply chain risk (ASI04). Check whether the agent follows them, including when the instruction is phrased as normal business text, as in EchoLeak.
2. Tool permissions and excessive agency
OWASP traces excessive agency (LLM06:2025) to three root causes: excessive functionality, excessive permissions, and excessive autonomy. Look for read integrations that also carry write or delete scopes, one shared service account acting for every user, and high-impact actions that run without human confirmation. Confirmation steps should be enforced by the application, not requested politely in a prompt.
3. Authorization at the tool layer, not in the prompt
A system prompt is not an access control. Every tool call should enforce the end user's permissions on the server side. Test whether one user can coax the agent into retrieving another user's records, and whether tenant boundaries hold inside retrieval systems such as vector stores, a leakage risk OWASP lists under Vector and Embedding Weaknesses (LLM08:2025).
4. MCP authorization, tokens, and third-party servers
MCP authorization is optional and defined for HTTP-based transports, so first confirm that a remote server enforces it at all. Where it does, the current specification (opens in a new tab) (version 2026-07-28) builds on the OAuth 2.1 draft, requires servers to accept only tokens issued specifically for them, and forbids passing a client's token through to upstream APIs. The security best practices (opens in a new tab) add tests for confused deputy attacks, server-side request forgery, state handle hijacking, local server compromise, and over-broad scopes. Review community MCP servers like any other dependency: OWASP cites a malicious npm package that impersonated a Postmark MCP server and secretly copied emails to an attacker.
5. Output handling and exfiltration paths
Model output is untrusted data, a risk OWASP calls Improper Output Handling (LLM05:2025). Check whether rendered Markdown or HTML can load external images or links that carry data out, whether output reaches a shell, a SQL query, or a template, and whether content security policy allow-lists include redirectors or proxies that defeat their purpose, as the Teams URL did in EchoLeak.
6. Logging, limits, and a way to stop
Confirm you can reconstruct what the agent did: the request, the retrieved content, each tool call, and each approval. Check rate limits and spending caps, which OWASP tracks as Unbounded Consumption (LLM10:2025). Make sure someone can disable a single tool quickly without taking the product offline, and rehearse that call in an incident readiness tabletop exercise before you need it.
What should an AI agent test report include?
Good findings describe a full path: the input, the agent's decision, the tool it called, and the business impact, with steps to reproduce. Because model behavior varies between runs, a single screenshot is weak evidence. Report how often each attack succeeded across repeated attempts.
The fixes are usually architectural: narrower tool scopes, per-user authorization, enforced confirmation steps, output sanitization, and restricted network egress. Retest after those changes land, and again whenever you add a tool, switch models, or connect a new data source. A test reports what was found within the agreed scope and time window. It does not guarantee that other weaknesses are absent.
Because agents reach users through web applications and APIs, their endpoints and integrations can be included in a penetration testing scope alongside the rest of the application.
Questions leadership should answer before launch
- Can we name every tool and data source this agent can reach, and who approved each one?
- Which actions can it take without a human confirming?
- Does it act with each user's own permissions, or with a broader shared identity?
- Where does its output render, and can that output trigger outbound requests?
- If it misbehaves on a Friday night, who can switch it off, and what records will they have?
- If a customer, auditor, or regulator asks how this agent was tested and governed, what evidence can we show?
If several are hard to answer, the gap is broader than one feature. A security assessment can place AI use within your overall exposure and sequence the work so testing goes where it changes decisions. To talk through an agent you plan to launch, start a conversation.
References
- OWASP, Top 10 for LLM Applications 2025 (opens in a new tab)
- OWASP GenAI Security Project, Top 10 for Agentic Applications for 2026 (opens in a new tab)
- Model Context Protocol, Authorization (2026-07-28) (opens in a new tab)
- Model Context Protocol, Security Best Practices (2026-07-28) (opens in a new tab)
- NIST National Vulnerability Database, CVE-2025-32711 (opens in a new tab)
- Aim Labs, Breaking down EchoLeak (opens in a new tab)
- Simon Willison, The lethal trifecta for AI agents (opens in a new tab)
