What this MCP server security checklist covers, and what it does not

An approved refund can time out after the payment system commits; the agent retries, and the customer is paid twice. This checklist covers writes, and the MCP gateway, proxy or action server that executes them for the agent.

Prompt injection, tool poisoning, supply chain and server hardening are separate problems; these requirements only limit what a manipulated or mistaken agent can do. For hardening and protocol-level attacks, start with the MCP project’s Security Best Practices and OWASP’s guide to secure MCP server development.

The guide to human-in-the-loop agent actions over MCP covers how to design the layer. The enterprise AI platform reference architecture builds agent actions last, once identity and the decision log exist; this checklist confirms they do.

The six requirements at a glance

Requirement What to ask A good answer
1. Exact approval Who approves, seeing what? A person approves the final parameters, outside the agent’s reach. Any change voids it.
2. Write classification Who decides a tool writes? Your versioned manifest in the gateway. Unknown tools are denied.
3. Single execution What does a retry do? One approval, one idempotency key, at most one execution.
4. Provable log What can the log prove? Append-only records in storage nobody can rewrite.
5. The user’s rights Whose permissions does the call carry? The user’s, via token exchange. No broad agent account.
6. Failure behaviour What if approval is down? Writes stop; pending writes have an owner.

1. Approval of the exact action, by a person, outside the agent

Approval means a person saw the exact call and accepted it: the tool, the target record and every parameter that will be sent. A model-written summary (“update the customer’s bank details”) does not count.

The MCP specification’s tools page says there SHOULD always be a human in the loop with the ability to deny tool invocations, and that clients SHOULD show tool inputs before calling. The specification overview adds that MCP cannot enforce its security principles at the protocol level. Approval is a property of your gateway.

The approval is bound to the parameters (a hash of tool, manifest version and arguments, say), so any change voids it. The gateway issues and checks it, in a channel the agent cannot write to. The approver sees the target system and the identity the call will run as.

For high-risk AI systems only, Article 14 of the EU AI Act asks for overseers who can override the output or stop the system, and who stay aware of automation bias. Most enterprise agents are not automatically high-risk; a reviewer clicking through prompts is still not oversight.

Test it: approve a write, change one parameter, resubmit. It must be refused.

2. The gateway decides what is a write, not the server’s own labels

MCP servers can label tools with annotations such as readOnlyHint and destructiveHint, which help an interface and fail as a control. The tools page says clients MUST consider annotations untrusted unless they come from trusted servers, and the schema reference calls them hints that may not describe the tool faithfully. Nothing in the protocol stops a server labelling a delete as read-only.

Keep the classification in the gateway: a versioned manifest per tool, reviewed in your organisation, saying whether it writes and which checks apply. I would do this even for your own servers, because the server that performs a write should not be the one declaring it harmless.

Deny any tool missing from the manifest. Tool lists can change over time; a new or changed tool stays blocked until it is reclassified.

Test it: relabel a write tool as read-only. It must still need approval.

3. A retry or a double click must run once

An approved write can run twice. The downstream API commits and the response times out; the agent proposes again; the approver double-clicks. One refund becomes two.

Stripe’s idempotent requests are a good public model, though not a standard. Stripe saves the status code and body of the first request with a given Idempotency-Key, success or failure, returns them for repeats, and errors if a repeat’s parameters differ. Old keys are pruned; reusing one afterwards starts a new request.

In an agent gateway the key belongs to the approval, not the attempt: one approval, one key, at most one execution. Keep keys at least as long as an approval stays valid, or pruning reopens the gap. A re-proposal brings a new approval and a new key, so only a precondition stops it: the state the action assumes (“the invoice is still unpaid”), checked at execution.

A gateway cannot make a downstream API idempotent. If the target takes no key, the gateway must read its state before retrying, and some timeouts will need a person.

Test it: send one approved call twice, then with a parameter changed, then have the agent re-propose it. Expect one execution and two refusals.

4. A log that can prove what happened

A log proves something only if it can reconstruct the decision and nobody could have changed it since. Record each proposal, approval, rejection and execution: proposer, tool and manifest version, inputs as executed, checks run, approver and what they saw, identity used, outcome.

An append-only table an administrator can truncate is a convention, so export the log to write-once storage. On AWS, S3 Object Lock in compliance mode means no user, including the root user, can overwrite or delete a locked object version, change its retention mode or shorten its period. Governance mode yields to anyone holding s3:BypassGovernanceRetention. AWS says it can help meet WORM requirements, not that it makes a log compliant.

That lock collides with erasure. A compliance-mode retention period cannot be shortened, and write parameters are often personal data, like the bank details above. Where the design allows, the locked export should hold references and hashes, not personal data in clear. Settle retention and erasure with your data-protection officer before you lock anything.

Test it: replay one approved write from the log, then ask who can delete that record.

5. The call carries the user’s rights, not the agent’s

The system of record should see the person the agent acts for and enforce their permissions. OWASP on excessive agency names excessive permissions as a root cause and recommends acting in the specific user’s context, with authorisation enforced downstream, not by the model.

Authorisation is OPTIONAL in MCP, so first ask whether there is any. Over HTTP, servers must accept only tokens issued for them (Authorization specification) and must not pass the received token on to the APIs they call (security considerations). The other shortcut, a service account broad enough for every user, is worse: the agent could do what no single user can.

The standard mechanism is OAuth 2.0 Token Exchange. The gateway presents the user’s token as the subject_token, optionally its own as the actor_token, and receives a token for the target system. Under delegation the act claim then names the gateway as the acting party, if your identity provider issues it: the target enforces the user’s rights and logs who acted for them.

Test it: have someone without refund rights ask the agent for a refund. The target system must refuse.

6. What happens when approval or the gateway is down

The gateway should fail closed. If approval, policy checks or the log store are unavailable, writes stop and the agent is told why. A gateway that writes what it cannot log has made the log optional.

Approvers are unavailable too. Approvals should expire, because a morning approval judges the morning’s state, and pending writes need a queue with a named owner. A timeout never approves.

The gateway is one more service to run and own, and its outage stops every agent’s writes. Approval on every write also limits throughput: relax it per tool, only where the manifest marks the risk as low (an internal note on a ticket, say), as a reviewed manifest change. I would not relax it per agent: “trusted agent” cannot be tested.

Test it: stop the approval service and ask the agent to write. Nothing should execute.

Which gateways do this?

I do not compare products: features move too quickly. Ask the vendor to show you one write being approved, then retried, then replayed from the log. If any of the three is a slide instead of a demo, you do not have the control.

Trade-offs and limits

Approval on every write is slow by design, and a poor fit for high-volume, low-risk automation.

A gateway that passes every test can still relay a poisoned instruction. The approver is then the last human check, and the user’s rights still cap the damage.

Token exchange needs an identity provider and targets that support it; a legacy system behind one shared service account cannot enforce the user’s rights, so the gateway must, which is weaker. A scheduled agent with no user behind it needs its own narrow identity, and that is where agents with their own privileges creep back in.

Common questions

Do MCP tool calls need human approval?

The MCP specification says there SHOULD always be a human in the loop with the ability to deny tool invocations. It is a recommendation the protocol cannot enforce. For tools that write, I would require approval of the exact parameters, enforced by the gateway rather than the agent.

Can I trust MCP tool annotations such as readOnlyHint?

Not as a control. The specification says clients must treat annotations as untrusted unless they come from a trusted server, and calls them hints that may not describe the tool faithfully. Keep your own versioned classification of every tool in the gateway, and deny tools it does not know.

How do I stop an agent doing more than the user is allowed to do?

Exchange the user’s token for one issued for the target system with OAuth 2.0 Token Exchange, never forward the token the gateway received, and give the agent no broad service account. The target then refuses what the user could not do.

What should an MCP gateway log?

Every proposal, approval, rejection and execution: who proposed it, the tool and manifest version, the inputs, the checks run, who approved and what they saw, the identity used, the outcome. Keep it append-only, and export references and hashes, not personal data in clear, once your data-protection officer has agreed retention and erasure.

Which MCP gateways support human confirmation before writes?

I do not rank products; the market changes faster than any list. Ask each vendor to show one write approved, retried and replayed from the log, then a parameter changed after approval, then the approval service stopped. Treat whatever they cannot demonstrate as missing.

In practice

On the ENSO platform, I specified agentic actions through a single MCP action server with versioned tool manifests, human validation on every write, idempotency and precondition checks, and an append-only decision log that makes every AI decision replayable, with a WORM export to S3 Object Lock. The AI inherits the user’s permissions through OAuth2 Token Exchange (Keycloak, OIDC) and IRSA workload identities. See the ENSO case study for the full architecture and the AI agents governance page for the service.

Sources