Skip to main content
← View All Insights
MCPArchitecture

Six Things the MCP Spec Should Fix

Published 2026-04-24·Updated 2026-09-29·7 min read

The observations in this post came from an eleven-server, 90+ tool production snapshot across eleven external APIs. The protocol's limitations become sharp. MCP excels at connecting agents to tools. It does not yet help agents understand what those tools mean, or safely act on what they understand.

These are the six changes that would matter most, drawn from what broke in production.

1. Smarter tool discovery

Today, every connected MCP server dumps all its tools into the agent's context at once. For a single server with 10 tools, that's fine. For a developer in an IDE with 10 servers connected simultaneously? That's 100+ tool descriptions competing for context.

This has a chilling effect on metadata investment: why write 200 lines of rich domain knowledge per tool if the client is going to flood the context with everything?

The spec needs a discovery layer: DNS for tools. Let servers declare tool groups with short summaries. The client presents groups to the model, the model selects the relevant group, and only those tool definitions are injected. Resolve first, load second.

Claude Code has since shipped a client-side version of this: tools connect deferred, by name only, with a search tool that fetches full schemas on demand. It works, and it's proof the diagnosis was right, but it's one client's fix, not a protocol primitive. A server still can't declare its own tool groups; every other client has to solve discovery on its own.

2. Extensible metadata without context flooding

All metadata currently goes into tool descriptions and input schema descriptions, both consuming context tokens on every call. There's no mechanism for metadata the agent can request on demand.

A tool could declare 200 lines of field-level documentation, cross-reference tables, and query strategy guides as structured metadata. The client loads the tool's name and short description by default, but injects the full metadata only when the model indicates it wants to call that tool.

A version of this has started showing up as a convention: a server points the agent at its detailed guidance via MCP Resources under a skill:// URI, with instructions that make the read mandatory for specific triggers, plus a fallback tool for clients that can't read resources at all. It works, not because Resources became spontaneously requested, but because the instructions force the read. That's a workaround, not the protocol guarantee this recommendation is asking for.

This would eliminate the trade-off between rich documentation and context efficiency, a trade-off that currently forces server authors to compress critical domain knowledge into artificially short descriptions.

Updated September 2026: I have since measured how short. On Claude Code only the first 2,048 characters of a tool description reach the model; the rest is dropped, and the server gets no signal that it happened. On my richest tool, 74% never arrived (Q7). Raising the client's limit delivers it, but adds 23.7% in tokens to every tool on every request (Q15b). So two cheap asks go with this one: hosts should publish the limit and mark a cut description as cut. Nobody publishes it today, and it is not one number. Cowork cuts at 4,096, and claude.ai chat, ChatGPT and Codex deliver the description whole (HD). I only know because I measured each host.

3. Resources that actually work

MCP Resources are the spec's answer to reference documentation: static or dynamic content that agents can request to inform their reasoning. In theory, perfect for domain guides and field dictionaries.

In practice, no tested client reliably surfaces resources to agents. Claude Desktop lists them in a sidebar but agents don't request them. Claude Code and Cursor ignore them entirely. I built resources, tested them, and removed them after validation showed zero agent-initiated requests.

For resources to work, clients need to either: automatically inject relevant resource content when a related tool is selected, allow servers to mark resources as required context for specific tools, or give the model an explicit get_resource primitive that it's trained to use.

4. First-class feedback loops

The current spec has no concept of agent-to-server feedback. Every interaction is request-response: call a tool, get data, move on. No standard way for agents to report confusion, flag data quality issues, or indicate which calls were unhelpful.

I solved this with a custom report_problem tool and queryIntent parameters on every input schema. This works, but it's a workaround. The spec should support feedback as a first-class primitive: a standardized way for agents to annotate tool calls with intent, satisfaction, and issues encountered. Server authors could subscribe to feedback events and improve metadata iteratively, closing the loop that currently requires custom infrastructure.

5. Make output schema model-visible, or say plainly that it isn't

I used to think this was a training gap: models just haven't been trained on rich MCP interactions, so of course they don't parse a WHEN TO USE block specially or read the .describe() on an output field. Then I measured it directly. Across Claude Code, claude.ai, and OpenAI's Responses API, the tool definition a model actually receives is the description and the input schema. None of them forward a single output-schema field annotation. The spec's own line pointing at LLMs there is non-normative, and no host I tested implements it.

That's not a training gap, it's a channel that's dead on arrival. Output schema still does real work: it validates the response and drives UI rendering. But on the hosts I tested, a model reasoning about what a tool returned never sees a word of the schema that documents it, no matter how the model was trained. The fix belongs in the spec or the clients: either start forwarding output-schema annotations to the model, or state plainly that it's validation-only so server authors stop burying interpretation knowledge somewhere most hosts never deliver. I've since moved that knowledge to where it demonstrably lands: the rules that hold for every record into the head of the tool description, and the rules for this one record into the response itself (Q11, Q15).

Updated September 2026: two clients turned out to be the exception. Codex and ChatGPT Work do pass the output schema on, Codex as a TypeScript return type. No Claude host does, so the advice stands for anyone whose users are on Claude (HD).

6. Scoped write permissions for MCP Apps

As servers evolve from read-only to interactive applications, write operations become inevitable. But the spec offers no mechanism to restrict a write tool to a specific interaction context. If a server exposes a POST tool, any connected agent can call it at any time.

I designed a workaround: the WriteIntent pattern. The agent calls a model-visible "open" tool that mints a server-side intent. The actual mutation is performed by a separate "commit" tool registered with visibility: ["app"], hidden from the agent, callable only by the MCP App form. The intent is user-bound, resource-scoped, one-shot, time-limited, and validated with typed schemas.

The spec should formalize this: a way to scope write tools to specific MCP apps, so that a tool marked as write-only-via-app can only be invoked through a validated form interaction, not through freeform agent reasoning.

The spec is built for transport, and the harder problems are knowledge delivery and safe interaction. Fixing these six at the protocol level would raise the floor for every server in the ecosystem, including the ones whose authors will never spend 40 hours on metadata engineering.


This post is one piece of a longer argument. Production MCP: A Practitioner's Guide puts all of them in order, from understanding your data through to identity-bound deployment.

Go deeper: read the full practitioner report, The Missing Layer, or explore the mcp-metadata-demo server, an open-source extract of these patterns, since the production servers run on private business data.