How to Write MCP Tool Descriptions: The 8-Block Pattern from 11 Production Servers
An academic analysis of 856 tools across 103 MCP servers (MCP Tool Descriptions Are Smelly! (arXiv:2602.14878, 2026)) found that 97.1% of tool descriptions contain at least one "smell": unstated limitations, missing usage guidelines, opaque parameters. That's not a fringe problem. That's the baseline. The sample was public, published servers, not private, internal ones like the eleven behind this post, so there's no current read on how that number looks inside a typical company today.
In The Six Levels of MCP Servers, I described what each maturity level looks like. Here's what the jump from Level 1 to Level 4 actually takes.
A tool description is not a sentence
The MCP specification's own best-practice proposal considers this a good description: "Read the contents of multiple files simultaneously. More efficient than reading files individually when analyzing or comparing multiple files." One sentence. That's the community's own example of good, and it's nowhere close to what an agent actually needs. The proposal wasn't rejected, it sat open for five months and was closed as dormant for lack of a sponsor.
After building more than 90 tools across eleven production servers, I arrived at a different standard. A tool description is an operational manual, structured into blocks, each added because the agent failed without it.
Updated September 2026: the blocks still hold, but I have since measured where they have to go. On Claude Code only the first 2,048 characters of a description reach the model, so part of each block now lives in the tool's response instead. The sections below say which, with the eval behind each claim (evidence register).
The eight blocks
| Block | Purpose | Without it... | Where it goes |
|---|---|---|---|
| RETURNS | Agent knows which fields come back, and which ones to ask for | Can't determine if this tool has the data it needs | Description head |
| WHEN TO USE | Agent knows when this tool fits | Picks wrong tool or misses this one | Description head |
| WHEN NOT TO USE | Prevents wrong tool selection | Tries this tool for queries that belong elsewhere | Description head |
| QUERY STRATEGY | Teaches summary-first, then drill down | Fetches full records when a summary would suffice | Description head and input schema |
| INTERPRETATION | Cross-field rules, type-to-field mappings | Returns raw numbers without conclusions | Rules for every record in the head; rules for this record in the response |
| RELATED TOOLS | Agent chains queries via join keys | Stops after first tool call | Description head |
| FEEDBACK | Agent reports friction | Issues go undetected, descriptions never improve | Description head or server instructions |
| ALERTS | Agent surfaces server-generated warnings | Ignores domain-specific warnings in the response | A pointer in the head; the alerts in the response |
These blocks are not theoretical. Each was added in response to a specific failure mode observed in production. The agent picked the wrong tool, so add WHEN NOT TO USE. The agent fetched 2,000 records instead of a summary, so add QUERY STRATEGY. The agent ignored that a status code meant something entirely different for a different record type, so add INTERPRETATION.
What actually works as a metadata channel
Not everything you write reaches the agent. My first answer came from testing across Claude Desktop, Claude Code, and Cursor, and it put tool descriptions at the top as the channel that always works. The evals I ran afterwards on a public reference server corrected that. What they measured on Claude Code:
| What the server ships | Reaches the model? |
|---|---|
| Field names | Always. The one piece of meaning guaranteed to arrive |
| Tool description | Up to character 2,048, and no further. Keep the head under about 1,800 (Q7) |
| Server instructions | Also cut at 2,048, where a client injects them at all. Other hosts differ (see below) |
Input schema, .describe() included |
In full, but re-sent every turn, so every word is paid on every call (IS) |
| The tool's response | Yes, under about 25k tokens (Q9) |
| Output schema | Not on Claude Code. It validates the response and drives UI rendering (Q11) |
The cut is silent. On my own richest tool, 74% of the description never reached the model, and nothing told me. Two things I had written off came back as well. A guidance tool works when the pointer to it is a requirement: I had removed my meta-tools because agents didn't call them, and with the pointer worded as "REQUIRED: call this first" Haiku called it 10 times in 10, against 0 in 10 for a soft hint (Q8b). MCP Resources are still unused; I built them, deployed them and removed them, and nothing since has changed that.
That table is Claude Code's. I have since measured the others (HD). Cowork cuts a description at 4,096, and claude.ai chat, ChatGPT and Codex deliver it whole. Server instructions vary the most. ChatGPT Chat cuts them at 512, and claude.ai chat and ChatGPT Work never deliver them. Codex and ChatGPT Work do pass the output schema on, and they drop every .describe() in an input schema once it passes 5,000 characters. A server that has to work on all of them keeps its instructions within 512 characters and each input schema under 5,000, and repeats every instruction line in a description head or the response.
The key insight is delivery, not channel. A delivered sentence works about as well in the description head as in the response: 20 of 20 in both places, and 1 of 20 when it sat past the cut (Q15). So what the agent needs before the call goes in the head and the input schema. What it needs to read a value goes in the response, next to the record it applies to.
Before and after
The difference in practice. A building profile tool, same API, same data:
Level 1: "Look up building information by postcode and house number." Two untyped parameters. No field documentation. The agent guesses what an energy label means, whether the building year is reliable, and which postcode format the API accepts.
Level 4, as I first built it: Structured WHEN TO USE / WHEN NOT TO USE blocks. A QUERY STRATEGY that warns the agent not to trust smart meter registration addresses as physical building addresses. An INTERPRETATION block that explains three different energy certification standards: NTA 8800 returns kWh/m2, Nader Voorschrift returns MJ total building energy. Without that block, the agent compares them as if they're the same unit and gives confidently wrong advice. Input schema with regex validation. The output schema validates the response shape; null patterns, cross-tool join keys, and benchmark-specific surface-area guidance live in the description and returned data, where the model can read them.
That version had a flaw I only found by measuring it. It was about 5,100 characters long, so most of it was cut, and one plausible line compared a calculated label figure with the Paris Proof target, which is defined on measured energy. It made 59 of 60 answers wrong. One sentence saying the figures are calculated, not measured, took the same question to 59 of 60 right (BT). The rebuilt version keeps a 1,800-character head and returns an interpretation block first in every response, with the rules and computed values for that one building.
Same tool. Different category. The Level 1 version exposes an API. The Level 4 version teaches the agent how a domain expert thinks about this data.
An open-source extract of this tool, with the metadata patterns intact, lives in the mcp-metadata-demo repo, at three tiers side by side. Read get-building-profile-best.ts for the rebuilt tool, and the wire view for exactly what the model receives from it.
This post is one piece of a longer argument. Production MCP: A Practitioner's Guide puts all of them in order, from understanding your data through to identity-bound deployment.
Go deeper: read the full practitioner report, The Missing Layer, or explore the mcp-metadata-demo server, an open-source extract of these patterns, since the production servers run on private business data.