---
title: "Your MCP Server Should Get Smarter Every Week"
description: "A number that sometimes means hours and sometimes kilometres, summed into a confident wrong total. No confidence score catches that. Production telemetry does."
canonical: "https://davidgolverdingen.nl/en/insights/mcp-server-smarter-every-week"
last-updated: "2026-09-29"
---

# Your MCP Server Should Get Smarter Every Week

Writing good tool descriptions is the start. Keeping them good is the real work.

The description blocks I covered in [a previous post](https://davidgolverdingen.nl/en/insights/97-percent-mcp-tool-descriptions-broken) are the initial knowledge layer. But how do you know what's missing? How do you find the gaps the agent silently works around without telling you? That's where the feedback architecture comes in.

I run sixteen MCP servers in production now, across 103 tools. The feedback loop under them has three layers, plus one design choice that ties them together. I've ordered the layers by how much each earns its keep. The ordering I'll defend. The exact split I won't; it's a working estimate, not a measurement.

## You can't observe understanding. You can observe confusion.

That is the whole idea. Nothing in the logs tells you the agent understood a tool. Plenty tells you it didn't.

And confusion shows up in three ways, not one.

Usually it's friction, not error. Something is underspecified, the agent takes a wrong turn, the user asks a follow-up, there's a retry or two, and together they arrive at the right answer. The answer ends up correct. What it cost was turns. Missing meaning usually doesn't produce a wrong answer. It produces extra work.

Sometimes someone notices something is off. The agent surfaces the friction in its answer, or the user feels it directly, and a person decides to flag it. The people using these servers are decent detectors, for a specific reason: they were onboarded on data they already know, so a wrong shape stands out. That's a design choice, not luck.

And rarely, the answer is just wrong and nobody notices. My worst case: a field called `Aantal` (Dutch for "quantity"), typed as a number, where a sibling enum decides whether the number means hours, kilometres or pieces. Sum the column and you get a confident, wrong total. No error, no retry, no complaint.

That last case is the hole in confidence-based triage. Route the uncertain findings to a domain expert for review, which is the right instinct, and the fields the model is confident about and wrong about never reach anyone. They pass by construction. `Aantal` wasn't caught by any score. A domain expert walking the high-stakes fields by hand caught it.

So three kinds of confusion, three instruments. Logs for the friction. A report when someone notices something is off. A domain expert for what nobody notices. The rest of this post is those instruments, and then what happens once knowledge you already validated starts to age.

## Layer 1: Tool call logs

Log every tool call: request parameters, response summary, session ID, timestamp. No conversation text, no user questions, no agent reasoning. Just the MCP layer.

This is the foundation, and it's the instrument for the friction case. Three consecutive calls with tightening filters on the same table means the agent is struggling with something the metadata should have covered. You don't need the agent to tell you it's confused. The call pattern tells you.

In the production implementation, every tool call writes to a persistent store with: tool name, user, queryIntent, filter fields and operators (not values, for privacy), summaryOnly flag, row count, duration, and error type. Session IDs correlate multi-tool sequences. The store keeps two years, and on the HR server, which carries special-category health data under GDPR article 9, the intent string is redacted before it's ever written.

## Layer 2: Pattern analysis

Weekly analysis of the raw logs. Look for repeated call sequences, redundant calls, unused tools that should have been used, tools called in unexpected order. This is the one layer that costs recurring human time. It sits inside the hour a week or so that twelve servers take to keep current, and it is the first thing to lapse when the week gets busy.

Session IDs make this possible. You can trace an entire multi-tool interaction from start to finish and see where the agent took a detour: a call to tool A when the answer was in tool B, a retry with different filters because the first call didn't return what it expected.

This is also where some of the silent-wrong cases surface, when a detour leaves a trace even though the agent never flagged anything. Not all of them. The ones that leave no trace at all are why a human stays in the loop.

## Layer 3: QueryIntent and the report tool

A queryIntent parameter on nearly every tool call: one sentence describing the business question being answered. Plus a report tool for the moments a person decides a problem is worth writing down.

The same three-call sequence, with and without queryIntent:

**Without:**

> `09:14:03 | get_service_records | summaryOnly=true, project=P-7056`
>
> `09:14:07 | get_service_records | type=13, project=P-7056`
>
> `09:14:09 | get_service_records | type=13, project=P-7056, completed=true`

Three calls. The last one added a filter. Why? Unknown.

**With:**

> `09:14:03 | intent: Overview of all service records for project P-7056`
>
> `09:14:07 | intent: Which maintenance orders exist for this project`
>
> `09:14:09 | intent: Only completed orders, previous call also returned open items`

Now you see exactly what happened: the agent expected only completed items but the metadata didn't make clear that a completion filter is needed. That's a targeted metadata fix requiring minutes to implement.

The report tool, `report_problem`, is the second instrument. A person decides a report is worth filing; the agent composes it. One detail matters more than any other: its most important input is not the complaint, it's the reasoning. The field asks the agent for its chain of thought, and the schema flags it as the "most valuable field." The report captures what the agent thought it was doing when it got stuck, not just that it got stuck. That's why the reports are worth reading.

So far, 82 have come in, 75 of them already worked through. The categories are what you'd expect from a system that mostly produces friction rather than errors: missing data, ambiguous semantics, pagination struggles, cross-tool mismatches, misleading descriptions. The most-reported tools are the two richest ones, `get_project` and `get_nacalculatie`. At least eight of those reports trace cleanly from a filed problem to a shipped fix. One of them traces to a "won't fix": an architectural limit of the connector-per-entity design that's honest to leave in place rather than paper over.

## Nudge patterns: making feedback happen

Agents don't self-report reliably. So the production implementation injects contextual nudges into tool responses:

- **Empty results:** "No records matched your filters. If unexpected, consider calling report_problem."
- **Pagination detected:** "You're paginating (skip=200). If you're doing this to manually aggregate data, call report_problem. We may be able to add that summary dimension."
- **Every response:** A feedback reminder as the last alert: "If this query took multiple attempts or returned confusing data, call report_problem before your next step."

These nudges are injected by the server's handler factory, not by the agent's judgment. They target specific friction patterns and make that friction visible in the response. But they don't close the loop on their own, and that turned out to be the more interesting finding. Even with the nudge sitting right there in the response, the agent almost never files on its own. It relays the friction, or the user feels it, and a person decides to report. Agents don't take initiative on meta-tasks, even when you ask them to. The nudge surfaces the moment. A human still makes the call.

## Validated knowledge isn't frozen

Here's the part the first version of this post skipped. Once a field has been validated and encoded, what re-checks it when the source system changes underneath?

Validation isn't a one-time event. Production keeps testing it. Three kinds of drift, and they are not equally hard.

**Structural drift is cheap to catch.** A supplier renamed its entire route surface on a pre-release API. Four of five tools broke, and the way they broke is the point: old payload keys came back as empty arrays that looked like real data. The failure wasn't an error. It was plausible emptiness. The fix was to pin the routes in one place and add a probe that asserts the live contract, so a silent rename fails loudly next time.

**Schema drift gets caught at runtime.** The handler emits a drift alert whenever the source returns fields the tool doesn't declare. A concrete one: `get_dossier_reacties` had a field typed as a number that the ERP actually returns as a text label. Every record quietly failed output validation until the schema was changed to a string. That is exactly the confident-and-wrong category: a validated assumption that was wrong, caught in production rather than in review.

**Semantic drift is the hard one, and only use surfaces it.** Everything still works technically. Every field validates. The understanding is just wrong, or has aged. No probe catches this. What catches it is friction in the logs, an explicit report, or a domain expert. That's the class `Aantal` belongs to.

I'll name the honest limit here, because it matters. When I wrote this, there was no eval framework. What existed was the ability to re-run the discovery loop against the live data and the current description. That catches a different class than an eval would: a frozen test set with recorded fixtures would never have caught the renamed routes; a re-run against live data would. What was missing was the behavioural half. Whether a description change actually made the agent act better was unmeasured.

The next step I wanted was a small set of questions whose answers a domain expert has already verified, re-run periodically. Not a benchmark. A regression test built out of validated truth. It's the only thing that would catch the silent-wrong class on a schedule instead of by luck.

*Updated September 2026:* I built it, on the public reference server rather than on production data: registered questions with known answers, a prediction written down before each run, more than 3,700 scored runs, and the measured versions frozen so a later edit can't quietly change them ([eval register](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md)). It found things no log would have shown me. Only the first 2,048 characters of a tool description reach the model on Claude Code, and one plausible line in my own richest tool made 59 of 60 answers wrong. The production servers don't have their own eval set yet. That is the next step now, and the live re-runs stay, because they still catch what fixtures can't.

## The loop doesn't just improve the docs. It evolves the interface.

The framing so far has been metadata: production finds a gap, you sharpen a description. That's the common case, but it undersells what actually happens. Sometimes the recurring intent in the logs doesn't point at a missing sentence. It points at a missing capability, and the answer is to change the agent-facing interface.

| Recurring production intent | Capability that emerged |
|---|---|
| "I only need part of this object" | Server-side field projection (`select`) |
| "I need to understand this large dataset" | Typed summaries with breakdowns |
| "Tell me if anything important is happening" | Conditional runtime alerts |
| "What does this combination of facts mean?" | Derived values |

I want to be honest about where these came from, because "they all emerged from telemetry" would be a tidy lie. Summaries were core design from the start. Projection came from efficiency pressure. The alerts have mixed origins, one of them an incident. Derived values like a charging status or an open-absence flag deliver interpretation as a value instead of asking the agent to reconstruct it from prose. What the telemetry gave me wasn't the ideas. It was the evidence that these particular ones recurred enough to be worth building.

The payoff is concrete. On the connectors that got projection, most full-record calls now request only the fields they need. One ticket read the same object twice, 24 seconds apart: the projected read was 3,969 bytes, the full body 265,418. That's roughly 67 times cheaper for the same answer.

## Why this compounds

Each resolved issue makes the system permanently smarter. A metadata update isn't a patch. It eliminates an entire category of uncertainty for every future session, every connected agent. When the fix is a new capability rather than a new sentence, the same is true a level up: the interface itself gets better at being used correctly. The metadata layer grows richer over time, the interface evolves, and the reasoning layer stays the same. Cheap, permanent, compounding progress.

---

This post is one piece of a longer argument. [*Production MCP: A Practitioner's Guide*](https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide) puts all of them in order, from understanding your data through to identity-bound deployment.

Go deeper: read the full practitioner report, [*The Missing Layer*](https://davidgolverdingen.nl/en/the-missing-layer), or explore the [mcp-metadata-demo](https://github.com/DaveGold/mcp-metadata-demo) server, an open-source extract of these patterns, since the production servers run on private business data.
