---
title: "The Missing Layer"
description: "Practitioner report on why enterprise AI can fail and the Rich Domain MCP Server pattern."
canonical: "https://davidgolverdingen.nl/en/the-missing-layer"
last-updated: "2026-09-25"
---

# Rich Domain MCP Servers, The Missing Layer

**Why AI on Enterprise Data Might Fail, and a Pattern That Could Fix It**

**David Golverdingen** · Senior Engineer & MCP Architect
*Living document, first published April 2026. The current revision always lives at
[davidgolverdingen.nl/en/the-missing-layer](https://davidgolverdingen.nl/en/the-missing-layer); the
"Updated" date in the page header is authoritative.*

> *Only about 5% of enterprise AI pilots achieve rapid revenue acceleration, according to MIT's NANDA initiative, the vast majority stall. Not because AI doesn't work. Not because the data is bad. But because enterprise implementations fail to adapt AI to their specific context, rarely does anyone tell the AI what the data **means** at the point where it reasons.*

*Built on a documented production snapshot of eleven MCP servers and 91 tools. Among them: a mid-market ERP system (11 tools), a BIM data explorer (Autodesk AEC Data Model, 6 tools), an energy monitoring platform (a mixed server aggregating a smart meter data provider, a building automation platform, a meteorological API, and two government building registries into a single domain interface, 13 tools), a construction standards database covering both building and product classification (7 tools), a pre-order calculation system (5 tools), a multi-million-item product catalog (6 tools), a visualization server (6 tools), a games server (2 tools), and two internal operations servers for data exploration and ticketing (23 tools). One cross-server feedback tool serves the whole estate. The estate grew to sixteen servers and 103 tools by the end of September, eight of those tools MCP Apps. Rich domain metadata, a three-layer feedback architecture, server-rendered interactive MCP Apps, and a deployed secure write pattern. Serving a Dutch installation company with roughly 350 employees. The systems named in this paper include AFAS Profit (ERP), Gilde Pro (pre-order calculation), Energiepartners and Priva Cloud (energy monitoring), Autodesk AEC Data Model (BIM), Ketenstandaard (NL-SfB and ETIM construction standards), Compano Select (product catalog), and the Dutch Kadaster BAG and EP-Online registries (building data).*

*Since September 2026 the design rules are also measured. A public reference server on Dutch building data runs three tiers side by side, and more than 3,700 scored eval runs show which of the knowledge actually reaches the model [28][29].*

---

## Abstract

Enterprise AI can fail, not because models lack capability, but because the domain knowledge rarely reaches the AI at the point where it reasons, the tool interface. This paper introduces **Introspective Context Engineering for MCP** (ICE), a five-step loop where AI **examines** a data source, **flags** every pattern it finds with a confidence level, a domain expert **validates** the uncertain ones, the confirmed knowledge is **encoded** into the parts of the interface the model actually receives (field names, the head of the tool description, the input schema and the tool's response), and production telemetry drives the next **iteration**. Evals run inside that loop: they measure whether what was encoded reaches the model and changes what it does. The result is a distinct category of AI-data integration: the **Rich Domain MCP Server**. Unlike RAG pipelines, Text-to-SQL engines, vendor-built copilots, or client-side skill files, this approach places domain intelligence directly in the tool interface where the LLM reasons over it at call time. The paper proposes a six-level maturity model for MCP servers, from bare API mappers through self-teaching domain servers to interactive MCP Apps with secure write operations, and presents six concrete recommendations for the MCP framework specification. The documented snapshot covers eleven MCP servers totaling 91 tools and eleven external APIs, across domains including ERP, pre-order calculation, BIM, energy monitoring, construction standards, and product data; by the end of September, the production estate had grown to sixteen servers and 103 tools. It serves a roughly 350-person Dutch installation company. The first server took one developer roughly 40 hours with no prior MCP experience; subsequent servers took 1–2 days by reusing the established patterns.

**Scope and evidence.** The pattern comes from production. The evidence for its design rules now comes from a public repository anyone can run [28]. It puts a thin, a rich and a best tier of the same building-data server on three live endpoints, and publishes the registered questions, the prediction made before each run, every scored run, and an evidence register that gives each rule a status: settled, direction, null or reversed [29]. Where this paper states a design rule, it cites the register entry behind it, for example ([Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). Those results have four standing limits. One author wrote the metadata, the questions, the ground truth and the scoring. They cover one domain family (Dutch building and weather data), one host for every effect on answers (Claude Code, with what arrives also measured on five other clients [HD](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)) and one model family (Claude Haiku, Sonnet and Opus). What the repository cannot reproduce is marked as production: the estate, its telemetry, its build times and its expert reviews. Those are observations from one company, not measurements. The production estate has not yet been migrated to everything the evals found, and where the two differ the paper says so.

---

## Contents

1. [The Context Gap](#01-the-context-gap), Why $37B in enterprise AI investment keeps falling short
2. [MCP Isn't Dead. It's Empty.](#02-mcp-isnt-dead-its-empty), Challenging the CLI/skills narrative
3. [The Missing Layer](#03-the-missing-layer), What is rarely built and why
4. [The Pattern: Introspective Context Engineering for MCP](#04-the-pattern-introspective-context-engineering-for-mcp), The five-step loop with implementation detail
5. [Before and After](#05-before-and-after), Level 1 vs the eval-derived reference in four layers
6. [Why It Works](#06-why-it-works), What the evals measured, feedback in practice, compound improvement, agent-agnostic
7. [Honest Limitations](#07-honest-limitations), What this pattern does not solve
8. [Recommendations for the MCP Framework](#08-recommendations-for-the-mcp-framework), Six changes that would raise the floor
9. [How to Apply This Monday Morning](#09-how-to-apply-this-monday-morning), A practical starting point
10. [The Bigger Picture](#10-the-bigger-picture), Why correct metadata is future-proof

---

## 01 The Context Gap

Enterprise AI investment tripled in a single year to $37 billion in 2025 [18]. The results don't match the spend. MIT's NANDA initiative analyzed 300+ deployments and found that only about 5% achieve rapid revenue acceleration [1]. S&P Global reports that 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before [2]. RAND concludes that over 80% of AI projects fail, twice the failure rate of non-AI IT projects [3].

The industry blames data quality. Informatica names it the number one obstacle at 43% [4]. Gartner predicts 50%+ of generative AI projects stall in the pilot phase for the same reason [5]. But MIT identifies something deeper: not a data problem, but an adaptation failure, generic AI tools excel for individuals, but stall in enterprise use because they don't learn from or adapt to specific workflows [1]. Companies build the highway but forget the navigation system.

All of the figures above describe 2025. Enterprise AI adoption moves quickly, and by the time you're reading this the picture may already look better than these numbers suggest; the pattern this paper addresses, the gap between having data and having a system that understands it, doesn't depend on the exact failure rate.

Clean data is not the same as understood data. Consider a service management ERP. A maintenance record has an "assigned technician" field. The AI reads it, finds it empty, and tells the user no technician was assigned. In reality, that field is never populated, the actual technician can only be found through time-booking records filtered by a specific work-order type. The data is perfectly clean. The AI simply doesn't have the interpretation layer to know that this field is a dead end. This is a data interpretation problem, not a data quality problem.

The market response follows a consistent pattern: solve the transport problem. RAG embeds documents in vector stores, useful for policy questions, useless for live operational queries [6]. Text-to-SQL translates natural language to syntax, but doesn't know that a 'type' field needs a specific numeric code or that certain schema fields are dead ends [7]. Vendor-built AI (SAP Joule, Microsoft Copilot for Dynamics 365, Oracle NetSuite AI) has intimate platform knowledge but mostly serves standard workflows, not free natural language access across domain-specific data models [8]. Enterprise data integration platforms (Informatica, K2view) connect the pipes but don't understand what flows through them [9]. In all cases, knowledge lives in infrastructure layers (vector stores, prompt templates, governance dashboards) rather than in the tool interface where the LLM actually reasons at call time.

Meanwhile, the Model Context Protocol (MCP), adopted by Anthropic, OpenAI, Google, and Microsoft, is becoming the universal standard for connecting AI to tools and data [10]. With over 97 million monthly SDK downloads and 10,000+ active servers as of December 2025 [10], the ecosystem is growing rapidly. But transport alone is not enough.

The author's analysis of 50+ publicly available MCP servers, conducted in March 2026, reveals a consistent maturity distribution:

| Level | Characteristics | Examples |
|---|---|---|
| **1. API Mapper** (~70%) | One-sentence descriptions. No domain context. AI must figure everything out alone. | Postgres, Filesystem, most community servers, generic ERP connectors, data platform wrappers |
| **2. Functional** (~20%) | Grouped tools, better organization. Still no domain knowledge or cross-references. | GitHub Official, Salesforce DX, Microsoft Dynamics 365 ERP MCP |
| **3. Metadata-Rich** (~8%) | Knowledge graphs, glossaries, lineage. Metadata typically curated manually by dedicated teams over weeks or months. | OpenMetadata, DataHub, Actian, Oracle Analytics Cloud |
| **4. Self-Teaching** (<2%) | AI-discovered domain knowledge, cross-tool references, query strategies, fallback routes, uncertainty markers. Metadata generated and maintained by AI with human validation. | Introspective Context Engineering for MCP (this paper). Possibly emerging in internal enterprise tools not publicly visible. |
| **5. Interactive App** (emerging) | Server-rendered interactive UI components (charts, tables, maps, forms) returned directly in the conversation [26]. The agent coordinates; the server controls presentation. Data tools + visualization tools in the same ecosystem. | Production implementation in this paper: render_chart, render_table, render_map, start_duurzaam (energy analysis intake). |
| **6. Secure Write App** (frontier) | Graduated write operations: agent-initiated bounded writes (feedback, scores) through standard tools for low-risk mutations; user-initiated secure writes through validated MCP App interactions for business-critical data. WriteIntent pattern: agent opens the app, user drives the mutation, server enforces every boundary (one-shot, user-bound, time-limited, typed schema). | Production: the WriteIntent pattern is deployed, `start_graph_comments` / `commit_graph_comments` commit user-authored annotations onto energy data through an app-only commit tool, alongside agent-initiated bounded writes (report_problem, submit_tetris_score) and role-gated ticket writes. Designed: ERP data corrections through the same form pattern. |

The progression across levels follows a clear arc: expose data (Level 1) → organize tools (Level 2) → understand the domain (Level 3) → learn from data and feedback (Level 4) → present through interactive apps (Level 5) → act through secure writes (Level 6).

Even the MCP specification's own best-practice proposal holds up this as a good tool description: "Read the contents of multiple files simultaneously. More efficient than reading files individually when analyzing or comparing multiple files." One sentence. That is the official example of good, and it falls well short of what an agent needs to use a tool correctly. The proposal itself was not rejected on its merits, it sat open for five months and was closed as dormant in January 2026 for lack of a sponsor on the core maintainer team [11].

**The difference is not subtle.** A Level 1 server says "query data from the ERP." A Level 4 server says "always start with a summary call, active projects have thousands of records. Type codes determine which fields are populated. Use the budget tool for planned costs, this tool for actuals. Report friction via the feedback tool." Not the same product. Not the same category.

## 02 MCP Isn't Dead. It's Empty.

A growing discourse argues that MCP is dead, that CLIs and skill files make the protocol unnecessary [19][20]. The frustration is real, and CLIs are genuinely simpler for developers already living in a terminal. The counterarguments are equally real: MCP, skills, and CLIs solve different problems [21][22]. Both sides are right about transport, and arguing about the wrong layer.

The CLI/skills argument breaks down for enterprise data. A product manager asking about project margins won't `pip install` a CLI; a service coordinator won't write a `CLAUDE.md`. These users need a one-click OAuth and tools that just appear, exactly what MCP provides. And even with flawless transport, an estimated 70% of servers would still produce unreliable answers because they contain little to no domain knowledge. Wrapping the same hollow APIs in a CLI doesn't fix that.

### What CLIs and Skills Cannot Replace

Four things only the server can provide: **server-side access control** (per-tool scoping verified against JWT claims, instead of whatever access the binary has); **structured tool metadata** (rich descriptions and typed input schemas, parsed by the model at tool-selection time, not stale `--help` output); **runtime feedback loops** (`report_problem`, `queryIntent`, session-correlated logs; reinventing these in a CLI is reinventing MCP, badly); and **enterprise governance** (audit trails, centralized auth, session correlation).

### Skills: A Client-Side Patch for a Server-Side Problem

Anthropic's skills guide [25] diagnoses the problem correctly, users connect an MCP server but don't know what to do next, and prescribes a client-side "knowledge layer" of markdown recipes. The diagnosis is right; the prescription treats the symptom. Skills exist because MCP servers are Level 1: if the server already taught the agent what the data means, when to use which tool, and how to chain calls, there would be nothing left for the skill to teach. The kitchen has no recipes, so they hand out cookbooks at the door.

Distributing knowledge client-side reintroduces the problems skills claim to solve: it goes stale when the server updates, fragments across clients and versions, and only helps the one client that loaded it, a skill uploaded to Claude.ai doesn't help the same server when connected to Cursor, Copilot, or a custom agent. The natural home for domain expertise about a service is inside that service. When the server owns the knowledge, every connected agent benefits automatically.

To be precise about the objection: it is about *where domain knowledge lives at call time*, not about the format. A skill is a poor container for what the agent needs while reasoning over a tool, and a good container for what a *builder* needs while writing that tool. The method in this paper is itself distributed as a skill, it teaches discovery and encoding, and ships no domain facts whatsoever.

MCP isn't dead. It's empty. The protocol works; the servers are hollow. The fix isn't a different transport or a client-side knowledge layer, it's filling existing servers with domain intelligence.

## 03 The Missing Layer

The transport problem is solved. What remains is the knowledge layer: metadata that tells the agent how the domain works, what the codes mean, which fields are reliable and which aren't, how to cross-reference between tables when the obvious path doesn't work, what patterns indicate problems versus normal operations.

### Why is it rarely built?

**Developer mindset.** Most MCP servers are built by developers wrapping APIs. Their mental model: "I make the API available, the LLM is smart enough to figure out the rest." For well-known domains, think GitHub, Slack, file systems. This works. The LLM knows what a pull request is from its training data. But mid-market ERP systems with industry-specific configurations? Equipment type codes, maintenance scheduling rules, subcontractor assignment patterns? None of that is in any model's training data.

**No complex domain.** Most MCP servers don't have this problem because their domain is universally understood. A GitHub MCP server doesn't need to explain what a commit is. An ERP MCP server does need to explain what a service ticket (type 10) means versus a maintenance order (type 13) versus a service planning (type 48), and how they relate hierarchically.

**No feedback loop.** The build-vs-buy analysis for MCP servers focuses almost entirely on transport capabilities [17]. Most MCP server builders ship to a package registry, publish, and move to the next project. They don't have five power users hitting limits daily, revealing which metadata is missing and which interpretations lead to wrong answers. Without that feedback, you don't know what's missing.

**New technology, unknown possibilities.** MCP was introduced by Anthropic in late 2024 and is still maturing as a protocol. Most developers and organizations are still learning what MCP is, let alone what it can do beyond basic API wrapping. When a technology is this young, exploration tends to focus on getting it to work at all, not on pushing the boundaries of what metadata it can carry. The idea that an MCP tool could carry domain knowledge, query strategies, and cross-tool references simply hasn't occurred to most builders yet. The builders who did try it tended to write that knowledge as long description prose, and the evals later showed that most of a long description never reaches the model (see [Which metadata channels actually work](#which-metadata-channels-actually-work)). The ecosystem needs time to discover its own possibilities, and to measure them.

The idea to enrich metadata is logical. But logical and obvious are different things. It is also logical to write unit tests and maintain documentation, not everyone does. McKinsey's 2025 AI Survey found that AI high performers are nearly three times more likely to have fundamentally redesigned workflows before selecting AI techniques [16]. The innovation here is not in the insight. It is in the degree of enrichment, the method (AI discovers domain and writes back), and the feedback mechanism to keep improving. The difference is in execution and systematization.

## 04 The Pattern: Introspective Context Engineering for MCP

Introspective Context Engineering for MCP (ICE) is a five-step loop:

| Step | What happens | Who does it |
|---|---|---|
| **Examine** | Point an AI at the real data, not the schema, and direct it through differentiated slices | AI, directed |
| **Flag** | Every pattern is tagged with a confidence level and the question it raises | AI |
| **Validate** | The uncertain findings are confirmed, corrected, or killed, and a small eval measures whether the model uses what was encoded | Domain expert, plus an eval |
| **Encode** | Validated knowledge is written into what the model receives: field names, the description head, the input schema, the response | AI, reviewed |
| **Iterate** | Production telemetry exposes the next gap and the loop runs again | The system |

The loop inverts the traditional metadata-curation pipeline. Instead of asking humans to write down everything they know, which does not scale past the second tool, the AI asks the questions and the human approves the answers, which does. Two one-time setup steps precede the first pass, and one step follows the last: the evals written along the way are kept as regression tests, so a later edit that breaks what the model used to get right fails a test instead of reaching a user. Everything in between repeats. In production, three or four passes per tool has been normal before the surprises stop; that is practice, not a measured figure.

The following sections detail each step with concrete implementation patterns drawn from a production deployment. Where previous descriptions of this pattern were necessarily abstract, these guidelines are grounded in direct experience across four different data domains (ERP financial data, BIM building models, energy monitoring, and construction standards) and documented in an operational metadata guide [23].

### Before the loop, research and guidelines

Two things are established once per organisation, not once per tool: which metadata channels the target clients actually honour, and the house style every tool description will follow.

#### Which metadata channels actually work

Have the AI research best practices for MCP tool design and metadata standards. Study what makes tool descriptions effective for AI reasoning. Analyze existing high-quality MCP servers for patterns worth adopting.

**What specifically to research.** The research phase should produce a clear understanding of which metadata delivery channels actually work, and the word that matters is *delivered*. A server can bind knowledge to a tool without the host ever showing it to the model. Earlier versions of this paper ranked the channels from production testing across several Claude clients [23]. The public evals have since measured them, first on Claude Code, and the table below is that measurement [29]. A second table further down covers the other hosts. They differ, so treat both as the method for finding out, and re-run the check against the hosts and models you actually deploy to.

| What the server ships | Reaches the model? | Evidence |
|---|---|---|
| Field names in the response | Always | the one carrier that is guaranteed to arrive |
| Tool description | Only the first **2,048 characters**, re-sent on every request. Keep the head under about 1,800 | [Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery) |
| Server instructions | Only the first 2,048 characters | [Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery) |
| Input schema, with its `.describe()` annotations | On Claude Code in full (measured up to about 21,700 characters per tool), re-sent every turn, so every word is paid on every call | [IS](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery), [Q25](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19) |
| The tool's response | Yes, below about 25k tokens. A larger result is replaced by a "saved to file" notice, and the guidance inside it is lost to any model without file tools | [Q9](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery) |
| Output schema | Not on Claude Code, and not on Cowork or claude.ai chat. Codex and ChatGPT Work do deliver it. It validates the response and drives UI rendering | [Q11](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery) |
| A separate guidance tool | Only when the pointer to it is worded as a requirement | [Q8, Q8b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery) |
| MCP Resources | Not requested by agents in production, and not measured in the eval set | production [23] |

The cut is silent. In the rich tier of the demo, 74% of the `get_building_profile` description never arrived, and nothing told the author. On Claude Code, a sentence that sits past character 2,048 does not exist for the model, however carefully it was written.

Other hosts cut somewhere else, or not at all. I measured five more in September 2026 with the same rich-tier server ([HD](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)):

| Host | Tool description | Server instructions | Output schema |
|---|---|---|---|
| Claude Code | Cut at 2,048 | Cut at 2,048 | Not delivered |
| Cowork (web and desktop app) | Cut at 4,096 | Whole | Not delivered |
| claude.ai chat | Whole, after a tool search | Not delivered | Not delivered |
| ChatGPT Chat | Whole | Cut at 512 | Sometimes (1 of 3) |
| ChatGPT Work | Whole | Not delivered | Delivered |
| Codex CLI | Whole, once the model looks it up | Prepended to every tool | Delivered |

Server instructions are the weakest surface. Codex and ChatGPT Work also drop every `.describe()` in an input schema once the serialized schema passes 5,000 characters. A server that has to work everywhere therefore keeps its instructions within 512 characters, each line repeated in a description head or the response, and each input schema under 5,000.

**Key validated insight: delivery, not channel.** Once a sentence is delivered, where it sits hardly matters. The same one-line rule scored 20 of 20 at character 330 of the description and 20 of 20 in the response, against 1 of 20 when it sat past the cut ([Q15](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). An earlier reading of these evals, that the response simply beats the description, turned out to be the cut and not the channel, and was reversed. Two practical consequences follow. What the model needs *before* it calls a tool goes in the description head and the input schema. What it needs to *read the result* goes in the response, next to the record it applies to, where it is sized to that record and cannot be cut off by a character budget. Raising a client's cap is no fix either: it delivered the block but added 23.7% tokens to every tool on every request ([Q15b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)).

A guidance tool, one the agent calls to fetch a recipe before it acts, is a working channel after all. The production estate had removed its meta-tools because agents rarely called them. The evals show why: a soft pointer to the tool was followed 0 times in 10 by Haiku, while the same pointer worded as "REQUIRED: call this first" was followed 10 times in 10, at a cost of under 1% in tokens ([Q8b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). It works better as a tool of its own than as a mode of the tool it guides, 110 of 120 against 78 of 120 ([Q25c](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)). MCP Resources are a different matter: agents in production never requested them, and the evals did not test them.

Two additional spec features were evaluated and set aside. Sampling (server-initiated LLM reasoning requests) is unnecessary when delivered metadata already guides the agent's reasoning, the interpretation the server returns and the input schema's `.describe()` annotations do what sampling would do, but through metadata that's always present rather than a runtime request. Elicitation (server-initiated user input requests) is unsupported by key clients including Claude Desktop; MCP Apps provide a more capable alternative with full interactive forms, typed validation, and multi-step workflows. Both cases follow the same principle: work with what clients actually support today, not what the spec promises.

#### The guidelines document

Have the AI create a comprehensive guideline document: what types of metadata to include, how to structure cross-references, what format to use for query strategies and fallback routes. This becomes the blueprint for every Encode step that follows.

**Description block structure.** The guidelines should define a standard set of blocks for every tool description. This block structure is not adopted from an existing standard, it emerged iteratively during production implementation, driven by observed agent failure modes. Each block was added because the agent failed without it. Recent academic research independently validates the need: an analysis of 856 tools across 103 MCP servers found that 97.1% of tool descriptions contain at least one "smell", unstated limitations, missing usage guidelines, opaque parameters [24]. That sample was drawn from public, published servers; no comparable study yet exists for private, internal deployments, so the current baseline inside a typical company is unknown. The block structure below addresses exactly these deficiencies. In the production implementation, each connector uses these blocks [23]. The evals changed where two of them go. Every block still has to arrive, and a description only arrives up to its cut, so the blocks the agent needs before the call sit in the description head, and the blocks about reading a result are split: the rule that holds for every record stays in the head, the part that depends on this record travels in the response.

| Block | Purpose | Without it... | Where it goes |
|---|---|---|---|
| **RETURNS** | Agent knows which fields come back, drives both tool selection and which fields to request | Agent can't determine if this tool has the data it needs | Description head, with the literal field names |
| **WHEN TO USE** | Agent knows when this tool fits | Picks wrong tool or misses this one | Description head |
| **WHEN NOT TO USE** | Prevents wrong tool selection; points to the correct tool | Agent tries this tool for queries that belong elsewhere | Description head |
| **QUERY STRATEGY** | Teaches summary-first, then drill down | Agent fetches full records when a summary would suffice | Description head, and the input schema's `.describe()` |
| **INTERPRETATION** | Cross-field behavioral rules, type→field-group mappings | Returns raw numbers without conclusions | Universal rules and a "read `interpretation` first" pointer in the head; this record's rules in the response, under `interpretation.notes` |
| **RELATED TOOLS** | Agent chains queries automatically via join keys | Stops after first tool call | Description head |
| **FEEDBACK** | Agent reports friction to report_problem | Issues go undetected, descriptions never improve | Description head or server instructions |
| **ALERTS** | Tells agent to expect and surface server-generated warnings from the response | Agent ignores domain-specific warnings (unapproved records, poor equipment condition, partial summaries) | A pointer in the head; the warnings themselves in the response, under `interpretation.alerts` |

Blocks are earned, not templated: a block belongs in a description because an agent failed without it, not because the template has a slot for it. The canonical block names cost nothing to keep; the same content under the eight names and in their order showed no regression ([Q19h](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)).

**Language strategy.** LLMs internally process through English representations, even for non-English input. The production implementation uses an English-with-glosses bridge: tool descriptions, block headers, and .describe() annotations are written in English, but domain-specific terms from the source system are kept in their original language with an English gloss, e.g., "Medewerker (employee ID)" or "Geaccordeerd (approved)." Tool titles stay in the user's language for UI display; descriptions in English for agent reasoning. The agent responds in the user's language automatically [23].

**A note on examples.** The production system uses Dutch field names and tool names throughout; the ERP is a Dutch product serving a Dutch organization. Examples in this paper have been translated to English for readability (e.g., get_employee was originally get_medewerker, EmployeeName was MedewerkerNaam). The pattern and the language bridge apply regardless of source language.

**Factory patterns.** The guidelines should specify shared patterns that eliminate boilerplate. In the production implementation, every connector shares the same input schema factory (filters, pagination, sorting, summaryOnly, queryIntent) and the same handler factory (permission check → build params → call connector → sanitize dates → transform → summarize → validate output → return). Each connector only supplies its field metadata, an optional transform callback for derived fields, and an optional summarize callback for domain-specific aggregation. The same factory pattern was implemented for both REST (ERP, energy) and GraphQL (BIM) backends, the handler abstraction is API-style agnostic [23].

### Examine, point the AI at the real data

Have the AI iterate over the actual rows rather than the schema definition. But this is not a passive scan. You actively direct the exploration: instruct the agent to apply specific filters, sort by different fields, compare subsets. Crucially, force the agent to look beyond the obvious. The default behavior of any AI is to fetch the first page of results and draw conclusions from the top 10 rows, but the top 10 rows are rarely representative. A table sorted by date shows only the most recent records. Sorted by ID, only the oldest. The interesting patterns live in the edges: the records where fields are unexpectedly null, where status codes don't match the expected flow, where cross-references break down.

**This is an iterative process.** The first pass reveals surface-level patterns. The second pass, informed by what you learned, goes deeper. By the third or fourth pass, you're discovering business rules that aren't documented anywhere, because the people who know them never thought to write them down. In the production implementation, the discovery phase was repeated across multiple sessions over several days. Each iteration refined the understanding: early passes caught null patterns and obvious code meanings; later passes uncovered conditional field population rules, cross-table join reliability issues, and exception patterns that only appeared when filtering for specific record types.

**Direct the agent explicitly.** Don't just say "examine the data." Give it specific instructions: filter by a particular type or status, sort by a numeric field descending, skip the first 50 records and look at what's underneath. Then change the filter and compare results. Then look at records where a key field is unexpectedly null and ask what other fields those records have in common. Each directed query teaches the AI something specific, and each answer informs the next question. You are guiding a structured exploration, not requesting a report.

This is where the AI finds things no API documentation contains. In a production ERP deployment, the AI discovered that the "assigned technician" field on maintenance location records was never populated despite being in the schema, that the actual technician could only be found through time-booking records filtered by a specific work-order type, and that equipment condition scores did not correlate with age, a 2001 installation could score better than a 2008 one due to maintenance quality. None of these discoveries came from the first query. They emerged from directed, repeated exploration of differentiated data slices.

**What the AI should look for:**

-   **Null patterns:** which fields are consistently empty for which record types? In the production system, the AI discovered that EmployeeName is always null for article bookings (TypeItem='Mat') and TicketId is null for project bookings but populated for service bookings. These null patterns belong in the tool description, where they prevent the agent from expecting data where there is none.

-   **Code-to-meaning mappings:** what do the values actually mean? Status codes, type codes, category codes; the AI reads real data and documents what each value represents in practice, not in theory.

-   **Cross-table join paths:** which ID fields connect which tables, and are those connections reliable? The AI discovers that TicketId in the actual-costs table links to the service ticket table, that EmployeeId is a join key to the employee table, and that TicketLineNumber on maintenance orders maps 1:1 to subscription line numbers.

-   **Business rules hidden in data:** the AI finds that service bookings on numeric-only project IDs (e.g. 6600xxx) are contract-covered and don't generate per-visit invoices, that project ID letter prefixes encode project types, and that the invoicing flow follows Approved → Completed → Invoiced sequentially.

-   **Cross-tool relationships:** when your server exposes multiple tools, instruct the agent to explore how they relate. Have it fetch a record from one tool, then use an ID field from that record to query another tool. Does the join work? Is the ID always populated? Does the related record add context that the first tool alone couldn't provide? In the production implementation, the AI discovered that the planned-vs-actual-vs-invoiced triangle (budget → actual costs → invoice lines) shared ProjectId as a reliable join key, but that the service ticket link only worked for service bookings, not project bookings. These cross-tool relationships become the RELATED TOOLS block in tool descriptions; that model-visible guidance enables the agent to chain tool calls instead of stopping after the first one.

### Flag, tag every pattern with a confidence level

When the AI encounters patterns it can observe but not explain (a status code that appears in unexpected contexts, a field that's populated for some record types but not others, a value distribution that doesn't match the schema documentation), instruct it to write a TODO marker with a confidence level immediately. Not after examination, *during* it: "When you find something, tag it: HIGH if confirmed across many records, MEDIUM if observed but possibly incomplete, LOW if inferred from limited data. Always add a `[CONFIDENCE: <level> — <observation>. TODO: DOMAIN EXPERT — <question>]` note."

Flagging after the fact does not work. By then the ambiguity has already been resolved with the model's best guess, and the guess reads exactly like an observation, for enterprise-specific domain knowledge it is often wrong, and invisibly so. The marker is what keeps the distinction between *observed* and *assumed* legible to the next reader, human or agent.

Every marker carries the same three parts: the confidence level, what was observed, and a specific question for a domain expert.

> [CONFIDENCE: HIGH — observed across thousands of records.
>
> TODO: DOMAIN EXPERT — is the 0.17-8 hour range for labor
>
> bookings correct? Are there legitimate bookings outside it?]

> [CONFIDENCE: MEDIUM — observed flow, but the ERP may allow
>
> skipping steps. TODO: DOMAIN EXPERT — is the invoicing
>
> flow Approved → Completed → Invoiced strictly
>
> sequential, or can steps be skipped?]

> [CONFIDENCE: LOW — only observed in test environment.
>
> TODO: DOMAIN EXPERT — is this flag a test artifact,
>
> or does it occur in production?]

One discipline governs the level: sample size. A pattern seen in fewer than a hundred rows of a single slice is LOW however convincing it looks. That threshold is a working rule from production; nobody has measured it. And if another query could settle the question, run the query, only genuine business-knowledge questions should reach a human.

### Validate, the domain expert confirms or kills

Flagged findings are hypotheses. What turns them into knowledge is someone who works with the data daily, and this is the step that is easiest to skip and most expensive to skip. In the production deployment (not reproducible from the public repository), domain-expert review has been completed across the ERP connectors: roughly **90% of the AI-discovered metadata held up under review, with the remaining ~10% corrected**. That 10% is precisely the part that would otherwise have shipped as fluent, confident and wrong, a bad description produces a plausible answer, not an error, so nothing downstream catches it.

The inversion is what makes this affordable. Asking a domain expert to dictate 500 words of operational guidance per tool produces dry, incomplete text and stalls after the second tool. Asking that same expert to *review* patterns the AI has already found, each with its evidence attached, converts authoring into judgement, and judgement is fast. It needs an expert who can sit with you for an afternoon, not a governance committee.

Three practices decide whether the session is worth their time:

**Batch, and lead with LOW.** One prepared list, not a question a week. The confidence level tells the expert where to focus: LOW markers are the most valuable to resolve because they represent the largest knowledge gaps, while HIGH markers are usually confirmations and can be dropped if time runs out.

**Ask closed questions with the evidence attached.** *"29 of 31 retail projects have no fiscal year booked, while every general project does, legacy-import artefact, or does the retail flow genuinely not book one?"* is answered in seconds. *"What does this field mean?"* buys a five-minute explanation of something already known.

**Treat a contradiction as a new finding.** When the expert says a field is always populated and the data says it is 40% null, neither is lying. That gap is configuration drift, an environment difference, or a process that changed without the data changing, and it is worth more than either answer alone.

Each resolved marker eliminates an entire category of potential misinterpretation, permanently. Resolved markers are then removed from the tool descriptions, so what remains inline at any given time is the open edge, not the total work done. An "I don't know" is a legitimate outcome: it means the knowledge does not exist in the organisation, and the description should say the meaning is unconfirmed rather than invent one.

**The expert confirms the knowledge is true; an eval confirms it works.** Those are different questions, and the evals showed how far apart they sit. A validated fact can be encoded where the model never sees it, or in a form it misreads. So Validate has a second half: a small eval. Five to ten questions whose right answers are already known, the tool before and after the change, ten runs per question, the runs interleaved in one batch, every run checked against the server's own call log. Read the result with a variance bar. At ten runs, a gap of eight or more is a size, four to seven is a direction, and less than four is noise ([V](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#method)). Register the prediction before the run. In the demo, 13 of the first 23 registered predictions were wrong, and every one of those would have shipped as a confident design rule without the run ([P](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#method)).

### Encode, write it back into the tool

Apply validated discoveries as operational metadata in the parts of the tool the model receives. This is where the loop diverges most sharply from traditional approaches. The knowledge doesn't go into a wiki, a data catalog, or a RAG index. It goes into the tool itself, and the evals settle where in the tool. What the agent needs before the call (when to use the tool, when not, how to form the call, what comes back) goes in the description head and the typed input schema. What it needs to read a value correctly goes in the response, as the first key, for the record it applies to. Field names do part of the work on their own, because they are the only semantics guaranteed to arrive. Output schema field annotations validate the response, but the model never sees them ([Q11](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)).

**Wrongness outranks volume.** The largest effect in the whole eval set was not missing knowledge but one plausible wrong line. The rich tier's description said `ep1 … Paris Proof kantoor: 70 kWh/m²`, putting a calculated label figure next to a target defined on measured energy. It made 59 of 60 answers wrong. One sentence stating that the label figures are CALCULATED, not MEASURED, and what they may be compared with, took the same question from 0 of 60 to 59 of 60 and removed all 22 fabrications ([BT](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)). Volume, by contrast, did not hurt: one rule among a hundred in a response was found as easily as the same rule alone ([Q17](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#volume-form-and-cost)). So audit what you already ship before adding more, and for every quantity say whether it is calculated or measured.

**Ship the fact with the instruction.** An instruction whose trigger is a fact the model does not have is inert. With both the fact and the instruction, 30 of 30; with only the instruction, 10 of 30 ([Q4](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)).

**Refuse, don't alert, when a fix is mandatory.** An alert attached to a successful call is read as a note. A render tool that warned about misleading table headers after drawing the table got the headers fixed in 2 of 10 runs, and was re-rendered in none. The same check as a refusal, with the fix in the error message, got them right in 10 of 10 ([Q22b, Q22c](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)).

**Six types of knowledge get encoded.** These map directly to the description blocks defined in the guidelines: domain interpretation populates the INTERPRETATION block, cross-tool references populate RELATED TOOLS, query strategies populate QUERY STRATEGY, fallback routes populate WHEN NOT TO USE, and uncertainty markers populate TODO annotations throughout. The RETURNS, WHEN TO USE, and FEEDBACK blocks provide the structural scaffolding that organizes this knowledge for the agent. The ALERTS block serves a different role: it doesn't define alert content in the description, it instructs the agent to expect server-generated warnings in the alerts array of every response and to always surface them to the user. Interpretation that depends on the record travels the same way, in the response, and the description keeps only the rules that hold for every record.

**1. Domain interpretation**, what codes and statuses mean in the business context.

> TypeItem: 'Lab' = labor hours, 'Cost' = costs,
>
> 'Mat' = physical articles/materials
>
> ChargeStatus: derived from 3 boolean flags into one readable label

**2. Cross-tool references**, join keys and resolution paths between tools, documented in the description's RELATED TOOLS block.

> EmployeeId, join key to get_employee
>
> TicketId, join key to get_service_ticket and get_actual_costs
>
> VisitNumber = TicketId, for invoice verification
>
> Triangle: planned costs → actual costs → invoice lines

**3. Query strategies**, performance guidance and multi-phase patterns.

> ALWAYS start with summaryOnly=true. Active projects
>
> accumulate thousands of records. Summary returns byMonth,
>
> byEmployee, byTypeItem, approval counts, no records.
>
> Then drill down with date range + type filters.

**4. Fallback routes**, alternatives when the expected path doesn't work.

> MaintenanceLocationTechnician is NOT populated in practice.
>
> To find the technician:
>
> 1. get_actual_costs with TicketType=52 → EmployeeName
>
> 2. get_ticket_responses → technicianFeedback

**5. Hierarchical relationships**, parent-child structures.

> Type 48 (service planning) → parent type 10 (servicemelding)
>
> Type 13 (onderhoud opdracht) → parent type 32 (onderhoud locatie)
>
> Type 52 (onderhoud planning) → parent onderhoud locatie

**6. Uncertainty markers**, the flagged findings that survived Examine but not yet Validate, encoded inline so the open edge is visible at the point of use rather than parked in a backlog. Format and examples are in [Flag](#flag-tag-every-pattern-with-a-confidence-level) above; what matters here is that they live *in the tool*, where the next agent and the next maintainer both read them.

**Implementation pattern: Two-Layer Data Enrichment**

A critical write-back technique is making data self-documenting rather than explaining codes in tool descriptions. It starts one step earlier than the two layers below, with the field names themselves. A name is the one piece of meaning that always reaches the model, so it should say what the quantity is, its scope and its unit. A name that implies the wrong quantity overrides the prose next to it, and a readable name without a unit gets a unit invented for it ([N2, N5](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#naming)). The production `Aantal` case described under Iterate is exactly this failure. When the API is not yours, rename in the server's transform step. The production implementation then uses two layers [23]:

**Layer 1: Source-side description joins.** The ERP's data connectors are configured to JOIN code fields with their human-readable descriptions. Every code gets a companion field: InstallationTypeCode → InstallationTypeDescription, ItemCode → ItemDescription. The agent reads "Service-specialist" directly from the data instead of looking up a mapping table. This eliminates \~2,650 tokens of code tables from tool descriptions.

**Layer 2: Server-side derived fields.** For interpretations that don't fit as simple JOINs, the MCP server computes derived fields before returning data. The production implementation derives a single ChargeStatus label from two boolean flags (a line-level "is charged" flag combined with either a labor or cost "is charged" flag, depending on the record type). The result is one readable label per row: "Fully charged", "Line-only charged", "Internal with cost tracking", "Cost allocation", "Internal", chosen from a closed enum the agent can reason about directly. This replaced \~200 tokens of boolean-combination explanation with one readable label in the returned data itself, self-evident enough that the tool description barely has to explain it.

The evals put numbers on both layers. A figure the server computed scored 20 of 20 on the value and on how it was derived; every arm that left the same arithmetic to the model scored 2 of 60 combined ([L1](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)). So compute what is determinate, and return it with its unit, its basis and whether it was calculated or registered, or return `null` with the reason it could not be computed. Where a rule needs data the payload does not have, ship the data rather than a pointer to it. A weather-correction rule that sent Haiku off to fetch a reference period scored 2 of 20; the same rule with the reference figure computed by the server scored 15 of 20. Sonnet and Opus were right either way, but needed 86% fewer calls ([Q16, Q16b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)).

**Implementation pattern: Annotation-driven field documentation**

Every output field carries a .describe() annotation, the field-level documentation for that column. These are not generic type descriptions, they encode domain knowledge. They are written for the maintainer, the validator and the client UI, not for the agent: no Claude host passes the output schema to the model ([Q11](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)); only Codex and ChatGPT Work do ([HD](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). The production estate still carries many of them, and the migration is to make sure each fact the agent needs also lives somewhere it is delivered:

> EmployeeName: z.string().nullable().describe(
>
> 'Employee full name. Null for Mat records.
>
> Use for display, only call get_employee
>
> for department/email.')
>
> ProjectId: z.string().nullable().describe(
>
> 'Project code, letter prefixes encode project
>
> types (G=general, K=small-scale, etc.),
>
> numeric-only = service contracts.
>
> Null for records without project allocation.')

Cross-tool join keys are documented in those annotations too: EmployeeId, join key to get_employee. For the agent, the same join has to appear where it is delivered, in the RELATED TOOLS block of the description head and in the input schema of the tool it leads to. That is how the agent learns to chain tool calls. The annotation keeps the fact next to the field for the next maintainer, and so does the reason each rule exists: an agent asked to improve a server cited a rule's history in 16 of 16 edits when the source carried its provenance, and in 0 of 16 when it did not ([Q18](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#volume-form-and-cost)). Provenance stays in the source and is never sent to the model.

### Iterate, production telemetry finds the next gap

The first four steps create the initial knowledge. Iterate keeps it improving by changing the loop's input: instead of your directed probes, the next round of findings comes from how real users and real agents actually behave. It runs on three layers. The concept of self-improving AI systems is not new, OpenAI's Self-Evolving Agents cookbook [13] and open-source projects like the Ultimate MCP Server's Autonomous Refiner [14] explore similar ideas. But these focus on prompt optimization and development-time simulation. The approach here operates on production runtime data, improving metadata rather than prompts.

**The governing observation: understanding cannot be observed, confusion can.** Nothing in the logs establishes that an agent understood a tool. A great deal establishes that it did not. That asymmetry is what makes production telemetry usable at all, and confusion arrives in three distinct forms, which is why three instruments are needed rather than one.

Most of it is friction rather than error. Something is underspecified, the agent takes a wrong turn, the user asks a follow-up, a retry or two follows, and together they reach the right answer. The answer is correct; what it cost was turns. Missing meaning usually does not produce a wrong answer, it produces extra work. Less often, a person notices something is off and decides to flag it. Rarely, the answer is simply wrong and nobody notices.

That last category is the blind spot in confidence-based triage, and it qualifies the Flag step described above. Routing the uncertain findings to a domain expert is the right instinct, but it means the fields the model is confident about *and wrong about* never reach a reviewer: they pass by construction. A production example: a field named `Aantal` (Dutch for "quantity"), typed as a number, where a sibling enumeration determines whether the number means hours, kilometres, or pieces. Summing the column yields a confident, wrong total, with no error, no retry, and no complaint. No confidence score caught it. A domain expert walking the high-stakes fields by hand did.

The three layers below map onto those three forms of confusion: logs for the friction, a report for what someone notices, a domain expert for what nobody notices. They are ordered by how much each earns its keep. The ordering is defensible; a precise split of the value between them is not, and is deliberately not claimed here.

A fourth instrument now sits beside them, and it is the one aimed at the class nobody notices. Telemetry only sees confusion. An eval sees correctness, because its questions have answers a domain expert has already verified. Kept as a regression suite and re-run after every change and on a schedule, it turns "a domain expert happened to walk the high-stakes fields" into a check that runs whether or not anyone remembers. In the public repository it caught what no log would: a correcting sentence past the description cut, an alert that was confidently wrong, a stale deploy, and an upstream quota the server exhausted by itself ([Method](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#method)). The production estate does not run one yet.

**Layer 1: Tool Call Logs (the foundation).** Log every tool call: request parameters, response summary, session ID, timestamp. No conversation text, no user questions, no agent reasoning, just the MCP layer. This is the primary data source, and the instrument for the friction case. Three consecutive calls with tightening filters on the same table means the agent is struggling with something the metadata should have covered. The agent does not need to report confusion; the call pattern reports it.

In the production implementation, every tool call, whether success, validation fallback, or error, writes to a persistent store with: tool name, connector ID, user, queryIntent, filter fields and operators (not values, for privacy), summaryOnly flag, row count, hasMore, duration, and error type. Session IDs correlate multi-tool sequences [23]. The store retains two years. On the HR server, which carries special-category health data under GDPR article 9, the intent string is redacted before it is written at all.

**Layer 2: Pattern Analysis (the intelligence).** Periodic (weekly) analysis of raw logs. Detect repeated call sequences, redundant calls, unused tools that should have been used, tools called in unexpected order. Session IDs enable correlation. This is the only layer that costs recurring human time, and in practice it is the first thing to lapse when the week gets busy. It surfaces some of the silently-wrong cases, those where a detour leaves a trace even though the agent flagged nothing. Not all of them. The ones that leave no trace are why a human stays in the loop.

**Layer 3: QueryIntent + Report Tool (the accelerator).** A queryIntent parameter on every tool call: one sentence describing the business question being answered. Plus a report tool for the moments a person decides a problem is worth writing down. The report tool's most valuable input is not the complaint but the reasoning: the schema asks the agent for its chain of thought and marks that field as the most valuable one, so a report captures what the agent thought it was doing when it got stuck, not merely that it got stuck. Academic research on LLM metacognition shows mixed results, models can sometimes calibrate confidence but not reliably [12], and emerging work on metacognitive state vectors suggests this may improve [15]. This layer is designed as a bonus signal, not the foundation.

In production (not reproducible from the public repository), 82 reports have been filed and 75 worked through. The categories are what a system that mostly produces friction rather than errors would predict: missing data, ambiguous semantics, pagination struggles, cross-tool mismatches, misleading descriptions. The most-reported tools are the two richest ones. At least eight reports trace cleanly from a filed problem to a shipped fix; one traces to a deliberate "won't fix", an architectural limit of the connector-per-entity design that is more honest to leave in place than to paper over.

**Validated knowledge is not frozen**

Encoding a validated fact is not the end of its lifecycle. Once a field has been validated and written into a description, something has to re-check it when the source system changes underneath. Production keeps testing the validation, and drift arrives in three kinds that are not equally tractable.

*Structural drift is cheap to catch.* A supplier renamed its entire route surface on a pre-release API. Four of five tools broke, and the manner of breaking is the point: old payload keys returned empty arrays that looked like real data. The failure was not an error, it was plausible emptiness. The fix was to pin the routes in one place and add a probe that asserts the live contract, so a silent rename fails loudly the next time.

*Schema drift is caught at runtime.* The handler emits a drift alert whenever the source returns fields the tool does not declare. One concrete case: a reactions connector had a field typed as a number that the ERP actually returns as a text label, and every record quietly failed output validation until the schema was corrected to a string. This is the confident-and-wrong category again, a validated assumption that was wrong, caught in production rather than in review.

*Semantic drift is the hard one, and only use surfaces it.* Everything still works technically, every field still validates, and the understanding is simply wrong or has aged. No probe catches this. What catches it is friction in the logs, an explicit report, or a domain expert. The `Aantal` case belongs to this class.

**The critical design: session ID as join key**

When an agent reports an issue, the session ID links to the complete tool call sequence that preceded it. You see what the problem was and how the agent tried to solve it. That's the recipe for a targeted metadata fix.

**Nudge patterns: teaching agents to self-report**

The production implementation embeds contextual nudges in tool responses rather than relying on agent initiative [23]:

-   **Empty results:** "No records matched your filters. If unexpected, consider calling report_problem with category 'unclear_filter'."

-   **Pagination detected:** "You're paginating (skip=200). If you're doing this to manually aggregate data the summary doesn't provide, call report_problem with category 'pagination_struggle'; we may be able to add that summary dimension."

-   **Incomplete summary:** "Summary is based on the first 500 records only; there are more. Tighten your filters for a complete summary."

-   **Every response:** A feedback reminder as the last alert: "If this query took multiple attempts, returned confusing data, or required workarounds, call report_problem before your next step."

These nudges are injected by the handler factory, not by the agent's own judgment. They target specific friction patterns observed in production and make that friction visible in the response.

They do not, however, close the loop on their own, and that turned out to be the more interesting finding. Even with the nudge sitting in the response, the agent almost never files a report unprompted. It relays the friction, or the user feels it directly, and a person decides to report. Agents do not take initiative on meta-tasks, even when explicitly invited to. The nudge surfaces the moment; a human still makes the call. Any design that assumes agent-initiated reporting as its primary signal should account for this.

The evals explain part of it. An invitation is weaker than it looks. A soft pointer to a guidance tool was followed by Haiku 0 times in 10; the same pointer worded as a requirement, 10 times in 10 ([Q8b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). And a note attached to a call that already succeeded reads as information, not as a reason to act ([Q22b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)). A nudge to report friction is both at once. Whether wording it as a requirement would change the reporting rate has not been measured, and a feedback tool the agent is required to call on every friction would bring its own noise.

## 05 Before and After

A concrete comparison using a universally understandable domain: building data. The same tool, looking up a building by address, at Level 1 and in the form the evals left standing. Both run live in the public reference repository, over the same public registers, with nothing between them but the interface [28]. The Level 1 version is its `thin` tier. The second is its `best` tier: [the tool](https://github.com/DaveGold/mcp-metadata-demo/blob/main/src/tools/get-building-profile-best.ts), and [byte for byte what the model receives](https://github.com/DaveGold/mcp-metadata-demo/blob/main/docs/wire/best.md).

### Level 1: Typical MCP Server

The entire tool definition:

> name: get_building_profile
>
> description: Look up building information by postcode and house number.
>
> parameters: {
>
> postcode: { type: string },
>
> huisnummer: { type: integer }
>
> }

No field documentation. No explanation of what gets returned. No cross-tool references. The agent has to guess what an energy label means, whether the building year is reliable, which postcode format the API accepts, and how to use the data for any follow-up analysis. The neighboring tools in this domain (`get_meters`, `get_usages`, `get_weather_context`) are equally hollow one-liners. The transport works; the interpretation layer is missing.

### What an earlier version of this section showed

Until September 2026 this section showed a longer description as the "after": about 5,100 characters of well-organised blocks, the `rich` tier of the same repository and the version presented at MCPCon Europe 2026 [30]. The evals found two things wrong with it, and both are instructive. First, Claude Code delivered only its first 2,048 characters, so 74% of it never reached the model, including the one sentence that said the label figures are calculated rather than measured. Second, its interpretation block carried a plausible line, `Paris Proof 2040 targets, kantoor: 70 kWh/m²`, and a matching runtime alert that compared a calculated label figure with that target. On the question of how a building stacks up against Paris Proof, that tier was wrong in 59 of 60 answers. Removing the false alert was not enough: Haiku still scored 0 of 10 and Sonnet 1 of 10. Moving the correcting sentence from character ~3,380 to character 766, and changing nothing else, took both to 10 of 10 ([BT, Q20, Q21](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)). More metadata was not the problem. Metadata that was wrong, or that never arrived, was.

### After: what the evals left standing

The same tool, with four layers the agent receives and acts on. Each solves a different reasoning problem, and each sits where the host actually delivers it.

#### Layer 1: A description head that fits the delivered budget

The whole description is 1,800 characters, inside the cut, so every word of it arrives:

```text
WHEN TO USE: what building is at a Dutch address — label, bouwjaar, area, and
label-based estimates (CO₂, space-heating gas, heat-pump readiness, overheating).

WHEN NOT TO USE: this server has NO metered energy consumption — no meter readings,
no actual gas or electricity use. Not for addresses outside the Netherlands.

RELATED TOOLS: get_weather_context(latitude/longitude = coordinaten.lat/lon) for
degree days and weather correction.

QUERY STRATEGY: postcode = 4 digits + 2 capitals, no space ("3543AR"). huisnummer =
integer only; a letter goes in huisletter (28A → 28 + "A"), an addition in
toevoeging. Several units match? Retry with one from candidates.

RETURNS: interpretation, derived, candidates, then BAG facts (bouwjaar,
oppervlakte_bag_verblijfsobject_m2 = ONE unit, coordinaten) and the EP-Online label
(energielabel, berekeningstype, *_berekend_* figures per m² of
gebruiksoppervlakte_thermische_zone_m2).

INTERPRETATION — read `interpretation` FIRST: this record's computed values and
reading rules. Quote them; do not recompute. For every record:
- CALCULATED vs MEASURED: every EP-Online energy figure is calculated by the label
  method, never measured. Paris Proof and other metered benchmarks are defined on
  measured final energy, so where a question asks for that comparison, say it
  cannot be made from this data and why, rather than producing a ratio.
- Totals use gebruiksoppervlakte_thermische_zone_m2, never the BAG area.
- A null field was not produced by that label method. Never substitute an estimate
  or a different field; say the data does not contain it.
- No registered label means no label is known. Do not infer one from bouwjaar or
  building type.

ALERTS: interpretation.alerts — computed verdicts and this record's branch (not
found, several units, no label).
```

The blocks are the same eight names, minus FEEDBACK, which the demo has no tool for. What changed is what they hold. The head carries only what the agent needs before the call, plus the handful of reading rules that hold for every record. The calculated-versus-measured sentence, the one whose absence made 59 of 60 answers wrong, is now the first rule under INTERPRETATION, well inside the budget.

#### Layer 2: An input schema that can express every valid call

```text
postcode   (string, required)   Dutch postcode: 4 digits + 2 capital letters, no space. Example: "3543AR".
huisnummer (integer, required)  House number, integer only. For "28A" pass 28 here and "A" as huisletter.
huisletter (string)             House letter, e.g. "A" for 28A.
toevoeging (string)             House-number addition, e.g. "bis", "I", "II".
queryIntent (string)            The business question this call answers. Used for observability.
```

The typed schema is for expressibility and validation, not for meaning. Where the right answer needs a parameter the Level 1 schema does not have, Level 1 scored 0 of 18 and the typed schema 18 of 18, and a postcode pattern stopped silent lookups of the wrong building ([L3](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)). The input schema is delivered in full, but re-sent on every turn, so it should hold what forming the call needs and no more.

#### Layer 3: Field names that cannot be misread, and values the server computes

Behind this one tool the server calls two separate government APIs, Kadaster BAG (building characteristics: year, floor area, intended use, GPS coordinates) and EP-Online (energy performance: label, certification method, CO2 emissions) [27]. The agent calls one tool with a postcode and house number and never needs to know two APIs were involved.

The fields it gets back are renamed so that the name carries the quantity, its scope and whether it was calculated: `ep1_energiebehoefte_berekend_kwh_m2`, `energieverbruik_berekend_niet_gemeten_kwh_m2`, `oppervlakte_bag_verblijfsobject_m2`, `gebruiksoppervlakte_thermische_zone_m2`. Anything determinate is computed on the server and returned with its unit, its basis and its provenance, or as `null` with the reason. From the live response for Gustav Mahlerlaan 10, Amsterdam:

```json
"derived": {
  "totalCo2KgPerYear": {
    "value": 6318764,
    "unit": "kg CO₂/year",
    "basis": "co2_emissie_berekend_kg_m2 53.47 × gebruiksoppervlakte_thermische_zone_m2 118174",
    "provenance": "calculated"
  },
  "spaceHeatingGasM3PerYear": {
    "value": null,
    "reason": "computed for residential (Woningbouw) only; utility heating systems vary too much for a gas-boiler assumption"
  }
}
```

The agent never multiplies a per-m² figure by the wrong area, because the server already did the multiplication with the right one and said which one it used.

#### Layer 4: Interpretation first in the response, for this record

Every response opens with an `interpretation` key: the alerts and reading notes that apply to *this* building, and the constants the rules need. The same response, trimmed:

```json
"interpretation": {
  "alerts": [
    "Overheating risk: NONE — indicator 0 (TOjuli/GTO, unitless; not °C, not hours). Thresholds: 0 none, up to 1.5 minor, above 1.5 significant.",
    "Total calculated CO₂: ~6318764 kg/year (co2_emissie_berekend_kg_m2 53.47 × gebruiksoppervlakte_thermische_zone_m2 118174). Calculated by the label method, not measured."
  ],
  "notes": [
    "CALCULATED vs MEASURED: every EP-Online energy figure here is CALCULATED by the NTA 8800 method, not a meter reading; nothing in this response is MEASURED. Paris Proof and other metered benchmarks are defined on MEASURED final energy, so none of these figures can be ranked against such a target — same unit, different quantity. …",
    "Two area scopes, not two measurements: oppervlakte_bag_verblijfsobject_m2 66581 m² (BAG, one unit) vs gebruiksoppervlakte_thermische_zone_m2 118174 m² (the zone the label covers; 1.77×). Per-m² label figures are per thermal-zone m²."
  ]
}
```

Ask this server how the building stacks up against the Paris Proof 2040 office target of 70 kWh/m² and it says the comparison cannot be made from this data, and why. That is the right answer, and the `best` tier gave it in 20 of 20, 9 of 10 and 10 of 10 runs across Haiku, Sonnet and Opus, where `rich` gave it in none ([Q19](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19)). The response carries rules sized to the record, so an office building never gets the residential gas estimate and a building without a label never gets told it has one. The agent does not compute the verdicts; it reads and quotes them.

### The Same Pattern Applied to Complex Enterprise Data

The building profile example demonstrates the pattern using universally understood data, addresses, building years, energy labels. But the same four-layer structure works on deeply domain-specific data that only industry insiders understand.

In the production ERP deployment, the same block structure is applied to a financial actuals tool (get_actual_costs) with domain knowledge that includes: type codes that determine which fields are populated ('Lab' populates employee name, 'Mat' has null employee), a three-flag charge status derived into a single readable label, contract-covered work detection (numeric-only project IDs indicate subscription billing, not missing invoices), a 13-status service flow, cross-tool join keys connecting actuals to budgets to invoices, and a summaryOnly pattern that aggregates thousands of records server-side to avoid context overflow. The metadata is entirely different. The structure (WHEN TO USE, WHEN NOT TO USE, QUERY STRATEGY, INTERPRETATION, RELATED TOOLS, FEEDBACK, ALERTS) is identical. The pattern transfers. The production descriptions were written before the cut was measured. Checking each one against the delivered budget, and moving what sits past it into the head or the response, is the migration still to do.

The Level 1 server exposes an API. The rebuilt server teaches the agent how a domain expert thinks about this data, what to query first, what the values mean, when to chain tools, and what it must not conclude, and it does so in the places the model actually reads.

## 06 Why It Works

### What the evals measured

Earlier versions of this paper argued from production experience alone and said so. Since then the claims have been put to a registered eval programme on the public reference server: more than 3,700 scored live runs across Claude Haiku, Sonnet and Opus. Each question's prediction was written down before its run, and each run was checked against the server's own call log, because the agents' own account of which calls they made turned out to be wrong ([A](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#method)). Six results changed the design most:

| Result | Evidence |
|---|---|
| **Wrong is worse than missing.** One plausible line put a calculated label figure next to a target defined on measured energy and made 59 of 60 answers wrong. One sentence saying the figures are calculated, not measured, took the same question from 0 of 60 to 59 of 60. | [BT](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship) |
| **Delivered is what counts.** Claude Code sends only the first 2,048 characters of a tool description and never the output schema. Moving one correcting sentence inside the cut took a failing question from 0 of 10 to 10 of 10. | [Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery), [Q21](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19) |
| **Ship the data, not just the rule.** A rule that sent the model off to fetch history: Haiku 2 of 20. The same rule with the server-computed reference figure: 15 of 20, and Sonnet and Opus needed 86% fewer calls. | [Q16, Q16b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship) |
| **The schema decides what can be asked.** Where the call needs a parameter the Level 1 schema lacks, 0 of 18 against 18 of 18 with a typed schema. | [L3](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship) |
| **Compute what is determinate.** A server-computed value scored 20 of 20 on value and derivation; every arm that left the arithmetic to the model, 2 of 60 combined. | [L1](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship) |
| **Refuse, don't alert, when a fix is mandatory.** An alert after a successful render fixed 2 of 10 and triggered no redo; refusing the call with the fix in the message, 10 of 10. | [Q22b, Q22c](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19) |

The differences between models matter less than the layer. At the top of the ladder every model saturates, so tailoring metadata per model buys at most a couple of percent in tokens ([M1](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#model-differences)). A stronger model does not recover a constant the payload lacks: where the answer needed one, Opus scored no better than Haiku ([L2](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)). And noticing is not acting. Opus named the calculated-versus-measured trap in 7 of 10 runs and still gave the forbidden verdict in 10 of 10, until the correcting sentence was delivered ([M3](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#model-differences)).

The programme also kept score on itself: 13 of the first 23 registered predictions were wrong, including several of the author's own design rules, and each wrong one changed the implementation. The limits stated at the top of this paper apply to every row: one author, one domain family, one host for the effects (delivery measured on six), one model family. The production estate has not yet been rebuilt on these results. The eval set measures the pattern on public data, and production is where the pattern came from.

### What the Feedback Architecture Reveals in Practice

To see why the system works, consider a concrete example. The same three-call sequence, with and without the queryIntent parameter:

**Without Intent:**

> 09:14:03 | get_service_records | summaryOnly=true, project=P-7056
>
> 09:14:07 | get_service_records | type=13, project=P-7056
>
> 09:14:09 | get_service_records | type=13, project=P-7056,
> completed=true

Three calls. The last one added a filter. Why? Unknown from the log alone.

**With Intent:**

> 09:14:03 | intent: Overview of all service records for project
> P-7056
>
> 09:14:07 | intent: Which maintenance orders exist for this
> project
>
> 09:14:09 | intent: Only completed orders, previous call also
>
> returned open items

The third call reveals exactly what went wrong: the agent expected only completed items but the metadata didn't make clear that a completion filter is needed. That's a targeted metadata improvement requiring minutes to implement. This is the system working as designed, the feedback architecture identifies the gap, and a metadata update fixes it permanently.

### The Underlying Properties

Introspective Context Engineering for MCP produces a two-layer intelligence system. The metadata layer provides domain knowledge discovered through data exploration. The reasoning layer (the LLM) fills remaining gaps using that knowledge as foundation. This creates several powerful properties.

**Graceful degradation.** Imperfect metadata does not crash the system; it gives the LLM less to work with. Every improvement eliminates an entire category of uncertainty, freeing reasoning capacity for harder problems.

**Compound improvement.** Each resolved TODO marker makes the entire system permanently smarter. The metadata layer grows richer over time while the reasoning layer stays the same. This is cheap, compounding progress, though "permanent" is the wrong word for it: as the drift discussion above makes clear, encoded knowledge has a maintenance cost, and the compounding holds only while something keeps re-checking it.

**The loop evolves the interface, not only the documentation.** Framing the loop purely as metadata undersells what it does. The common case is that production finds a gap and a description gets sharpened. But sometimes the recurring intent in the logs does not point at a missing sentence, it points at a missing capability, and the answer is to change the agent-facing interface itself. Four capabilities in the production estate emerged or were confirmed this way:

| Recurring production intent | Capability |
|---|---|
| "I only need part of this object" | Server-side field projection (`select`) |
| "I need to understand this large dataset" | Typed summaries with breakdowns |
| "Tell me if anything important is happening" | Conditional runtime alerts |
| "What does this combination of facts mean?" | Derived values |

Honesty about provenance matters here, because "all four emerged from telemetry" would be too tidy. Summaries were core design from the start. Projection came from efficiency pressure. The alerts have mixed origins, one of them an incident. Derived values deliver an interpretation as a value rather than asking the agent to reconstruct it from prose. What the telemetry supplied was not the ideas but the evidence that these particular ones recurred often enough to be worth building.

The payoff is measurable in production (not reproducible from the public repository). On the connectors that received projection, most full-record calls now request only the fields they need. In one logged instance the same object was read twice, 24 seconds apart: the projected read was 3,969 bytes against 265,418 for the full body, roughly 67 times cheaper for the same answer.

**Transferable pattern.** The loop works identically regardless of data source, API style, or domain. The comparison below covers four servers, deliberately, not exhaustively. They were chosen because they span the widest range the estate contains: REST and GraphQL, a five-API aggregation hidden behind one unified interface, flat relational records and hierarchical classification trees, transactional data and time-series telemetry. Seven more servers followed on the same foundation and are omitted here because they add scale, not new evidence, a further REST server with well-annotated tools tests nothing the first four did not already test.

| | ERP | BIM (Autodesk AEC) | Energy monitoring (mixed) | Construction standards |
|---|---|---|---|---|
| **API style** | REST | GraphQL | REST (5 APIs) | REST |
| **Tools** | 11 data | 6 data | 13 (10 data + 3 app) | 7 data |
| **Data shape** | Relational, flat records | Hierarchical, spatial | Time-series, building profiles, weather data, cadastral records | Hierarchical classification |
| **Key discoveries** | Invoice triangle (budget → actual → invoiced), null patterns per record type, technician fallback path, contract-covered work detection, ItemCode groupings per discipline | Hub → project → model → element navigation, HVAC system classifications, pipe/duct specifications, API pagination quirks (silently returns 0 at limit=100+) | Smart meter vs building automation as separate tools with metadata-guided selection, weather context enrichment via meteorological API, building profile from two government registry endpoints, Priva variable discovery across 0–60K sensors per building (zero is a legitimate state, not a lookup error: the tenant knows the asset but no monitoring is configured on it) | NL-SfB classification hierarchy, code → description resolution, ETIM product classes and typed features |
| **Domain metadata** | Financial interpretation, cross-tool join keys, charge status derivation, 13-status service flow, 30-status quote lifecycle | Spatial relationship navigation, element classification hierarchies, Revit category mappings, custom parameter filtering limitations | Energy analysis methodology, building profile enrichment from two government APIs, sustainability scoring, five-API routing logic hidden behind unified tools | Construction standard lookups, code validation |
| **Cross-system links** | — | Project ID mapping to ERP | search_projects → ERP project ID | NL-SfB codes → BIM element classification |
| **Confidence markers** | 59 (9 HIGH, 24 MEDIUM, 26 LOW) | Verified project mappings, API limitation flags | Data source routing validated | Code hierarchy validated |
| **Build time** | ~40 hours (first server, from scratch) | ~1–1.5 days (reusing patterns) | ~1 day (energy domain discovery) | ~0.5 day (classification mapping) |

Additionally, a **Utility server** provides three interactive rendering tools ([render_chart](https://github.com/DaveGold/mcp-metadata-demo/blob/main/src/tools/render-chart.ts), [render_table](https://github.com/DaveGold/mcp-metadata-demo/blob/main/src/tools/render-table.ts), and render_map) that every data tool in the ecosystem can use for presentation, plus an interactive form tool (render_form). A **Games server** demonstrates MCP Apps with Firebase-backed leaderboards. Four further servers (pre-order calculation, product catalog, internal data exploration, and ticketing) were built on the same foundation after this table was compiled.

Feedback is centralized rather than replicated: a single `report_problem` tool on the Utility server serves the whole estate, and each server points agents at it with its own identifier. The earlier per-server tool was consolidated once the estate grew, eleven copies of the same definition competed for context without adding anything, which is the tool-discovery problem of §08.1 in miniature.

Each domain required entirely different metadata, but the five-step loop, the description block structure, the handler factory pattern, and the feedback architecture were identical across every server. The BIM server operates on a GraphQL API rather than REST; the energy server is a mixed-backend server aggregating five separate APIs (a smart meter data provider, a building automation platform, a meteorological API, and two government building registries) behind a unified tool interface, separate tools exist for smart meter data (`get_usages`) and building automation data (`get_priva_usages`), and the metadata's WHEN TO USE / WHEN NOT TO USE blocks guide the agent to the correct tool for each building. The agent doesn't need to know the backend architecture; it reads the tool descriptions and picks the right one. The standards server maps a hierarchical classification system. The same introspective approach, discovering what the data means through actual exploration and writing it back as operational metadata, applies to any structured data source.

**Agent-agnostic by design.** Because the domain intelligence lives in the MCP server, not in the client, a single Rich Domain MCP Server can serve multiple AI agents simultaneously. Claude.ai, Microsoft Copilot agents, Cursor, custom-built assistants: any MCP-compatible client benefits from the same metadata without duplicating the knowledge investment. Build the domain layer once, connect it to every agent your organization adopts. As the ecosystem of MCP clients grows, so does the return on every hour spent writing metadata. One caveat the evals added: hosts differ in what they deliver, and the 2,048-character cut is Claude Code's alone. Put correctness where every host has to deliver it, in field names, a short description head and the response. Raising one client's cap does deliver the rest, but it added 23.7% in tokens to every tool on every request ([Q15b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)).

**Enterprise-grade security.** Every server shares a common security layer: Microsoft Entra ID OAuth 2.1 authentication with JWT validation on every tool call, audience and domain restrictions, and role-based access control (RBAC) scoped per tool. Only authenticated employees within the organization can access any tool. The security infrastructure (OAuth proxy, JWT verification, RBAC checks, session correlation) is implemented once in a shared factory and reused by every server, adding zero security engineering per new deployment. Tool call logging to Firestore provides a complete audit trail: tool name, user identity, query parameters, session ID, and timing, without capturing conversation content.

**No new infrastructure required.** The upgrade from Level 1 to Level 4 needs no infrastructure beyond what any MCP server already deploys. It does need more than better text. Several of the fixes that measured best are server code: computing determinate values, shipping the reference data a rule needs, refusing a call that must be corrected. They live in the handler the server already has, not in a new system beside it.

**Server-side summarization.** The production implementation moves analytical computation from the agent to the server through a `summaryOnly` pattern. When set to true, the server paginates through up to 5,000 matching records internally and returns pre-aggregated domain-specific breakdowns (by month, by employee, by project, by item type, by installation type, by condition score) plus financial totals and approval counts. Each connector supplies its own `summarize` callback that produces different aggregation dimensions relevant to its domain, along with contextual alerts (e.g., "23% of records are unapproved, approval backlog?" or "summary is based on the first 500 records only, tighten your filters"). One `summaryOnly` call answers most analytical questions ("who worked on this project?", "how are hours distributed?", "what was the busiest month?") without returning a single individual record. This is server-side intelligence that no amount of prompt engineering on a Level 1 server can replicate; the server does the computation, the agent does the interpretation.

**Metadata as instruction set.** The metadata serves a dual purpose: it documents the domain for human maintainers and simultaneously instructs the AI agent at call time. There is no separate "AI training" step: the rules a new developer reads in the source are the rules the agent receives at call time. The two readers do not see the same bytes, though. The agent gets the field names, the description head, the input schema and the interpretation the server returns. The maintainer also gets what the agent never sees: output-schema annotations, the text past the cut, and the provenance of each rule. Keep one canonical copy of each rule in the source and render the agent's version from it, and every connected agent behaves better after the next deploy. This eliminates the duplication inherent in client-side approaches like skill files, where domain knowledge must be maintained in two places, the server and the instruction file, and kept in sync manually.

## 07 Honest Limitations

Intellectual honesty requires acknowledging what this pattern does not solve.

**Domain expert dependency.** The 40 hours are not 40 hours from zero. They are 40 hours from someone who already works with the data daily and can validate what the AI discovers. The AI generates the metadata; a human confirms it. Without someone sitting on top of the data daily, the error margin is higher and the TODO markers take longer to resolve.

**Evals are now inside the loop, and they have limits.** Earlier versions of this paper called the missing eval framework the sharpest gap in the method. That gap is closed on the public reference server: a small set of questions whose answers are verified in advance, arms that differ by one variable, predictions registered before each run, every run audited against the call log, and the arms that were measured frozen so a later edit cannot silently change them [29]. It answers the question the method could not answer before, whether a given change made the agent act better. Three limits remain. The evidence is narrow: one author, one domain family, one host for the effects (delivery measured on six), one model family, and the results should be re-run before they are trusted anywhere else. The production estate does not have its own eval set yet; its regression suite, built from questions its domain experts have already answered, is the next step. And frozen fixtures still miss what only live data shows. A frozen test set would never have caught the renamed routes described above, while a re-run against live data does, so the evals complement the live re-runs rather than replace them.

**Single-tenant depth.** The current implementation serves one company, one ERP configuration, one domain. The same status code can mean different things at different companies running the same ERP. Scaling to multi-tenant requires per-tenant metadata variants, not a vector database, but more than one set of tool descriptions.

**From data to applications (Levels 5 and 6).** Levels 1–4 address understanding, how well does the server help the agent interpret data? But the production implementation has already moved beyond understanding into two additional capabilities that differ in kind rather than degree.

**Level 5: Interactive MCP Apps (read).** The production deployment includes server-rendered interactive UI components, charts (9 types including sankey flows), sortable/filterable data tables (8 cell types including badges, icons, currency formatting), geographic maps with typed markers, and a guided energy analysis intake form. Each MCP App is built as a standard Angular + Tailwind component, compiled by Vite into a single self-contained HTML file that the MCP server serves directly, no separate hosting, no iframe URLs, no build infrastructure beyond what any frontend developer already uses. These are not client-side widgets; they are MCP App tools served by the server, rendered directly in the conversation. The agent selects the right visualization and parameters; the server controls the presentation. Even a perfect Level 4 server with flawless domain knowledge still relies on the agent to present 847 rows of data as text. A Level 5 server renders that data as an interactive, properly-formatted table, directly from the server. The effect is combinatorial: each new visualization tool multiplies the presentation capabilities of every data tool in the ecosystem.

**Level 6: Secure Write Apps (graduated deployment).** Because every tool call is already authenticated via Entra ID, authorized via RBAC, and logged to Firestore, the write security concern is narrower than it first appears: the question is not "can someone unauthorized do this?", that's already solved, but "can the agent trigger a mutation without the user intending it?" This focused threat model makes bounded writes practical.

The production implementation already includes agent-initiated bounded writes: a cross-server `report_problem` tool writes structured friction reports to Firestore (typed Zod schema, session-correlated, categorized by friction type), and a game server writes leaderboard scores to Firebase. These are low-risk, bounded writes through standard model-visible tools, the agent calls them directly, the schema constrains what can be written, and every call is logged.

For business-critical writes (ERP corrections, invoice approvals, status updates), a more restrictive pattern is deployed: the WriteIntent architecture. It first shipped on energy data, where `start_graph_comments` opens the form and an app-only `commit_graph_comments` persists user-authored annotations against a specific building and time range. The agent opens an MCP App form (model-visible tool), the server mints a one-shot WriteIntent bound to the user and scoped to a specific resource, the user makes corrections in the form, and the form calls a commit tool that is hidden from the agent (`visibility: ["app"]`). The server validates the intent (same user, not expired, not consumed) before executing the write. Security properties: agent cannot see the commit tool, agent cannot obtain the intentId (stored in `structuredContent`, not `content`), intent is user-bound, resource-scoped, one-shot, time-limited (15 minutes), RBAC re-checked at commit, and both open and commit calls are logged to Firestore. The blast radius is bounded even if every host-layer guarantee fails, each write requires two separate tool calls, each intent is single-use, and the trail is noisy and auditable.

The graduated approach (agent-initiated writes for low-risk operations, user-initiated writes for business-critical mutations) reflects a practical security model: match the write pattern to the risk level. This would allow Rich Domain MCP Servers to evolve into Rich Domain MCP Apps, not just reading and interpreting data, but safely writing it back through validated, schema-enforced interactions at the appropriate trust level.

**Agent metacognition limits.** Academic research shows mixed results on whether LLMs can reliably assess their own knowledge boundaries [12]. The feedback architecture accounts for this: tool call logs (always reliable) are the foundation, queryIntent (sometimes useful) is a bonus, and the report tool (rarely used without strong prompting) captures only genuine blocking issues. The system is designed for the agent's actual behavior, not ideal behavior.

**Client fragmentation.** Not all MCP clients handle metadata equally. Server instructions are injected by some clients but not others. Output schemas with structured content are supported by some but ignored by others. Descriptions are cut, and silently: on Claude Code at 2,048 characters, on Cowork at 4,096, on claude.ai chat not at all. No host publishes its limit; I found these by measuring ([HD](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). A server author cannot see the cut from the server side. The only way to find it is to measure what the model actually receives, which the public repository does in its wire views ([Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). The implementation maintains both structured and text fallback responses for every tool call, and pre-validates output before the SDK's own validation to maintain control over error handling [23].

These limitations are real, but they are bounded. None of them invalidate the core pattern.

**A note on evidence.** The paper rests on two kinds of evidence, and it keeps them apart. The design rules are measured: before-and-after comparisons on answer accuracy, derivation, fabrication, tool calls, tokens and wall time, on a public server anyone can run and with every result published [28][29]. That moves the case from "this works in practice" to "this works measurably better, on this data, on this host". The production claims are not measurements. The detailed inventory covers eleven servers and eleven external APIs, and the estate had grown to sixteen servers and 103 tools by the end of September. Build times, report counts and review rates come from that one company. What is still open is breadth: other domains, other hosts, other model families, other authors writing the metadata, and eventually other organizations. The production codebase is an in-progress system, with confidence markers still being resolved with domain experts and metadata continuously improving through the feedback architecture, which is exactly how the pattern is designed to work.

## 08 Recommendations for the MCP Framework

The limitations described above are not inherent to the protocol; they are gaps in the current specification and client ecosystem. Based on production experience building a Level 4 Rich Domain MCP Server, these are the changes that would make the biggest difference.

### 1. Smarter Tool Discovery

Today, every connected MCP server dumps all its tools into the agent's context at once. For a single dedicated agent connecting to one domain server, this is fine, 8-10 tools with rich descriptions fit comfortably in context. The problem emerges at the client level: a developer in an IDE with 10 MCP servers connected simultaneously faces 100+ tool descriptions competing for context. This is not a server-side problem, it's a client-side discovery problem. But it has a chilling effect on metadata investment: why write 200 lines of rich domain knowledge per tool if the client is going to flood the context with descriptions from every connected server? The specification needs a discovery layer: a way for clients to query available tools by category, domain, or capability before loading them into context. Think of it as DNS for tools, resolve first, load second. A simple approach: let servers declare tool groups with short summaries. The client presents groups to the model, the model selects the relevant group, and only those tool definitions are injected. This is not hypothetical, it mirrors how the skills pattern already works, and it is one of the things skills genuinely do better than MCP today.

**Update.** Claude Code has since shipped a version of this at the client level: tools connect deferred, by name only, and a dedicated search tool fetches full schemas on demand, closing most of the context cost described above. That confirms the diagnosis, this was solvable, but it is still client-specific. Nothing in the protocol lets a server declare its own tool groups; the fix depends on each client building its own discovery layer rather than servers being able to rely on one that works everywhere.

### 2. Extensible Metadata Without Context Flooding

The current specification forces all metadata into tool descriptions and input schema descriptions, both of which consume context tokens on every call. There is no mechanism for metadata that the agent can request on demand. Resources were designed for this, but in practice agents never request them autonomously (validated across Claude Desktop, Claude Code, and Cursor [23]). The framework needs a middle ground: a metadata field on tool definitions that clients can surface selectively. For example, a tool could declare 200 lines of field-level documentation, cross-reference tables, and query strategy guides as structured metadata. The client loads the tool's name and short description into context by default, but injects the full metadata only when the model indicates it wants to call that tool. This would eliminate the current trade-off between rich documentation and context efficiency, a trade-off that currently forces server authors to compress critical domain knowledge into artificially short descriptions.

**Update.** A version of this has started appearing as a convention rather than a spec feature: an MCP server points the agent at its detailed guidance via MCP Resources under a `skill://` URI, with instructions that make the read mandatory for specific triggers, and a fallback tool for clients that can't read resources at all. It works, not because Resources became spontaneously requested, but because the instructions force the read and a fallback covers the clients that still ignore it. That is a workaround built on convention, not the protocol guarantee this recommendation is asking for.

**Update, September 2026.** The trade-off this recommendation describes has since been measured, and it is sharper than it looked. On Claude Code only the first 2,048 characters of a tool description reach the model, the rest is dropped without any signal to the server, and 74% of one carefully written description never arrived ([Q7](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). Raising the limit on the client delivers the rest but charges it to every tool on every request, +23.7% in tokens ([Q15b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). Two asks follow, and both are cheap. Hosts should publish the limit and mark a cut description as cut, so an author can see it. And the specification should give servers a place for detail that is loaded when the tool is used rather than paid for on every turn. Until then the working answer is the one the evals found: a short head for the decision to call, and the rest in the response.

### 3. Resources That Actually Work

MCP Resources are the specification's answer to reference documentation, static or dynamic content that agents can request to inform their reasoning. In theory, this is the perfect channel for domain guides, field dictionaries, and cross-reference tables. In practice, no tested client reliably surfaces resources to agents. Claude Desktop lists them in a sidebar but agents don't request them. Claude Code and Cursor ignore them entirely. The production implementation built resources, tested them, and removed them after validation showed zero agent-initiated requests [23]. For resources to fulfill their design intent, clients need to either: (a) automatically inject relevant resource content when a related tool is selected, (b) allow servers to mark resources as required_context for specific tools, or (c) give the model an explicit get_resource primitive that it is trained to use. Without at least one of these, resources remain a specification feature that does not work in the real world. The closest working substitute today is a guidance tool the description points at as a requirement. A soft pointer to it was ignored, and "REQUIRED: call this first" was followed in 10 of 10 runs ([Q8b](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). That is option (b) built by hand, and it is exactly the kind of convention the protocol should make unnecessary.

### 4. First-Class Feedback Loops

The current specification has no concept of agent-to-server feedback. Every interaction is request-response: the agent calls a tool, gets data back, and moves on. There is no standard way for agents to report confusion, flag data quality issues, or indicate which tool calls were unhelpful. The production implementation solved this by building a custom report_problem tool and injecting queryIntent parameters into every tool's input schema [23]. This works, but it is a workaround. The specification should support feedback as a first-class primitive: a standardized way for agents to annotate tool calls with intent, satisfaction, and issues encountered. A feedback field on tool responses. A report method on the protocol itself. Server authors could subscribe to feedback events and use them to improve metadata iteratively, closing the loop that currently requires custom infrastructure.

### 5. Make Output Schema Model-Visible, or Document That It Isn't

This was assumed to be a training gap: current foundation models are not trained on rich MCP interactions, so surely they simply hadn't learned to read per-field output-schema annotations the way they read a tool description. Direct measurement disproves that. Across Claude Code, claude.ai, and OpenAI's Responses API, the tool definition a model actually receives was the description and the input schema; none of them forwarded a single output-schema field annotation. That check was made on the production estate; the public evals have since confirmed it on Claude Code by token accounting, where a 7,659-character difference in output schema changed the request by only 181 to 336 tokens, far too little for the schema to have been sent ([Q11](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). The specification's own note pointing at LLMs here is non-normative, and no Claude host implements it. Later measurements found the exception: Codex and ChatGPT Work do deliver the output schema, Codex as a TypeScript return type ([HD](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery)). Keep any meaning a model needs out of it all the same, because the Claude hosts drop it. Output schema does real work, it validates the response and drives client UI rendering, but on a Claude host a model reasoning about a tool's result never sees the schema that documents it, regardless of how the model was trained [23].

That reframes the fix. It is not model training; it is the specification and the client implementations. Either hosts start forwarding output-schema annotations to the model, even a trimmed form, field names and one-line descriptions, would close most of the gap, or the specification should say plainly that output schema is validation-only, so server authors stop investing documentation effort into a channel that no Claude host delivers to its reader. Until one of those happens, the correct practice is what the evals settled: keep output schema for validation, and put what the agent needs to interpret a response where it is delivered. That means the rules that hold for every record in the description head, inside its budget, and the rules for this record in the response itself.

### 6. Scoped Write Permissions for MCP Apps

As MCP servers evolve from read-only data sources toward interactive applications, write operations become inevitable. But the current specification offers no mechanism to restrict a write tool to a specific interaction context. Today, if a server exposes a POST tool, any connected agent can call it at any time, the only protection is either generating short-lived tokens server-side or adding an isApprovedByUser boolean parameter that the agent can simply set to true without actual user approval. Neither is real security.

The production implementation has designed and validated a workaround: a two-tool pattern where the agent calls a model-visible "open" tool that mints a server-side WriteIntent, and the actual mutation is performed by a separate "commit" tool registered with `visibility: ["app"]`, hidden from the agent, callable only by the MCP App form. The intent is user-bound, resource-scoped, one-shot, time-limited, and validated with typed Zod schemas per write operation. Both calls are logged. This creates a real security boundary despite the specification not supporting one natively.

The specification should formalize this pattern: a way to scope write tools to specific MCP apps, so that a tool marked as write-only-via-app can only be invoked through a validated form interaction, not through freeform agent reasoning. The `ext-apps` extension already supports `visibility: ["app"]`, but this is advisory, the underlying `tools/call` endpoint remains callable regardless. Making write scoping a protocol-level guarantee rather than a host-level hint would let servers safely expose bounded write operations, creating a time booking, approving an invoice, updating a status, without giving the agent unbounded write access. The schema enforces what can be written; the app scope enforces when and how. This is the missing primitive that would allow Rich Domain MCP Servers to evolve into Rich Domain MCP Apps.

These six recommendations share a common theme: the MCP specification is designed for transport, but the real challenge is knowledge delivery and safe interaction. The protocol excels at connecting agents to tools. It does not yet help agents understand what those tools mean, or safely act on what they understand. Addressing these gaps at the framework level would raise the floor for every server in the ecosystem, not just the ones built by teams willing to invest 40 hours in metadata engineering.

## 09 How to Apply This Monday Morning

This is not a framework to install. It is a method to follow.

**Prerequisites, stated honestly.** The build times quoted below assume TypeScript proficiency, direct access to a domain expert who knows the data daily, AI-assisted development throughout, and willingness to iterate on metadata for weeks after the initial deploy. Without these, especially the domain expert, the pattern still works but takes longer, produces lower-confidence metadata, and compounds more slowly. The 40-hour first build is what one developer achieved; it is not a benchmark.

1.  **Pick your data source.** Your company's ERP, database, or API, anything with structured data where the domain is complex enough that AI can't figure it out from the schema alone. The more obscure or industry-specific the domain, the higher the value of this pattern.

2.  **Build a basic MCP server.** Use the MCP TypeScript SDK. Expose your data through standard tools. This commodity layer takes hours, not days. An [open-source reference server](https://github.com/DaveGold/mcp-metadata-demo) provides a standalone starting point on public Dutch building and weather data. It runs the same tools at three tiers, `thin`, `rich` and `best`, on live endpoints you can connect to, plus the interactive chart/table/map MCP App tools. It leaves out the production infrastructure; the metadata patterns, and the evals behind them, are all there [28]. Start from `best`.

3.  **Do the two setup steps once.** Have the AI research which metadata channels your target clients actually honour, then generate the guidelines document, the house style every tool description will follow. That document becomes your team's reference for consistent metadata quality across all tools, and you write it once, not once per tool. The demo repository ships this step as an [agent skill](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/SKILL.md), `rich-domain-mcp-server`, which encodes the loop, the description blocks, the confidence markers, the delivery rules, an audit for servers that already exist, and the evals [28]. Run the audit first on a server you already have: the richest tool in the reference repository turned out to deliver 28% of its own description. The method remains a method; the skill is a way of handing it to an agent rather than re-deriving it per project.

4.  **Run the loop: Examine, Flag, Encode.** Point the AI at the real data and direct it through differentiated slices, filter, sort, skip, compare. Make it tag every pattern with a confidence level and the question it raises. Write what it confirms where the model receives it: what the agent needs before the call in the description head (under about 1,800 characters on Claude Code) and the typed input schema, what it needs to read a value in the response, and meaning that fits in a field name in the field name. Output schema field annotations reach no model on a Claude host. Compute what is determinate on the server. Expect three or four passes per tool before the surprises stop; that figure is practice, not a measurement.

5.  **Validate with domain experts.** Review the TODO markers, batched, LOW first, closed questions with the evidence attached. Each one is a question for someone who knows the business. Each answer makes the system permanently smarter, and in production roughly one in ten corrected something the AI got confidently wrong.

6.  **Measure.** Before you trust a change, run a small eval: five to ten questions whose answers you already know, the old and the new version, ten runs each, in one batch, checked against your own call log. Write the prediction down first. At ten runs, a gap of eight or more is real, four to seven is a direction, less than four is noise ([V](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#method)). Keep the questions: they become the regression suite that stops the next edit from quietly undoing this one.

7.  **Iterate.** Deploy to power users, log every tool call with a queryIntent parameter, and read the logs weekly. Where agents struggle is where the metadata is missing, and each fix takes minutes, not a design cycle. Re-run the evals after each fix. The system compounds.

> *In the author's production implementation (not reproducible from the public repository), the first server (ERP, 11 tools), from blank MCP server to domain-intelligent assistant, took one developer roughly 40 hours, with no prior MCP server experience. The 40 hours were achieved with TypeScript proficiency, domain access, and AI-assisted development throughout, the same AI-first approach the pattern itself advocates. The rough breakdown: ~1 day to set up a proof-of-concept on a single tool, ~1 day to implement the Entra ID security layer with OAuth flow, ~1 day to add more tools and enrich the approach across connectors, and ~2 days to improve metadata in collaboration with domain experts resolving TODO markers.*
>
> *The pattern accelerates with reuse. The second server (BIM, 7 tools) took 1–1.5 days including metadata improvement, the handler factory, description block structure, and feedback architecture carried over directly. The third server took half a day, building on all the foundation already in place. Subsequent servers (energy monitoring, construction standards, visualization, games) each took 0.5–1 day. The first seven servers, 52 tools, were built by one developer in approximately 8 weeks of cumulative effort, with each server benefiting from the patterns established by its predecessors; four more followed on the same foundation, bringing the estate to eleven servers and 91 tools. Five more servers brought the estate to sixteen and 103 tools by the end of September. The intelligence lives in the tool interface, not in infrastructure. The investment is time and domain expertise, not budgets and vendor contracts.*

## 10 The Bigger Picture

MCP is becoming infrastructure. The protocol tripled its enterprise adoption in twelve months, with 76% of AI solutions now purchased rather than built internally [18]. As MCP servers become standard integration points, the competitive advantage will not be having a server; it will be having a server that understands your domain.

The production implementation described in this paper is no longer a set of disconnected tools; it is a complete agentic AI platform. By the end of September, sixteen servers with cross-system joins, shared authentication, interactive visualizations, and a common feedback architecture created a unified interface across more than a dozen backend systems. Non-technical employees (service coordinators, project controllers, energy advisors) use it daily to query operational data, visualize trends, and make decisions informed by AI-interpreted data. A service coordinator asks about maintenance schedules and sees interactive tables with condition scores and equipment details. A project controller compares budgeted versus actual costs and sees the margin visualized in a chart. An energy advisor assesses a building's sustainability profile by combining government building data with smart meter consumption and weather context, across three APIs, in one conversation. These are not demos. They are production workflows that already influence how the organization allocates resources, prioritizes maintenance, and advises clients.

The Rich Domain MCP Server sits between expensive infrastructure and cheap wrappers. It is cheap to build, fast to deploy, and agent-agnostic. But there is a stronger reason to invest now: correct metadata is future-proof. Better models reason more effectively over structured metadata, follow cross-references more reliably, and handle multi-tool sequences with less confusion, and the evals show the gap closing at the top: with the full layer delivered, Haiku, Sonnet and Opus all saturate ([M1](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#model-differences)). What a smarter model does not do is recover what the interface never said. Where the answer needed a constant the payload lacked, Opus did no better than Haiku ([L2](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#content--what-to-ship)), and a stronger model that noticed a trap still walked into it until the correcting sentence was delivered ([M3](https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#model-differences)). Wrong metadata is not future-proof either; it misleads every model generation equally. A Level 1 API wrapper gains the least from a smarter model, there's nothing for the reasoning to latch onto. A Level 6 Rich Domain App, correct and measured, gains the most.

> *The pattern is open. The method is documented. What matters now is who applies it first.*

## References

[1] Challapally, A. et al. The GenAI Divide: State of AI in Business 2025. MIT Project NANDA, 2025. Based on 150 interviews, 350 employee surveys, and 300+ public AI deployment analyses.

[2] S&P Global Market Intelligence. 2025 Enterprise AI Survey. Survey of 1,000+ enterprises across North America and Europe. Reports 42% of companies abandoned most AI initiatives in 2025.

[3] Ryseff, J., De Bruhl, B.F., and Newberry, S.J. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. RAND Corporation, RR-A2680-1, 2024.

[4] Informatica. CDO Insights 2025 Survey. Identifies top obstacles to AI success: data quality and readiness (43%), lack of technical maturity (43%), shortage of skills (35%).

[5] Gartner. Generative AI Pilot Forecast 2025. Predicts 50%+ of generative AI projects abandoned at pilot stage due to poor data quality.

[6] Smink, W. et al. Designing and Evaluating a Generative AI-Powered Chatbot for Enterprise Software Support. Journal of Systems and Software, January 2026.

[7] Peliqan. Text-to-SQL for ERP Data. Product documentation, 2025.

[8] Microsoft. Copilot for Dynamics 365. 2025. SAP. Joule: Generative AI Assistant. 2025. Oracle. NetSuite AI. 2025.

[9] K2view. GenAI Data Fusion Platform. Technical documentation, 2025.

[10] Anthropic. Model Context Protocol Specification. modelcontextprotocol.io, 2024–2026. Donated to the Linux Foundation under the Agentic AI Foundation, December 2025.

[11] Model Context Protocol. SEP-1382: Documentation Best Practices for MCP Tools. GitHub specification proposal, opened 2025, closed as dormant January 2026.

[12] Steyvers, M. and Peters, M.A.K. Metacognition in Large Language Models: A Survey. Sage Journals, 2025.

[13] OpenAI. Self-Evolving Agents Cookbook. OpenAI documentation, 2025. Post-hoc evaluation and prompt optimization from production traces.

[14] Dicklesworthstone. Ultimate MCP Server: Autonomous Refiner. GitHub, 2025. Development-time tool simulating agent failures.

[15] Courchaine, C. et al. Metacognitive State Vectors for LLMs. The Conversation, February 2026.

[16] McKinsey & Company. The State of AI: How Organizations Are Rewiring to Capture Value. March 2025.

[17] PubNub. Build vs Buy for MCP Servers. Technical analysis, 2025.

[18] Menlo Ventures. 2025 State of Generative AI in the Enterprise. December 2025.

[19] Holmes, E. MCP is dead. Long live the CLI. ejholmes.github.io, 2026. Argues CLIs provide equivalent value to MCP with less friction for developer workflows.

[20] Various. Skills as MCP alternatives. Developer community posts, Medium, 2026. Skills as self-contained markdown instruction sets that load on-demand.

[21] Cramer, D. MCP, Skills, and Agents. cra.mr, 2026. Counterargument that MCP, skills, and CLIs solve fundamentally different problems.

[22] jngiam. MCPs, CLIs, and Skills: when to use what? Blog post, 2026. Balanced analysis of complementary tool primitives.

[23] Golverdingen, D. MCP Server Metadata Guide, Rich Domain Server Patterns. Internal technical documentation, 2026. Operational metadata guide documenting an eleven-server, 91-tool production snapshot, among them AFAS Profit (11 tools), Autodesk AEC Data Model (6 tools), Warmtebouw Duurzaam (13 tools), Ketenstandaard (7 tools), Compano Select (6 tools), Gilde Pro calculation (5 tools), Utility (6 tools), Warmtebouw Games (2 tools), Firestore Explorer (5 tools), and WBTickets (18 tools). The estate grew to sixteen servers and 103 tools by the end of September. Cited here for production observations only; the design rules are cited to [29].

[24] Wen, Y. et al. Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions. arXiv:2602.14878, February 2026. Analysis of 856 tools across 103 MCP servers finding 97.1% contain description "smells" including unstated limitations, missing usage guidelines, and opaque parameters.

[25] Anthropic. The Complete Guide to Building Skills for Claude. https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf, 2026. Introduces skills as a client-side knowledge layer for MCP integrations, using markdown instruction files to teach Claude workflows and best practices for connected tools.

[26] Model Context Protocol. MCP Apps: Bringing UI Capabilities to MCP Clients. blog.modelcontextprotocol.io, January 26, 2026. First official MCP extension enabling servers to deliver interactive UI components (dashboards, forms, visualizations) directly in the conversation.

[27] Kadaster. Basisregistratie Adressen en Gebouwen (BAG) API and EP-Online Energielabel API. Public government APIs providing building characteristics (year, floor area, intended use, coordinates) and energy performance data (label, certification method, CO2 emissions). bag.basisregistraties.overheid.nl and ep-online.nl.

[28] Golverdingen, D. Rich Domain MCP, reference implementation and evals. GitHub, 2026, v2.0.0. https://github.com/DaveGold/mcp-metadata-demo. The same building and weather tools over public Dutch data (Kadaster BAG, EP-Online, Open-Meteo) at three tiers: `thin` (the raw API as a tool), `rich` (the server as presented at MCPCon Europe 2026, with only its measured defects fixed) and `best` (rebuilt from the evals, the reference), each on a live endpoint, alongside the interactive chart/table/map MCP App tools. Includes wire views that show byte for byte what each tier delivers to the model, a readable `queryIntent` call log, and the method as an agent skill, `rich-domain-mcp-server`, covering the ICE loop, an audit for existing servers, the description blocks, confidence markers, the delivery rules and the evals. Leaves out the production handler factory and security layer.

[29] Golverdingen, D. Rich Domain MCP: eval register. GitHub, 2026. Narrative: https://github.com/DaveGold/mcp-metadata-demo/blob/main/evals/README.md. Registered questions and predictions: https://github.com/DaveGold/mcp-metadata-demo/blob/main/evals/open-questions.md. Raw results: https://github.com/DaveGold/mcp-metadata-demo/tree/main/evals/results. Evidence register, one row and one status per design rule: https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md. More than 3,700 scored live runs across Claude Haiku, Sonnet and Opus on Claude Code, September 2026.

[30] Golverdingen, D. Most MCP servers are empty. AGNTCon + MCPCon Europe 2026, Amsterdam, September 2026. Slides and a slide-by-slide page with what the evals changed since: https://github.com/DaveGold/mcp-metadata-demo/tree/main/talks.

---

### About the Author

David Golverdingen is a software engineer with 18+ years of experience connecting complex systems to the people who use them. He started in industrial automation, programming PLCs, SCADA, and robotic controls for manufacturing and horticultural machinery, before moving into software engineering and data visualization in 2013, where his real-time monitoring work for energy and manufacturing clients earned a "Best in Show" award at the OSIsoft Visualization Hackathon. Seven years of freelance consulting followed, delivering senior frontend and full-stack engineering for de Volksbank, KLM/Air France, Allianz, Rijkswaterstaat (via Technolution), Nederlandse Spoorwegen, and Pro-Fa Automation. The common thread: turning complex business requirements into applications that work at scale.

In September 2025 he joined Warmtebouw, a roughly 350-person Dutch installation company, as Senior Engineer & MCP Architect, leading the development team. There, the gap between what AI could theoretically do with enterprise data and what it actually did became the central problem. The MCP servers described in this paper are the result: production infrastructure that lets entire teams query live ERP, BIM, energy, and building data through natural language. The pattern emerged not from AI research but from 18+ years of translating business problems into software solutions and turning industry trends into systems that people actually use.

**About this paper.** Written with the assistance of Claude (Anthropic) and curated, reviewed, and validated by the author. Implementation patterns and production observations reflect direct hands-on experience; AI assisted with research synthesis, structural organization, and prose. The domain knowledge and conclusions are the author's own.

**David Golverdingen** · MCP Architecture & Enterprise AI Engineering
LinkedIn: https://www.linkedin.com/in/davidgolverdingen/
