Back to articles
AI Agents

Cartograph Makes MCP Tool Discovery More Efficient for AI Agents

3 min read

Why this matters

The Model Context Protocol gives AI agents a standard way to discover and call external tools. That convenience becomes expensive when an agent is connected to a large catalog: loading every name, schema, and description consumes context before the agent has even started solving a task. Cartograph addresses this discovery bottleneck by placing a federated proxy between the agent and MCP servers.

How Cartograph works

  • Progressive disclosure: In the reported deployment, 22 servers offered 374 tools, but the agent saw three proxy tools rather than all 374 definitions. The proxy first identifies promising servers and then searches for tools within that smaller set.
  • Operator-attested capability cards: Cartograph uses Ed25519-signed cards whose descriptions are generated under the control of the deploying operator. This shifts attention away from publisher-ranked copy and gives the retrieval layer a verifiable description source. The system also records which provenance was used for each query.
  • Rift analysis: Rift combines density clustering, query-margin analysis, and token diagnosis to locate tools that may be confused with one another. It found 49 confusable clusters in the reported evaluation, including four marked HIGH risk among bootstrap-generated cards.
  • Smaller discovery exchanges: A measured top-five discovery exchange used 475 tokens, compared with 42,450 tokens under the paper’s full-catalog accounting. The point is not that tool execution becomes free, but that the model need not ingest the entire catalog before retrieval begins.

What the evaluation shows

On a 49-query benchmark constructed by the authors, Cartograph reached R@5 of 0.816, compared with 0.592 for a Jaccard keyword baseline. An exploratory comparison involving 119 LLM-generated descriptions removed an observed zero-distance cluster. However, mixing cards produced under different generation regimes could reduce R@5, suggesting that description consistency matters as much as description fluency.

The gateway added an average of 5 milliseconds, or 0.8% in the reported comparison, relative to direct stdio MCP calls across ten trials. These numbers are encouraging, but they come from one 22-server deployment and an author-created query set. They should therefore be read as an initial systems result, not as a general guarantee for every MCP ecosystem.

Broader implications

Cartograph’s main contribution is architectural and operational. It turns tool discovery into a staged, inspectable retrieval process, while linking ranking decisions to signed descriptions controlled by the deploying operator. That combination could help organizations manage large internal API inventories, plugins, and third-party MCP services without flooding an agent’s context window.

The approach is complementary to code-execution systems. Code execution governs how an agent composes capabilities, while Cartograph governs which tool descriptions are surfaced first and where those descriptions came from. Future validation will need broader datasets, independent benchmarks, and production workloads. For now, the paper shows a plausible way to make MCP discovery more scalable without treating the full catalog as mandatory context.

Source: arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles