# Your agent does not need every tool in the room.

Separate the tool catalog from the model's working set. A practical design for discovery, schema loading, authorization, and context-budget tests.

Published: 2026-10-09 | Updated: 2026-10-09
Author: /init editorial
Method: AI-assisted, source-checked
Canonical: https://init.news/notes/mcp-tool-discovery-context-budget/

## The answer

Keep a complete authorized catalog at the client or gateway, then give the model a small, task-relevant working set. Measure whether it can discover a missing capability before narrowing that set further. Catalog size, prompt size, and permission are different decisions: selecting fewer schemas is a context strategy, while the execution boundary must still enforce authorization.

## Keep four different questions separate

A tool can exist, be available to a caller, be visible to a model, and be permitted for a particular action. Treating those as one switch makes context management brittle. Our recommendation is to maintain the full catalog outside the prompt, expose a concise discovery interface, and load complete schemas only when a task needs them. This is an application design, not an MCP requirement.

Layer | Question it answers | Record to keep
--- | --- | ---
Registry | What capabilities exist, and who owns them? | Stable server identity, tool name, owner, schema version, and descriptive purpose.
Authorized catalog | What may this caller discover? | Principal and scope boundaries, catalog version, and expiry or refresh behavior.
Model working set | Which tools help this task now? | Selection reason, loaded schema identifiers, and measured serialized token count.
Execution gate | Is this invocation allowed now? | Validated arguments, current permission decision, confirmation where required, and result receipt.

## Respect the protocol while narrowing the prompt

The MCP tools specification dated 28 July 2026 supports paginated tool listing. The available set can depend on the authorization supplied with a request, but must not change merely because a different connection is used or another request happened on that connection. It recommends deterministic ordering and describes tool-list change notifications. Client-side working-set selection therefore should not be implemented as an undocumented, conversation-dependent mutation of tools/list. [1]

The same version's caching specification defines TTL and public/private cache hints. Its TTL is a freshness hint, not a promise that underlying state cannot change. Keep catalog caching separate from a permission check at execution time, and include the relevant identity and authorization boundary in your application cache design. Check the negotiated protocol version before relying on these fields. [2]

## Build a discovery path that can recover

- Give each tool a short purpose statement, required inputs, output shape, and clear side-effect description. Distinguish similarly named tools with a stable server identifier.
- Let discovery return candidate identifiers and concise descriptions. Load each selected tool's complete schema before generating an invocation; do not guess arguments from a search snippet.
- When discovery finds no adequate match, return an explicit capability gap. Permit a broader search or a human decision instead of substituting a nearby write tool.
- Validate inputs and outputs at the boundary. Restrict oversized responses through explicit projection or a separate paginated read, while preserving the evidence needed for the task.
- Log the selected catalog and schema versions with the invocation so a later failure can be reproduced against the interface actually shown.

## Measure the trade before building a gateway

Start with a fixed set of tasks that includes a familiar tool, an obscure tool, overlapping names, a denied action, and an unavailable capability. Compare the current full-schema prompt with your proposed discovery path using the same model and task inputs. Record task completion, wrong-tool selections, discovery round trips, schema tokens, total latency, and authorization failures separately.

Include a catalog update between discovery and execution. The expected behavior is a controlled refresh or actionable error, not silent argument repair. Also test two principals with different permissions against the same cache. A context reduction that leaks a private catalog or increases unresolved tasks is not a successful optimization.

Choose a threshold before running the comparison. If the existing curated set is small and reliable, a static working set may remain the better design. Add dynamic discovery when the measured context or maintenance cost justifies another subsystem. No benchmark or universal tool-count limit is claimed here; the useful result is a decision you can revisit as the catalog grows.

## Sources

- [MCP specification 2026-07-28: Tools](https://modelcontextprotocol.io/specification/2026-07-28/server/tools) — [1] Tool listing, pagination, authorization-dependent catalogs, deterministic ordering, and change notifications; inspected 9 October 2026.; checked 2026-10-09
- [MCP specification 2026-07-28: Caching](https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/caching) — [2] Cacheable methods, TTL semantics, cache scope, and compatibility considerations; inspected 9 October 2026.; checked 2026-10-09
