Skip to main content

MCP Tool Token Optimization Patterns

18 min read

Overview​

Model Context Protocol (MCP) servers perform upfront loading of all tool definitions upon connection. Each tool consumes 300–1,000+ tokens in JSON Schema, and a configuration with 10 servers × 20 tools occupies 100,000 tokens in the context window before user input. This document quantifies token overhead and presents four optimization techniques: Progressive Discovery, tool compression, Code Execution, and prompt cache alignment.

Document Location

Background: Quantifying the Problem​

Upfront Loading Cost​

MCP servers return all tool metadata (name, description, JSON Schema) when list_tools is called. Clients include this in the system prompt sent to the LLM, so the more tools, the higher the initial context window occupancy.

Measured Cases (Sources Cited)​

  • StackOne Analysis: 10 MCP servers × 20 tools × average 500 tokens = 100,000 tokens preempted before user input (Source)
  • Atlassian Measurement: GitHub MCP server (94-tool) without compression = 17,600 tokens (Source)
  • Anthropic Case: 10,000-row spreadsheet exposed as 5 rows via code execution = 150,000 → 2,000 tokens (98.7% reduction) (Source)

Compound Cost​

Token overhead increases costs across three dimensions.

DimensionImpactQuantitative Example
Input Token CostTool definitions transmitted per requestClaude Sonnet 4.5: $3/M tokens → 100k tool definitions = $0.30/request
Context Window ExhaustionUser conversation length constraint100k preemption out of 200k window → 50% effective availability
Prompt Cache Hit Rate DegradationCache invalidation on tool list changesRetransmission on every dynamic tool addition/removal

Architecture: 4 Optimization Techniques​

Expand the diagram to read labels and follow connections.
Diagram

Technique 1: Progressive Discovery​

Concept​

Tool definitions are lazy loaded as needed. On initial connection, only tool names and one-line descriptions are transmitted; detailed schemas are retrieved when the LLM selects a specific tool.

3-Step Flow​

MCP official client best practices recommend the following steps.

  1. Catalog (Search): search_tools(query="file operations") → Returns tool name list
  2. Inspect (Schema Retrieval): get_tool_schema(tool_name="read_file") → Returns JSON Schema
  3. Execute (Invocation): invoke_tool(tool_name="read_file", args={...}) → Actual execution

Hybrid Threshold​

MCP official documentation suggests a 1–5% of context window threshold. A hybrid approach that switches to Progressive Discovery when tool definition tokens exceed the threshold is practical.

# pseudo-code: Threshold-based loading strategy
def load_tools(mcp_servers: list, context_window: int):
threshold = context_window * 0.05 # 5%
total_tokens = 0
loaded_tools = []

for server in mcp_servers:
tools = server.list_tools()
for tool in tools:
tool_tokens = estimate_tokens(tool.schema)
if total_tokens + tool_tokens < threshold:
loaded_tools.append(tool) # Upfront loading
total_tokens += tool_tokens
else:
loaded_tools.append({
"name": tool.name,
"description": tool.description,
"schema": "lazy" # Lazy loading
})
return loaded_tools

Trade-offs​

AdvantagesDisadvantages
Removes detailed schemas from initial context (reduction magnitude varies by tool set composition)Added round-trip for schema retrieval per tool call increases latency
Frees context window spaceLLM cannot see the full tool list at a glance
Improves prompt cache stabilityRepeated retrieval possible in multi-step reasoning

Technique 2: Tool Compression Proxy​

Atlassian mcp-compressor​

Atlassian Labs open-sourced a proxy that wraps existing MCP servers to compress tool descriptions. It provides three APIs.

  1. list_tools: Compressed tool list (name + minimal description)
  2. get_tool_schema: Detailed schema for specific tool
  3. invoke_tool: Delegates invocation to original server

Performance by Compression Strength​

Atlassian measurement for GitHub MCP server (94-tool):

Compression StrengthToken CountReduction RateNote
No Compression17,6000%Original
Low3,90078%Main parameters retained
Medium3,30081%Optional parameters removed
High2,20087%Only required parameters
Extreme50097%Name + one-line description

The proxy approach allows adoption without modifying original MCP servers or agent code; compression strength is controlled via proxy settings.

Suitable Scenarios​

  • Large Tool Sets: Agents using 50+ tools
  • Static Tool Configuration: Environments where tool lists rarely change
  • Token Cost Optimization Priority: When cost is more important than latency

Technique 3: Code Execution / Programmatic Tool Calling​

Concept​

Tools are exposed as programming APIs (e.g., TypeScript file tree) instead of JSON Schemas, and the LLM writes and executes code in a sandbox to invoke tools. Intermediate results are filtered within the execution environment, so they do not pass through the model context.

Anthropic Case​

Anthropic published the following results for a 10,000-row spreadsheet processing scenario.

  • Conventional Approach: Entire data passed to context → 150,000 tokens
  • Code Execution: Filtered via Python code → Only final 5 rows to context → 2,000 tokens (98.7% reduction)

(Source)

Cloudflare Code Mode​

Cloudflare introduced "Code Mode," which runs MCP servers in Workers sandboxes. Instead of tool definitions, it provides TypeScript APIs, and LLM-generated code executes in an isolated V8 runtime.

(Source)

Trade-offs​

AdvantagesDisadvantages
Up to 98.7% token reduction (Anthropic measurement)Requires sandbox infrastructure (Cloudflare Workers, Lambda, etc.)
Intermediate result filtering enables large data processingSecurity and resource isolation costs
Converts tool definition tokens → code execution tokensDepends on LLM code generation capability

Security Considerations​

Code Execution permits arbitrary code execution, so sandbox isolation is mandatory. Refer to the Tool Allow-list and Scoped Token sections in AI Gateway Guardrails to restrict executable API scope.


Technique 4: Prompt Cache Alignment​

Problem​

When MCP servers dynamically add/remove tools, the tools array changes, invalidating the prompt cache. Entire tool definitions are retransmitted per request, losing cache benefits.

Placement Strategy​

Fix the static tool list before the cache breakpoint, and append dynamic tools after the cache.

# pseudo-code: Cache-friendly tool placement
system_prompt = f"""
You are a customer support agent.

# Static Tools (cacheable)
{json.dumps(static_tools)}

<cache_breakpoint />

# Dynamic Tools (per-session variation)
{json.dumps(dynamic_tools)}

User Request: {user_query}
"""

Static vs. Dynamic Classification Criteria​

Tool TypeExamplePlacement Position
Staticsearch_kb, create_ticket, get_weatherBefore cache breakpoint
DynamicUser-specific custom actions, session temporary toolsAfter cache breakpoint

Effect​

Anthropic Prompt Caching reduces input token cost by 90% on cache hits (regular $3/M → cached $0.30/M). Caching 100k static tool tokens saves $0.27 per request.


Deep Dive: Gateway-Level Integration​

Relationship with Agent Data Plane​

The Tiered Gateway Architecture separates model request routing in Tier 1–2 from the Agent Data Plane's handling of tool calls, MCP/A2A connections, and stateful sessions. MCP/A2A can also use HTTP as a transport; the distinction is the responsibility each component owns.

Token optimization applies at the following layers.

LayerOptimization ResponsibilityImplementation Method
Agent Data Plane (agentgateway)MCP server discovery and schema retrieval (Progressive Discovery)Provides search_tools / get_tool_schema APIs
Tier 2 ② LLM API Gateway (Bifrost/LiteLLM)Prompt cache alignment, tool compression proxy integrationSeparate static/dynamic tool placement, mcp-compressor wrapping
Client SDKCode Execution sandbox invocationExposes TypeScript/Python APIs, filters execution results

Governance Integration​

Tool Allow-list, MCP server Fingerprint, and Scoped Token policies refer to the "§5.2 Tool Allow-list + Scoped Token" section in AI Gateway Guardrails. Token optimization addresses efficiency, while Guardrails addresses security. Both perspectives are independent and should be applied simultaneously.


Conclusion​

First measure where tokens are being used. If tool definitions account for a large share, consider Progressive Discovery, which loads definitions when needed, or a tool-compression proxy. If tools return large intermediate results, Code Execution can reduce that data in the execution environment before sending the necessary result to the model. This approach requires operating a sandbox.

Account for prompt-cache billing savings separately from reductions in the number of tokens sent in context. The cited examples use different tools and data. Compare input-token counts, actual cost, and tool-call accuracy on the same request set before choosing a combination.


References​

Official Documentation​

Technical Blogs​