Observed Signal · Aug 9, 2026 · Technical Analysis · Source: DEV Community · Impact: 2/5 · Sentiment: Neutral
MCP tool lists cost median 3,150 tokens
The author measured the token cost of MCP servers' tool discovery payloads by probing a random sample of 60 servers from their MCP directory. 27 servers did not respond; 33 returned tool lists. Using a chars/4 approximation for tokens, the median returned tool list contained 11 tools and cost ~3,150 tokens, with a wide distribution (min 174, max 19,923 tokens). Tool count is a poor proxy for token cost because schema verbosity (long descriptions, nested parameters, inline enums, examples) drives size. Because tool definitions are sent on every agent turn, costs compound quickly across multiple servers and multiple turns. Recommended mitigations include pruning unused servers, client-side allow-lists, and loading tool definitions on demand; the author describes implementing a load-on-demand approach via fetchbean.
Practical technical analysis relevant to developers building LLM agents: it quantifies hidden token costs from tool discovery and gives concrete mitigations. Useful to engineers but not industry-shifting.
Track Real-Time Large Language Models (LLM) & AI Signals & Market Shifts
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Author probed 60 MCP servers; 27 did not respond and 33 returned tool lists.
- Median tool list among responders: 11 tools, approximately 3,150 tokens (chars/4 approximation).
- Token cost distribution among 33 responders: min 174, p25 913, median 3,150, p75 14,113, max 19,923 tokens.
- Tool count poorly predicts token cost; schema verbosity (long descriptions, nested parameters, examples) can increase token cost (example: 84 tools → 7,569 tokens; 54 tools → 19,923 tokens).
- Mitigations recommended: prune unused servers, use client-side allow-lists, or load tool definitions on demand (author built fetchbean: 56 providers, 748 endpoints, exposing 4 meta-tools).
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
MCP Servers Add Significant Per-Turn Token Overhead
A developer measured how MCP (Modular Capability Protocol) servers increase Claude Code session input tokens by injecting tool-definition blocks into the system prompt on every turn. Tool definitions are roughly 80–150 tokens each and compound across multi-turn sessions. In experiments, a minimal 3-tool server added ~180 tokens/turn (~3,600 tokens over 20 turns), an official filesystem server added ~640 tokens/turn (~12,800 over 20 turns), and the official GitHub MCP server added ~3,100 tokens/turn (~62,000 over 20 turns). At scale (e.g., 2,000 turns) the GitHub server can add ~6.2M tokens of overhead (~$18.60 at $3/MTok for Sonnet 4 input pricing). The author recommends project-scoped configs, preferring servers with fewer tools, and concise tool descriptions to reduce costs.
MCP Tool Surfaces Create a Repeated Token Bill
A DEV Community post by Boris Pan (published 2026-06-10) warns that MCP servers which expose tools to LLM agents incur an often-overlooked cost: every exposed tool (its name, description and full JSON input schema) is included in the model context on every single agent turn. Large or verbose tool menus can burn thousands of tokens before any useful action and can degrade tool-selection accuracy. To make that cost visible, the author built a small CLI that reports the per-call token bill caused by an MCP tool surface. The post includes practical observations for developers building MCP servers and exposes a prompt-engineering / tooling design trade-off.
MCP Tools Cut Agent Token Costs and Hallucinations
A developer describes building 24 custom Model Context Protocol (MCP) tools (organized into six categories) to power an autonomous agent managing six production services. The post argues MCP tool schemas act as "native prompts," reducing system-prompt length, lowering hallucination risk, and enforcing constraints in code rather than prose. In an example workflow, structured MCP calls reduced a three-turn, ~809-token interaction to a single ~250-token turn (≈3.2x token savings). Additional optimizations include dynamic per-execution mini-configs (load only relevant servers/tools) and session-based context retention, which compound token and cost savings. Implementation details: a single-file Python MCP server using the MCP SDK, claude -p (Claude Code CLI) as the agent runtime (Max plan), tool-level blocklists and hard-coded constraints, and practical cost comparisons across subscription and per-token pricing models.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
