There is a cost column missing from every MCP server directory: what connecting the server does to your context window. Every tool a server exposes ships a definition (name, description, JSON schema for the parameters) and that definition is injected into your model’s context on every single turn of the conversation. Connect enough servers and you have spent a meaningful slice of your context, and a meaningful slice of every request’s token bill, before the user has typed a word.
This is the real argument in the “one server vs many” debate, and it is worth doing properly, because the answer is not “consolidate everything” either.
The cost nobody itemizes
Tool definitions are not free in three distinct ways.
Tokens. Each definition is a few hundred tokens once you count the description and parameter schemas. A server exposing 40 tools can burn thousands of tokens of context per turn. Multiply by every turn in a long agent session and it becomes a real line on the model bill.
Attention. Models pick tools better from short menus. This is a qualitative claim but a widely observed one: as the tool list grows, models get worse at choosing correctly between near-duplicate options (“search_web” from one server, “web_search” from another) and start hallucinating parameter shapes. Two connected servers that both expose a search tool is a recipe for nondeterministic tool choice.
Operations. Every server is a connection to maintain, an auth flow, a thing that can be down. Ten local servers means ten processes; ten remote servers means ten OAuth dances and ten status pages.
The case for consolidation
If several tools share a shape (call a capability, get a normalized result back), they belong behind one server. This is the design we chose for the route.tools remote MCP server: a single server exposing 9 tools, one per category (search, scrape, parse, image, video, speech, transcription, embeddings, code execution). The consolidation is not just packaging. Behind each single tool definition sits the whole provider roster for that category, with failover and routing handled below the protocol line, invisible to the model.
Compare the alternative: connecting a Serper server, a Brave server, an Exa server, and a Firecrawl server to get search coverage. That is four connections and four overlapping tool definitions competing for the model’s attention, and the model, not your infrastructure, ends up deciding which vendor gets your traffic on any given call. Moving provider choice out of the context window and into configuration is, in my opinion, strictly better: the model decides what it needs (a search), the routing layer decides who serves it. I made the broader argument in MCP servers explained for developers.
The token math follows directly. Nine definitions covering nine categories is a small, stable context cost that does not grow as providers are added behind them. Nine servers exposing overlapping per-vendor tools is a context cost that scales with your vendor count.
The case for many servers
Consolidation has limits, and pretending otherwise would be selling.
Unique capabilities do not consolidate. Your Postgres server, your GitHub server, your internal admin tools: these are not substitutable providers behind a category, they are distinct capabilities with distinct schemas. A router-style server makes sense for commodity categories with multiple interchangeable vendors. It makes no sense for your own database.
Normalization has a floor. A single tool definition per category means a common schema, and a common schema is a lowest common denominator. If you need a deeply provider-specific feature, a dedicated server for that provider exposes it natively. (Our escape hatch is include_raw, which returns the provider’s untouched response alongside the normalized one, but an escape hatch is not the same as a native interface.)
Blast radius. One server down is one capability down. If that one server fronts nine categories, its availability matters more. Worth weighing, though I would note the consolidated server can be the more reliable path per category, since each tool call gets up to three providers’ worth of failover behind it rather than one.
The heuristic I actually use
Count your tools two ways: how many capabilities your agent needs, and how many definitions it currently sees. If definitions meaningfully exceed capabilities, you have duplication, and duplication in the context window is pure cost. Consolidate the commodity categories behind as few servers as possible, keep dedicated servers for genuinely unique systems, and be ruthless about disconnecting servers your agent has not called in weeks.
A reasonable end state for a working agent looks like: one server for commodity tool categories, one or two for the proprietary systems that define your product, and nothing else. Under ten servers’ worth of definitions, most of them earning their context rent daily.
For the commodity half of that stack, our MCP server gives you all nine categories, roughly 30 providers, and automatic failover through one connection; setup takes a few minutes via the MCP docs, and pricing per category is on pages like the search API comparison.