blog

April 14, 2026

The hidden cost of 50 tools

Every tool you register is a recurring tax: input tokens for its definition, output tokens for the reasoning to choose it. Here's the math, measured.

Your agent has 40 tools wired in. It works. But every call feels a little slower than it should, the model reaches for the wrong tool more often than you’d like, and the bill keeps creeping up. You’re paying a tax — and it never shows up as a line item.

The good news: it’s not mysterious. It’s arithmetic.

The tax you can see

When you register a tool, its definition goes into the model’s context: the name, the description, and the full JSON schema for its arguments. A single definition runs roughly 50–150 tokens. Ten tools is about 1,000 tokens. Fifty is 5,000 or more — and those tokens are in context on every call, every turn, every retry. They’re not a setup cost. They’re rent.

Put real numbers on it. Picture a mid-size agent wired to six servers, ~50 tools total, at ~100 tokens each — 5,000 tokens of definitions. A real task rarely finishes in one turn; say 20. That catalogue rides along on all 20:

5,000 tokens of tool definitions
× 20 turns
= 100,000 input tokens spent describing tools

And across that whole task, the model probably called six or seven of those 50 tools. The other forty-odd were described, in full, twenty times over — and never used.

The tax you can’t see — the bigger one

Before the model can call a tool, it has to choose one. That means weighing the whole list and reasoning about which fits — and reasoning is output tokens, which cost roughly 3–5× what input tokens cost. More tools, more weighing, on every decision.

Then the failure mode nobody budgets for. Show a model 50 tools — several with overlapping descriptions, three “create issue” variants across three trackers — and it will sometimes pick the wrong one. A wrong pick is a full wasted round trip: the bad call, the error, the model reading the error, the recovery, the retry. Every step re-sends the 5,000-token catalogue and burns more reasoning. The two costs don’t add — they feed each other.

The fix: make the list stop growing

Praxec exposes exactly two tools — praxec.query for every read and praxec.command for every write — no matter how many capabilities sit behind them. The model stops scanning and starts searching: it calls praxec.query with a plain-language query and gets back one scored result with a title and a link, not 50 definitions. The response to a command carries the legal next moves as pre-filled links; it follows one instead of reasoning about which of 50 tools comes next.

There’s a subtler win in the schema. In a flat list you pay every schema, for every tool, on every call. With the two-tool surface, the model fetches one schema on demand — only when it’s about to use that capability. The 49 it isn’t using cost nothing.

Where this honestly doesn’t pay off

A search call plus a start call is two round trips before the real action. If your agent has five tools and always knows which one it wants, that indirection is overhead, not savings. And the search dispatch is lexical, not semantic by default — vague metadata produces vague results. The two-tool surface earns its keep when you have many capabilities, or want governance around them.

You don’t have to take this on faith. Open a recent trace, count the definitions, multiply by 100 tokens and by turn count, then count how many tools actually got called. The gap is what you’re spending to describe tools that never got used.

← All posts