Your agent has 12 tools and it’s fine. So you’ve never thought of the tool list as a design problem. It’s just a list. Lists are free. You need a new capability, you append it.
That instinct is about to cost you — not because the list gets expensive (a separate post covers the token bill), but because a flat list stops being a usable structure long before you notice it has.
A flat list has no idea what’s legal
Start with the deepest problem, the one that has nothing to do with size. A flat tool list is stateless. Every tool in it is always “available,” always callable, right now.
But your actual system has order. You can’t issue the refund before the charge settles. You don’t deploy before the tests pass. The PR gets created after the branch exists, not before. A list can’t express any of that — it has no slot for “this comes after that” or “this isn’t legal yet.”
So that knowledge has to live somewhere else, and “somewhere else” is your system prompt, or the model’s guesses. At 12 tools you paper over it with a few lines of “always check X before Y,” and it mostly holds. At 200 tools across a dozen subsystems, there is no prompt long enough. Every “do X only after Y” rule is really a tiny state machine — the legal order your system already enforces in code. With a flat list, none of that lives anywhere the model can check, so it’s left improvising your state machine on every turn, with no way to know when it’s wrong.
A flat list has no notion of relevance
The second problem is navigation. The model needs one tool out of 200. The list offers exactly one way to find it: consider all 200. There’s no “search the tools.” There’s only “scan the tools.” And the scan gets harder as the list grows in the most natural way — you’ve wired in three issue trackers, so now there are three create_issue variants with near-identical descriptions, and the model has to disambiguate lookalikes on every relevant turn.
What a kernel adds that a list can’t
Praxec replaces the list with structure. The workflow’s current state declares which moves are legal — order and preconditions are data the model can read, not prose it has to remember. And discovery is search, not scan: the model asks praxec.query for what it needs and gets a small, scored answer, then follows the links each response hands back.
The tool list stays two entries whether there are five capabilities behind it or five hundred. Your capability count stops being a context-budget decision and stops being a navigation problem. You add the tool because the work needs it — the structure absorbs the rest.