blog

May 19, 2026

Your LLM is a generator, not a calculator

An LLM does one kind of work: generation. It can't compute — by design, not by immaturity. Match each job to the tool built for it, and both shine.

Your coding agent told you the tests passed and shipped. They hadn’t. It scrolled the log, saw something that looked like a green run, and decided it had passed — instead of checking. No error, no exception. Just a confident, wrong “all green,” already deployed.

It’s the same failure as the small one: ask the model to total a few numbers or compare two dates, and it answers instantly, with total confidence — and it’s wrong. The instinct is to think the model isn’t smart enough yet and wait for the next version. That instinct is the bug. A smarter model will not fix this. You handed a generator a calculator’s job.

An LLM does exactly one thing

An LLM generates. That is the whole tool. Every output it produces — a sentence, a decision, a function, a yes-or-no — is generated: assembled, plausibly, from patterns. It is not computed. Ask the same question twice and you can get two different answers, because generation is nondeterministic by nature. That isn’t a defect waiting on a patch. It is the definition of the thing.

So “the LLM makes a decision” and “the LLM writes code” are not two different talents. They’re the same act — generation — aimed at different outputs. Which means it cannot compute, and “did the tests pass?” is a computation: there is one correct answer, and it’s knowable by checking, not by generating a plausible-sounding verdict.

Match the job to the tool

The fix isn’t a better model. It’s putting the deterministic work — the checking, the rules, the gating — back into code that runs the same way every time, and keeping the agent to the judgment calls it’s genuinely good at.

That’s what Praxec is. “Did the tests pass?” becomes an evidence record the kernel checks, not a claim the model asserts — the tests_passed record only exists because a deterministic step actually ran the suite and it came back green. “Should we deploy?” — a genuine judgment call — is where the model is invited in. The kernel runs the computations; the model supplies the judgment; neither is asked to do the other’s job.

Give the generator generation and the calculator arithmetic, and both are excellent. Ask either to do the other’s work, and you get a confident “all green” on a red build.

← All posts