Checking your workflows
Prove a config loads, behaves, and holds its declared properties with praxec check, fuzz, and test -- no model, no network, CI-ready.
Workflows put your agent on a state machine where the wrong move isn’t blocked – it’s not there to call. That guarantee is only as good as the machine you wrote. And you can’t eyeball a state machine: a dozen states and fifty transitions with guards and branches is past the point where reading the YAML proves anything.
praxec ships three commands that prove it for you. Each answers a stronger question than the last.
| Command | Question | What it does |
|---|---|---|
praxec check |
Does it load? | Static: schema, reachability, dead-ends |
praxec fuzz |
Does it behave? | Graph walk + per-transition isolation fuzzing + integration smoke |
praxec test --scenarios |
Does it hold its properties? | Asserts declared reaches / never_reaches / final_state / outcome_met |
All three are deterministic and exit non-zero on failure. Wire them into CI.
check – does it load?
praxec check is the static pass. It reads the config without running anything and proves the structural facts:
- the schema is valid
- every transition target points at a real state
- every state is reachable from the initial state
- no non-terminal state is a dead end (a state you can enter but never leave)
praxec check --config gateway.yaml
This catches typos and dangling pointers – the bugs that make a config wrong on paper. It’s the cheapest check you have and should be the first thing in CI. But a config that loads cleanly can still misbehave the moment it runs.
See the validation check reference for the full rule list.
fuzz – does it behave?
praxec fuzz proves the config behaves, with no real model and no network. It’s deterministic and exits non-zero on a failure.
praxec fuzz --config gateway.yaml
What fuzz checks
Fuzzing does not traverse the graph at random. That space explodes – every branch multiplies every path. Instead it splits the work in two.
Graph walk (structure). One pass over each workflow’s states and transitions confirms every state is reachable and nothing is orphaned. The shape is sound.
Per-transition isolation fuzzing (behavior). Each transition is tested alone, against the things that can go wrong with it:
- it fires on an input that satisfies its guard
- it rejects an input that violates the guard
- it handles the executor failing
- it resolves its output
- it never panics
- where a transition has human-gate branches, each branch is tested
Then a thin one-run integration smoke runs the whole flow once, end to end, to confirm the pieces fit together and not just in isolation.
The coverage model
The reason this is cheap enough to run on every commit:
If every transition is correct for all the values it can branch on, and the graph it sits in is sound, then the whole machine is sound.
So you don’t walk every path. You walk the graph for structure and fuzz each transition for behavior. Prove the parts and prove the wiring, and the composition follows. Cost is linear in the number of transitions, not exponential in the branching – fifty transitions is about fifty units of work, not fifty factorial.
test – does it hold its properties?
check and fuzz prove the machine is well-formed and well-behaved. They don’t know what you meant. praxec test --scenarios lets you declare the properties a workflow must hold and asserts them.
praxec test --config gateway.yaml --scenarios scenarios.yaml
A scenario file is a list of named assertions, each pinned to a workflow:
tests:
- name: happy path deploys
workflow: deploy_pipeline
expect:
reaches: [build, ready_to_deploy]
never_reaches: [rolled_back]
final_state: [deployed]
outcome_met: [shipped]
- name: refund needs approval
workflow: refund_flow
expect:
never_reaches: [refunded_without_approval]
reaches: [awaiting_approval]
The four properties you can assert:
reaches– the workflow reaches each of these statesnever_reaches– the workflow never reaches any of these statesfinal_state– the workflow ends in this stateoutcome_met– this declared outcome is satisfied
These are the claims you’d otherwise hold in your head – “the refund flow can never skip approval” – written down where a tool checks them every time the config changes.
Honest limitations
The fuzzer reports what it can prove, not a 100% green banner. Two patterns will show as un-driveable, and they’re expected:
- Evidence-gated human moves report as un-driveable. A transition behind a human gate that requires evidence a person produces (an approval record, a signed-off review) can’t be driven by the fuzzer – there’s no real actor to produce that evidence. It reports honestly rather than faking the evidence.
- Deterministic chains into capabilities may report
CHAIN_FAILED. A deterministic chain that runs into a capability which doesn’t yet resolve will reportCHAIN_FAILEDon that transition – a real signal that capability composition has a gap there, not a false negative to suppress.
When we ran all three over the cognitive-architectures library (28 workflows), the walk found 0 orphaned states and per-transition fuzzing verified 41 of 69 transitions driveable – the remaining 28 pinpointed exactly where capability composition still has gaps. That’s the intended output: a transition-by-transition map of what’s proven and what isn’t.
Wire it into CI
All three commands are deterministic and exit non-zero on failure, so a CI step is three lines:
praxec check --config gateway.yaml
praxec fuzz --config gateway.yaml
praxec test --config gateway.yaml --scenarios scenarios.yaml
The next time someone asks whether your state machine is right, you don’t say “it should be.” You point at the green.