Intent is not delivery.
A generated action can be structurally reasonable while the surrounding system drops, rewrites, or misreads it.
Fixed-scope agent failure audit
One reproducible workflow. One failure family. A fault tree that separates model output from tool envelope, state, validator, adapter load, template, tokenizer, and environment.
The page sells a diagnostic, not a miracle. The buyer gets a boundary: where the first decisive failure appears, what was ruled out, and which next repair is rational.
A generated action can be structurally reasonable while the surrounding system drops, rewrites, or misreads it.
A verifier that checks requested intent instead of observed state makes failure look solved and sends teams toward the wrong spend.
An independent probe separates real environment defects from model or harness blame, so the report can stop cleanly.
Freeze expected and observed state, preserve the raw model response, follow the action through transport and execution, then read the destination independently. If one layer is clearly responsible, fix it without buying the audit.
Keep the raw assistant segment, exact tool envelope, timestamps, and final destination state. Re-tokenized or reconstructed traces can move the apparent failure.
Separate missing model intent from dropped transport, failed execution, stale state, and a validator that measures the wrong thing.
More evidence has no value once it no longer changes the repair. The checklist ends at the narrowest evidence-backed next intervention.
Included
No client logos. No invented customer quotes. This public sample shows the shape of the deliverable without claiming a buyer result.
The audit starts by freezing the desired post-action state and the observed failure, so later tests cannot quietly redefine success.
If it cannot be reproduced in the agreed environment, the result is still useful: exactly what was tried and what evidence is missing.
The report separates the model's output from the transport, parser, state, validator, and environment layers.
The buyer receives a runnable check or a clearly bounded hypothesis they can keep as a recurrence guard.
Good fit
Bring this audit when a model proposes plausible work but the final state, test, or tool result does not match the intent.
Not included
No promised uplift. No production operation, model training, legal advice, security certification, or transfer of protected material.