Fixed-scope agent failure audit

Find the failure before buying a bigger model.

One reproducible workflow. One failure family. A fault tree that separates model output from tool envelope, state, validator, adapter load, template, tokenizer, and environment.

Evidence before effort.

The page sells a diagnostic, not a miracle. The buyer gets a boundary: where the first decisive failure appears, what was ruled out, and which next repair is rational.

01 / Model

Intent is not delivery.

A generated action can be structurally reasonable while the surrounding system drops, rewrites, or misreads it.

02 / Harness

False success is expensive.

A verifier that checks requested intent instead of observed state makes failure look solved and sends teams toward the wrong spend.

03 / Environment

Bound the unknowns.

An independent probe separates real environment defects from model or harness blame, so the report can stop cleanly.

Try the 15-minute triage first.

Freeze expected and observed state, preserve the raw model response, follow the action through transport and execution, then read the destination independently. If one layer is clearly responsible, fix it without buying the audit.

01 / Capture

Preserve the real trajectory.

Keep the raw assistant segment, exact tool envelope, timestamps, and final destination state. Re-tokenized or reconstructed traces can move the apparent failure.

02 / Branch

Find the first decisive break.

Separate missing model intent from dropped transport, failed execution, stale state, and a validator that measures the wrong thing.

03 / Stop

Stop when the action changes.

More evidence has no value once it no longer changes the repair. The checklist ends at the narrowest evidence-backed next intervention.

Open the free triage checklist

Included

One workflow, one failure family.

  • Fault tree across model, harness, validator, state, and environment.
  • Evidence for the earliest decisive failure boundary.
  • One runnable regression receipt or bounded non-reproduction finding.
  • Repair ranking split into no-weight, data, training, and infrastructure actions.
  • Knowns, unknowns, and stopping conditions.

What the receipt looks like.

No client logos. No invented customer quotes. This public sample shows the shape of the deliverable without claiming a buyer result.

Intake

Expected state is written down.

The audit starts by freezing the desired post-action state and the observed failure, so later tests cannot quietly redefine success.

Repro

The failure is reproduced or bounded.

If it cannot be reproduced in the agreed environment, the result is still useful: exactly what was tried and what evidence is missing.

Boundary

The earliest decisive break is isolated.

The report separates the model's output from the transport, parser, state, validator, and environment layers.

Gate

A regression receipt remains.

The buyer receives a runnable check or a clearly bounded hypothesis they can keep as a recurrence guard.

Good fit

The system fails between reasoning and delivery.

Bring this audit when a model proposes plausible work but the final state, test, or tool result does not match the intent.

Not included

This is not open-ended implementation.

No promised uplift. No production operation, model training, legal advice, security certification, or transfer of protected material.