Showing version 4 bot legacy · api · 2026-08-13T07:10:22Z

Trust but Verify

Trust but Verify

The defining hazard of generative coding is plausible-but-wrong output. Diff looks right, names match the convention, tests pass — and a spec requirement is missing, or a contract is silently broken.

Trust-but-verify is the workflow's answer: an independent adversarial reviewer between implement and ship.

The principle

  • Hallucinations don't announce themselves. You can't ask the same agent "did you do this correctly?" and trust the answer.
  • An independent reviewer is cheap. One Opus invocation against a finished diff costs a fraction of a re-implementation.

Every workflow that ships AI-written code needs a verify phase. Not "we should code-review more" — a phase, in the pipeline, with a structured output and a pass/fail gate.

The phase

flowchart LR
    Impl["/speckit.implement/"] --> V["/speckit.agentic.verify/"]
    Spec & Plan & Tasks & Diff --> V
    V --> Findings[BLOCKER / WARN / NOTE]
    Findings --> Refute{Survives refutation?}
    Refute -->|no| Drop[Dropped]
    Refute -->|yes| Verdict{PASS / FAIL?}
    Verdict -->|FAIL| Fix[Back to implement]
    Verdict -->|PASS| Retro["/speckit.agentic.retrospective/"]

Inputs: spec, plan, tasks, diff. Output: structured findings + verdict.

Seven dimensions

# Dimension What it catches
1 Spec divergence Requirements not implemented
2 Missing tasks Tasks marked [X] without matching diff content
3 Security OWASP Top 10:2025 — auth, injection, secrets, scopes, supply chain (A03), failing open (A10)
4 Correctness Off-by-one, races, edge cases, swallowed errors, false idempotency claims
5 Contract breaks Schema, struct, signature, response-shape changes that break callers
6 Necessity Abstraction with one caller, a dependency replacing three lines, unrequested config
7 Velocity illusions Unread "working" code, edge-blind tests, PR text that doesn't match the diff

Spec divergence catches the most common failure. Contract breaks catch the most expensive. Necessity catches the one nobody else is looking for — code that works and shouldn't exist is still a defect, and no test will ever flag it.

Dimension 7 is NOTE-only. It rarely blocks anything; it predicts where the next three bugs land.

Hunt until dry, then try to kill it

One pass finds the obvious problems and stops. Two protocols wrap the dimensions:

Loop until dry. Sweep, then sweep again on whatever produced findings. Stop when two consecutive rounds surface nothing new. Deduplicate against every finding seen — including refuted ones, or a rejected candidate resurfaces each round and the loop never terminates. A fixed target ("find 10 bugs") either stops early or pads to reach itself; convergence does neither.

Refute before reporting. For each finding above NOTE, try to kill it. Default to killed when uncertain — the burden of proof sits on the finding. Ask whether it reproduces, whether the path is reachable, whether a caller already guards it. Where independent subagents are available, run several refutation attempts and keep the finding only if a majority fail.

Findings that are plausible but wrong cost more than findings you miss, because they get acted on.

Severity

Severity Effect
BLOCKER Must fix. Verdict = FAIL.
WARN Real but non-blocking. Verdict still PASS if no BLOCKERs.
NOTE Observation. Informational.

Verdict math: BLOCKER == 0 → PASS.

The reviewer's prompt

The single most important property: look for what's missing, not validate what's present.

"You are an adversarial reviewer. Find missing items, contract breaks, security regressions, correctness bugs. Quote line ranges for every finding. Do not praise correctness — that is not your role."

Without this framing, the reviewer falls into the same trap as the implementer: reads code, sees plausible patterns, approves. Same model, same diff, different prompt → different catch-rate.

Example finding

### Spec divergence — FR-005 (hard delete on file removal)
**[BLOCKER]** Spec FR-005 requires that when a feature folder disappears,
all rows for that slug are removed. The diff implements file-level deletion
in `ScanFeature`, but `ScanAll` does not call `DeleteAEFeature` for slugs
that have disappeared. Quote: importer.go:50-60.

Composition

Verify is the last line. Earlier gates do the heavy lifting:

  • Plan review catches design-level errors before code is written.
  • Pre-fetched context prevents drift into unrelated files.
  • The decision ladder keeps unnecessary code from being written at all — necessity is a verify dimension, not just correctness.
  • Constitution rules block recurring mistakes by category.

Verify catches the residual.

All seven dimensions read the diff. For anything with a UI, that leaves a gap
nothing above closes: the artifact itself is never observed running. A build can
satisfy every requirement in the spec, break no contract, and still throw on
first click or take six seconds to paint — none of which is visible in a patch.
Driving a real browser against the built thing is the complement, not a
substitute: see QA and QoE Testing with Playwright MCP for the mechanics and
QA and QoE for why the experience half needs measuring separately from the
correctness half.

Anti-patterns

  • "We have tests, we don't need verify." Tests check the test's claims. Verify checks the spec's claims. Different.
  • Skip when in a hurry. Same trap as The Illusion of Speed. Verify is faster than the bug-hunt cycle it replaces.
  • Same model as implement. Lower yield. Opus for verify even when implement ran Sonnet.
  • Verified the diff, never opened the page. Seven green dimensions and nobody loaded the app. For UI work the browser pass is part of verification, not a follow-up.

Hallucinations don't announce themselves. The verifier is the announcement.