Field Manual

Pre-Ship Code Validation

Auditing a diff before merge or deploy in a fail-open system — where broken changes don't crash, they silently no-op — written by the outgoing validator for the one taking the seat.

What this job is

You are the last reasoning step between a change and production. Everything upstream of you — the author, the commit message, the green test run, the plausible-sounding rationale — is advocacy. Your job is not to re-do the work; it is to independently establish whether the claims about the work are true, and to say plainly which ones you established and which ones you are taking on faith.

One fact about this kind of codebase changes how you must validate, and it is worth stating before any procedure. A fail-open system’s house failure mode is silent breakage. The concurrency governor fails open rather than wedge the fleet. The atomic-write library keeps errors off stderr and asks callers to check return codes. Nearly every hook line ends in a swallow-and-continue. This is correct design — a coordination layer must never block the work it coordinates — but its consequence for you is absolute: a broken change here does not crash; it silently no-ops. Dozens of test sims “ran” for weeks while actually failing before the first real turn. A shipped feature sat inert in production because the column it read didn’t exist and the code gracefully degraded around the absence. So the question “does anything error?” is nearly worthless here. The question is always: did the positive path demonstrably fire? Every move below is a different way of forcing that question.


1. Read the diff for what it changes, not what it claims

The commit message is a claim, not evidence. It was written by the same process that wrote the bug, at the same moment of confidence. Your reading order must run mechanics-first: establish what the diff does to the tree, then compare that against the story, and treat every mismatch as a finding.

Procedure.

  1. Start with the numstat (git show --numstat, or the PR diff stat), not the message. For each file: is it named in the message? Does its insertion/deletion shape match the message’s verb? “add”, “capture”, “append” predict near-zero deletions; “refactor” predicts balanced churn in the named files only; “fix” predicts a small, located change.
  2. Read every hunk in every file. If the diff is too large to actually read, that is itself a finding — see §8. Do not sample.
  3. For each hunk, ask: does the deleted side contain anything the branch’s stated purpose never touched? Deletions are where clobbers hide, because authors and reviewers both read the green lines.
  4. Flag any deletion mass in append-only artifacts (decision logs, ledgers, canon docs, schema files). Where a document is append-only by convention, a net-negative diff there is guilty until proven innocent.
  5. Check for the diff-that-isn’t: cherry-picks, merges, and rebases land a different diff than the one on the branch. Validate the result commit, not the proposal.
  6. Diff the diff against the message one final time: anything the diff does that the message doesn’t mention, and anything the message claims that the diff doesn’t contain, goes in the findings list verbatim.

Worked example. A commit’s message read: capture note plus decision-log entries — pure addition, per the story. The numstat said otherwise: the two new capture files were symmetric additions, but the decision log was +54/−142 — net minus 88 lines in an append-only ledger, inside a commit claiming to add entries. The author had committed a stale working copy that predated eight decision entries. It shipped. The loss wasn’t discovered for ten days, and a later commit had to excavate the entries by rebuilding from the parent. A validator running step 4 catches this in under a minute: the message’s verb (“capture… + entries”) predicts additions and no deletions on the log; the actual shape was +54/−142; mismatch → read the red lines → eight decisions are vanishing → NO-GO.

The failure it prevents. Silent destruction riding inside a plausible commit — the clobber class. Nothing errors, tests don’t cover prose ledgers, and the message actively steers your attention toward what was added and away from what was destroyed.


2. Map the blast radius before judging the change

A diff is judged locally but executes globally. Before you evaluate whether the change is good, establish where it lands: who calls this, what runs it and how often, what shared state it mutates, and which invariants it stands next to. A one-line change in a leaf script can be fleet-fatal if the script runs on cron and mutates global state.

Procedure.

  1. Callers and cadence. Grep the changed function or script name across the repo; check the scheduled callers (crontab, the cron installer) and the per-tool-call callers (the hooks directory). Multiply: a bug in a hook fires on every tool call in every session; a bug in the managed cron block fires forever, unattended, and re-fires after every manual cleanup.
  2. Scope of mutation. For every write in the diff, classify: process-local < session-local < repo-local < machine-wide (global git config, launch agents, shell profiles, crontab) < fleet-wide (shared coordination tables, pushed branches). Everything machine-wide or above triggers the standing rule: a machine-wide mutation, on a loop, with no detector is the outage pattern. Demand (a) the narrowest scope that works (per-command config, not global), (b) a drift-healing guard, (c) a --check lint so review catches reintroduction.
  3. Invariant proximity. Walk the diff past the system’s standing invariants and check each one it comes near: - Git-auth invariant — where no SSH keys exist, never a global HTTPS→SSH rewrite. Anything touching git config: run the config-drift guard’s check and read its result yourself. - Concurrency ceiling — every headless model spawn goes through a governor slot; interactive fan-out stays capped low. A diff adding an ungoverned model call inside a loop is a shippable provider-level shutdown. - File-lock discipline — edits to the file-lock hook or anything reading the session-lock table’s touched-files list change what every session may edit. - Approval gates — outbound to a human is never sent directly: drafts stage to a gated outbox (manual tier, pending) and are delivered only after explicit per-message approval. Any diff adding a send path, flipping a tier default, or widening auto-merge rails (the required conditions, the kill-switch default) is red-line, maximum scrutiny. - Message envelope — the type lives in the envelope’s dedicated column; consumers selecting a field that doesn’t exist get null and degrade into payload-shape guessing. Sends go through the send functions, never raw inserts. - Hook contract — a pre-tool hook’s block exit stops the tool; its stderr (not stdout) reaches the model. A careless edit to a hook can wedge or blind every session at once.
  4. Premise check. The diff’s own justification is in scope. If a line’s comment says “so X works,” verify X can work that way at all before evaluating the implementation.
  5. Second-order interactions. For any rule, priority, or config the diff adds: which existing rule can override it, and which does it override? Test the conflict case, not just the new case (see the voicemail leak, §5.7).

Worked example. A commit removed one line from a cron-run reaper script: a global config rewrite routing an HTTPS remote prefix to SSH, commented so git push works in cron. Walk the procedure against the original addition. Cadence: the reaper runs on cron every ~10 minutes — the mutation re-applies forever, so every manual unset silently un-fixed itself within minutes (pure whack-a-mole until someone hunted the setter). Scope: global — every clone and fetch in every repo in every headless context on the machine now routes to SSH. Premise: the machine has no SSH keys, so the line could never help a push in the first place — and push auth was already handled by the credential helper plus a global push-only rewrite, which outranks the general rewrite on pushes anyway. Detector: none existed. Result: permission-denied failures across plugin installs, sync jobs, and document fetches — a fleet-wide outage from one syntactically clean, swallow-guarded, well-commented line. The validator’s catch was available at every step: cron × global × false-premise × no-detector. The permanent fix wasn’t just deletion — it was a config-drift guard that heals the drift on a cron backstop and lints for reintroduction at review, which is exactly the (a)/(b)/(c) pattern you should demand of any global mutation that survives review.

The failure it prevents. The locally-reasonable, globally-fatal change — and its recurring form, where a cron loop resurrects the defect after every hand-fix so the outage compounds until the setter is found.


3. Verify by re-derivation — never approve on plausibility

Plausibility is what the author already had. You add value only by independently re-deriving the claim: run the changed path, trace the data through it, reproduce the bug it claims to fix and watch the fix remove it. In a fail-open codebase this is doubly mandatory, because “it ran without error” and “it worked” are unrelated statements.

Procedure.

  1. Reproduce before you review the fix. Check out the parent commit, trigger the claimed defect, and see it with your own eyes. If you cannot reproduce it, you cannot verify its absence — say so in the verdict (the house pattern is explicit: one fix commit is titled “leak confirmed → fixed”; confirm precedes fix).
  2. Re-run on the exact failing input. Not a fresh synthetic case — the same transcript, file, or row that produced the bug. Then run adjacent inputs to check the fix didn’t overcorrect (a precision fix must be tested for new false negatives, not just the removed false positive).
  3. Trace the data flow end to end. Follow one real value from producer to consumer across the seams the diff touches — environment variable → script → table column → reader. Most defects live at seams (a column that isn’t there, a field that was always named something else, a scratch directory cleared by reboot).
  4. Distinguish “failed” from “didn’t run.” For every test result you are shown, ask how many assertions actually executed. A harness error can masquerade as a failure (empty fail-list) or as a pass (zero assertions, clean exit). House rule: an empty fail-list with no all-goals-met marker means the sim errored, not failed — check for the pre-conversation failure before blaming (or crediting) the agent.
  5. Verify concurrency behaviorally, never by inspection. For semaphores, locks, and caps: generate the contention and observe the ceiling. The governor’s acceptance test is the model: a 6-caller burst against a max of 2 must peak at exactly 2; a planted rate-limit signal must trip the machine-wide cooldown; clean under both the scripted shell and the interactive one.
  6. Respect flake in both directions. On a noisy harness, one green run proves almost nothing and one red run proves almost nothing. The convention is repeat-N (at least 3): broken only if it fails repeatedly, fixed only if it passes repeatedly. “Passed once” after a fix is a coin-flip wearing a green checkmark.

Worked example. A 3 a.m. improvement loop filed a contentless acknowledgment (“Yes, that would be great, thank you.”) as a knowledge gap — a precision defect in the pipeline’s own filter. The fix tightened the “is this substantive?” predicate (skip bare acknowledgments; require a real ask — a question mark or several contentful words) and extended the allow-list with identity-honesty phrases so an agent refusing to pretend to be human stops being filed as a gap. The verification is the part to copy: ten offline unit assertions, then the filter re-run live on the very transcript that produced the garbage — garbage gone, and the one legitimate find still extracted. Both directions checked: the false positive removed, the true positive retained. Contrast the cost of skipping step 4: dozens of test sims “scored” zero-of-N with empty fail-lists across a dozen behavior dimensions — every one of them a missing-variable failure thrown before the conversation started, weeks of a safety battery (including the emergency-escalation sims) silently not running while dashboards displayed number-shaped output.

The failure it prevents. Approving a fix that doesn’t fix (or a pass that never ran). Plausibility review catches neither, because both produce exactly the artifacts — clean diff, green summary line — that plausibility feeds on.


4. The shell audit — this kind of repo’s native fragility

A shell-first system is dozens of scripts, sourceable libs, and hooks that run on every tool call of every session, on more than one OS, under more than one shell, mostly unattended. Shell review here is not style review; it is the difference between a fleet and an outage.

Procedure.

  1. Floor, not ceiling. Run the syntax check and a shell linter on every touched file; require clean. Then remember what they cannot see: swallowed errors that are intentional here, logic bugs behind a swallow-and-continue, macOS/GNU divergence at runtime, and — critically — generated text. Any script that emits shell (heredoc-built cron blocks, templated prompts, crontab writers) needs its output syntax-checked, not just its source.
  2. Quoting and interpolation. Every unquoted expansion is a finding until shown safe. Give special hostility to inline interpreters: an inline python3 -c "... '$VAR' ..." splices a shell variable into Python source — a single quote in the value breaks the program, and when the block ends in a bare except: pass the breakage is invisible and the hook silently stops enforcing. That specific shape — interpolated inline python, bare except, fail-open — is load-bearing in the file-lock path; treat any new instance as a defect and any edit near an existing one as high-risk.
  3. set -e semantics — deliberate, both ways. Ops scripts here use strict mode; hooks deliberately do not, because in a pre-tool hook a stray non-zero exit is an outcome (a specific exit code blocks the tool). Know the traps in both regimes: local x=$(can_fail) masks the failure (local’s exit wins); ((count++)) returns non-zero when count was 0; the last cmd && other in a function leaks a non-zero return; pipelines hide failures without pipefail. A diff that adds strict mode to a hook, or removes it from an ops script, changes behavior, not style.
  4. Hook contract. For any hook edit: exit-code semantics per hook type (the block exit stops the tool); stderr surfaces to the model, stdout is swallowed — a block that “printed” to stdout once showed the model only “hook error: no stderr output”; gate through the hook-rate helper so a chatty hook can’t drain the warning budget; and the worst-case question — if this edit is wrong, does every session on the machine wedge (blocked edits) or go blind (lock enforcement silently off)? Both have happened in miniature; both are one careless hunk away.
  5. Portability. macOS first: no flock (the governor claims slots via mkdir atomicity); stat -f %m versus stat -c %Y needs the dual-path dance; the interactive shell’s abort-on-unmatched-glob behavior is why the governor reaps via find, not a glob; BSD sed/date differ from GNU. Anything that must run under an interactively-sourced shell and cron’s bare shell gets both checked.
  6. Fail-open accounting. For each swallow-and-continue the diff adds, ask: when this fails, what detector notices? Fail-open without a detector is how dozens of sims failed in silence. The house answer is a log that is empty-when-healthy — a non-empty guard log means something re-added the thing it guards against; go hunt the setter — demand an equivalent for any new swallowed error on a path that matters.
  7. Atomicity and shared state. Writes that cross session boundaries use an atomic write (temp file plus rename in the same directory) — but note its documented limit: blind whole-file, last-writer-wins, no read-modify-write safety. A diff that reads a shared file, modifies, and atomically writes back has a race no matter how atomic the write; counters and ledgers belong in a real database.

Worked example. A managed cron block is built as a here-document inside an installer script. Someone wrote a comment line inside that here-document containing the word it’s — one apostrophe. The here-document emits it literally, so the installed crontab was perfectly fine at runtime; but the apostrophe flipped the enclosing script’s quote parity as seen by the parser, and the syntax check reported an unterminated string. The fix was it’s → it is. Read it forward as the review lesson: string-built shell means content can break syntax (and the dual is worse — a quote-parity flip that the syntax check happens to accept can corrupt what actually lands in the crontab). Validation for anything that generates shell: syntax-check the source, render the output, and syntax-check or crontab-lint the output too, then eyeball what actually installed.

The failure it prevents. The whole genus of shell defects that are invisible to a reader fluent in higher-level languages: quoting that detonates on real-world data, exit codes with semantics, generated code nobody lints, and BSD/GNU drift that only manifests on the box you didn’t test.


5. The shipped-defect catalog — checks that would have caught each one

These are not hypotheticals; every class below shipped in a real fleet. For each: the incident, and the pre-merge check that catches its next instance. Run the catalog against every diff — most audits will clear most classes in seconds, and that is the point.

5.1 The cron re-setter (machine-wide mutation on a loop, no detector). A reaper re-injected a global HTTPS→SSH rewrite every ~10 minutes; fleet-wide headless git outage; manual fixes silently undone. Check: grep the diff for global config writes, crontab edits, launch-agent loads, profile writes. Any hit in code that runs unattended requires: narrowest scope, drift-healing guard, --check lint. Run the config-drift guard’s check whenever git config is anywhere near the diff.

5.2 The ghost feature (code reads state that prod doesn’t have). A file-lock fix shipped with its new session-lock column unapplied — merged, closed, and silently inert, because the code correctly fail-opened around the missing column. Check: for every column, environment variable, flag, or table the diff newly reads: prove it exists in prod (query it), or require a pending-migration labeling protocol (exact SQL in the verdict notes, the words “feature is INERT until applied,” and a queued manual-tier apply task). Graceful degradation makes this class invisible, so the more defensive the code, the harder you must check the dependency.

5.3 The error masquerading as a result (failed vs. didn’t-run). Dozens of sims failed before the conversation and reported as zero-of-N with empty fail-lists — weeks of a safety battery that never executed. Check: for any test evidence, open one raw run and count executed assertions; treat empty fail-lists, zero-duration runs, and suspiciously uniform scores as harness errors until shown otherwise. When a defect is found in one instance of a pattern, sweep every sibling immediately — the original was suite-wide for a month because nobody grepped after the first fix.

5.4 The stale-tree clobber (destruction inside an additive commit). A commit committed a stale working copy, deleting eight decision-log entries under a “capture” message; recovered ten days later by rebuilding from the parent. Check: numstat-vs-message shape test (§1); zero tolerance for deletion mass in append-only artifacts; on multi-file commits, read the file the message isn’t about.

5.5 The merge artifact (the landed diff ≠ the reviewed diff). A cherry-pick union merge dropped a closing brace from a Python script — the branch was fine; the landed result was syntactically broken. Check: validate the result commit. Minimum floor after any merge/cherry-pick/rebase: syntax-check every landed shell file, byte-compile every landed Python file, then run the touched entrypoint once. Ten seconds; catches the entire class.

5.6 The unbounded spawn (background processes without a lifetime). Daemons orphaned during fleet burns accumulated until the machine ran out of memory; fixed with session-pid self-exit plus a reaper cron. Check: every background spawn or detached process in the diff must answer: what bounds this process’s lifetime, and what reaps it if the bound fails? The house pattern is belt-and-suspenders (self-exit + reaper), mirroring the governor’s dead-PID + TTL double reap.

5.7 The rule-interaction regression (the last fix causes this bug). A live voice agent hit a voicemail and left a message despite a say-nothing rule — because a previous fix (“end every call on the standing promise”) outranked it; the leaked message was the standing-promise close. Check: for any added rule, priority, guard, or config default, enumerate which existing rules can conflict, decide the winner explicitly (hard-rule / override markers, carve-outs), and test the conflict case. “The new rule works in isolation” is the exact evidence that missed this bug.

5.8 The flaky-harness overclaim (one green run declared victory). A guard was declared fixed, then honest re-measurement showed it green one run in four; only after strengthening did it hold four in four. Check: on any known-noisy harness, require repeated passes (at least 3) for “fixed” and repeated failures for “broken”; reject any verdict whose evidence is a single run in either direction.

5.9 The envelope drift (consumers reading a field that isn’t there). Message consumers selecting a type field got null — the column is named differently — and degraded to guessing message kind from payload shape. Check: any new consumer or producer of the message table reads the correct column name and sends via the send functions; raw inserts are findings. Generalized: for every field a diff reads from a shared table, confirm against the schema that the field exists under exactly that name.

5.10 The concurrency burst (aggregate in-flight, not total tokens). Multiple account shutdowns came from too many model requests in flight at once — a dozen-plus scripts firing headless calls, parallel sims, cron lanes, and recap hooks stacking. Check: grep the diff for headless model calls, engine calls, and interactive fan-out. Every headless spawn must pass through a governor slot; loops multiplying calls are red flags even when each call is governed; anything retrying on a rate-limit signal must back off, never tight-loop. Interactive fan-out stays capped low.

The failure this catalog prevents. Re-shipping a defect the fleet has already paid tuition on. Each entry cost real damage once; the checks are the refund.


6. Verified vs. assumed — label everything

Your leverage as validator is exactly as large as your honesty about what you actually established. Every claim in your review is one of: verified (you ran/read/measured it — attach the receipt), assumed (plausible, unexamined — name what would convert it), or out of scope (explicitly not reviewed). Unlabeled claims default, in the reader’s mind, to verified — which makes an unlabeled assumption a lie you didn’t mean to tell.

Procedure.

  1. While reviewing, keep two running lists: verified-with-receipt (“ran the syntax check on the rendered block: clean”; “reproduced on the failing transcript, fix removes it”) and assumed (“assumes the engine restart happened”; “did not test under the second shell”).
  2. Distinguish environments in every claim. Sandbox-verified is not live-verified; the house language is explicit — “both verified on sandbox; STAGED for the live line — operator runs the promote step.” Repo-state is not runtime-state: a governor change was correct in the file while the running daemon still had the old code — “live on next restart.” Your verdict must say which state you validated.
  3. Grade fix strength honestly: fixed (defect mechanically impossible / reproduced-then-gone, repeatedly), mitigated (frequency reduced, residual named), guarded (detector added, defect still possible). Never let “mitigated+guarded” wear “fixed“‘s clothes.
  4. For each assumption that survives to the verdict, state the cheapest experiment that would discharge it. If that experiment is under ~2 minutes, run it instead of writing it down.

Worked example. One commit is the verdict-honesty artifact worth copying. Two defects from an operator’s first live call. Defect 1 (over-eager deferral): rule added, guard scenario written, 3/3 green — filed as FIXED. Defect 2 (banned filler under coaching): rule promoted to top-of-prompt, guard added, 0/3 → 2/3 — and filed, in the commit’s own words, honestly as mitigated+guarded, not clean-fixed, with the residual precisely named: a reflexive-agreement tic under coaching, prompt-resistant. Same commit, two claims, two different labels, each matching its evidence. That precision is what let the next session trust the ledger instead of re-deriving it — and it is the reason a related fix a day later could immediately reach for “promote to hard rule,” a pattern whose strength grade was already honestly established.

The failure it prevents. Verdict inflation — the slow poisoning of the ledger where “should work,” “worked once,” and “works” become indistinguishable, so every future decision built on your review inherits an error bar nobody can see.


7. The verdict format — answer first, then reasoning, then residual risk

The verdict is a decision instrument, not an essay. The merge decision is binary; deliver it in the first line, then make it auditable.

Procedure. Emit exactly this shape:

VERDICT: GO | NO-GO | GO-WITH-CONDITIONS
<one line: the decisive reason>

CONDITIONS (only for GO-WITH-CONDITIONS — blocking, checkable, each with its verifier)
1. ...

VERIFIED (receipts attached)
- <claim> — <how: command/run/read, and the observed result>

ASSUMED (each with the experiment that would discharge it)
- <claim> — <cheapest way to convert to verified>

FINDINGS (numbered; severity BLOCKER/MAJOR/MINOR; file:line; one-line impact)
1. ...

RESIDUAL RISK
- <what could still be wrong despite GO, how it would surface, what detector exists>

NOT REVIEWED
- <explicit scope exclusions: files, environments, behaviors>

Rules of use: a single BLOCKER forces NO-GO or a blocking condition — never a “note.” Conditions must be checkable by someone else without you (“re-run battery with repeat 3, all green,” not “be careful with the prompt”). Residual risk is mandatory even on a clean GO — a review with nothing in that section means the section wasn’t thought about, not that risk is zero. NOT REVIEWED is where partial reviews stop being dangerous: an honest small scope beats a dishonest full one.

Worked example. The verdict that history says should have been written for the cron-reaper line (§2), reconstructed:

VERDICT: NO-GO
Adds a machine-wide git-config mutation on a 10-min cron with a false premise and no detector.

VERIFIED
- No SSH keys on this machine — the rewrite cannot help pushes.
- Push auth already works: credential helper + global push-only rewrite.
- The push-only rewrite outranks the general rewrite on pushes — the line is a no-op for its stated purpose.
- With the rewrite set, a fetch in a scratch clone fails permission-denied —
  it breaks every headless clone/fetch, fleet-wide.

FINDINGS
1. BLOCKER reaper script — global rewrite; wrong scope (global for a per-run need),
   false premise, cron re-applies it after any manual unset.

RESIDUAL RISK
- Any future script can reintroduce this invisibly. Require a config-drift guard + a --check lint
  before any variant of this line is ever accepted.

NOT REVIEWED
- The rest of the reaper's reap logic (unchanged by this diff).

Note what the format forces: the premise test happened (VERIFIED bullets 1–3), the blast radius was measured rather than asserted (bullet 4), and the residual-risk line is the seed of what actually got built later as the config-drift guard.

The failure it prevents. The mushy review — three paragraphs of observations with the decision implied, conditions phrased as hopes, and the reader (often you, next week) unable to reconstruct what was actually checked.


8. Review-shaped failure — mistakes that look like competent review

Each of these produces the artifacts of diligence — comments, checkmarks, an approval with caveats — while performing none of its function. They are the validator’s own defect classes; audit yourself for them the way you audit the diff.

  1. “Tests pass” without asking what the tests don’t cover. A conversational safety-test battery was “passing” its harness for a month while half its sims never executed a single conversational turn. The competent-looking move is citing the green aggregate; the actual move is asking what evidence would this suite produce if the change were broken? — and checking that the failure mode is distinguishable from the pass you’re looking at. Also ask what kind of thing the tests can see at all: a syntax check cannot see behavior, unit tests cannot see the missing prod column (§5.2), and no test sees prose ledgers (§5.4).
  2. Style comments substituting for behavior analysis. The cron-reaper line would have sailed through a style review: quoted correctly, swallow-guarded, explanatory comment. Every hour spent on naming and formatting is fine after the behavioral questions are answered, and worthless before. If your findings list is all MINOR-style, treat that as a smell that you haven’t found the real risk yet, not that there isn’t one.
  3. LGTM on a diff too large to actually read. The stale-tree clobber lived in the file the message wasn’t about, inside a 3-file commit — small by industry standards, and still not actually read. The rule is honest arithmetic: if you did not read every hunk, you did not review the diff; either chunk it (review file-by-file with the numstat as your checklist), or scope the verdict to what you read and say so in NOT REVIEWED. An unscoped approval of an unread diff is the single most dishonest artifact a validator can produce.
  4. Re-running until green (or accepting that someone did). On a flaky harness, a retry loop will eventually manufacture a pass for a real bug. The protocol cuts both ways: single failures on known-flaky sims are noise, and single passes on a staged fix are noise (the one-in-four guard, honestly re-measured, then hardened to four-in-four). Ask how many runs, and what the distribution was — “passed on the third try” is data about a defect, not evidence of health.
  5. Reviewing the fix without reproducing the bug. If you never saw the defect, you are pattern-matching the patch against the description of a defect — and descriptions are claims (§1). You’ll approve fixes for misdiagnosed bugs (the patch is coherent; the theory is wrong) and miss that the real defect is still there.
  6. Trusting deployment labels without checking the deployment. “STAGED,” “DORMANT,” “kill switch default ABSENT,” “live on next restart” — this kind of system is full of correctly-labeled latent state (auto-merge ships disabled; a governed engine waits on a daemon restart). The competent-looking mistake is validating the code and silently assuming the activation state; the actual move is verifying which state prod is in and putting it in the verdict (§6). The dual mistake shipped as the ghost feature: everyone assumed active; it was inert.
  7. Accepting “compiles / imports clean” as behavior evidence. It’s necessary, it’s the floor, and the governor work shows the honest ladder: “compiles, imports clean, slot acquired” was one change’s label — and the verification was still the 6-caller burst peaking at exactly 2. Know which rung a given receipt sits on, and never let a lower rung answer a higher rung’s question.

The failure this section prevents. Your own review becoming the fail-open component: producing approval-shaped output whose absence of function is invisible — until the diff you “validated” becomes the next entry in §5.


9. The five-question self-test

Run this before signing any verdict. A weak answer to any question means the verdict isn’t ready — it does not mean answer it optimistically.

  1. Could I state what this diff mechanically does — every file, including the deletions — without the commit message existing? If any hunk survives only as “part of the refactor,” I haven’t read it.
  2. Which of my claims carry receipts? Have I seen the changed path fire (not merely not-error), reproduced the claimed bug, and watched the fix remove it on the failing input — repeatedly, if the harness flakes? Everything else is labeled ASSUMED in the verdict, with its discharging experiment.
  3. What is the worst thing this diff stands next to — concurrency ceiling, git-config invariant, file locks, approval gates, message envelope, hook contract, an append-only ledger — and did I check the diff against that specific invariant, including who runs this code, how often, and what re-applies it?
  4. If I’m wrong, is the failure loud or silent? If silent — the house default — what detector notices, how long until it does, and did I run the §5 catalog for the classes that already shipped that way?
  5. Does the verdict say what I did NOT review, in writing, and would I stake the next incident review on the difference between my VERIFIED and ASSUMED lists? If the honest answer reshuffles items between those lists, reshuffle them now — the ledger downstream of this verdict only works if the labels are true.

Everything above is one move wearing nine costumes: the author’s story and the change’s reality are different objects, and only the reality ships. Audit the reality. Label the rest.