Guides / Stop Claude Code marking work done without evidence

Stop Claude Code from marking work done without evidence (the Stop hook pattern)

2026-10-11 · 744 words · Pack: Evidence Gate

An agent's final message describes what it intended to do. "Implemented the limiter, added tests, wired the router — done." Sometimes every word is true. Sometimes the tests were never run, because the agent knows what they would say. From the outside the two messages are identical, and the difference only surfaces when the next person builds on the result.

The fix is not to ask the model to be more careful. It is to make "done" a state the harness checks. Claude Code's Stop hook is the place to do it.

The Stop hook, briefly

Per the hooks documentation, a Stop hook runs every time Claude is about to end its turn. It receives JSON on stdin — session_id, cwd, stop_hook_active, and last_assistant_message — and can prevent the stop in two equivalent ways: exit code 2 with the reason on stderr, or exit 0 with {"decision":"block","reason":"..."} on stdout. Either way Claude receives the reason and keeps working. The docs also warn about loops: stop_hook_active is true when the current turn exists because a previous Stop hook blocked, and Claude Code caps consecutive continuations at eight.

So a Stop hook can hold the door shut, but it has to be able to say precisely why — a block without a reason is just a stuck session.

It also has to be cheap and deterministic. A Stop hook runs on every turn, so it cannot call a model, hit the network, or take more than a second or two. Reading one file and applying a regex is the right weight class. That constraint shapes the whole design: the hook does not judge whether the evidence is good, only whether it exists in the place the convention says it must. Judgement stays with the human reading the checklist; the hook enforces that there is something to read.

Make the checklist the source of truth

The pattern needs something concrete to check. A checklist file in the project root works well: ISA.md, CHECKLIST.md, TODO.md, whatever your team already uses, with criteria as checkboxes and the proof under each one:

status: in-progress

## Criteria
- [x] Sliding window limiter implemented
  evidence: src/limiter.ts read back, exports allow(), 61 lines
- [x] Unit tests cover burst + steady state
  evidence: `bun test src/limiter.test.ts` → 6 pass, 0 fail
- [ ] Router returns 429 with Retry-After

Two things matter here. Every checked item has an evidence: line naming something observed — a command and its output, a test count, a file read back. And the file carries a status. The gate should only engage when the agent claims completion by flipping status: in-progress to status: complete; a hook that blocks every stop while work is merely unfinished would be intolerable.

Parsing it

The parser is small. Find bullet checkboxes, then look at the indented continuation lines beneath each one for an evidence marker:

const ITEM_RE = /^(\s*)[-*+]\s+\[( |x|X)\]\s+(.*)$/;
const EVIDENCE_RE = /\bevidence\s*:/i;
const COMPLETE_RE = /^\s*(?:[-*]\s*)?\**(?:status|phase)\**\s*[:=]\s*\**\s*(complete|completed|done)\b/im;

for (let i = 0; i < lines.length; i++) {
  const m = ITEM_RE.exec(lines[i]);
  if (!m) continue;
  const indent = m[1].length;
  let hasEvidence = EVIDENCE_RE.test(m[3]);
  for (let j = i + 1; j < lines.length && !hasEvidence; j++) {
    const nm = ITEM_RE.exec(lines[j]);
    if (nm && nm[1].length <= indent) break;          // next sibling item
    if (EVIDENCE_RE.test(lines[j])) hasEvidence = true;
  }
  criteria.push({ line: i + 1, text: m[3].trim(), checked: m[2] !== " ", hasEvidence });
}

The completion regex tolerates the ways people actually write it — status: complete, Status: done, - phase = completed. Being strict here only produces false negatives.

Deciding, and saying why

When the file declares completion, every criterion must be checked and evidenced. Anything else is a named problem:

if (declaredComplete) {
  for (const c of criteria) {
    if (!c.checked) problems.push(`line ${c.line}: UNCHECKED — ${c.text}`);
    else if (!c.hasEvidence) problems.push(`line ${c.line}: NO EVIDENCE — ${c.text}`);
  }
}

The output is the whole point. Instead of "not done", Claude sees:

evidence-gate: ISA.md declares completion but 2 criteria are unproven.
  - line 8: NO EVIDENCE — Unit tests cover burst + steady state
  - line 9: UNCHECKED — Router returns 429 with Retry-After
Each criterion needs `- [x]` AND an `evidence:` line describing the observed proof ...

That is a to-do list, not a scolding. The agent runs the tests, pastes the count as evidence, finishes the router work, and the next stop passes.

Emit both block signals

To block, the hook prints the JSON form and exits 2:

if (event === "Stop") console.log(JSON.stringify({ decision: "block", reason }));
console.error(reason);
process.exit(2);

Printing both costs nothing and survives harness versions that honour only one of them. When stop_hook_active is true the message is shortened — the agent has already seen the full explanation once — and the eight-continuation cap means a genuinely stuck checklist cannot lock a session forever.

Catch it earlier with PostToolUse

A Stop hook fires at the end of the turn. The claim of completion happens earlier, when the agent edits the checklist. Registering the same script as a PostToolUse hook on Edit|Write|MultiEdit, and having it act only when the edited file is the checklist, surfaces the problem at the moment status: complete is written. PostToolUse cannot undo the edit — the docs are clear that exit 2 there does not block, since the tool already ran — but stderr is still shown to Claude, which is all that is needed.

{
  "hooks": {
    "Stop": [{ "matcher": "", "hooks": [{ "type": "command", "command": "bun \"${CLAUDE_PROJECT_DIR}/.claude/hooks/evidence-gate.ts\"", "timeout": 20 }] }],
    "PostToolUse": [{ "matcher": "Edit|Write|MultiEdit", "hooks": [{ "type": "command", "command": "bun \"${CLAUDE_PROJECT_DIR}/.claude/hooks/evidence-gate.ts\"", "timeout": 20 }] }]
  }
}

Where it earns its keep

Use this on work someone will rely on: deploys, migrations, anything touching money, long unattended sessions where nobody is watching the transcript. Skip it for a one-line fix. The discipline of writing evidence lines costs a minute per task; the cost of a confident false "done" on a migration is measured differently.

Install

The Evidence Gate pack is the hook above with the checklist parser, README and settings fragment, tested across Stop and PostToolUse payloads for block, allow, in-progress silence and malformed stdin.

curl -fsSL https://hookcrate.com/install | bash   # Windows: irm https://hookcrate.com/install.ps1 | iex
hookcrate login hc_YOUR_KEY
hookcrate install evidence-gate

$9/month or $99 lifetime; one license covers every pack in the catalog. (bunx hookcrate is coming soon on npm.)

The pack

Evidence Gate Hook
Refuses to let a task be marked complete while checklist criteria are unchecked or lack an evidence line.
View pack

$9 / month or $99 lifetime — one license covers every pack. Get access.