Guides / Reproduce before you fix: bug workflow for coding agents

Reproduce before you fix: a bug-report workflow for coding agents

2026-10-11 · 791 words · Pack: Reproduce First

Give a coding agent a bug report and it will have a theory in ten seconds. "CSV export is empty" → it opens the exporter, spots a suspicious await, and starts editing. The theory is plausible, which is the problem: plausible fixes to unreproduced bugs are how you end up with three commits titled "fix export" and a bug that still happens on the user's machine.

Humans learned this the hard way and encoded it as a rule — reproduce first — that most of us break under time pressure. An agent will break it every time unless the workflow makes reproduction a prerequisite rather than a suggestion.

The rule, stated so a machine can follow it

No Edit or Write to source files until repro/<slug>.md exists with a captured failure.

A "captured failure" is one of three things: a command and its verbatim output including the exit code; a failing automated test with its assertion message; or, for UI bugs, a screenshot path plus the steps that produced it and a console or DOM excerpt. "I read the code and I see the problem" is explicitly not on the list.

That rule lives in a skill file the agent loads when a bug report appears — any message with "broken", "fails", "error", "not working", "regression", a pasted stack trace. The skill's job is to sequence the work: name the bug, capture it, fill the checklist, then diagnose.

A capture helper that mirrors exit codes

Capturing by hand means copying terminal output into a markdown file, which agents do inconsistently. A small helper removes the friction. capture.ts runs a command, records its stdout, stderr and exit code into the record under a timestamped heading, and then exits with the same code so it composes with && and ||:

const proc = Bun.spawn(opts.command, { cwd, stdout: "pipe", stderr: "pipe", stdin: "ignore" });
const [stdout, stderr] = await Promise.all([
  new Response(proc.stdout).text(),
  new Response(proc.stderr).text(),
]);
exitCode = await proc.exited;
block += "```\n$ " + opts.command.join(" ") + "\n";
if (stdout.trim()) block += clip(stdout.trimEnd()) + "\n";
if (stderr.trim()) block += "[stderr]\n" + clip(stderr.trimEnd()) + "\n";
block += `[exit ${exitCode}]\n` + "```\n";
appendFileSync(recordPath, block);

Usage is one line:

bun .claude/skills/reproduce-first/scripts/capture.ts --slug csv-export-empty -- bun src/cli.ts export --format csv --out /tmp/out.csv

The --phase after flag writes the second block once the fix is in, so the record ends with a before/after pair produced by the identical command. --note records things a command cannot — "screenshot: repro/modal-overlap.png, Chrome 130, 1280×800" — for UI bugs.

The checklist that forces the questions

The record starts from a template whose Before fixing section cannot be skipped:

## Before fixing (all required)
- [ ] Exact steps to trigger (numbered)
- [ ] Capture attached below (command + output, failing test, or screenshot path)
- [ ] Expected vs actual stated in one line each
- [ ] Reproduced on the same branch/data the reporter used
- [ ] Scope noted: does it affect one path or many?

"Expected vs actual" is the item agents most want to skip and most need to write. Stating what correct output looks like — a header row plus one line per record, 4.8 KB not 0 bytes — is what turns "the file is empty" into something a test can assert.

When the capture does not fail

This happens more than you would think, and it is the most valuable outcome the workflow produces. If the command the agent ran succeeds, the bug is not reproduced. The skill says so, asks the reporter for the exact steps, data and environment, and does not touch code. Compare that with the alternative: an agent that "fixes" something in a code path the user never hits, closes the ticket, and the real bug stays.

Intermittent bugs are the sharp version of this. A capture that fails one time in five is still a capture — the record shows four [exit 0] blocks and one [exit 1], which already tells you it is timing or state dependent.

Diagnose with the capture open

Only now does the agent read code. The record gets a Diagnosis section — hypothesis, evidence for it, root cause once confirmed — and the discipline is to write the hypothesis before editing. If the hypothesis is "the writer flushes before the stream ends because await writer.end() was dropped in a refactor", the fix is obvious and the regression test writes itself: assert the output length is greater than zero after export.

Prove it with the same command

The fix is done when the same capture, re-run with --phase after, shows the expected behaviour. Not a different command, not a unit test alone — the user's flow, as captured. The final message quotes the pair:

Fixed csv-export-empty: before 0 bytes → after 4,812 bytes with the same command; regression guarded by csv.test.ts › writes all rows before closing. Record: repro/csv-export-empty.md.

That sentence contains everything a reviewer needs and nothing they have to take on faith.

Anti-patterns the workflow is designed to stop

Editing first because the cause is "obvious". Declaring a fix without re-running the original failing command. Reproducing on a different branch or dataset than the report and calling it the same bug. Fixing a proxy — a symptom visible in a unit test — while the user's actual flow still fails. Each of these is a shortcut an agent will take under a vague instruction and will not take under a checklist.

Install

The Reproduce First pack is the skill file, the record template and capture.ts, tested end to end (capture mirrors exit codes, before/after blocks, slug validation).

curl -fsSL https://hookcrate.com/install | bash   # Windows: irm https://hookcrate.com/install.ps1 | iex
hookcrate login hc_YOUR_KEY
hookcrate install reproduce-first

$9/month or $99 lifetime; one license covers every pack in the catalog. (bunx hookcrate is coming soon on npm.)

The pack

Reproduce First Skill
Forces a captured reproduction (command output, failing test or screenshot) before any code edit on a bug report.
View pack

$9 / month or $99 lifetime — one license covers every pack. Get access.