Guides / Auditing your CLAUDE.md: cut, keep, sharpen

Auditing your CLAUDE.md: cut, keep, sharpen

2026-10-11 · 752 words · Pack: Bitter Pill Audit

A CLAUDE.md starts as ten lines about the package manager and the test command. Six months later it is three hundred lines, and the agent has started ignoring the parts you care about most. Nothing in it is wrong exactly. It is just that every incident added a rule and no incident ever removed one, and now the file contains instructions the model follows anyway, instructions that say the same thing twice, and at least one pair that contradict each other.

The uncomfortable fact is that instruction files have a budget — of attention, not just tokens — and every rule spends it. A vague rule does not merely fail to help; it dilutes the sharp rule next to it. So the audit's default posture is to cut, and a rule has to earn a KEEP.

The five questions

Ask them of every rule in order. The first "yes" is the verdict.

  1. Does the model already do this? "Write clean, maintainable code." "Think step by step." "Be concise." A capable model does these unprompted. CUT.
  2. Does it contradict another rule? "Always commit directly to main" in one section, "never commit to main without a PR" in another. The model picks one at random per session. RESOLVE — usually by scoping each to its real trigger.
  3. Is it a duplicate? "Use bun, never npm" and "Prefer bun over npm where possible" are one rule with a hedge bolted on. MERGE into the clearer one.
  4. Is it a one-off that generalised badly? A rule written after a single incident that now fires everywhere. CUT or scope it.
  5. Is it vague? No observable behaviour; hedge words — "appropriate", "as needed", "properly", "try to", "best practices". SHARPEN into a trigger and a concrete do/don't.

What survives is a minority, and it tends to be the same kinds of thing: facts the model cannot know (the test command, the deploy steps, which port), hard prohibitions with a concrete trigger, routing rules ("for X use skill Y"), verification requirements, and specific examples.

A mechanical first pass

Judging three hundred rules by hand is tedious, so the pack ships audit.ts to do the mechanical part. It splits the file into rules — bullets, numbered items, table rows, imperative sentences under headings, with fenced code ignored — and flags each one:

if (HEDGES.some((h) => t.includes(h))) r.flags.push("vague");
if (!ACTION_VERBS.test(r.text) && !/[`:]/.test(r.text)) r.flags.push("no-action");
if (DEFAULTS.some((d) => d.test(r.text))) r.flags.push("default");
if (r.tokens > 60) r.flags.push("long");
for (let j = 0; j < i; j++) {
  if (similarity(r.text, rules[j].text) >= threshold) { r.flags.push(`dup:${rules[j].n}`); break; }
}

The duplicate check is token-set Jaccard similarity over the rule text with stop words removed, threshold 0.6 by default. The conflict check looks for pairs of always/never rules that share three or more content words and disagree on polarity. The hedge list and the "default behaviour" patterns are the opinionated part:

const HEDGES = ["appropriate", "as needed", "where possible", "properly", "try to",
                "best practices", "be careful", "high quality", "reasonable", "ideally", ...];
const DEFAULTS = [
  /\b(write|produce|create)\b[^.]{0,40}?\b(clean|good|readable|maintainable|quality)\b[^.]{0,30}?\bcode\b/i,
  /\bbe (helpful|concise|clear|accurate|professional|polite)\b/i,
  /\bthink (step[- ]by[- ]step|carefully|before (you )?(act|answer|respond))\b/i,
  ...
];

Tokens are estimated at roughly four characters each — accurate enough to compare before and after, which is all the number is for. The output is a table:

Rules: 87  Tokens: ~2140 → ~1180 if suggestions applied (−45%)
Suggested: KEEP 31 · CUT 19 · SHARPEN 28 · MERGE 9

| #  | line | section   | rule                                                      | tokens | flags         | suggested |
| 3  | 9    | Principles| Always write clean, maintainable code following best pr…  | 16     | default vague | CUT       |
| 13 | 25   | Tooling   | Prefer bun over npm where possible                        | 8      | vague dup:12  | MERGE     |
| 58 | 131  | Git       | Always commit directly to main                            | 7      | conflict?:61  | SHARPEN   |

The suggestions are heuristics, and the skill says so in its last line. They are there to put the five questions in front of you with the evidence attached, not to decide for you.

The judgement pass

Walk the table. For each flagged rule write the final verdict and, for SHARPEN and MERGE, the replacement. The rewrite shape is deliberately boring:

<TRIGGER>: <DO this> — <not that>. e.g. <one example>

"Make sure tests are run appropriately before finishing" becomes "Before claiming any task is done: run bun test and quote the pass/fail count — never 'tests should pass'." Shorter, observable, and it carries its own verification.

For the contradiction, read the context. In the example above, "commit to main" sat under Personal repos and "never without a PR" under Work repos. Neither was wrong; both were unscoped. The resolution is one rule with two triggers.

Calibration checks

If more than 70% of rules survive as KEEP, you were too kind — go back to question one. If a rule's only defence is "it can't hurt", it can; cut it. If a rewrite is longer than the original, you are adding scaffolding, not sharpening. And the finished file should read like a runbook, not a personality sheet.

Keep it read-only

The audit produces a report and a proposed file, then stops. The instructions file is the owner's written content; overwriting it is a separate, explicit step. In practice the proposed file gets a human read-through, a few verdicts get reversed, and then it goes in — typically a third to a half shorter than it was.

Install

The Bitter Pill Audit pack is the skill with the five-question workflow and audit.ts, tested against a fixture file containing default, vague, duplicate and conflicting rules.

curl -fsSL https://hookcrate.com/install | bash   # Windows: irm https://hookcrate.com/install.ps1 | iex
hookcrate login hc_YOUR_KEY
hookcrate install bitter-pill-audit

$9/month or $99 lifetime; one license covers every pack in the catalog. (bunx hookcrate is coming soon on npm.)

The pack

Bitter Pill Audit Skill
Audits a CLAUDE.md or instructions file for redundant, vague and contradictory rules with keep / cut / sharpen verdicts and token estimates.
View pack

$9 / month or $99 lifetime — one license covers every pack. Get access.