Packs / Research Fleet

Research Fleet

Skillv1.0.0

Decomposes a research question into sub-questions, runs parallel subagents, cross-checks claims and verifies every URL before reporting.

Install

# once per machine
curl -fsSL https://hookcrate.com/install | bash
# then
hookcrate install research-fleet

Get access

One license covers all packs.

What it does

An orchestration skill for questions that deserve more than one search. Claude splits the question into 3–5 sub-questions, launches one subagent per sub-question in parallel, merges their findings into a claim table, cross-checks each claim across sources, runs verify-urls.ts so that no dead or hallucinated link reaches the report, and emits a report where every claim carries a [HIGH] / [MED] / [LOW] / [CONFLICT] confidence tag. Includes the SKILL.md orchestration workflow, the report template, and the URL verifier script.

Usage: bun .claude/skills/research-fleet/scripts/verify-urls.ts report-draft.md --json

When to use this

Use Research Fleet for questions where being wrong is expensive and one search cannot settle the matter: choosing a database for a new system, checking whether a library is still maintained, comparing pricing across vendors, or finding out what actually changed in a dependency's last three releases. A single search returns one source's framing; an agent summarising it will sound confident regardless of how thin the evidence is. The fleet pattern fixes both problems. The question is split into three to five independent sub-questions, one subagent runs each in parallel, and the results are merged into a claim table where every claim is tagged by how well it is supported. The verify-urls.ts step then fetches every link in the draft and fails the report if any is dead or points at a page that does not mention the claim — the fix for hallucinated citations is to test them like code. Use it when the answer will drive a decision, a document someone else will read, or a purchase. Do not use it for a single documented fact or an API signature; that is one search and done. Expect a run to take a few minutes and produce a report you can hand to someone without caveats.

When it triggers

  • User asks to research, investigate, compare, find out, or 'what's the state of' a non-trivial topic
  • User invokes /research-fleet

Preview — a sample session

Illustrative transcript showing the pack in use.

User: Research whether SQLite is a sane choice for a multi-tenant SaaS backend in 2026. I want a real answer, not a blog post.

Claude: Non-trivial, contested topic → research-fleet. Decomposing first.

Sub-questions
1. What write-concurrency limits does SQLite have today (WAL, BEGIN CONCURRENT)?
2. Which production SaaS products run SQLite per-tenant, and at what scale?
3. What do the replication/backup options (Litestream, LiteFS, rqlite, Turso) actually guarantee?
4. Adversarial: strongest documented evidence of teams that moved OFF SQLite and why.
● Agent(explorer-1: write-concurrency limits)   ● Agent(explorer-2: production per-tenant users)
● Agent(explorer-3: replication guarantees)     ● Agent(explorer-4: who left SQLite and why)
  ⎿  4 agents running in parallel …
  ⎿  explorer-1: 6 claims, 2 NOT FOUND · explorer-2: 5 claims · explorer-3: 7 claims · explorer-4: 4 claims, 1 conflict

Claude: 22 raw claims → 14 after dedup. Cross-checking: C3 (single-writer limit) is confirmed by the official docs and two independent engineering posts → [HIGH]. C9 ("one vendor runs 40k tenant DBs") appears in three places but all cite the same talk → [MED]. C12 conflicts: one source says LiteFS is deprecated, another dated later says it is maintained → [CONFLICT], both dates kept. Now verifying every link before writing.

● Bash(bun .claude/skills/research-fleet/scripts/verify-urls.ts draft.md --expect "BEGIN CONCURRENT=https://sqlite.org/cgi/src/doc/begin-concurrent/doc/begin_concurrent.md")
  ⎿  | ok | status | url | final | title/keyword |
     | PASS | 200 | https://sqlite.org/wal.html | | Write-Ahead Logging |
     | PASS | 200 | https://sqlite.org/cgi/src/doc/begin-concurrent/… | | BEGIN CONCURRENT: found |
     | FAIL | 404 | https://example-blog.dev/sqlite-at-scale-2023 | | Not Found |
     …
     13/14 passed, 1 FAILED — remove or replace before publishing.

Claude: One dead link — the claim it supported (C7) also had a second source, so it drops from [HIGH] to [MED] and the dead URL is removed. Writing the report.

● Write(research/sqlite-multi-tenant.md)

Claude: Answer: SQLite is a defensible choice for database-per-tenant designs with modest write rates per tenant (C1, C3 [HIGH]), but not for a shared single-file multi-tenant store under concurrent writes (C3, C5 [HIGH]). Replication tooling gives you durable backups, not synchronous multi-writer HA (C10 [HIGH], C12 [CONFLICT] on LiteFS status). Teams that left cite fleet operations at >10k DBs, not raw performance (C13 [MED]). Full evidence table, conflicts and the 2 things I could not find are in research/sqlite-multi-tenant.md; all 13 links verified.

FAQ

How many subagents does it launch?

One per sub-question, three to five, launched in a single message so they run concurrently. An extra adversarial sub-question — what is the best evidence the common answer is wrong — is recommended for any topic with a popular narrative.

What do the confidence tags mean?

[HIGH] is two or more independent sources with at least one primary. [MED] is one credible source or several that share an origin. [LOW] is a single weak source or your own inference. [CONFLICT] means credible sources disagree and both sides are shown with dates.

What happens to a claim whose URL fails verification?

The URL is removed and the claim drops one confidence level; if it had no other source it is removed entirely. A dead or wrong link never reaches the final report.