Skillshop

26 skills and 196 lenses.

shape it harden & verify it ★ recommended

docs

0

recommended

fix the instructions first

13 lenses
0.5

recommended

apply the audit's fixes

structure

1

recommended

settle structure

data

2

recommended

lock the data layer

13 lenses

correctness

3

financial correctness review

14 lenses
4

recommended

security audit

16 lenses
5

the account lifecycle as one machine, on whitehat's control inventory

5 lenses
6

what is stored about people, on the same data map

5 lenses
7

concurrency correctness, before anything is measured

9 lenses
8

recommended

clocks, timezones and expiry, beside race-cop

8 lenses

perf

9

recommended

performance audit

8 lenses
10

fit against the target machine, on the paths perf-cop costed

quality

11

recommended

broad quality sweep on the settled structure

15 lenses

contract

12

settle the contract before anything else binds to it

6 lenses
13

the browser's half of the same contract, beside api-cop

5 lenses
14

what the interface does once the user acts, on the exchange web-cop settled

6 lenses
15

settle the settings an operability review depends on

5 lenses

ops

16

resilience and operability, as a design review

10 lenses
17

each dependency against down, slow, flapping and back

5 lenses
18

whether the signals compose, on ops-cop's deploy shape

5 lenses

adversarial

19

what an untrusted caller can spend, where fuzz then probes

5 lenses
20

recommended

adversarial hardening

tests

21

recommended

validate tests last

14 lenses

prose

22

recommended

prose the user reads

10 lenses
23

recommended

prose the operator reads

7 lenses
24

recommended

prose the maintainer reads

12 lenses

Outside the run

Audit general-agent prompts (persona, scope, guardrails, tool descriptions, examples) — LLM-agnostic

Write git commits

Set up or modernize Python projects (uv, ruff, pyproject.toml)

Audits this collection as a set — reviewer lists against each reviewers/ directory, orchestration-file parity, a lens that lives in two skills, cross-references, the review order, run-log pointer parity.

Picks which review skills and which lenses to run next in the project you are in, and names every skill that would find nothing there.

Guidelines for authoring general agent system prompts, personas, guardrails, tool descriptions, and the skill / output-style / command files a harness loads as instruction (LLM-agnostic)

Changelog

412 changes.

2026-08-28

abuse-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
accountant
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
api-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
audit-agent-docs
  • Step 4 now carries a lens that did not finish its coverage into the report, naming the area it left unread.
  • A sub-agent the auditor spawns has to return before the report, cannot be passed name:, and leaves the area it did not read named in the report.
audit-agent-prompt
  • A sub-agent the auditor spawns has to return before the report, cannot be passed name:, and leaves the area it did not read named in the report.
authflow-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
codehealth
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
comment-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
config-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
dba
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
degraded-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
follow-the-trace
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
fuzz-my-stuff-up
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
log-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
muscle-memory
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
ops-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
perf-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
privacy-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
race-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
should-i-abstract
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
string-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
test-my-tests
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
time-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
web-cop
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
whitehat
  • Distill now receives which reviewers reported coverage they did not complete, and prints the area each one left unread.
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.
will-it-run
  • A reviewer that spawns a sub-agent now has to wait for it, is barred from passing name:, and has to name any dispatched scan that did not return along with the area it left unread.

2026-08-27

abuse-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
accountant
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
api-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
audit-agent-docs
  • The structure lens counts lines as a trigger to inspect for overlapping and stale rules, rather than as a token budget. Long windows and prompt caching made the per-session token argument for >200 lines cost nothing; competing rules still cost adherence.
  • The domain deep-dive area question now carries a No deep-dive option, so full scope minus that one lens is a listed answer rather than an off-menu one. Step 3 skips the lens on that answer.
authflow-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
codehealth
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
comment-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
config-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
dba
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
degraded-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
follow-the-trace
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
fuzz-my-stuff-up
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
log-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
muscle-memory
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
ops-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
perf-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
privacy-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
race-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
should-i-abstract
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
string-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
test-my-tests
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
time-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
web-cop
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
whitehat
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
will-it-run
  • The focus-area question now carries a No focus area option, so declining a focus is a listed answer rather than an off-menu one.
write-agent-docs
  • Token Budget is now Rule Density. The 200/300-line marks prompt a read for overlapping and stale rules instead of a split, and the refusal to add content past 300 lines is gone: long windows and prompt caching made the token argument cost nothing, while a needed rule was being withheld over length.

2026-08-22

skillshop
  • Dropped the note that skill-reviews/runs/ is gitignored: the reports are tracked now, so a clone has the files its INDEX rows point at.
test-my-tests
  • Added four reviewers for the failure modes of a model-written suite: orphan-tests (the subject was deleted, or the runner never collects the file), suppressed-tests (skip, xfail, swallowed failure, eroded assertion), tautological-tests (the expected value comes from the implementation) and coverage-theater (trivial and duplicated tests).
  • Added the Slop mode: those four reviewers alone, for a suite an agent wrote or grew.
  • Added items to four existing reviewers: permissive doubles (mock-debt), assertion roulette and name-body mismatch (assertion-quality), over-broad raises (error-paths), magic numbers (data-realism).
  • Widened distill's Red tier to cover coverage that only appears to exist, and split removal-shaped findings by what the removal leaves behind.
  • Put removals outside distill's 35-point cap, in their own section capped at 60 entries, collapsing to one entry per file when a whole file goes.
  • Raised the per-agent cap from 12 to 30 for orphan-tests, suppressed-tests and coverage-theater, and made agent.md's cap overridable by the assignment.

2026-08-20

audit-agent-docs
  • Added Step 6, Record the Run: the report and an skill-reviews/INDEX.md row now go to disk under RUN-LOG-PROTOCOL.md, so skillshop can see the audit ran. Presenting and applying moved to Step 7.
  • The Findings Summary table gained a Done box per row, the resume state a later session reads out of the written report.
codehealth
  • Added a scope boundary against should-i-abstract, which rules on whether an abstraction should exist. duplicates and extract-logic were returning the same consolidation with no file naming the owner.
  • The description now opens its neighbour clause with should-i-abstract, so a session choosing between the two reads the split before loading either.
skill-fitness
  • Check 8 expects a run-log pointer in audit-agent-docs as well, which now writes one. audit-agent-prompt and skill-fitness stay out of the set.
  • Check 6's order() now accepts the * bullet that marks an essential step, so the ten starred entries stop reading as missing from the review order.
  • Check 2 skips a row whose skill has no reviewers/ directory instead of diffing its names against nothing. fuzz-my-stuff-up gained a twenty-name list in its README row and was producing twenty-two lines a run.
  • Dropped the claim that perf-cop runs twice at steps 7 and 10. It appears once, at step 9.

2026-08-19

audit-agent-docs
  • Moved the 13 lens definitions out of SKILL.md and into reviewers/, one file per lens, matching every other skill that launches an agent per lens. Step 3 now reads the file into {lens_instructions}.
  • Renamed the Redundancy lens to doc-redundancy: redundancy is already a reviewer name in string-cop, and two lenses under one name read as one in a findings table.
  • Replaced the lens numbers in cross-lens boundary notes with lens names, which survive a file moving.
codehealth
  • Corrected the step numbers in the test-my-tests boundary: this skill runs at 11 and test-my-tests at 21, not 9 and 14.
  • Added two lenses, taking the skill from thirteen to fifteen. optional-discipline covers absent values consumed unchecked, truthiness presence tests, boundary Any and type suppressions; failure-cleanup covers state left half-written when one call raises partway through.
  • Drew the new boundaries: error-gaps owns the producer of an absence and optional-discipline the consumer; type-structs owns a value's shape and optional-discipline the type layer around it; a swallowed exception is error-gaps while a correct exception over half-written state is failure-cleanup. Against the ops skills, failure-cleanup keeps the half-write one call leaves behind — ops-cop caps any finding that cannot name a deploy shape at Low, and one call raising over local state names none.
dba
  • Added the write-durability lens — fsync=off, SQLite synchronous=OFF, UNLOGGED tables, a data directory on ephemeral storage, replicas read as committed — taking the skill from twelve lenses to thirteen and Quick mode from five to six. transaction-gaps owns whether a write is wrapped; write-durability owns whether the commit survives power loss, and stops at the database's acknowledgement, where ops-cop/crash-recovery picks up.
  • Widened the scan to server config the repo carries — postgresql.conf, my.cnf, *.conf under docker/ — plus the compose and Kubernetes manifests that set the server's flags and mount its data directory.
follow-the-trace
  • Named the three-way overlap on a single span attribute: cardinality owns its value space, privacy-cop/pii-in-logs owns whose data is in it, and whitehat/egress-payload owns the payload leaving for a third party.
muscle-memory
  • Added the skill with six lenses: dead-click, interrupt-budget, keeps-my-place, path-length, sync-status, undo-not-confirm. Ships a references/precedents.md alongside the usual agent.md, distill.md and scan-steps.md — the only reviewer skill with a references directory.
perf-cop
  • Added the tail-latency lens — the spread between median and p99, per-caller size, miss paths, queueing — taking the skill from seven reviewers to eight. The boundary against blocking: blocking owns whether a wait exists and how the resource is sized, tail-latency owns the distribution across it. A fix that speeds up every request equally belongs to neither.
privacy-cop
  • Extended pii-in-logs to tracing: it now reads span attributes, the SDK's span processors and each exporter's endpoint, and treats the collector's attribute processor as that sink's redaction filter. The lens also states where it stops — follow-the-trace/cardinality reads the same span attribute for its value space, not for whose data is in it.
race-cop
  • Added the backpressure lens — unbounded queues and spawn loops, and the work dropped or lost when the backlog is cut — taking the skill from eight reviewers to nine. The boundaries came with it: perf-cop/blocking owns the same queue as waiting, backpressure owns it as loss, and against task-lifecycle it owns the ceiling rather than the owner of each task. The scan now records every queue and channel with its maxsize, or unbounded where the constructor takes none.
skillshop
  • Added a muscle-memory probe for controls a user acts on, and a note that the grep is not the input: the same form submitted forty times a shift and twice a year hit the same pattern and are different reviews.
  • Corrected the description and catalog.md: the catalog is read out of the skills tree itself, and the three sources now rank README.md first, then each reviewers/*.md H1, then SKILL.md frontmatter — frontmatter fills the purpose only for a skill with no README.md row.
  • Added the skill: SKILL.md, catalog.md and probes.md.
time-cop
  • Gave dst-arithmetic a second half with no zone precondition: the day the target month does not have. Every calendar library clamps 31 January + 1 month to 28 February and none says so, and clamping does not compose — two one-month steps from 31 January land on 28 March, one two-month step on 31 March. Added the grep step for dates derived from a stored day of month, which pushed the lens to eleven steps.

2026-08-18

git-commit-craft
  • Split the .dogcats rule in two, because the old blanket "never run git operations on it" also banned the staging the tracker needs. Git may never change tracker content — no checkout, restore, stash, reset --hard, no hand-edited JSONL — but git add .dogcats after every dcat write is now required, so the status a commit records is the status that commit produced.
  • Made steps 1 and 6 check git status --short .dogcats, since a dcat write between them leaves an MM line with one hunk staged and one not. dcat diff --staged now runs as its own shell call: chained, its exit code is hidden and an unquoted === separator aborts the line under zsh.

2026-08-16

abuse-cop
  • Added the skill with five lenses: enumeration, outbound-amplification, quota-and-lockout, rate-limit-coverage, work-per-request, plus agent.md, distill.md and scan-steps.md.
accountant
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
api-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
authflow-cop
  • Added the skill with five lenses: concurrent-flows, flow-exit, session-on-state-change, state-machine-holes, token-crossover, plus agent.md, distill.md and scan-steps.md.
codehealth
  • Named the boundary against test-my-tests: when both skills run, test-gaps reports nothing and hands its list over, since the review order runs this skill at step 9 and test-my-tests at step 14. Run alone, test-gaps keeps the lens.
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
comment-cop
  • Renamed the llm-slop lens to machine-prose, across the description, the mode lists, the boundaries and distill.md. The finding is about the prose, not about who wrote it.
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
config-cop
  • Added the skill with five lenses: missing-key-failure, naming-and-precedence, sample-drift, secret-handling, unsafe-defaults, plus agent.md, distill.md and scan-steps.md.
dba
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
degraded-cop
  • Added the skill with five lenses: dependency-map, flapping, hard-vs-soft, recovery, slow-not-down, plus agent.md, distill.md and scan-steps.md.
follow-the-trace
  • Added the skill with five lenses: cardinality, correlation-keys, sampling, span-hygiene, trace-continuity, plus agent.md, distill.md and scan-steps.md.
fuzz-my-stuff-up
  • Renamed fuzzer-agent.md to agent.md, matching every other reviewer skill.
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
log-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
ops-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
perf-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
privacy-cop
  • Added the skill with five lenses: data-inventory, erasure, minimization, pii-in-logs, retention, plus agent.md, distill.md and scan-steps.md.
race-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
  • Added the skill with eight lenses: async-misuse, cancellation, check-then-act, lock-discipline, reentrancy, shared-mutable-state, task-lifecycle, unsafe-sharing, plus agent.md, distill.md and scan-steps.md. Every finding must name the concurrent execution it assumes, and the orchestrator builds one codebase snapshot that all agents share.
should-i-abstract
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
skill-fitness
  • Added check 8, run-log pointer parity: every meta-table skill plus should-i-abstract and will-it-run must point at RUN-LOG-PROTOCOL.md. checks.md went from eight blocks to nine.
  • Added the skill: SKILL.md and checks.md.
string-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
test-my-tests
  • Renamed test-agent.md to agent.md, matching every other reviewer skill.
  • Added the sibling boundaries to the description and the skill: the source under the tests is codehealth's, a clock bug that also breaks production is time-cop's, and where both run, this skill owns the suite while codehealth/test-gaps steps back.
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
time-cop
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
  • Added the skill with eight lenses: boundary-definition, clock-choice, dst-arithmetic, expiry-math, ordering-skew, parse-format, schedule-drift, tz-storage, plus agent.md, distill.md and scan-steps.md.
web-cop
  • Added the skill with five lenses: error-coverage, form-state, no-js, post-redirect-get, status-codes, plus agent.md, distill.md and scan-steps.md.
whitehat
  • Added two lenses, taking the skill from 14 reviewers to 16. credential-scope covers how far one of your own credentials reaches — wildcard scopes, prod keys in non-prod, personal tokens as service identities. egress-payload covers data the code sends to a third party on purpose: error-tracker bodies, analytics traits, model prompts, webhook dumps. Both came with boundaries against token-lifecycle, authz, secrets-at-runtime and error-disclosure, grep patterns in the scan, and a skip rule for repos with no IAM policy or no outbound client.
  • Added the "Record the Run" step: the report goes to skill-reviews/runs/ and a row goes to skill-reviews/INDEX.md in the reviewed repository.
will-it-run
  • Added the "Record the Run" step: the skill now writes its report to skill-reviews/runs/ and appends a row to skill-reviews/INDEX.md in the reviewed repository, so a later session can pick the review back up.

2026-08-15

accountant
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
api-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path, omitted when nothing was dropped.
  • Added the skill with six lenses: breaking-changes, cli-ergonomics, exit-codes, help-accuracy, output-contract, versioning, plus agent.md, distill.md and scan-steps.md.
codehealth
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
comment-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
dba
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
log-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path, omitted when nothing was dropped.
  • Added the skill with seven lenses: dead-logs, format-consistency, level-abuse, missing-context, silence, spam, unactionable-warnings, plus agent.md, distill.md and scan-steps.md.
ops-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path, omitted when nothing was dropped.
  • Added the skill with ten lenses: config-source, crash-recovery, health, idempotent-reruns, local-state, log-destination, observability, partial-failure, retry-policy, shutdown, plus agent.md, distill.md and scan-steps.md.
perf-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path, omitted when nothing was dropped.
string-cop
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
test-my-tests
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path.
whitehat
  • Raised the report cap from 25 action points to 35, and added a Below the cap section — one line per theme with a count and one example path, omitted when nothing was dropped.
will-it-run
  • Added the skill: SKILL.md, agent.md and scan-steps.md, checking whether a codebase fits the machine it is meant to run on. No reviewers — the checks live in the scan steps.

2026-08-14

audit-agent-docs
  • Renamed the agent's "Self-Verification" section to "Output Contract" and the skill's Step 5 to "Report Conformance". Both check the report's shape, which is not the same thing as an instruction to verify.
audit-agent-prompt
  • Made the self-verification lens direction-dependent on the target model: where the model already checks its own work, a named soft-check is the finding; where it skips the check, absence is. Naming how to verify — the command, the file to diff against — always stays, since the model cannot guess it. Step 1 now asks which model the prompt runs on.
comment-cop
  • Added reveal openers to the slop lens: "Here's the thing", "It turns out" — the run-up implies the reader believed something else a moment ago.
  • Gave transient the planning-artifact case: per the M3 spec, see the migration plan, T12 covers the rest. The document is written to be superseded, and plan, spec and draft have day jobs in a query planner or a parser, so the reference is flagged and never the noun.
  • Made rambling sort candidates by size first — blocks over three content lines, comment lines over 100 characters — so the findings land before attention runs out.
string-cop
  • Added reveal openers to llm-slop: "Here's the thing", "Here's why", "It turns out" promise a turn the string never takes and cost the width the fact needed.
  • Added the compressed form of self-justification to its lens — "Kept here for backwards compatibility", "left in place" — with the fix being a limit ("Read-only; edit on /settings") rather than the history.
write-agent-docs
  • Split the self-verification section in two. What belongs in the doc is the affordance — the command, the file to diff against, the expected output — because that is environment knowledge the agent cannot guess. An instruction to verify was dropped: current models already check their work, so the line cost tokens and changed nothing.
write-agent-prompt
  • Made the "don't restate defaults" rule name its moving target: the default set shifts with each model generation, so it gets rechecked at every migration. Self-verification is the current example — frontier models check their own work, and instructing it again costs tokens for nothing.
  • Told the writer to check the target model before adding self-check instructions: for Claude Opus 5 and its generation, delete them rather than reword, since they add cost and can trigger over-verification.
  • Marked the pre-ship checklist as the writer's own; none of it ships inside the prompt.

2026-08-12

audit-agent-docs
  • Dropped disable-model-invocation: true, so the skill can load on its own trigger again.
comment-cop
  • Added two lenses: counting (a count of a set the prose does not own, a lead-in numbering its own list, "step 5"-style positional references) and settled-history (prose that is true and about the past — removals, migrations, fixed hazards, rejected alternatives, edit annotations). Both came with boundaries against contradicts-code, doc-drift, transient, dead-comments and the slop lens.
  • Struck "actually", "genuinely" and "clearly" from the skill's own prose across SKILL.md, distill.md, seven reviewers and the scan steps.
string-cop
  • Struck "actually", "genuinely", "simply" and "truly" from the skill's own prose, across SKILL.md, agent.md, distill.md, seven reviewers and the scan steps. The skill grades that vocabulary in other people's strings.

2026-08-10

accountant
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5: one priming agent in the foreground, then a five-wide window refilled on each completion.
  • Added the background-agent contract and the default-agent-type rule, and replaced the caching estimate with the 5-agent measurement.
codehealth
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5, and added the background-agent contract and the default-agent-type rule.
dba
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5, and added the background-agent contract and the default-agent-type rule.
fuzz-my-stuff-up
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5, and added the background-agent contract and the default-agent-type rule.
git-commit-craft
  • Stopped assuming the default branch is main. The skill reads it from git symbolic-ref --short refs/remotes/origin/HEAD, falls back to the current branch, and rebases against $DEFAULT_BRANCH.
perf-cop
  • Replaced 1+Parallel with Rolling 5, then with 1+Rolling 5. The window refills on each completion notification instead of running waves, because a wave leaves every finished slot idle until its slowest agent returns; a sixth concurrent agent is never allowed, since a 429 mid-run wastes the work of every agent that already finished.
  • Made the first agent run alone in the foreground. A cache entry becomes readable only once the request writing it starts streaming, so five agents launched together all miss it and all pay the 1.25× write. Agents now inherit the default model and the default agent type — either override changes the system prompt or the tool definitions and invalidates the shared entry.
  • Replaced the "agents share nothing" measurement with numbers from a 5-agent run: first agent read 0 and wrote 16,713 tokens, each later agent read 5,994 and wrote ~10.9K.
  • Split the background-agent contract out: foreground agents return findings in the tool result, background agents return an id and deliver findings in the completion notification. Never call TaskOutput on a subagent and never read its .output file — that symlink is the full transcript.
string-cop
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5: one priming agent in the foreground, then a five-wide window refilled on each completion.
  • Added the default-agent-type rule alongside the default model — either override invalidates the shared cache entry — and the background-agent contract, including never calling TaskOutput on a subagent.
  • Replaced the caching estimate with the 5-agent measurement: first agent read 0 and wrote 16,713 tokens, each later agent read 5,994.
test-my-tests
  • Replaced 1+Parallel with Rolling 5, then 1+Rolling 5: one priming agent in the foreground, then a five-wide window refilled on each completion.
  • Added the default-agent-type rule alongside the default model, the background-agent contract, and the 5-agent cache measurement.
whitehat
  • Replaced 1+Parallel with Rolling 5, then with 1+Rolling 5: the window refills on each completion notification rather than running waves, and never holds a sixth agent, since a 429 mid-run wastes the work of every agent that already finished.
  • Made the first agent run alone in the foreground. A cache entry is readable only once the request writing it starts streaming, so a simultaneous burst misses it and every agent pays the 1.25× write. Agents now inherit the default model and the default agent type.
  • Replaced the caching numbers with a 5-agent measurement: first agent read 0 and wrote 16,713 tokens, each later agent read 5,994 and wrote ~10.9K.
  • Split the background contract out: foreground findings arrive in the tool result, background findings in the completion notification. Never call TaskOutput on a subagent, and never read its .output symlink.
write-agent-docs
  • Widened the hand-off to write-agent-prompt to name skill, output-style, slash-command and subagent files, which are prompt text rather than repo docs.
write-agent-prompt
  • Extended the skill to the files a Claude Code harness loads as instruction — SKILL.md, output-styles/*.md, .claude/commands/, .claude/agents/ — and added the trigger phrases for them. The frontmatter description decides whether the body is ever injected, so both are written under these rules.

2026-08-09

accountant
  • Stopped defaulting to Full mode and Sequential when the user names neither.
audit-agent-docs
  • Added the hardcoded-count check: a number in prose mirroring a set defined elsewhere — files in a directory, rows in the table below, entries in .claude/rules/. The fix is deleting the number, not correcting it, since a corrected count drifts again on the next addition.
audit-agent-prompt
  • Added the count-versus-enumeration check to the weak-language lens: "five tools available" above six definitions. Both halves sit in the prompt, so the audit counts the list, and recommends dropping the number rather than correcting it.
codehealth
  • Stopped defaulting to Full mode and Sequential when the user names neither.
comment-cop
  • Gave doc-drift the count case: a number mirroring a set that lives elsewhere is right when written and wrong the moment the set grows. The fix is deleting the number, not correcting it, unless the number itself is the rule.
dba
  • Stopped defaulting to Full mode and Sequential when the user names neither.
fuzz-my-stuff-up
  • Stopped defaulting to Full mode and Sequential when the user names neither.
perf-cop
  • Stopped defaulting to Full mode and Sequential when the user names neither. Silence is not a choice: the skill asks and waits unless the invocation already named a mode or strategy.
  • Renamed statusline-command.py to statusline_command.py in the worked examples.
  • Added the skill with seven lenses: allocations, blocking, caching-wins, hot-loops, io-batching, payloads, startup, plus agent.md, distill.md and scan-steps.md. Every finding must name the workload it costs.
string-cop
  • Stopped defaulting to Full mode and Sequential when the user names neither.
  • Gave contradicts-view the hardcoded-count case: "12 checks run on every commit" beside a registry of thirteen. The fix is the count-free wording, not the corrected number, because a corrected literal drifts again.
test-my-tests
  • Stopped defaulting to Full mode and Sequential when the user names neither.
whitehat
  • Stopped defaulting to Full mode and Sequential when the user names neither: the skill asks and waits unless the invocation already named one.

2026-08-04

accountant
  • Rewrote the description into a triggerable one — reach for it whenever code stores, calculates, aggregates or displays money, and before shipping anything a user acts on financially — and replaced the inert frontmatter with argument-hint and disable-model-invocation.
  • Retracted the "~90% cheaper input" caching claim; the --- divider is a section divider, not a cache boundary.
  • Added the errata contract for a brief found wrong mid-run, with the matching distill step.
codehealth
  • Rewrote the description into a triggerable one — technical debt, duplication, dead code, what to refactor, and what a structural change left behind — with the boundaries against accountant, dba and whitehat. Replaced the inert frontmatter with argument-hint and disable-model-invocation.
  • Retracted the "~90% cheaper input" caching claim, and added the errata contract with its distill step.
comment-cop
  • Rewrote the description into a triggerable one, naming what the skill protects (why-comments carrying a gotcha) and where the boundaries are (string-cop for user-facing strings, codehealth for the logic).
  • Retracted the "~90% cheaper input" caching claim, and added the errata contract with its distill step.
dba
  • Rewrote the description into a triggerable one, naming the boundary against codehealth's single query-smells reviewer and stating that NoSQL is out of scope. Replaced the inert frontmatter with argument-hint and disable-model-invocation.
  • Retracted the "~90% cheaper input" caching claim, and added the errata contract with its distill step.
fuzz-my-stuff-up
  • Rewrote the description into a triggerable one — stress-test, harden, find crash inputs, run after the happy path works — and replaced the inert frontmatter with argument-hint and disable-model-invocation.
  • Retracted the "~90% cheaper input" caching claim; the --- divider is a section divider, not a cache boundary.
  • Added the errata contract with its distill step.
python-bootstrap
  • Replaced eslint with oxlint in the web templates, and added oxfmt as its own layer — oxlint does no formatting, so dropping eslint without it silently loses the indent, quote and semicolon rules. eslint.config.js gave way to .oxlintrc.json, .oxfmtrc.json and a pnpm-workspace.yaml that skips @parcel/watcher's build script, which otherwise halts pnpm install for approval.
  • Fixed the parallel-linter recipes. The trailing wait never propagated a failure — sh -c 'true & false & wait; echo $?' prints 0 — so the recipes now collect each job's pid and OR the results. djlint --reformat takes a || true, since it exits non-zero whenever it rewrites a file.
  • Moved @eslint/js and globals out of dependencies, where they had been misfiled and propagated into at least one bootstrapped project.
  • Replaced the inert args/user-invocable frontmatter with argument-hint.
should-i-abstract
  • Replaced the inert args and user-invocable frontmatter with argument-hint: "[area]" and disable-model-invocation: true. The skill runs only when the user asks for it.
  • Corrected what the 1-hour snapshot TTL is for: a staleness backstop, not a match to the prompt-cache window. Editing an already-dirty file leaves git status --porcelain unchanged, so the cache key alone goes stale. The cache is per-skill, and with one agent there is no shared prefix to prime.
string-cop
  • Rewrote the description into a triggerable one, and replaced the inert args/user-invocable frontmatter with argument-hint and disable-model-invocation: true.
  • Retracted the "~90% cheaper input" caching claim: agents share no prompt cache with each other, the --- divider is a section divider rather than a cache boundary, and snapshot size is the lever that moves cost.
  • Added the errata contract for a brief found wrong mid-run, and the matching distill step that drops earlier findings resting on a corrected claim.
test-my-tests
  • Rewrote the description into a triggerable one — it grades whether the suite would catch a real bug, and it reads tests without running them — and replaced the inert frontmatter with argument-hint and disable-model-invocation.
  • Retracted the "~90% cheaper input" caching claim; the --- divider is a section divider, not a cache boundary.
  • Added the errata contract and the matching distill step for findings resting on a corrected claim.
whitehat
  • Rewrote the description into a triggerable one — what it does, when to reach for it, how it differs from /security-review, and what belongs to dba and fuzz-my-stuff-up — and replaced the inert args/user-invocable frontmatter with argument-hint and disable-model-invocation: true.
  • Retracted the "~90% cheaper input after the first agent" claim. Agents share no prompt cache with each other: the Agent tool takes one prompt string, so the shared snapshot and the per-agent assignment land inside the same cached unit and can never match across agents. Launch order does not change cost; snapshot size does. The --- divider is a section divider, not a cache boundary.
  • Added the errata contract for a brief that turns out to be wrong mid-run. The resolved template freezes at the first launch; corrections go into an errata list appended to every later agent, and distill drops earlier findings that rested on the corrected claim, counting them as stale-brief.
  • Corrected the snapshot TTL note: a staleness backstop, not a prompt-cache window, and per-skill rather than shared across meta-skills.

2026-08-03

accountant
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
audit-agent-docs
  • Added the spawn contract and the background-agent rules.
audit-agent-prompt
  • Pinned the spawn contract — run_in_background: false, no name: — and made the findings go into the reply verbatim.
codehealth
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
comment-cop
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
dba
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
fuzz-my-stuff-up
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
should-i-abstract
  • Pinned the spawn contract: run_in_background: false and no name: — a named agent becomes a mailbox teammate and the findings never come back. The findings now go into the reply verbatim, since only the reply is rendered.
string-cop
  • Added the spawn contract: never pass name:, since a named agent becomes a mailbox teammate whose findings never come back. The distill output goes into the reply verbatim.
  • Dropped the references to the removed /sweep and /impeccable skills; a cramped page is now "a design concern", named as out of scope.
test-my-tests
  • Added the spawn contract: never pass name:, and paste the distill output into the reply verbatim.
whitehat
  • Added the spawn contract: never pass name: — a named agent becomes a mailbox teammate whose findings never come back and whose run_in_background: false is ignored. The distill output now goes into the reply verbatim, since only the reply is rendered to the user.
  • Added the skill with 14 lenses: authn, authz, crypto-misuse, dep-supply-chain, error-disclosure, file-perms, input-trust, network-exposure, secrets-at-runtime, secrets-in-git, subprocess-shell, symlink-path-safety, timing-side-channels, token-lifecycle, plus agent.md, distill.md and scan-steps.md.

2026-08-02

string-cop
  • Added the skill with ten lenses: contradicts-view, cross-screen, empty-and-error, lecture, llm-slop, reassurance, redundancy, scaffold-filler, self-justification, widget-narration, plus agent.md, distill.md and scan-steps.md. The scan extracts strings rather than dumping files, byte-for-byte, and only contradicts-view can raise a Critical.

2026-07-31

comment-cop
  • Rewrote transient's ticket-id section: the id belongs in the commit message, and in the file it resolves to nothing without tracker access and to a closed ticket about moved code with it. Two shapes — a trailing citation to delete, and an id standing in for the fact, where the fact has to be recovered from the code. Grade the spread, not the instance: one stray id is Low, a house habit across dozens of files is one High naming the prefix and the count.
  • Added "one rationale re-argued at every call site" to rambling, and the rule that length is not proof of care: comment volume tracks how hard the decision felt, not how surprising the code reads.

2026-07-30

audit-agent-docs
  • Added the prose-tics lens: mechanical greps first (curly quotes inside a command, unfilled placeholders, tool artifacts, chat residue), then sentence shapes — copulative avoidance, negative parallelism, rule of three, uniform rhythm, participle tails, elegant variation, promotional register, changelog voice. It ranks below contradictions, gaps and the stale-instruction check, and never launders a false claim by rewriting the prose around it.
  • Carved out bullet density: dense bullets are the correct form for CLAUDE.md, AGENTS.md and .claude/rules/, so the density check applies to prose docs only and never proposes turning a rule list into paragraphs.
  • Capped each lens at 12 findings, batched systemic patterns into one counted row, and replaced the bare severity list with a scale that says what each level means.
  • Added the what-not-to-flag list to actionability: a hedge carrying a real condition is a scoped rule, a definite claim is a commitment.
audit-agent-prompt
  • Added the prose-tics lens: copulative avoidance, negative parallelism, rule of three, uniform rhythm, participle tails, elegant variation, editorial filler, chat residue, promotional register, plus mechanical greps for curly quotes and tool artifacts (oaicite, utm_source=chatgpt.com). It ranks below contradictions and platitudes, requires a cluster rather than a single hit, and never speculates about authorship.
  • Capped the report at 20 findings and batched systemic patterns into one row with a count — forty rules sharing an opening formula is one finding.
  • Added the what-not-to-flag list to weak language: a hedge carrying a real condition is a scoped rule, and a definite claim is a commitment.
comment-cop
  • Added the llm-slop lens — vocabulary tics, the antithesis flourish, em-dash density, assistant narration idioms — and its boundaries: rambling owns volume, restates-code redundancy, noise decoration.
  • Ranked style below truth, in the lens, in the boundaries and in distill.md: when the same prose trips contradicts-code or doc-drift, the truth finding leads and the style note becomes a clause inside it. A rewrite must never make an unverified claim more convincing.
  • Added uniform rhythm as the strongest tell, with the carve-out that bullet density is correct form in agent-facing docs and never flagged there.
  • Replaced "load-bearing" with "carries a fact" throughout, since the skill's own slop lens flags the phrase.
write-agent-docs
  • Added the adjective test: does the word name a fact or a mood? "Drops the column before the backfill runs" is checkable; "this migration is critical" is not.
  • Added the one-name-per-thing rule — rotating "deploy script", "release tool" and "publish step" breaks the mapping between the doc and the command.
  • Ranked polish below correctness: verify a rule still holds before rewording it, since a smoothly phrased stale instruction is more convincing than the clumsy one.
write-agent-prompt
  • Added the adjective test: "the lock is held across the retry" is a fact, "the lock is critical" is a mood, and moods carry nothing the agent can apply.
  • Added the one-name-per-thing rule — rotating "payload", "request body" and "incoming data" breaks the mapping to the field the agent must emit. Padding rules to equal length loses the caveat that resisted the mold.
  • Added the polish-below-correctness check: verify a rule is still true before rewording it.

2026-07-13

codehealth
  • Set the reviewers' reporting stance: coverage rather than pre-filtering, with honest severities, since distill validates every finding.
dba
  • Set the reviewers' reporting stance: coverage, not pre-filtering, with honest severities, since distill validates every finding.
fuzz-my-stuff-up
  • Reframed the agent from attacker to defender: this is a hardening review of our own code, so a finding names the weakness, a concrete triggering input, and the fix — no working exploits, weaponized payloads or attack tooling. The "Attack" step became "Probe the defenses" and the "Attack Narrative" became an "Exposure Summary".
  • Set the reporting stance: coverage rather than pre-filtering, honest severities, since distill validates every finding.
should-i-abstract
  • Tightened the agent's deliverable from "reasoning" to a cited framework test: a finding that names no rule is incomplete.
test-my-tests
  • Set the reviewers' reporting stance: the distill step validates every finding, so report each genuine gap within the cap rather than pre-filtering, and mark severity honestly.

2026-07-03

accountant
  • Collapsed every reviewer's bespoke prose block into the shared ## Findings Summary table. Fourteen lenses had fourteen output formats, each with its own labelled fields; the table now carries Severity and File:Line plus whatever domain columns the lens defines.
  • Corrected the agent count in the scan step from 12 to 14.
audit-agent-docs
  • Split the .claude/agents/*.md boundary: frontmatter and tool scoping are this skill's, the prompt prose belongs to audit-agent-prompt.
  • Replaced the loose summary with a mandatory findings table sorted by file and line, with severity and flagging lenses per row.
audit-agent-prompt
  • Split the .claude/agents/*.md boundary: the prompt prose is this skill's, the frontmatter description and tool scoping are audit-agent-docs'.
  • Added input delimiting to the framing lens — user input, retrieved documents, examples and schemas wrapped in semantic tags.
  • Added over-hardening as a finding: hardening every rule is bloat, and the signal is lost.
codehealth
  • Collapsed every reviewer's bespoke prose block into the shared ## Findings Summary table. Thirteen lenses had thirteen output formats.
  • Added the who-does-what split between orchestrator and reviewer agents, and the severity remap: a "Critical" that is only maintainability or style becomes "High" before tiering.
comment-cop
  • Added the who-does-what split between orchestrator and reviewer agents, and collapsed every reviewer's bespoke table into the shared ## Findings Summary format.
  • Added the severity remap: a reviewer's "Critical" that does not mislead on a safety-critical property becomes "High" before tiering.
dba
  • Collapsed every reviewer's bespoke prose block into the shared ## Findings Summary table. Twelve lenses had twelve output formats.
fuzz-my-stuff-up
  • Wrote down why the fuzzers share one methodology instead of per-lens criteria files: fuzzing is open-ended, and a fixed checklist would narrow it. Each fuzzer differs only by its attack angle.
  • Added universal severity definitions as the distill baseline: easy exploitability bumps a finding up a tier, and a "Critical" needing implausible conditions drops to "High".
  • Added the who-does-what split between orchestrator and fuzzer agents, and widened the scan's git log from 15 commits to 20.
git-commit-craft
  • Added the direct-to-main rule: this repo commits to its default branch by design, no branches.
  • Handled a failing or file-modifying pre-commit hook — show the output, re-stage, retry once — and allowed git commit --amend only before a push.
  • Fixed the venv line to source $PROJECT_FOLDER/.venv/bin/activate &&.
python-bootstrap
  • Added a GitHub Actions CI template: one ubuntu-latest job that installs uv and just, syncs the locked environment, then runs just lint-all and just test-all.
  • Added [project.scripts] for CLI entry points, and three standing decisions: templates target the single pinned Python version (a library needing more tests them in a CI matrix), and pre-commit stays out, since the justfile and CI already cover formatting and linting.
  • Loosened the pinned dev dependencies to floors (pytest-cov>=7.1.0, pytest-xdist[psutil]>=3.8.0), and swapped .tox out of .gitignore for dist/, .ruff_cache/, .pytest_cache/ and *.egg-info/.
  • Fixed the SCSS stylelint globs to match .scss rather than .css, and the eslint globs to match .js rather than js.
should-i-abstract
  • Fixed the user-invokable typo in the frontmatter.
test-my-tests
  • Added the who-does-what split between orchestrator and reviewer agents, and universal severity definitions — a reviewer's "Critical" that is only a test-quality gap is remapped to "High" before tiering.
  • Normalized every reviewer file's headings from ## to #.
write-agent-docs
  • Pointed system-prompt, persona and tool-description work at write-agent-prompt.
write-agent-prompt
  • Dropped the model version from the reasoning-model caveats: "Claude with extended thinking" rather than a pinned release.

2026-07-02

comment-cop
  • Added the skill with nine lenses: contradicts-code, dead-comments, doc-drift, docstring-gaps, missing-why, noise, rambling, restates-code, transient, plus agent.md, distill.md and scan-steps.md. The scan reproduces files byte-for-byte with comments intact, and flagging a good why-comment counts as the worst error a reviewer can make.

2026-05-28

git-commit-craft
  • Dropped the co-author trailer. The commit carries no Co-Authored-By and no other attribution.

2026-05-24

codehealth
  • Added the caching lens — stale or leaky caches, unbounded growth, missing invalidation or memoization — with the boundary that query-smells owns the query underneath and caching the cache around it.

2026-04-30

accountant
  • Rewrote distill.md as a standalone agent prompt in two passes: validate, classify and dedupe mechanically, then tier and rank by judgment. Distillation moved to a fresh Sonnet agent that gets the findings tables and not the snapshot.
  • Added the high-severity validation checks — is the float drift visible, is the sign convention actually inconsistent, is the transfer filtered earlier, is the is_ignored filter applied in a parent query — and a one-tier downgrade for clear false positives.
  • Added auto-skip: no money-active language aborts the run, a single-currency codebase drops currency-mixing, and no transfer concept drops double-counting.
  • Added the snapshot cache under .claude-cache/, capped each reviewer at 12 findings, and replaced the token-count snapshot limit with wc -c at ~1,250,000 bytes.
  • Retitled the scan's grep list: the patterns select files for the snapshot, and the agents do the diagnosing.
  • Simplified the dcat probe to running dcat list --agent-only, and switched the non-git file fallback from find to Glob.
codehealth
  • Rewrote distill.md as a standalone two-pass agent prompt — mechanical validate, classify and dedupe, then judgment tiering — run by a fresh Sonnet agent that receives the findings tables and not the snapshot.
  • Added auto-skip: no SQL or ORM drops query-smells, no manifest drops dep-hygiene, no test infrastructure drops test-gaps.
  • Added the snapshot cache under .claude-cache/, capped each reviewer at 12 findings, and replaced the token-count snapshot limit with wc -c at ~1,250,000 bytes.
  • Retitled the scan's grep list: the patterns select files, the agents judge severity.
  • Simplified the dcat probe to running dcat list --agent-only, and switched the non-git file fallback from find to Glob.
dba
  • Rewrote distill.md as a standalone two-pass agent prompt — mechanical validate, classify and dedupe, then judgment tiering — run by a fresh Sonnet agent that receives the findings tables and not the snapshot.
  • Added the high-severity validation checks: is the N+1 collection actually unbounded, is the table on a hot path, does the framework auto-wrap transactions.
  • Added auto-skip: no ORM drops orm-antipatterns, no migrations directory drops migration-safety and schema-drift, no GRANT or RLS drops privilege-scope, no pool config drops connection-mgmt.
  • Added the snapshot cache under .claude-cache/, capped each reviewer at 12 findings, and replaced the token-count snapshot limit with wc -c at ~1,250,000 bytes.
  • Retitled the scan's grep list: the patterns select files, the agents judge severity.
  • Simplified the dcat probe to running dcat list --agent-only, and switched the non-git file fallback from find to Glob.
fuzz-my-stuff-up
  • Rewrote distill.md as a standalone two-pass agent prompt, run by a fresh Sonnet agent that receives the findings tables and not the snapshot.
  • Added auto-skip for six fuzzers with no target patterns: no concurrency primitives, no network code, no locale code, no path handling (drops both filesystem-edge and path-traversal), no query construction, no timezone arithmetic.
  • Added the snapshot cache under .claude-cache/, capped each fuzzer at 12 findings, and replaced the token-count snapshot limit with wc -c at ~1,250,000 bytes.
  • Simplified the dcat probe to running dcat list --agent-only, and switched the non-git file fallback from find to Glob.
should-i-abstract
  • Added the Iron Law to the scan: files go into the snapshot byte-for-byte and conclusions stay out. It names what it blocks — layer maps, "key excerpts" headings, (76 lines, key parts) digests, inline commentary, counted findings, thematic grouping, ... elisions — with an excuse-versus-reality table and a red-flag list. A snapshot that has already labeled code as duplicated turns the review into a ratification of the orchestrator's guess.
  • Added the snapshot cache: key from skill, path, git rev-parse HEAD, git status --porcelain and the language list, stored under .claude-cache/, reused for an hour.
  • Replaced the token-count size limit with wc -c over the selected files at ~1,250,000 bytes, and said how to fit: drop whole leaf modules, never abridge a file. Added redaction stubs for .env*, key and credential files, and a final check that re-reads the snapshot for anything outside a ### file: block.
  • Simplified the dcat probe to running dcat list --agent-only and reading the error, and switched the non-git file fallback from find to Glob.
test-my-tests
  • Rewrote distill.md as a standalone agent prompt in two passes: validate, classify and dedupe mechanically, then tier and rank by judgment. Distillation moved to a fresh Sonnet agent that receives the findings tables and not the snapshot, which would have added ~200K tokens for nothing.
  • Added auto-skip for lenses with no target patterns — no mocking library drops mock-debt, no clock or randomness drops flaky-risks, no tests at all stops the run.
  • Added the snapshot cache under .claude-cache/, keyed on skill, path, HEAD, dirty state and languages.
  • Capped each reviewer at 12 findings, and replaced the token-count snapshot limit with wc -c at ~1,250,000 bytes; drop whole files rather than abridging one.
  • Simplified the dcat probe to running dcat list --agent-only, and switched the non-git file fallback from find to Glob.

2026-04-22

audit-agent-docs
  • Added the scope statement, the definitions block (@ import, .claude/rules/, hooks, Iron Law, rationalization table, Red Flags, lens), and the rule-hardening lens with its selectivity warning.
  • Added the code-derivable boundary table, the red flags that precede applying edits unasked, the anchoring loopholes an agent uses to peek at another lens, the worked example finding, the self-verification step, and what to do when the user declines the changes.
audit-agent-prompt
  • Replaced the parallel fan-out with a single agent applying every selected lens in one pass. The target is one small artifact, so per-agent startup and coordination dominated the cost.
  • Added the rule-hardening lens with the harvest procedures for each of its three defenses, and rewrote the pink-elephant exception with its own rationalization table and red flags.
  • Added the in-scope/out-of-scope statement with the verbatim redirect for a CLAUDE.md target, the agent's clarification policy and escalation conditions, the false-positive filters, and the severity scale.
write-agent-docs
  • Added the guardrails: the skill covers agent docs only, and never edits them without an approved plan — a CLAUDE.md edit cascades into every future session. Named the red flags that precede editing anyway ("the user will obviously want this", "absence of 'don't' is permission").
  • Added the glossary — Iron Law, rationalization table, Red Flags, baseline run, @path import, .claude/rules/, hooks — and the three hardening moves for rules that keep breaking, with a worked | Excuse | Reality | table.
  • Added "rules come from incidents, not imagination": every rule answers what incident it is the memorial for, and the agent's verbatim self-justification feeds the next rule's loophole list.
  • Turned the line budget into an action: over 200 lines, recommend splitting; over 300, refuse further content without an approved restructuring plan.
  • Added the source hierarchy — the repo's own docs win over this skill's advice, and a conflict is named to the user rather than silently resolved.
write-agent-prompt
  • Added the skill-priority order against write-agent-docs and claude-api, and a glossary for Iron Law, baseline run, rationalization table, Red Flags, long-context deployment and scripted redirect.
  • Added the three hardening moves, the harvest loop that supplies them (baseline, record, write, pressure-test, refactor), and the counterweight: when not to harden, since hardening everything makes hardening invisible.
  • Turned rule economy into an Iron Law with its own loophole list, excuse table and red flags, and mirrored critical rules at the tail for long-context deployments.
  • Replaced "never ask more than one question" with an ordering rule: resolve the highest-impact ambiguity first, circle back later, and re-read the user's message before asking anything.
  • Added worked refusal and clarification examples, the tool-boundary disambiguation pattern with a tie-breaker, this skill's own persona as a worked example, and a pre-ship self-check list.
  • Reframed example drift: when an example and a rule disagree, the example shows what the prompt actually produces, so fix the rule.

2026-04-21

audit-agent-docs
  • Split the lens list into core and full, and rewrote the rules to say why each holds — independent agents avoid anchoring, distill needs every finding in hand, misrouted findings move noise rather than reduce it.
audit-agent-prompt
  • Removed the injection-resilience lens the day it was added: token-level defenses are not a security boundary, and the fix belongs at the architecture layer.
  • Added the skill: SKILL.md and audit-agent.md, auditing a general agent's prompt through 20 lenses.
write-agent-docs
  • Replaced the one-line cross-tool consistency note with the four drift patterns that actually bite (test framework, lint command, safety rule, workflow step), and the fix: one file is the source of truth, and identical copies get a pre-commit diff.
  • Softened the sourcing of the self-verification claim from "Anthropic's own guidance calls this the single highest-leverage thing" to "one of".
write-agent-prompt
  • Added internal consistency (persona versus output shape, scope versus tools, clarification policy versus output shape), cold-start readability, and the reasoning-model caveats — no "think step by step", minimal few-shot, constrain only the final answer.
  • Added the redundancy rule and XML structuring for anything the model must parse, and sent strict formats to structured-output schemas instead of prose.
  • Added and then removed a prompt-injection section the same day: token-level defenses are not a security boundary, and the advice belonged at the architecture layer rather than in prompt prose.

2026-04-14

accountant
  • Added the skill with 14 lenses: currency-mixing, display-vs-store, division-hazards, double-counting, float-money, idempotency, off-by-one-period, overflow-limits, phantom-records, rounding-errors, sign-convention, sum-integrity, temporal-consistency, unit-confusion, plus agent.md, distill.md and scan-steps.md.

2026-04-12

git-commit-craft
  • Rewrote the skill from a persona ("expert Git practitioner and technical writer") into a numbered workflow, and folded the edge cases into the steps that hit them. Added a verify step, git log --oneline -1, and said why open issues stay out of the message: a grep for the ID would read the commit as the one that resolved it.

2026-04-11

audit-agent-docs
  • Fixed 1+Parallel to batch at most five agents, and dropped the arXiv ids from the pink-elephant citation.
codehealth
  • Fixed 1+Parallel to batch at most five agents, since a 429 mid-run wastes the work already done.
dba
  • Fixed 1+Parallel to batch at most five agents, since a 429 mid-run wastes the work already done.
fuzz-my-stuff-up
  • Fixed 1+Parallel to batch at most five agents, since a 429 mid-run wastes the work already done.
test-my-tests
  • Fixed 1+Parallel to batch at most five agents: Anthropic rate-limits large simultaneous bursts, and a 429 mid-run wastes the work already done.
write-agent-docs
  • Added the self-verification section: name the test and lint commands, the preview variant of any destructive action, and a canonical example file.
  • Allowed file structure back in where it is non-obvious — a monorepo map earns its place, a tree of src/components/ does not.
  • Sent deterministic enforcement (format-on-save, lint-on-commit) to hooks in settings.json rather than prose.
write-agent-prompt
  • Added the skill: one SKILL.md on the prose that defines a general agent — persona, scope, clarification policy, output shape, uncertainty, escalation, tool descriptions and examples.

2026-04-09

codehealth
  • Listed the twelve reviewer names in the Full-mode description instead of the count alone.
dba
  • Listed the twelve reviewer names in the Full-mode description instead of the count alone.
fuzz-my-stuff-up
  • Listed the twenty fuzzer names in the Full-mode description instead of the count alone.
test-my-tests
  • Listed the ten reviewer names in the Full-mode description instead of the count alone.

2026-04-07

audit-agent-docs
  • Fixed the scope-boundary example to name a real global-config import.

2026-04-06

audit-agent-docs
  • Added two .claude/rules/ checks: a paths: glob matching no file means the rule silently stopped loading, and a rules file with no paths: at all costs the same tokens as CLAUDE.md every session.
should-i-abstract
  • Added the skill: SKILL.md, agent.md and scan-steps.md, one agent reviewing in both directions — code that should be shared, and abstractions that should be inlined.

2026-04-04

write-agent-docs
  • Added the skill: one SKILL.md on writing CLAUDE.md, AGENTS.md, copilot-instructions.md and .claude/rules/ for agents that can already read the source.

2026-04-02

codehealth
  • Split the agent template at the --- divider into a shared prefix and a per-agent half, resolved once and reused. (The cache claim behind this was retracted on 2026-08-04.)
dba
  • Added the skill with twelve lenses: connection-mgmt, data-integrity, index-coverage, injection, migration-safety, n-plus-one, orm-antipatterns, privilege-scope, query-scatter, raw-perf, schema-drift, transaction-gaps, plus agent.md, distill.md and scan-steps.md.
fuzz-my-stuff-up
  • Split the agent template at the --- divider into a shared prefix and a per-agent half, resolved once and reused. (The cache claim behind this was retracted on 2026-08-04.)
test-my-tests
  • Split the agent template at the --- divider into a shared prefix and a per-agent half, resolved once and reused, with a marker comment in the template. (The cache claim behind this was retracted on 2026-08-04.)

2026-03-30

audit-agent-docs
  • Extracted the agent prompt into audit-agent.md and added the launch-strategy question, replacing "always run lenses in parallel".
codehealth
  • Added the launch-strategy question, replacing "run all twelve in parallel", and pointed test-gaps at /test-my-tests for deeper test-quality work.
  • Added the risk-pattern grep to the scan and a snapshot size limit.
fuzz-my-stuff-up
  • Moved the prescan out to scan-steps.md, and replaced "launch all twenty in one parallel batch" with a launch strategy the user picks. The old rule claimed parallel launching was needed to catch timing bugs; the agents read a static snapshot, so it was not.
  • Restructured fuzzer-agent.md so the shared half — snapshot, languages, ground rules, output format — precedes the per-fuzzer assignment.
test-my-tests
  • Added the skill with ten lenses: assertion-quality, boundary-values, coverage-gaps, data-realism, error-paths, flaky-risks, fragile-tests, happy-path-only, mock-debt, user-flows, plus test-agent.md, distill.md and scan-steps.md.

2026-03-28

codehealth
  • Absorbed twelve standalone skills as reviewer criteria files: complexity, dead-code, dep-hygiene, duplicates, error-gaps, extract-logic, hardcoded, naming, query-smells, simplify-code, test-gaps and type-structs each stopped being its own skill and became reviewers/<name>.md. The reviewer-to-skill-path table went with them.

2026-03-27

codehealth
  • Made the orchestrator prescan once and pass one snapshot to every agent, instead of twelve agents each scanning the same files. scan-steps.md became an orchestrator playbook, and the agent template a snapshot consumer.
fuzz-my-stuff-up
  • Made the orchestrator prescan once and pass one snapshot to every fuzzer, instead of twenty agents each scanning the same files. Agents keep Grep, Glob and Read for tracing a specific path, but no longer scan broadly.

2026-03-25

fuzz-my-stuff-up
  • Added the skill: SKILL.md, fuzzer-agent.md and distill.md, with twenty attack angles from empty inputs and unicode chaos to state-machine abuse and adversarial users.

2026-03-22

codehealth
  • Added the skill: SKILL.md, agent.md and scan-steps.md, then distill.md the same day.
  • Added the language prescan — group git ls-files by extension, confirm the list with the user, pass it to every agent — so a repo's smaller languages do not get skipped, plus error handling for a missing reviewer file, a non-git target and agents that return nothing.

2026-03-17

audit-agent-docs
  • Renamed the skill from audit-docs to audit-agent-docs.
  • Added the tool-restriction check: prose alone is not a security boundary, so a "never run kubectl delete" rule gets a positive rewrite and a note about whether any hook or deny rule backs it.

2026-03-16

audit-agent-docs
  • Added the structure, hygiene, guardrails and agent-quality lenses: token budget with a 200/300-line threshold, progressive disclosure through @ imports, path-scoped .claude/rules/, the skills boundary, secrets, stale instructions, hook-enforceable rules and lost-in-the-middle placement.
  • Added the framing principle to actionability — positive directives over prohibitions, with NEVER reserved for catastrophic irreversible actions — and the weak-modal, weasel-phrase and unmeasurable-quality patterns.
  • Widened discovery beyond Claude Code to AGENTS.md and copilot-instructions.md, with @ imports resolved recursively to five hops, a project-root scope boundary, and a report of which tools were found.
python-bootstrap
  • Split the skill into four files: SKILL.md plus core-templates.md, web-templates.md and migration-guide.md. SKILL.md had grown past 800 lines of inline templates.

2026-03-15

audit-agent-docs
  • Added the skill as audit-docs: one SKILL.md auditing CLAUDE.md and agent_docs/ for redundancy, contradictions, gaps, actionability and misplaced content.

2026-03-07

python-bootstrap
  • Widened the ruff select list — ASYNC, FA, FBT, PL, S, T20 — and made the ignore list carry its reasons: S101 for asserts in tests, PLR0913 and PLR2004 as too noisy, C901 as too rigid, G004 because f-strings in logging are fine. Added per-file-ignores for tests/**.
  • Gave the .bootstrap stamp a time as well as a date, from date '+%Y-%m-%d %H:%M:%S'.
  • Fixed the granian invocation: --factory for a factory function, and --host/--port as separate flags rather than --bind.

2026-03-06

python-bootstrap
  • Added granian as the WSGI server, replacing uWSGI and gunicorn — no C compiler in the Docker build, one binary for WSGI and ASGI.
  • Added the re-bootstrap flow and the .bootstrap / .bootstrap-skill stamp: a dated snapshot of SKILL.md, diffed on the next run to show what moved.

2026-03-03

python-bootstrap
  • Added dcat as the issue tracker, with a CLAUDE.md template carrying the workflow, and .issues.lock in .gitignore.

2026-02-26

python-bootstrap
  • Added the .gitignore template.

2026-02-21

python-bootstrap
  • Added the Dockerfile template: multi-stage Alpine on the official uv image, uv sync --no-dev --no-install-project for dependencies only, source loaded through PYTHONPATH.

2026-02-20

python-bootstrap
  • Added the whole web-application half — djlint, eslint flat config, stylelint, optional vitest, and the parallel fmt/lint recipes — and banned npm alongside black, isort, mypy, bandit, tox and setuptools.

2026-02-16

git-commit-craft
  • Added the skill: one SKILL.md that reads git diff --staged, drafts a plain English message, asks before committing, then rebases onto origin/main.
python-bootstrap
  • Added the skill: one SKILL.md for a uv, ruff, pyright, vulture and pytest stack, with a justfile and a src/ layout.