repo-standard/CLAUDE.md
Kjell Tore Guttormsen e00ed3340c feat(engine): VERIFY-COMMAND — the one command a stranger has instead of CI
The forge has no Actions runners, so this org publishes no CI badge; the
stated substitute is one command runnable from a clean clone. A repo with
something runnable and no such command in its README is a WARN — and the
finding names what the repo already has, so the remedy is one line.

The subject is MEASURED, never read off a class. Five of 21 clones have
nothing runnable at all and answer VERIFY-NONE at OK; they span plugin,
shared-asset AND standalone, so every class-level phrasing of this rule
would fail a correct repository somewhere. Measured: 10 document a
command, 6 do not, 5 have no subject.

Two things bound the rule. It adds no API call, so it has no SKIP at all
— copying the null-input guard from every check since PIN-DEAD would
print a false "not run". And it runs nothing, so its OK says documented,
never passing.

Not built, with distinct reasons recorded as invariants: RELEASE-ASSETS
is rejected permanently for having NO SUBJECT (0 of 21 READMEs mention an
asset download; the 18/18 fire rate is a proxy and must not be quoted as
the reason). TAG-SIGNED is BLOCKED ON AN OPERATOR DECISION, not rejected
— filing it with the rejections would read as settled when it is one
yes/no from acquiring its whole subject.

Also fixes this repo's own surface, which had drifted behind its engine:
four checks had shipped without a row in the README check table, and
Requirements still said "two network calls" after the third was added.

230 tests (from 213).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LwZeAZ8cHmGZofM9dryuT9
2026-08-12 23:21:34 +02:00

24 KiB

repo-standard

Per-repo gate for the open/ presentation standard, packaged as a marketplace plugin.

Context

Two components, one boundary:

  • Engine (scripts/repo-standard-check.mjs) — pure classifiers with all I/O resolved into their input, mirroring the marketplace's check-versions.mjs. Findings are ERROR/WARN/SKIP/OK; exit 1 on ERROR. Pinned by scripts/repo-standard-check.test.mjs (npm test).
  • Skill (skills/repo-standard/) — the judgement the script cannot encode. No checking logic lives here; it calls the engine.

register/repos.json is the single taxonomy register (name → class, per-class requirements, and the known non-repo names). Central by design: per-repo copies would recreate, in data, exactly the drift this plugin exists to remove.

Invariants

  • This repo has a PUBLIC remote. STATE.md is LOCAL-ONLY and gitignored from the first commit. Repos on a private remote track theirs — do not carry that habit across in either direction.
  • The gate sees ONE repo. Anything needing a view across the whole org (topic coverage, competing install forms, catalog-vs-forge divergence) does not belong here. It is measured where the org is enumerated.
  • It records, it does not fix. Findings first, remediation afterwards. Patching while measuring is how the inconsistency it detects was produced.
  • SKIP is never a pass — and never a severity either. A check that could not run says so and names why. But a skip is the ABSENCE of a verdict, so it cannot be the worst of a set that contains real ones: a repo's status is the worst judged finding, and SKIP only when nothing was judged (unregistered, or an empty finding set). It used to outrank OK, which made five repos with 0 ERROR, 0 WARN and a dozen OK headline as "skipped" in org-ops census 05 — okr, with the most OK in the org, read as unread. Fixing it in the status alone would have traded that for the worse direction, "skipped checks look clean", so coverage rides its own axis: notChecked in --json, a · N not checked qualifier on the summary line. Absent notChecked means an older engine and prints the pre-0.7.0 line — it is not zero.
  • A skip is one of TWO facts, and merging them says the same wrong thing more quietly. byDesign (the check saw it and declined — out-of-repo links, fixture paths; it can never become a verdict, nobody has an action) versus notRun (a re-run or an operator action resolves it). Measured across 19 clones, 35 of 39 skips were the first kind, which is why portfolio-optimiser — OK · 11 not checked read as eleven unread checks. The kind is set at the EMISSION site, never looked up from the code: VERSION-TAG is emitted at SKIP with no tags and at ERROR with the wrong one, so a code→kind map has to re-derive a reason the emission site already had. Untagged falls to notRun, the loud side. The summary line names only notRun; the deliberate ones keep their own NOT JUDGED heading, because the wanted side effect is exposure, not silence — landing the split in --json alone would have left the symptom untouched. The test that holds this is a scan of the engine SOURCE, not of a finding set: a corpus test only sees the sites it manages to trigger, and the next site added would take the reader's default invisibly.
  • "Nothing to judge" is an OK, not a skip — and it is a THIRD shape, not a third bucket. A check that ran, saw everything, and found no subject has produced a verdict: there is nothing here to be wrong. checkReadmeLanguage already answered that way for a README with no prose; VERSION-NONE (0.9.0, was VERSION-UNAVAILABLE/notRun) does for a repo that claims no version. The deferral that kept it a skip for one release was recorded in the engine as "re-levelling moves a repo's status" — measured false: an added OK cannot worsen the worst judged finding, and all three affected repos already read OK. Status moves only for a repo whose entire finding set was skips. The lesson is not about this code: a recorded decision that is wrong is worse than no record, because the next session re-derives the measurement instead of reading it. And the guard against OK blessing a real gap is that the gap is caught elsewhere — no class requires a version file, and a plugin missing its manifest is an independent FILE-MISSING ERROR. Without that, this would have been "SKIP is never a pass" one level up.
  • When a check fires many times in one repo, suspect the CHECK. The first link pass produced ~30 findings against llm-security and all were noise — regexes in code spans, file: URLs, relative paths resolved against the wrong directory. Strip code before scanning text; resolve links against the file they sit in. A gate that is wrong this often gets switched off.
  • Two axes on every finding. Level (ERROR/WARN/SKIP/OK) and bucket (broken/missing/weakening). They are independent — a weakening finding can be an ERROR.
  • Class is structural, traits are judgement. Class is read off the catalog and the remotes. A trait (security) says what the code does, which no remote knows, and the operator owns that list.
  • Who the reader is decides what is required. Contributor-facing documents are required by no class — solo-maintained, and the published stance says so. Consumer-facing ones (SECURITY.md, LICENSE, non-goals, limitations) are untouched by that. This is not a rule against having the others.
  • No CI badge, because there is no CI. The forge has zero Actions runners registered (measured). The substitute is one command from a clean clone, said plainly. A static badge asserting a run is the anti-pattern this gate flags — and an early draft of this README carried one.
  • A dead pin is not a dead link, and it is not a wrong manifest. PIN-DEAD asks whether the ref a README install command pins actually resolves — LINK-DEAD asks whether the repo exists, VERSION-TAG asks whether the MANIFEST's version was tagged. All three coincided on guard only because one wrong number was written in three places. It is an ERROR because a dead documentation link costs a stranger a 404 while a dead pin costs them the install. Resolved against the FORGE, never the clone: a local tag can exist unpushed, which portfolio-optimiser demonstrates. A pin at a branch or a sha is a byDesign skip — ls-remote --tags cannot answer it, and looseness is a different finding from deadness.
  • Three outcomes on references. "No match" and "match on a known non-repo" must stay distinct findings. Collapsing them hides real loss inside correct text — the exact defect class this gate exists to catch.
  • Three API calls per invocation, anonymous, with 429 retry. The org listing (description + topics) is one; the catalog's marketplace.json for INSTALL-TRUTH is the second (added after this used to say "one call" — that line went stale and stayed stale until a 13-repo shell loop trusted it and tripped the rate limiter at 26 requests); the repo's releases listing for RELEASE-CURRENT is the third, and is the only subject the cheaper git channel cannot cover, because a release is not a ref. This count is the line that has now gone stale twice — it is updated in the SAME commit as the check that changes it, never afterwards. All go through fetchWithRetry, which retries HTTP 429 rather than silently reporting SKIP. All are anonymous — no token, confirmed no different with one — so the gate works for any reader, not only someone holding one. A sweep across every repo still does not belong here: it needs the listing fetched once, not once per invocation, which is a different shape of caller (org-ops), not a flag on this engine. The arithmetic moved with the count: a full online sweep of 22 repos is now 66 API calls, not 44. Nobody has measured 66, and the only knee ever measured is the ~40 marked unusable below. A sweeping caller paces or fetches once — this sentence exists so the number is not re-derived from a stale "two calls", which is how that line failed twice. The "13 calls in a loop" explanation was incomplete (2026-08-04): the forge's nginx never sends Retry-After on its 429s (measured directly), so fetchWithRetry always falls back to exponential backoff — the Retry-After branch is live code with no live path yet. The limit is also smaller than "loop of 13" implied: 20 concurrent requests from one IP reproduced it directly, no loop needed, and a single well-formed 2-call invocation can still lose if something else on the same IP is calling the forge at the same moment (other repos' hooks, another session). The block is a leaky bucket, not a fixed ban — a 20-25 request burst took up to ~15s to fully drain. fetchWithRetry defaults to retries: 5 / maxDelayMs: 8000 (23s worst case) to cover that. Burst is not the only way in (2026-08-12): a steady 1 request/second tripped 429 at around request 40, so pacing alone does not buy immunity — only a smaller total does. Recovery measured at 35s drain plus 3s pacing: 12/12 through. That ~40 is unmarked and cannot be quoted as a forge number: the run did not record which endpoint class it hit, and org-ops measured ~110 sequential /raw/branch/main/ reads at 0.4s with zero 429 the same day. Either they are two different limiters — the likelier reading — or the ~40 is wrong, and nothing in the note distinguishes those. The one API-axis figure this repo can stand behind: 22 sequential /api/v1/ calls at 2s pacing, anonymous, zero 429 (the RELEASE-CURRENT baseline run). That is below the suspected knee, so it bounds nothing — it is a floor, not a ceiling. Every future measurement records its endpoint class.
  • Two acquisition channels, and refs are not one of the API calls. Four queued checks need state the local clone does not have, and answering the mechanism question once is the point — asked per check, it gets four inconsistent answers and the budget grows by accretion. Anything answerable from refs goes over the git protocol: git ls-remote --tags is anonymous, reports annotated-vs-lightweight for free (the peeled ^{} ref), and does not share the API's bucket — 20 rapid calls all returned, and an API call immediately after still got 200. Measured sequentially and only to 20, so it is "does not consume the API budget", not "unlimited". Releases have no git equivalent, so RELEASE-CURRENT is the third API call and the only new one. Derive the URL from the register (forge + org + name), never from origin. app-creator's origin is ssh://git@…, which needs the operator's key — it would work on this machine and fail for every other reader, which is the anonymity invariant failing silently rather than loudly. Signature status splits the same way: git cat-file tag reads the signature block locally, while the forge's "Verified" badge needs a key registered there — so TAG-SIGNED answers is it signed, never does the forge vouch for it.
  • A stale release page is a WARN, and "no releases" is not a finding at all. RELEASE-CURRENT compares the newest release the forge publishes against the newest tag the forge lists — both sides from the FORGE, never the clone. Comparing a local tag against a published release would report portfolio-optimiser as stale when the real defect is a tag that was never pushed (v1.0.0 local, v0.1.0 published) — REMOTE-SYNC's subject, not this one. Measured across all 22 registered repos (2026-08-12): 4 have no tags, 2 tag without ever publishing a release, 11 are current, 5 lag. Those 2 — ktg-plugin-marketplace and llm-security-commons — are why zero releases is an OK (RELEASE-TAGS-ONLY) and not a finding: nothing in a repo says which of the two legitimate conventions it follows, which is the measurement that rejected VERSION-DRIFT one check over. Lagging is a WARN because the remedy is safe — publishing a release for a tag that already exists moves no published ref, unlike the remedy TAG-ANNOTATED has to withhold. It is also WARN because this repo is the worst offender in the org (v0.9.0 tagged, v0.3.0 published), and an ERROR would have let the gate decide an operator question — whether this org publishes releases at all — by exiting 1 on its own author.
  • An unpushed tag is a version that exists for nobody, and it is VERSION-TAG's blind spot, not its duplicate. VERSION-TAG reads LOCAL tags, so a manifest claiming 1.0.0 against an unpushed v1.0.0 reads as a clean pass — portfolio-optimiser read OK until REMOTE-SYNC existed. One subject in the corpus is what got BRANCH-STALE rejected; the difference is that an unpushed tag is never one of two legitimate conventions the way tag-only releasing is, the remedy (git push origin <tag>) moves no published ref, and the finding recurs at every release rather than once. The reverse direction — a clone behind the forge — is deliberately not a finding: that is the reader's machine state, not the repository's, and failing correct repos on it is how gates get switched off.
  • Version order, in the measurement as well as in the code. The shell that measured this check's baseline sorted tags with sort -t. -k1,1V and put v0.9.0 above v0.10.0 — the exact defect compareTags already exists to fix, reproduced in the tool that was supposed to validate it. It named repo-mailbox's newest tag as v0.9.0 when it is v0.21.0. A baseline derived with different ordering than the rule cannot validate the rule. Derive "newest" with the engine's own comparator, or the measurement is fiction.
  • Codepoints, not bytes, not UTF-16 units. Use [...s].length. An em-dash exposes only the byte layer; astral characters expose the rest.
  • The reader decides a link's level, not just what is required. Root documents are the shop window — a dead link there is an ERROR. Below the root it is a WARN: that is where session plans, agent working files and path-traversal fixtures with deliberately invalid targets live. Measured, 30 of 43 findings were down there and all were ERRORs.
  • A fixture-path dead link is SKIP, not WARN — and never silently dropped. test/, tests/, fixtures/ (exact segment) and *golden* (substring) mark a path as presumed intentional; the finding still fires as LINK-INTERNAL-FIXTURE with its file and line, it just isn't judged. Grounded in nav-golden-escape/bundle/index.md's deliberate ../../../../etc/passwd escape: the deep .. pops the whole base path rather than resolving to null, so it read as a genuine WARN — third tool in the org to hit this exact pattern, which is the signal the check was at fault. Measured before shipping: 16 findings before, 16 after, across all 20 local clones — every one converted 1:1, none disappeared.
  • A repo's name is its remote, not its directory. catalog/ holds ktg-plugin-marketplace. The basename left it unregistered with zero checks run, against the one repo every catalog rule depends on.
  • A decision needs somewhere to live, or the gate repeats itself forever. The engine already held that a human README title is the operator's call — and still warned about it every round, because a YES could not be recorded. Six warnings, unchanged across censuses 05, 06 and 07. titles in the register is that record: set, the H1 matching it is OK; unset, the WARN stands. What the gate must never do is make "we decided this" and "nobody looked" the same output. The wanted side effect is exposure, not silence — ai-psychosis is deliberately unregistered so it stands alone.
  • An exemption is a finding, not a deletion. readme_desc_match: false turns off README-DESC equality for a class, and the check still emits an OK naming why. An exception nobody can see reads exactly like a check that silently stopped running.
  • Class rules live in the register, never as a class name in the engine. The org-profile exemption is a flag on the class, not if (klass === 'org-profile'). Per-repo copies of a rule are the drift this plugin exists to remove; a class name hardcoded in a classifier is the same defect one level up.
  • The version names a file; only the sha names the code. engineVersion was added because a stale cache served an old engine silently — but a feature and its version bump are two commits, so a worktree carries new behaviour under the old number for a window. Measured: a sweep stamped 18 raw files 0.4.0, four of them holding findings from a check that only exists in 0.5.0. engineCommit closes that, derived from the same checkout with no network call. It is present-and-null when underivable, never absent — an absent key means an older engine, null means this one ran without a HEAD.
  • A finding must name a remedy the operator can safely perform. A lightweight tag is movable without a trace, and the catalog pins plugins by tag — so it is a supply-chain property, not tidiness. But the levels come from a measurement: 155 tags across 19 clones, 14 lightweight, and only ONE repo whose newest tag is lightweight. The newest can be re-cut at no cost (ERROR); the older ones can only be "fixed" by force-moving an already published ref — the very act the check warns about — so they are exposed once, as a count (WARN), never as fourteen findings. Left unrecorded, that WARN can never be cleared, which is the titles problem again; no acceptance record is built until a repo actually needs one.
  • VERSION-DRIFT was approved, measured, and not built. The proposed rule — the manifest carries a released version string while HEAD sits N commits past that tag — fires on 15 of the 18 tagged repos, twelve of which are simply between releases (ms-ai-architect 131 commits, the catalog 262). A gate that fails a correct repository is the mechanism that gets gates switched off, and this one fails almost all of them. The two repos that motivated it are both already answered: guard's manifest claims 0.7.0 with no such tag, which is an existing VERSION-TAG ERROR; okf's case turns on behaviour-changing commits past the tag, which no classifier reads off git. Two legitimate conventions coexist here — bump-at-release and bump-first — and nothing in a clone says which one a repo follows. This says nothing about org-ops' PIN-DEAD, which is a different check on a different subject: VERSION-TAG reads the MANIFEST and asks whether that version was ever tagged, PIN-DEAD reads a README INSTALL COMMAND and asks whether the ref it pins resolves. They coincide on guard only because the same wrong number was written in both places; a README pinning a bad ref in a repo with a correct manifest is invisible to VERSION-TAG. Recorded rather than deferred: a decision that is wrong is worse than no record, and the next session should not re-derive this measurement.
  • RELEASE-ASSETS is rejected permanently, and the reason is NO SUBJECT — not the fire rate. All 18 releases in the org carry no assets, but that number is the weaker half of the measurement. The other half: 0 of 21 READMEs mention an asset download at all — no releases/download link, no tarball, nothing. Every install path in this org goes through marketplace add, plugin install, pip/uv, or a clone. So there is no reader who is harmed and no remedy anyone would perform. That is a stronger rejection than VERSION-DRIFT's, which was "fires on repos following the other legitimate convention"; this one has no convention on either side because it has no subject. Do not re-derive it from "18/18" alone — the fire rate is a proxy, and reasoning from the proxy is what would resurrect this check the first time one repo publishes an asset.
  • TAG-SIGNED is BLOCKED ON AN OPERATOR DECISION, not rejected — and filing it with the rejections is the error to avoid. 0 of 18 repos sign their tags today, which looks like the same measurement that killed VERSION-DRIFT and is not: if the operator decides this org signs tags, the check acquires its whole subject the same day. A rejection is permanent; this is one yes/no away. Recorded here so the next session reads a pending question instead of re-measuring a settled one. The scope is also fixed in advance: git cat-file tag reads the signature block from the clone, so the check can answer is it signed and must never claim the forge vouches for it — the forge's "Verified" badge needs a key registered there.
  • A gate whose subject varies per repo MEASURES the subject; it does not read it off the class. VERIFY-COMMAND fires only where something runnable exists (scripts.test, or a tracked executable test file), and answers VERIFY-NONE at OK where nothing does. The five subject-less repos — human-friendly-style, llm-security-commons, playground-design-system, portfolio-optimiser-commons, app-creator — span plugin, shared-asset AND standalone, so every class-level phrasing of this rule fails a correct repository somewhere. The rate is fine (6 of 21 WARN) for a reason that must travel with the rule: the org publishes exactly ONE stance on verification ("no CI badge, so one command from a clean clone"), so an undocumented suite is not a second convention the way bump-first vs bump-at-release was for VERSION-DRIFT. And the check runs nothing, so its OK says documented, never passing — a green line implying a passing suite would be a claim on the surface that nobody verified.
  • The check table in the README is part of the surface it measures. Four checks shipped without a row (PIN-DEAD, TAG-ANNOTATED, RELEASE-CURRENT, REMOTE-SYNC), and the Requirements section still said "two network calls" after the third was added — the same stale-count defect the invariant above warns about, one file over, on the README of the repo whose entire job is presentation. A new check is not shipped until its row exists.
  • No hook until the rule is precise. A blocking gate that fails a correct repository is the mechanism that gets gates switched off.

Commands

npm test                                                  # 230 tests
node scripts/repo-standard-check.mjs --dir "$PWD"         # gate one repo
node scripts/repo-standard-check.mjs --offline            # no network call
node scripts/repo-standard-check.mjs --json               # machine output
node scripts/repo-standard-check.mjs --refresh            # register vs. forge

Release

Run --refresh before every release. Register freshness is owned HERE, not by the sweeps that read the register. Twice running, a newly published repo was missing when a census ran, and the cost is not a gap — it is false ERRORs in a different repo: portfolio-optimiser earned three LINK-DEADs against a repo that existed, in the same round it fixed its three real ones, so its status line did not move even though the work was done. A stale register makes the gate lie about repos that are not even the stale one. One owner, no shared duty: consumers of the register are told not to check freshness themselves.

Polyrepo rule: a version bump is not finished until the tag vX.Y.Z is pushed and the catalog ref is bumped to it. Use release-plugin.mjs, never a hand-edited ref.

check-versions.mjs reads the catalog README's per-plugin label, and a missing entry is a silent null rather than an error — release-plugin.mjs rewrites an existing heading but cannot create one. A new plugin's catalog README entry has to be added by hand once.