repo-standard/CLAUDE.md
Kjell Tore Guttormsen 425e7e9bdf docs(invariants): the forge-state acquisition model, decided once
Four queued checks (REMOTE-SYNC, PIN-DEAD, RELEASE-CURRENT, TAG-SIGNED) need
state a local clone does not have. STATE.md carried the mechanism as an open
question per check, which is how a two-call budget grows by accretion — asked
four times, answered four ways.

Decided once, from measurement:

- Refs go over the git protocol. `git ls-remote --tags` is anonymous, reports
  annotated-vs-lightweight for free via the peeled `^{}` ref, and does not
  share the API's rate-limit bucket: 20 rapid calls all returned, and an API
  call immediately after still got 200. Stated as "does not consume the API
  budget", not "unlimited" — the run was sequential and stopped at 20.
- Releases have no git equivalent, so RELEASE-CURRENT is the third API call
  and the only new one.
- The URL comes from the register (`forge` + `org` + name), never `origin`.
  `app-creator`'s origin is `ssh://git@…`; using it would work on this machine
  and fail for every other reader — the anonymity invariant failing silently.
- TAG-SIGNED answers "is it signed" (`git cat-file tag`, local), never "does
  the forge vouch for it" — the Verified badge needs a key registered there.

Also corrects the rate-limiter note: a steady 1 request/second tripped 429 at
around request 40, so the burst description was not the whole limit. Pacing
alone does not buy immunity; a smaller total does.

VERSION-DRIFT is recorded as approved, measured, and NOT built. The rule fires
on 15 of 18 tagged repos, twelve of them merely between releases. Its two
motivating cases are already answered elsewhere or need judgement no classifier
has. Recorded rather than left open, so the next session reads the measurement
instead of repeating it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AkHEqTSr1k3HbeiHu1ggW
2026-08-12 22:27:03 +02:00

16 KiB

repo-standard

Per-repo gate for the open/ presentation standard, packaged as a marketplace plugin.

Context

Two components, one boundary:

  • Engine (scripts/repo-standard-check.mjs) — pure classifiers with all I/O resolved into their input, mirroring the marketplace's check-versions.mjs. Findings are ERROR/WARN/SKIP/OK; exit 1 on ERROR. Pinned by scripts/repo-standard-check.test.mjs (npm test).
  • Skill (skills/repo-standard/) — the judgement the script cannot encode. No checking logic lives here; it calls the engine.

register/repos.json is the single taxonomy register (name → class, per-class requirements, and the known non-repo names). Central by design: per-repo copies would recreate, in data, exactly the drift this plugin exists to remove.

Invariants

  • This repo has a PUBLIC remote. STATE.md is LOCAL-ONLY and gitignored from the first commit. Repos on a private remote track theirs — do not carry that habit across in either direction.
  • The gate sees ONE repo. Anything needing a view across the whole org (topic coverage, competing install forms, catalog-vs-forge divergence) does not belong here. It is measured where the org is enumerated.
  • It records, it does not fix. Findings first, remediation afterwards. Patching while measuring is how the inconsistency it detects was produced.
  • SKIP is never a pass — and never a severity either. A check that could not run says so and names why. But a skip is the ABSENCE of a verdict, so it cannot be the worst of a set that contains real ones: a repo's status is the worst judged finding, and SKIP only when nothing was judged (unregistered, or an empty finding set). It used to outrank OK, which made five repos with 0 ERROR, 0 WARN and a dozen OK headline as "skipped" in org-ops census 05 — okr, with the most OK in the org, read as unread. Fixing it in the status alone would have traded that for the worse direction, "skipped checks look clean", so coverage rides its own axis: notChecked in --json, a · N not checked qualifier on the summary line. Absent notChecked means an older engine and prints the pre-0.7.0 line — it is not zero.
  • A skip is one of TWO facts, and merging them says the same wrong thing more quietly. byDesign (the check saw it and declined — out-of-repo links, fixture paths; it can never become a verdict, nobody has an action) versus notRun (a re-run or an operator action resolves it). Measured across 19 clones, 35 of 39 skips were the first kind, which is why portfolio-optimiser — OK · 11 not checked read as eleven unread checks. The kind is set at the EMISSION site, never looked up from the code: VERSION-TAG is emitted at SKIP with no tags and at ERROR with the wrong one, so a code→kind map has to re-derive a reason the emission site already had. Untagged falls to notRun, the loud side. The summary line names only notRun; the deliberate ones keep their own NOT JUDGED heading, because the wanted side effect is exposure, not silence — landing the split in --json alone would have left the symptom untouched. The test that holds this is a scan of the engine SOURCE, not of a finding set: a corpus test only sees the sites it manages to trigger, and the next site added would take the reader's default invisibly.
  • "Nothing to judge" is an OK, not a skip — and it is a THIRD shape, not a third bucket. A check that ran, saw everything, and found no subject has produced a verdict: there is nothing here to be wrong. checkReadmeLanguage already answered that way for a README with no prose; VERSION-NONE (0.9.0, was VERSION-UNAVAILABLE/notRun) does for a repo that claims no version. The deferral that kept it a skip for one release was recorded in the engine as "re-levelling moves a repo's status" — measured false: an added OK cannot worsen the worst judged finding, and all three affected repos already read OK. Status moves only for a repo whose entire finding set was skips. The lesson is not about this code: a recorded decision that is wrong is worse than no record, because the next session re-derives the measurement instead of reading it. And the guard against OK blessing a real gap is that the gap is caught elsewhere — no class requires a version file, and a plugin missing its manifest is an independent FILE-MISSING ERROR. Without that, this would have been "SKIP is never a pass" one level up.
  • When a check fires many times in one repo, suspect the CHECK. The first link pass produced ~30 findings against llm-security and all were noise — regexes in code spans, file: URLs, relative paths resolved against the wrong directory. Strip code before scanning text; resolve links against the file they sit in. A gate that is wrong this often gets switched off.
  • Two axes on every finding. Level (ERROR/WARN/SKIP/OK) and bucket (broken/missing/weakening). They are independent — a weakening finding can be an ERROR.
  • Class is structural, traits are judgement. Class is read off the catalog and the remotes. A trait (security) says what the code does, which no remote knows, and the operator owns that list.
  • Who the reader is decides what is required. Contributor-facing documents are required by no class — solo-maintained, and the published stance says so. Consumer-facing ones (SECURITY.md, LICENSE, non-goals, limitations) are untouched by that. This is not a rule against having the others.
  • No CI badge, because there is no CI. The forge has zero Actions runners registered (measured). The substitute is one command from a clean clone, said plainly. A static badge asserting a run is the anti-pattern this gate flags — and an early draft of this README carried one.
  • Three outcomes on references. "No match" and "match on a known non-repo" must stay distinct findings. Collapsing them hides real loss inside correct text — the exact defect class this gate exists to catch.
  • Two API calls per invocation, anonymous, with 429 retry. The org listing (description + topics) is one; the catalog's marketplace.json for INSTALL-TRUTH is the other (added after this used to say "one call" — that line went stale and stayed stale until a 13-repo shell loop trusted it and tripped the rate limiter at 26 requests). Both go through fetchWithRetry, which retries HTTP 429 rather than silently reporting SKIP. Both are anonymous — no token, confirmed no different with one — so the gate works for any reader, not only someone holding one. A sweep across every repo still does not belong here: it needs the listing fetched once, not once per invocation, which is a different shape of caller (org-ops), not a flag on this engine. The "13 calls in a loop" explanation was incomplete (2026-08-04): the forge's nginx never sends Retry-After on its 429s (measured directly), so fetchWithRetry always falls back to exponential backoff — the Retry-After branch is live code with no live path yet. The limit is also smaller than "loop of 13" implied: 20 concurrent requests from one IP reproduced it directly, no loop needed, and a single well-formed 2-call invocation can still lose if something else on the same IP is calling the forge at the same moment (other repos' hooks, another session). The block is a leaky bucket, not a fixed ban — a 20-25 request burst took up to ~15s to fully drain. fetchWithRetry defaults to retries: 5 / maxDelayMs: 8000 (23s worst case) to cover that. Burst is not the only way in (2026-08-12): a steady 1 request/second tripped 429 at around request 40, so pacing alone does not buy immunity — only a smaller total does. Recovery measured at 35s drain plus 3s pacing: 12/12 through.
  • Two acquisition channels, and refs are not one of the API calls. Four queued checks need state the local clone does not have, and answering the mechanism question once is the point — asked per check, it gets four inconsistent answers and the budget grows by accretion. Anything answerable from refs goes over the git protocol: git ls-remote --tags is anonymous, reports annotated-vs-lightweight for free (the peeled ^{} ref), and does not share the API's bucket — 20 rapid calls all returned, and an API call immediately after still got 200. Measured sequentially and only to 20, so it is "does not consume the API budget", not "unlimited". Releases have no git equivalent, so RELEASE-CURRENT is the third API call and the only new one. Derive the URL from the register (forge + org + name), never from origin. app-creator's origin is ssh://git@…, which needs the operator's key — it would work on this machine and fail for every other reader, which is the anonymity invariant failing silently rather than loudly. Signature status splits the same way: git cat-file tag reads the signature block locally, while the forge's "Verified" badge needs a key registered there — so TAG-SIGNED answers is it signed, never does the forge vouch for it.
  • Codepoints, not bytes, not UTF-16 units. Use [...s].length. An em-dash exposes only the byte layer; astral characters expose the rest.
  • The reader decides a link's level, not just what is required. Root documents are the shop window — a dead link there is an ERROR. Below the root it is a WARN: that is where session plans, agent working files and path-traversal fixtures with deliberately invalid targets live. Measured, 30 of 43 findings were down there and all were ERRORs.
  • A fixture-path dead link is SKIP, not WARN — and never silently dropped. test/, tests/, fixtures/ (exact segment) and *golden* (substring) mark a path as presumed intentional; the finding still fires as LINK-INTERNAL-FIXTURE with its file and line, it just isn't judged. Grounded in nav-golden-escape/bundle/index.md's deliberate ../../../../etc/passwd escape: the deep .. pops the whole base path rather than resolving to null, so it read as a genuine WARN — third tool in the org to hit this exact pattern, which is the signal the check was at fault. Measured before shipping: 16 findings before, 16 after, across all 20 local clones — every one converted 1:1, none disappeared.
  • A repo's name is its remote, not its directory. catalog/ holds ktg-plugin-marketplace. The basename left it unregistered with zero checks run, against the one repo every catalog rule depends on.
  • A decision needs somewhere to live, or the gate repeats itself forever. The engine already held that a human README title is the operator's call — and still warned about it every round, because a YES could not be recorded. Six warnings, unchanged across censuses 05, 06 and 07. titles in the register is that record: set, the H1 matching it is OK; unset, the WARN stands. What the gate must never do is make "we decided this" and "nobody looked" the same output. The wanted side effect is exposure, not silence — ai-psychosis is deliberately unregistered so it stands alone.
  • An exemption is a finding, not a deletion. readme_desc_match: false turns off README-DESC equality for a class, and the check still emits an OK naming why. An exception nobody can see reads exactly like a check that silently stopped running.
  • Class rules live in the register, never as a class name in the engine. The org-profile exemption is a flag on the class, not if (klass === 'org-profile'). Per-repo copies of a rule are the drift this plugin exists to remove; a class name hardcoded in a classifier is the same defect one level up.
  • The version names a file; only the sha names the code. engineVersion was added because a stale cache served an old engine silently — but a feature and its version bump are two commits, so a worktree carries new behaviour under the old number for a window. Measured: a sweep stamped 18 raw files 0.4.0, four of them holding findings from a check that only exists in 0.5.0. engineCommit closes that, derived from the same checkout with no network call. It is present-and-null when underivable, never absent — an absent key means an older engine, null means this one ran without a HEAD.
  • A finding must name a remedy the operator can safely perform. A lightweight tag is movable without a trace, and the catalog pins plugins by tag — so it is a supply-chain property, not tidiness. But the levels come from a measurement: 155 tags across 19 clones, 14 lightweight, and only ONE repo whose newest tag is lightweight. The newest can be re-cut at no cost (ERROR); the older ones can only be "fixed" by force-moving an already published ref — the very act the check warns about — so they are exposed once, as a count (WARN), never as fourteen findings. Left unrecorded, that WARN can never be cleared, which is the titles problem again; no acceptance record is built until a repo actually needs one.
  • VERSION-DRIFT was approved, measured, and not built. The proposed rule — the manifest carries a released version string while HEAD sits N commits past that tag — fires on 15 of the 18 tagged repos, twelve of which are simply between releases (ms-ai-architect 131 commits, the catalog 262). A gate that fails a correct repository is the mechanism that gets gates switched off, and this one fails almost all of them. The two repos that motivated it are both already answered: guard's manifest claims 0.7.0 with no such tag, which is an existing VERSION-TAG ERROR (and the same root cause as org-ops' PIN-DEAD — the README pins @v0.7.0 too); okf's case turns on behaviour-changing commits past the tag, which no classifier reads off git. Two legitimate conventions coexist here — bump-at-release and bump-first — and nothing in a clone says which one a repo follows. Recorded rather than deferred: a decision that is wrong is worse than no record, and the next session should not re-derive this measurement.
  • No hook until the rule is precise. A blocking gate that fails a correct repository is the mechanism that gets gates switched off.

Commands

npm test                                                  # 182 tests
node scripts/repo-standard-check.mjs --dir "$PWD"         # gate one repo
node scripts/repo-standard-check.mjs --offline            # no network call
node scripts/repo-standard-check.mjs --json               # machine output
node scripts/repo-standard-check.mjs --refresh            # register vs. forge

Release

Run --refresh before every release. Register freshness is owned HERE, not by the sweeps that read the register. Twice running, a newly published repo was missing when a census ran, and the cost is not a gap — it is false ERRORs in a different repo: portfolio-optimiser earned three LINK-DEADs against a repo that existed, in the same round it fixed its three real ones, so its status line did not move even though the work was done. A stale register makes the gate lie about repos that are not even the stale one. One owner, no shared duty: consumers of the register are told not to check freshness themselves.

Polyrepo rule: a version bump is not finished until the tag vX.Y.Z is pushed and the catalog ref is bumped to it. Use release-plugin.mjs, never a hand-edited ref.

check-versions.mjs reads the catalog README's per-plugin label, and a missing entry is a silent null rather than an error — release-plugin.mjs rewrites an existing heading but cannot create one. A new plugin's catalog README entry has to be added by hand once.