feat(gate): two presentation checks the evidence actually supports

Adversarial deep research (25 sources, 124 claims extracted, 25 verified:
11 confirmed / 14 refuted) plus direct measurement of all 18 cloned open/
repos. The useful half of the result is what it REFUSED to support, so the
research is recorded in docs/presentation-research-2026-08-03.md rather than
being spent and forgotten.

- BADGE-COUNT (WARN) — past five badges. Trockman et al., ICSE 2018
  (n=294,941 npm packages) measured a non-linear relationship with popularity
  inflecting at five, motivated by surveyed maintainers calling over-badged
  READMEs cluttered and "trying too hard". WARN and never ERROR: the
  coefficient sits in an appendix with no CI or p-value. Counting deliberately
  uses a NARROWER rule than the existing claim check, so a screenshot or an
  architecture diagram is never counted as clutter. Fires on 8 of 18.
- README-LANGUAGE (WARN) — prose not in the language this repo's readers were
  declared to speak, via a new `locales` axis in the register. Class is
  structural, a trait is what the code DOES, a locale is who it is FOR — the
  standard's own "who the reader is decides what is required". English is the
  default; ms-ai-architect and okr are declared nb, named by the operator as
  Norway-only in audience. Stopword-frequency comparison over prose with code
  stripped: a Norwegian flag name in a shell example cannot decide the
  document. Fires on exactly those two, silent on all sixteen English repos.

One design correction found mid-implementation: the first version returned
SKIP when a README had too little prose to judge, which broke a passing
fixture and would have stopped any terse repo from ever reaching OK. SKIP is
for a check that could not RUN; this one ran, saw everything and found no
prose to be in the wrong language — the same shape as "no licence claim to
back". Insufficient prose is now OK, and evenly bilingual prose is the SKIP,
because there the question is live and unanswered.

Deliberately NOT built, because the evidence does not reach: any rule about
images, diagrams or terminal recordings (every such claim refuted 0-3); a
README length bound (no evidence-based target exists); a section count (would
fire on 9 of 18 — textbook "suspect the CHECK"); Mermaid source length and
#gh-dark-mode-only (zero occurrences, and the instance limit is not readable
via the API, so any threshold would be a guess).

Verified live against the operator's own forge (15.0.6+gitea-1.22.0): Mermaid
DOES render in README.md — two div.mermaid-block iframes carrying real SVG —
while #gh-dark-mode-only landed only in Gitea 1.26.0 and is unavailable here.

103 tests green, up from 92. No version bump: the catalog ref still trails at
v0.1.1 against 0.1.3, and starting a second release chain over that is the
operator's call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCAGKZT8h9F46ygzSkhEee
This commit is contained in:
Kjell Tore Guttormsen 2026-08-04 09:41:53 +02:00
commit dc386d4471
7 changed files with 474 additions and 2 deletions

View file

@ -4,6 +4,39 @@ All notable changes to this project are documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
versioning is [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
### Added
Two presentation checks, both from adversarially-verified research rather than
taste — `docs/presentation-research-2026-08-03.md` records what the evidence
supports and, more usefully, what it refuses to support.
- `BADGE-COUNT` (`WARN`) — more than five badges. Trockman et al., ICSE 2018
(n=294,941 npm packages) measured a non-linear relationship with popularity
inflecting at five, motivated by surveyed maintainers calling over-badged
READMEs cluttered and "trying too hard". `WARN` and never `ERROR`: the
coefficient sits in an appendix without CI or p-value, so it carries "more is
not better" and cannot carry a hard limit. Counting uses a narrower rule than
the existing claim check — a screenshot or architecture diagram must not be
counted as clutter. Measured: 8 of 18 `open/` repos are past it.
- `README-LANGUAGE` (`WARN`) — the prose is not in the language this repo's
readers were declared to speak, via a new `locales` axis in the register.
English is the default; `ms-ai-architect` and `okr` are declared `nb` as
Norway-only in audience. Detection is a stopword-frequency comparison over
prose with code stripped, so a Norwegian flag name in a shell example cannot
decide the document. Evenly bilingual prose is a `SKIP` — the question is
live and unanswered. No running prose is an `OK`: nothing claims a language,
and a thin README is `checkFirstScreen`'s business. Measured: fires on
exactly the two declared repos, silent on all sixteen English ones.
Deliberately **not** built, because the evidence does not reach: any rule about
images, diagrams, screenshots or terminal recordings (every such claim was
refuted 0-3); a README length bound (no evidence-based target exists); a section
count (would fire on 9 of 18 — textbook "suspect the CHECK"); Mermaid source
length and `#gh-dark-mode-only` (zero occurrences, and the instance's actual
limit is not readable via the API, so any threshold would be a guess).
## [0.1.3] — 2026-08-03
### Fixed