Article rule 1 ("give Claude judgement instead of rules") gets its register
entry, and deliberately no detector. `BP-JUDG-001` carries `lensCheck: null`.
The cut between deterministic prefilter and prose judge was the open design
decision. It was settled by measurement, and the measurement declined both
halves:
- 409 real CLAUDE.md files (38488 lines, 8689 prose blocks): the caging class
fires 7 times, and all 7 are false positives ("rendered prose-side",
"naming is a flag on the class"). Verified along an independent grep path
that bypasses block-splitting and sentence-splitting entirely, in both word
orders: 5 lines and 1 line, none an instruction.
- The narrow variant (absolute + form noun + numeric cap) fired 8 times —
one duplicated block seen seven times across plugin caches, precision 0 %.
The pre-committed rule required 90 % over 20 distinct fires.
- Where the shape does occur — 45 lines across 4755 skill/agent/command files
— it is the author's editorial policy (emoji, sentence length, slide
titles). Nothing in the text separates that from a vendor's over-tight
guardrail, and the article's reasoning does not transfer: the model is not
the author of a user's config.
So no CA-OPT-002; finding-codes keeps OPT next-free = 2. The numbers live in
the entry's own `note`, so the next session does not re-derive the question.
Two premises the chunk falsified. The brief justified a separate axis by
saying these blocks sit inside `floor-exclusion`'s floor — but the article's
own canonical line carries no floor marker at all, so "inside the floor"
cannot define the axis (the corpus tendency is 76 %, which is a tendency, not
a mechanism). And the fasit's own form-noun vocabulary was wrong: `name` and
`format` alone drove 97 % of fires.
Not folded into `--subtract`: a third "loosen instead of delete" verdict in
the subtraction judge is the AS#5 mixing STATE forbids, and with the corrected
vocabulary there are 0 collisions to arbitrate anyway.
Guards, both seen red against their own defect first: the entry must exist,
be confirmed, date its source and name NO lensCheck; and every lensCheck in
the register must be backed by a real detector.
No behaviour changed — no new finding, no output change, nothing consumes the
entry yet — hence `docs`, not `feat`. Suite 1701 -> 1703/0; frozen v5.0.0 and
default-output baselines 0 changed files. Fasit:
docs/b2-judgment-lens-fasit.local.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017mCkx9wGywqNQzsXkBMzJ1
228 lines
15 KiB
Markdown
228 lines
15 KiB
Markdown
# v5.14 Plan — Doctor Overlap, Model Routing, Effort Awareness, Dead References
|
||
|
||
> **Filename note:** kept as `v5.13-…` so existing references resolve; the target has been
|
||
> v5.14 since the `5.13.0` slot was consumed by the pipeline-hardening batch. Rewritten
|
||
> 2026-08-03 (session #53) to merge the `/doctor`-overlap decision (Oppgave A) and the
|
||
> context-engineering follow-ups (B1–B4) from `docs/v5.14-doctor-overlap-brief.md` into the
|
||
> pre-existing chunks. One plan, top-to-bottom, no relitigation.
|
||
|
||
## Inputs and their status
|
||
|
||
| Input | Status |
|
||
|---|---|
|
||
| `/doctor`-overlap measurement (session #53, fasit-first) | **DONE** — decisions binding, recorded in `docs/doctor-overlap-results.local.md` |
|
||
| B1 register freshness defect (evidence-age / supersededBy / sources[]) | **DONE** — landed as `d66035e` (TDD, 1477/1477 green) |
|
||
| Video-derived model/effort chunks (verified 2026-07-14) | Open — carried below unchanged |
|
||
| Open posts from dogfood sessions #45–#51 | Open — detail lives in `STATE.md`; referenced here by ID only |
|
||
|
||
## A. The `/doctor` verdict (measured 2026-08-03, CC 2.1.220 — binding)
|
||
|
||
`/doctor` (alias `/checkup`, v2.1.205+) is an **agent-driven, usage-data-backed, quota-priced,
|
||
non-reproducible** checkup: 10 checks incl. unused skills/plugins/MCP vs. context cost,
|
||
CLAUDE.md dedup/contradiction judgment, derivable-content trim (checked-in files only),
|
||
lazy-loading migration, hook latency, and a chars÷4 context manifest. It overlaps our
|
||
**judgment lenses and alarms — not the deterministic validators.** Full per-scanner table in
|
||
`docs/doctor-overlap-results.local.md`.
|
||
|
||
Binding outcomes (0 whole scanners removed; 2 measured function-duplicates removed;
|
||
5 surfaces repositioned):
|
||
|
||
| ID | Chunk | What |
|
||
|---|---|---|
|
||
| **D1** | GAP autoMode removal | Delete the «No autoMode classifier» gap dimension (`feature-gap-scanner.mjs:485`) — `/doctor` Sjekk 8 checks AND fixes it. Byte-stability: existing-scanner finding-removal variant of [[adding-scanner-byte-stability]]; humanizer entries for the removed title go too. |
|
||
| **D2** | SKL alarm re-scope | Remove CA-SKL-002's aggregate-over-budget **alarm** role (native startup warning v2.1.105/2.1.181 + `/doctor` Sjekk 6 both do it better-placed); re-position CA-SKL-001 as per-description **attribution** (author lens, repo scope); CA-SKL-003 untouched. Exact form (delete vs. re-scope 002) decided inside the chunk against the frozen-baseline cost. |
|
||
| **D3** | Positioning rewrite | **DONE in #53** — README section «config-audit vs. the built-in /doctor» (division-of-labor table + per-area split for SET/HKV/CNF/OPT/tokens/manifest, citing `/doctor`'s own referral to `optimize --subtract`) + CLAUDE.md invariant line. Remaining refinement (per-command copy in `commands/*.md`, if wanted) rides along with D1/D2. |
|
||
|
||
**Strategic line (constrains all copy):** our defensible identity = determinism/byte-stability,
|
||
all scopes incl. LOCAL, cross-repo campaign breadth, zero quota. `/doctor`'s unmatchable edge =
|
||
usage telemetry. Long-term (v5.15+ candidate, NOT this release): transcript/usage telemetry as a
|
||
scanner *input*, so the axes compose instead of competing.
|
||
|
||
## B. Context-engineering follow-ups (article 2026-07-24, brief §2)
|
||
|
||
- **B1 — DONE** (`d66035e`): `sources[]` + `published` + `supersededBy` + evidence-age rule
|
||
(green-but-outdated is now expressible and flagged; re-verifying the old source no longer
|
||
clears it). BP-SUB-001 carries the article as verified corroborating source.
|
||
**Deviation from brief, verified 2026-08-03:** the article contains NO mechanism-choice or
|
||
size-limit content → it was NOT added to BP-MECH-*/BP-SIZE-001 (that would be false
|
||
provenance). If a future read finds real coverage, add it then.
|
||
- **B2 — DONE, and the detector was DECLINED by measurement.** `BP-JUDG-001` ships as
|
||
register knowledge with `lensCheck: null`; **no `CA-OPT-002`** (OPT next-free stays 2).
|
||
Measured over **409 real CLAUDE.md files** (38 488 lines, 8 689 prose blocks): the caging
|
||
class fires **7 times, all 7 false positives**, confirmed along an independent grep path in
|
||
both word orders (5 lines / 1 line, none an instruction). Where the shape *does* occur — 45
|
||
lines across 4 755 skill/agent/command files — it is the author's **editorial policy** (emoji,
|
||
sentence length, slide titles), and nothing in the text separates that from a vendor's
|
||
over-tight guardrail: the article's reasoning does not transfer, because the model is not the
|
||
author of a user's config. Precision-first ⇒ silence. Two premises the chunk falsified: the
|
||
brief's «these blocks are inside the floor» (the article's own example carries **no** floor
|
||
marker; the corpus tendency is 76 %, which is not a mechanism), and the §4 form-noun
|
||
vocabulary (`name`/`format` alone were 97 % of fires). Full record:
|
||
`docs/b2-judgment-lens-fasit.local.md` §9.
|
||
- **B3 — two-layer duplication/contradiction (CNF extension):** article rule 4 + `/doctor`'s
|
||
measured Sjekk 2 catch (global CLAUDE.md model-policy ↔ agent frontmatter) prove the class
|
||
exists and is catchable. Extend `conflict-detector` with cross-layer checks where the pair is
|
||
structurable (CLAUDE.md/rule, rule/skill-description, CLAUDE.md-keyword ↔ frontmatter field).
|
||
This **supersedes the old "Explicitly rejected #5"** below — it now has evidence.
|
||
- **B4 — thinner gaps (backlog, after everything above):** rule 2 (skills/agents leaning on
|
||
examples where a parameter enum would do) and rule 6 (reference *form*: prefer in-code files;
|
||
`import-resolver` is the natural owner). Park until D/B2/B3 land.
|
||
- **Rules 3 and 5 need nothing** (verified in brief §B4): progressive disclosure = BP-LOAD-001..006
|
||
+ BP-MECH-003 + `token-hotspots`/`manifest`; router pattern + auto-memory covered.
|
||
|
||
## C. Carried chunks (verified 2026-07-14 — unchanged specs)
|
||
|
||
Source-verification table, rejected-claims list, and full chunk specs below are carried
|
||
verbatim from the 2026-07-14 revision; only numbering context changed.
|
||
|
||
### C1 — Register entries: model routing + effort (dogfoods `knowledge-refresh`)
|
||
|
||
- **BP-MODEL-001** (`model-fit`): mechanical/read-only subagents can pin a cheaper model via
|
||
`model:` frontmatter; orchestrator keeps the strong model. Source: code.claude.com/docs/en/sub-agents
|
||
→ `confirmed`.
|
||
- **BP-MODEL-002** (`model-fit`): reasoning effort tunable at five levels in five places;
|
||
default `high`; higher is not universally better. Source: code.claude.com/docs/en/model-config
|
||
→ `confirmed`.
|
||
- Schema per `scanners/lib/best-practices-register.mjs:42-102`. New sources SHOULD carry
|
||
`published` so the B1 evidence-age rule has teeth.
|
||
|
||
### C2 — fix-engine effort hygiene (tiny, TDD)
|
||
|
||
`fix-engine.mjs:26` `VALID_EFFORT_LEVELS` missing `xhigh` → nearest-match "fix" for `xhig`
|
||
corrects to `high`. Red test first: `findNearestEffortLevel('xhig') === 'xhigh'`. Align with
|
||
`settings-validator.mjs:75`.
|
||
|
||
### C3 — CA-CML dead prose references (new deterministic check)
|
||
|
||
Flag backtick-quoted relative file paths in CLAUDE.md prose that do not exist on disk (today
|
||
only `@import` targets are checked). Conservative v1: skip URLs, globs, placeholders,
|
||
absolute/`~/` paths. Severity low. New CA-CML-NNN (verify next free NNN at implementation).
|
||
Byte-stability per [[adding-scanner-byte-stability]] incl. humanizer step 7.
|
||
|
||
### C4 — feature-gap + inventory: model/effort awareness
|
||
|
||
New T3 opportunity check: authored agents where NO agent sets `model:`/`effort:` → routing
|
||
opportunity citing BP-MODEL-001/002. Fires only when authored agents exist; opportunity
|
||
framing, suppressable. `whats-active`/`manifest` surface `model`/`effort` per agent.
|
||
Humanizer step 7; verify via direct `scan()` ([[agent-commands-need-scanner-scoping]]).
|
||
Known tension with the operator's own Opus-for-everything policy stands as written 2026-07-14:
|
||
the check serves general users; on this machine it gets suppressed.
|
||
|
||
### C5 — planner-agent adversarial gate (AFTER DEL B 3.2 dogfood)
|
||
|
||
Required "Failure modes" section in `agents/planner-agent.md`'s action-plan contract.
|
||
Sequencing: only after the DEL B fasit pass that judges planner-agent, or the fasit target
|
||
moves mid-evaluation.
|
||
|
||
## D. Open posts from dogfooding (detail in STATE.md — not restated here)
|
||
|
||
C-SKL1 (#37) · M-BUG-26 · **M-BUG-28** (suppression-ID positional instability — ID-semantics
|
||
change touching all scanners + frozen snapshots, own chunk) · **M-BUG-41** (no scope-gate from
|
||
scan to write, two arms — design change, own chunk) · **arg-sluk CLI arm**
|
||
(`optimize-lens-cli` + `token-hotspots-cli`, `KNOWN_OPEN` in
|
||
`tests/scanners/cli-unknown-flag-rejection.test.mjs`) · **P6/M-BUG-44** (scanner-side stdout
|
||
with `--output-file`) · knowledge-refresh write-CLI · cleanup-invisible session files ·
|
||
web-poll candidates (nothing written; primary sources unread).
|
||
|
||
## Priority order for v5.14
|
||
|
||
Cheap-and-loud first, judgment-heavy later; D-chunks early because they DELETE code the rest
|
||
must not build on. **Re-ordered 2026-08-10 (operator decision, session #61):** the operator wants
|
||
to *use* the subtraction axis, so its write half — and the scope-gate it depends on — move ahead
|
||
of the remaining additive work. Everything below step 4 is unchanged in content, only in position.
|
||
|
||
1. ~~**C2**~~ ✅ · 2. ~~**Arg-sluk CLI arm**~~ ✅ · 3. ~~**D1 + D2**~~ ✅ · 4. **D3** (rest dropped,
|
||
`stop-at-meaningful-value`) · 5. ~~**M-BUG-28**~~ ✅ · 6. ~~**C1**~~ ✅ · 7a. ~~**C4**~~ ✅
|
||
8. **M-BUG-41** (scope-gate design) — **promoted from 10.** Prerequisite for anything that writes
|
||
outside the repo the session stands in, which subtraction-write does by definition
|
||
(`~/.claude/CLAUDE.md`).
|
||
9. **SUB-WRITE** (new) — the write half of `optimize --subtract`; see §C6 below.
|
||
10. **C3** (CA-CML dead prose references) — was 7b.
|
||
11. **B2** (new lens axis)
|
||
12. **B3** (CNF two-layer extension)
|
||
13. **P6/M-BUG-44**, knowledge-refresh write-CLI, cleanup glob (small batch)
|
||
14. **C5** (after DEL B 3.2), **B4** backlog last
|
||
15. Release-cut via `release-plugin.mjs` when the batch is coherent. **Level is MAJOR — v6.0.0:**
|
||
M-BUG-28 shipped as `fix(scanners)!` with a `BREAKING CHANGE:` footer, which outranks the
|
||
minor the D1/D2 removals plus C3/C4 additions would have implied. `release-plugin.mjs` does
|
||
not derive the level — pass `--version` explicitly.
|
||
|
||
### C6 — SUB-WRITE: the write half of `optimize --subtract`
|
||
|
||
`--subtract` proposes and never writes (`commands/optimize.md`), which is correct for a v1 whose
|
||
judge is an agent. The operator now wants the removal executed. Scope: apply an approved
|
||
subtraction candidate to the CLAUDE.md it came from, with backup and rollback.
|
||
|
||
Non-negotiable frames, all inherited rather than invented here:
|
||
|
||
- **The floor is not the judge's decision.** `floor-exclusion.mjs` runs deterministically before
|
||
anything is proposed, and that ordering must not migrate into the write path either.
|
||
- **`~/.claude` is git-tracked with a `.gitignore` of `*`** — archive by `mv` into `_archive/`,
|
||
never `rm`. Machine-side config writes need operator approval.
|
||
- **User level is mandatory in v1**: that is where the cost is (~4 300 tokens every turn in every
|
||
repo). Project level follows.
|
||
- **Honest sizing:** the #40 fasit measured deletable ≈1 400 always-loaded tokens, realistically
|
||
≈850 after tier-2 earn-backs, against a ≈4 300-token file — **≈20 %, not 80 %.** Do not let the
|
||
command's copy imply more.
|
||
- Verify the backup covers the file the write actually touched, not merely that a backup exists
|
||
(M-BUG-31's shape).
|
||
|
||
Open decision, to be settled in the chunk's fasit before code: whether removal is a `fix`-engine
|
||
action, a `plan`/`implement` step, or its own flag — decided against M-BUG-41's gate, not before it.
|
||
|
||
**Rejected 2026-08-10, do not revive:** a sibling `/repo-reinit` skill that rewrites a CLAUDE.md
|
||
from scratch. It would be a third implementation of one judgement (this axis, plus `/doctor`
|
||
Check 3) and fails the binding `/doctor` positioning; and regenerating destroys exactly the floor
|
||
— local facts, gotchas, policy invariants — that a mature repo's CLAUDE.md is most valuable for.
|
||
`repo-init` already owns the fresh-repo case.
|
||
|
||
## Source verification (done 2026-07-14 — carried)
|
||
|
||
| Claim from video | Verdict | Source |
|
||
|---|---|---|
|
||
| Orchestrator + cheaper worker models is supported/recommended | VERIFIED | code.claude.com/docs/en/sub-agents, /workflows |
|
||
| Effort tunable per settings/session/launch/agent-frontmatter/SDK; `low..max` | VERIFIED | code.claude.com/docs/en/model-config#adjust-effort-level |
|
||
| Leaked Fable 5 system-prompt principles | VERIFIED near-verbatim, provenance unconfirmed | github.com/asgeirtj/system_prompts_leaks |
|
||
| "Fable low ≈ Opus high" chart | **CONTRADICTED** | anthropic.com/news/claude-fable-5-mythos-5 |
|
||
|
||
## Explicitly rejected (unchanged unless noted)
|
||
|
||
1. "Fable low ≈ Opus high" framing — never encode.
|
||
2. Tool-call-count effort scaling as register entry — unconfirmed leak, not carried.
|
||
3. Cost/intelligence/"taste" routing-table generator — subjective, doesn't fit provenance-gated design.
|
||
4. "Fable mode" skill — out of plugin scope.
|
||
5. ~~CLAUDE.md prose contradiction detection~~ — **superseded by B3** (2026-08-03: article
|
||
rule 4 + `/doctor` Sjekk 2 measurement supplied the evidence the 2026-07-14 rejection lacked).
|
||
|
||
## Verification (per chunk, unchanged discipline)
|
||
|
||
- Full suite green (`node --test 'tests/**/*.test.mjs'`; baseline 2026-08-03: 1477/0), frozen
|
||
`tests/snapshots/v5.0.0/` untouched (`git status --porcelain` empty), red test before every
|
||
production change.
|
||
- **D1:** GAP fixture with autoMode absent → no finding; humanizer has no orphaned entries
|
||
(M-16/M-17 checks reversed for removal); frozen baselines untouched or consciously re-seeded
|
||
per [[adding-scanner-byte-stability]].
|
||
- **D2:** over-budget fixture → no 002-alarm (or re-scoped payload per in-chunk decision);
|
||
001 fires per oversized description with attribution copy; 003 unchanged.
|
||
- **D3:** README/CLAUDE.md name `/doctor` explicitly; `self-audit --check-readme` PASS.
|
||
- **C1–C5:** criteria as specified 2026-07-14 (C2 red-first nearest-match; C3 fixture
|
||
missing-path fires / URL-glob-placeholder silent; C4 authored-agent matrix; C5 failure-modes
|
||
section present).
|
||
- **B2 (superseded by the measurement above):** the criterion was written for a detector that
|
||
measurement declined. What was verified instead: `BP-JUDG-001` present, `confirmed`, primary
|
||
source dated `2026-07-24`, `lensCheck` **absent** — plus a guard that every `lensCheck` in the
|
||
register is backed by a real detector, so no later session can "complete" the entry by wiring
|
||
one. Both assertions seen RED against their own defect before landing.
|
||
- **B3:** fixture with same instruction in CLAUDE.md + rule → CNF finding; single-layer only →
|
||
silent ([[guard-can-be-green-on-its-own-defect]]: assert the blanket invariant).
|
||
|
||
## Key assumptions (test at implementation)
|
||
|
||
1. Per-agent `effort` frontmatter still official — re-fetch sub-agents + model-config pages.
|
||
2. Finding-type REMOVAL in an existing scanner leaves frozen v5.0.0 snapshots untouched only if
|
||
no frozen fixture carries the type — verify per D1/D2 before committing; re-seed consciously
|
||
if not.
|
||
3. Next free CA-CML/CA-GAP/CA-OPT NNN — grep tests + snapshots before assigning.
|
||
4. `/doctor`'s check set is version-fluid (2.1.205→220 changed it materially) — re-run the
|
||
overlap measurement cheaply (CLI + one in-session run) before executing D1–D3 if CC has
|
||
moved significantly past 2.1.220.
|