# R11 pilot results — measured, 2026-08-03 **The §10 acceptance measurement of `docs/r11-tiered-fix-design.md`, run against the live ledger. No KB file was edited and no ledger record was written.** Instrument: `scripts/kb-eval/lib/fix-op.mjs` (+ `tests/kb-eval/test-fix-op-classify.test.mjs`, 30 tests) driven by `scripts/kb-eval/classify-fix-ops.mjs`. The classifier **is** the O1 driver with writes disabled — it constructs the swap and checks the §4 invariant, so measurements 1 and 3 come out of the mechanism that would later touch the corpus, not out of a proxy heuristic. Artefact: `scripts/kb-eval/data/r11-pilot-classification.json` (untracked, regenerable; per-flag records so the run can be re-analysed without re-running). > **Two different 202s.** This population is 202 flags. §3's "202 flags whose > claim *and* quote contain a numeric token" is a different 202, measured over > the full 712-flag population. They are unrelated. --- ## 1. The headline **Nine provable, correct value swaps exist in the entire 776-flag `not_grounded` population — 1.2 %.** The machine half of the R11 tiering buys nine edits. Every other flag needs a human. | Population | Files | Flags | O1 admitted | O1 hand-verified correct | |---|---|---|---|---| | Pilot (`not_grounded` ≥ 7) | 24 | 202 | 2 | 2 | | Whole `not_grounded` corpus | 218 | 776 | 15 | **9** | This is the answer §10 asked for, and it is materially worse than the design assumed: *"If the split is materially worse than assumed, that is known after one session rather than after ten."* ## 2. §4 as written is not sufficient — measured, not argued §4 claims its invariant is *"deliberately stronger than human review at scale."* It is not. Run exactly as specified over the pilot, it admitted **6 swaps, of which 4 were wrong** — precision **2/6**: | Proposed swap | Why it is wrong | |---|---| | `30-dagers` → `24-dagers` | **Unit crossing.** The quote says 24 **hours**. | | `3000 requests/sekund` → `50` | **Metric crossing.** The quote is a *query* throttle per index; the claim is an *indexing* rate per replica. | | `Microsoft Agent 365` → `Agent 7` | **Identifier mutilated.** The `7` was harvested out of `E7`. | | `text-embedding-ada-002` → `ada-2` | **Identifier mutilated.** The `2` came from a dimensions column. | The defect is structural, not incidental. §4 constrains **where the new value came from** (verbatim in the cited quote) and **what the edit looks like** (one line, rest byte-identical). It constrains nothing about whether the two tokens **denote the same quantity**. Same-type-and-provenance is not same-referent. ### The added condition `contextCorresponds()` requires the token to sit under **the same label or the same trailing unit on both sides**. It is deliberately lexical, with **no translation table**: `dokumenter` is not taught to equal `documents`, because a synonym table introduces a new fact source and is an operator decision, not an engineering one. Consequence, measured: a swap is provable essentially only where the context is language-neutral — a URL, a code sample, a parameter key. ## 3. The condition is necessary but still not sufficient Corpus-wide the context condition cut 38 admissions to 15. Hand-verifying all 15 splits them cleanly by token type: | Token type | Proposals | Correct | Failure mode | |---|---|---|---| | `iso_date` | 9 | **9** | — every one is an `api-version=` bump in a URL or code sample | | `number` | 5 | 2 | `AI-900` → `AI-901`, `gpt-4o` → `gpt-5.1o` (×2) | | `version` | 1 | 0 | Java agent `3.7.5` → `3.4.0` — a downgrade | A matching identifier *prefix* (`AI-`, `gpt-`) satisfies the context condition while the digit is part of a **name**, not a quantity. **Only `iso_date` survives hand-verification**, and the report marks it as the sole recommended class (`o1_recommended`). `number` and `version` proposals must not be applied. ## 4. The four §10 measurements 1. **O1 / O2 / O3 split.** O1 = 2/202 on the pilot (9/776 corpus-wide, safe class only). O2 is **undetermined** — it does not exist as a class until §5 is ratified, so every non-O1 item is O3 by design. O3 ≥ 200/202. 2. **How much of R8 resolves as O2.** **Not answered, and not answerable by machine.** R8 is 87/202 on the pilot (366/776 corpus-wide) and yields **zero** O1. Of the pilot's 87, **46 are structural enumerations** — the O2 candidate shape. Which of them subtract cleanly turns on the judge's prose `reason`, and no regex reads prose. This needs prose classification (see §6). 3. **O1 abort rate: 99 % (200/202).** Typed, because "99 %" alone is not actionable: | Code | Pilot | Class | |---|---|---| | `MULTI_PART_CLAIM` | 96 (47.5 %) | intrinsic — not a value swap at all | | `MULTI_VALUE_TOKEN` | 29 | intrinsic | | `NO_VALUE_TOKEN` | 28 | intrinsic — the claim asserts prose | | `STATUS_SYNONYM` | 15 | **operator question** (§6.2) | | `NOT_VERBATIM` | 15 | intrinsic | | `LOCATOR_AMBIGUOUS` | 7 | **fixable engineering gap** | | `MULTI_REPLACEMENT` | 6 | intrinsic | | `CONTEXT_MISMATCH` | 4 | intrinsic — these are the 4 wrong edits above | **Only 7 of 200 aborts (3.5 %) are a fixable engineering gap.** More locator engineering cannot move the O1 number materially. 4. **Review throughput per class. NOT MEASURED.** It requires human review sessions, which have not happened. Recording it as measured would be false. ## 5. Two further findings **F1 — subtraction can leave a misleading remainder.** §5 argues O2 *"cannot introduce a new error, because it asserts strictly less."* True of the sentence, false of the reader's inference. Real case: *"Deep Research tool (o3-deep-research + Bing) er GA (juni 2025)"* where the source says the tool is **deprecated**. Subtracting `er GA (juni 2025)` leaves the tool standing in a list of available tools. Strictly less asserted; still misleading. O2 therefore still requires a human to look at the remainder — cheaper than O3 (no fact-finding) but not mechanical. **F2 — subtraction can destroy true information.** Real case: a list of seven prebuilt model IDs where the judge found six correct and `prebuilt-check` wrong — the real ID is `prebuilt-check.us`. Subtraction drops a model that **exists**; the correct fix is a swap. Subtraction is not the safe default everywhere. **F3 — `disposition` carries zero information.** It is `outdated` on **202 of 202** flags. `docs/r11-flag-format-2026-07.md` specifies `not_grounded → {outdated, wrong}` with *"the human assigns which at R11"*, but the pass hard-assigned `outdated`. Do not use it as a classifier signal. Spec/data divergence, recorded. **F4 — claims are not file text.** `claim` is an LLM-extracted, translated restatement: **0 of 202** match their file line verbatim, and 188 share no 40-char run with it. For table claims, `line` points at the **header**, not the value. This is why the locator exists at all, and why it searches the enclosing block rather than the line. ## 6. Open operator decisions 1. **Ratify O2 (§5)?** Until then O2 does not exist and everything is O3. F1/F2 above argue for ratifying it **with a remainder check** rather than as a blanket rule. 2. **Amend §4 with a ratified synonym table?** 15 pilot flags (54 corpus-wide) are `STATUS_SYNONYM`: the corpus writes `**Preview**` / `**GA**`, the source writes *"generally available"*. Either a small ratified equivalence table admits them, or they are permanently O3. 3. **Is O1 worth building at all?** Nine edits corpus-wide, all `api-version` bumps. The alternative is to drop the O1 driver and treat those nine as O3. ## 7. What this does not change The design's core reading survives: the expensive half (locating the source, reading it, extracting the deciding passage) was already paid for by the judge pass, and §7's freshness re-fetch per distinct URL is untouched. What the pilot falsifies is the assumption that a meaningful share of that evidence converts into machine-provable edits. It does not. R11 is a human review programme with a nine-item machine assist, and its leverage lies entirely in the O2 decision.