ms-ai-architect/docs/r11-pilot-results.md
Kjell Tore Guttormsen 41f769a166 docs(ms-ai-architect): footer-klassen handklassifisert — 58 medlemmer, en fjerde unnslippsform, og «alle avvik samme vei» falt
§9.16. Leveransen er en maaling og en scope-anbefaling for aapent spoersmaal
#17; ingen korpusfil er roert.

«>=48» var riktig som gulv. Handklassifisert: 58 ekte medlemmer i 58 filer.
Bøtter: 30 konsistent / 11 inkonsistent / 1 tvetydig / 16 ikke sjekkbar.
Hvert anker verifisert ordrett og unikt (0 avvik).

FJERDE UNNSLIPPSFORM: tellingen baaret av handlingsordet («3 soek»,
«4 search queries», «2 deep fetch»), ikke av «kall». Mitt eget nett var blindt
for den — domain-specific-prompt-optimization.md:589 ble kun reddet av
nabolinja, og model-versioning-registry-management.md var usynlig i sin helhet
for baade MCP- og verktoeynavn-nettet.

COMPLETENESS: i stedet for flere ordgjetninger ble hele etikettvokabularet i
proveniensblokkene enumerert — 68 distinkte etiketter / 141 linjer, alle
plasserbare i kall-telling, kilde-telling eller annet. Ingen femte form.
Den kjente residualen (alle nett krever et SIFFER) ble testet, ikke antatt:
21 tallord-treff, alle broedtekst, null medlemmer.

RETNINGSARGUMENTET RE-MAALT: 10 av 11 inkonsistente gaar oppgitt < sum, 1 gaar
motsatt (chain-of-thought-prompting.md, oppgitt 4 over enumerert 3). «Alle
avvik gaar samme vei» er falsifisert som absolutt og maa slutte aa siteres
slik; tendensen (91 %) overlever. «Null motsatt» var et artefakt av at
populasjonen var maalt til under halv stoerrelse.

FORMEN DEKKER IKKE KLASSEN: 46 av 58 er linjeform, 12 er blokkform der «slett
linja» er udefinert — og én blokk blander naboklassen (Unique sources) inn i
samme punktliste som kall-linjene.

ANBEFALING #17: én klasse-entry for de 42 uten aapen beslutning, 16 egne
entries for dem som baerer én. Beslutning #14 («én entry per fil») ble tatt da
klassen ble antatt aa vaere 14 filer; 42 nesten-identiske entries er nettopp
mekanismen §9.15 viste at lot en feilmaaling gjemme seg bak sine egne kopier.

Datasett: scripts/kb-eval/data/r11-footer-class-2026-08-11.json
Suite 1047/1047.
2026-08-11 21:52:19 +02:00

2098 lines
138 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# R11 pilot results — measured, 2026-08-03
**The §10 acceptance measurement of `docs/r11-tiered-fix-design.md`, run against
the live ledger. No KB file was edited and no ledger record was written.**
Instrument: `scripts/kb-eval/lib/fix-op.mjs` (+ `tests/kb-eval/test-fix-op-classify.test.mjs`,
49 tests after §4b) driven by `scripts/kb-eval/classify-fix-ops.mjs`. The classifier **is**
the O1 driver with writes disabled — it constructs the swap and checks the §4
invariant, so measurements 1 and 3 come out of the mechanism that would later
touch the corpus, not out of a proxy heuristic.
Artefact: `scripts/kb-eval/data/r11-pilot-classification.json` (untracked,
regenerable; per-flag records so the run can be re-analysed without re-running).
It holds the **pilot** run — `node scripts/kb-eval/classify-fix-ops.mjs --write`.
Every corpus-wide figure below is from `--threshold 1`, and the per-table
reproduce command is stated where it is used.
> **Two different 202s.** This population is 202 flags. §3's "202 flags whose
> claim *and* quote contain a numeric token" is a different 202, measured over
> the full 712-flag population. They are unrelated.
---
## 1. The headline
**Nine provable, correct value swaps exist in the entire 776-flag `not_grounded`
population — 1.2 %.** The machine half of the R11 tiering buys nine edits. Every
other flag needs a human.
| Population | Files | Flags | O1 admitted | O1 hand-verified correct |
|---|---|---|---|---|
| Pilot (`not_grounded` ≥ 7) | 24 | 202 | 2 | 2 |
| Whole `not_grounded` corpus | 218 | 776 | 15 | **9** |
> These are **numeric-path** figures, and they stay that way. §4b (the status
> synonym table) was implemented afterwards and adds a separate class with its own
> hand-verification — see §8. A run today prints O1 = 7 (pilot) and 23 (corpus)
> because the status proposals are included in the total; the numeric line above
> is unchanged and is still what `s4_as_written` compares against.
Reproduce: `node scripts/kb-eval/classify-fix-ops.mjs` (pilot) and
`node scripts/kb-eval/classify-fix-ops.mjs --threshold 1` (corpus). The nine are
enumerated with verdicts in appendix A — that hand-verification is the only thing
separating 9 from 15, so it is recorded rather than left in a session transcript.
This is the answer §10 asked for, and it is materially worse than the design
assumed: *"If the split is materially worse than assumed, that is known after one
session rather than after ten."*
> **The O2 half, measured afterwards (§9), does not rescue the number.** Of the 46
> R8 multi-part claims that have the O2 shape, 17 are candidates and **2 are clean
> subtractions on the available evidence**; all 29 non-candidates fail because the
> source supplies a *corrected value*, which makes them swaps or rewrites. O2's
> value in R11 is triage — it tells a human which 17 to look at first — not
> automation.
## 2. §4 as written is not sufficient — measured, not argued
§4 claims its invariant is *"deliberately stronger than human review at scale."*
It is not. Run exactly as specified over the pilot, it admitted **6 swaps, of
which 4 were wrong** — precision **2/6**:
| Proposed swap | Why it is wrong |
|---|---|
| `30-dagers``24-dagers` | **Unit crossing.** The quote says 24 **hours**. |
| `3000 requests/sekund``50` | **Metric crossing.** The quote is a *query* throttle per index; the claim is an *indexing* rate per replica. |
| `Microsoft Agent 365``Agent 7` | **Identifier mutilated.** The `7` was harvested out of `E7`. |
| `text-embedding-ada-002``ada-2` | **Identifier mutilated.** The `2` came from a dimensions column. |
The defect is structural, not incidental. §4 constrains **where the new value
came from** (verbatim in the cited quote) and **what the edit looks like** (one
line, rest byte-identical). It constrains nothing about whether the two tokens
**denote the same quantity**. Same-type-and-provenance is not same-referent.
### The added condition
`contextCorresponds()` requires the token to sit under **the same label or the
same trailing unit on both sides**. It is deliberately lexical, with **no
translation table**: `dokumenter` is not taught to equal `documents`, because a
synonym table introduces a new fact source and is an operator decision, not an
engineering one. Consequence, measured: a swap is provable essentially only where
the context is language-neutral — a URL, a code sample, a parameter key.
## 3. The condition is necessary but still not sufficient
Corpus-wide the context condition cut 38 admissions to 15. Hand-verifying all 15
splits them cleanly by token type:
Reproduce: `node scripts/kb-eval/classify-fix-ops.mjs --threshold 1`
(the persisted artefact is the **pilot** run — the corpus-wide tables in §1 and §3
come from this threshold-1 run). Per-proposal verdicts: appendix A.
| Token type | Proposals | Correct | Wrong | Unverified | Failure mode |
|---|---|---|---|---|---|
| `iso_date` | 9 | **9** | 0 | 0 | — every one is an `api-version=` bump in a URL or code sample |
| `number` | 5 | 0 | 3 | 2 | `AI-900``AI-901`, `gpt-4o``gpt-5.1o` (×2) |
| `version` | 1 | 0 | 1 | 0 | Java agent `3.7.5``3.4.0` — a downgrade |
A matching identifier *prefix* (`AI-`, `gpt-`) satisfies the context condition
while the digit is part of a **name**, not a quantity. **Only `iso_date` survives
hand-verification**, and the report marks it as the sole recommended class
(`o1_recommended`). `number` and `version` proposals must not be applied.
## 4. The four §10 measurements
1. **O1 / O2 / O3 split.** O1 = 2/202 on the pilot (9/776 corpus-wide, safe
class only). ~~O2 is **undetermined** — it does not exist as a class until §5 is
ratified, so every non-O1 item is O3 by design.~~ §5 was ratified 2026-08-03,
and O2 has since been measured over the class where it can exist at all — see
measurement 2 and §9. It remains **unmeasured outside R8 ∧
`MULTI_PART_CLAIM`**; the classifier still routes every non-O1 item to O3, so
O3 ≥ 200/202 stands as the machine's own partition.
2. **How much of R8 resolves as O2.** **MEASURED 2026-08-03 — see §9.** R8 is
87/202 on the pilot (366/776 corpus-wide) and yields **zero** O1. Of the
pilot's 87, **46 are structural enumerations** (R8 ∧ `MULTI_PART_CLAIM`) — the
O2 candidate shape. Which of them subtract cleanly turns on the judge's prose
`reason`, and no regex reads prose, so this was done by prose classification:
**17 of 46 (37 %) are O2 candidates, 29 are O3.** All 29 are foreclosed by
condition 3.
3. **O1 abort rate: 99 % (200/202).** Typed, because "99 %" alone is not
actionable:
| Code | Pilot | Class |
|---|---|---|
| `MULTI_PART_CLAIM` | 96 (47.5 %) | intrinsic — not a value swap at all |
| `MULTI_VALUE_TOKEN` | 29 | intrinsic |
| `NO_VALUE_TOKEN` | 28 | intrinsic — the claim asserts prose |
| `STATUS_SYNONYM` | 15 | **operator question** (§6.2) |
| `NOT_VERBATIM` | 15 | intrinsic |
| `LOCATOR_AMBIGUOUS` | 7 | **fixable engineering gap** |
| `MULTI_REPLACEMENT` | 6 | intrinsic |
| `CONTEXT_MISMATCH` | 4 | intrinsic — these are the 4 wrong edits above |
**Only 7 of 200 aborts (3.5 %) are a fixable engineering gap.** More locator
engineering cannot move the O1 number materially.
4. **Review throughput per class. NOT MEASURED.** It requires human review
sessions, which have not happened. Recording it as measured would be false.
## 5. Two further findings
**F1 — subtraction can leave a misleading remainder.** §5 argues O2 *"cannot
introduce a new error, because it asserts strictly less."* True of the sentence,
false of the reader's inference. Real case: *"Deep Research tool (o3-deep-research
+ Bing) er GA (juni 2025)"* where the source says the tool is **deprecated**.
Subtracting `er GA (juni 2025)` leaves the tool standing in a list of available
tools. Strictly less asserted; still misleading. O2 therefore still requires a
human to look at the remainder — cheaper than O3 (no fact-finding) but not
mechanical.
**F2 — subtraction can destroy true information.** Real case: a list of seven
prebuilt model IDs where the judge found six correct and `prebuilt-check` wrong —
the real ID is `prebuilt-check.us`. Subtraction drops a model that **exists**; the
correct fix is a swap. Subtraction is not the safe default everywhere.
**F3 — `disposition` carries zero information.** It is `outdated` on **202 of
202** flags. `docs/r11-flag-format-2026-07.md` specifies `not_grounded →
{outdated, wrong}` with *"the human assigns which at R11"*, but the pass
hard-assigned `outdated`. Do not use it as a classifier signal. Spec/data
divergence, recorded.
**F4 — claims are not file text.** `claim` is an LLM-extracted, translated
restatement: **0 of 202** match their file line verbatim, and 188 share no 40-char
run with it. For table claims, `line` points at the **header**, not the value.
This is why the locator exists at all, and why it searches the enclosing block
rather than the line.
## 6. Operator decisions — ALL THREE RATIFIED 2026-08-03
All three were put to the operator with the recommendations below and **all three
were accepted as recommended**. The contract text now lives in
`docs/r11-tiered-fix-design.md` §4a/§4b/§5; this section records what was asked
and what the answer was.
**None of the three is implemented yet.** The classifier still aborts
`STATUS_SYNONYM` and still routes every non-O1 item to O3. A later session builds
against the ratified contract — it must not assume the code already honours it.
1. **Ratify O2 (§5)?****RATIFIED, with the remainder check** (not as a blanket
rule), exactly as F1/F2 above argued. Contract: design doc §5, three
conditions, human-confirmed.
2. **Amend §4 with a ratified synonym table?****RATIFIED, narrow and closed.**
Contract: design doc §4b — four label rows, closed table, complete-label-only,
corpus-side value with the file's own markup preserved. Unlocks up to 54
corpus-wide flags.
3. **Is O1 worth building at all?****KEPT, locked to `iso_date`.** Contract:
design doc §4a condition 5 — a driver may apply `iso_date` proposals and must
never apply `number` or `version` ones. Nine edits corpus-wide.
## 7. What this does not change
The design's core reading survives: the expensive half (locating the source,
reading it, extracting the deciding passage) was already paid for by the judge
pass, and §7's freshness re-fetch per distinct URL is untouched. What the pilot
falsifies is the assumption that a meaningful share of that evidence converts into
machine-provable edits. It does not. R11 is a human review programme with a
nine-item machine assist, and its leverage lies entirely in the O2 decision.
---
## 8. §4b implemented — the status class measured, 2026-08-03
The ratified synonym table (`docs/r11-tiered-fix-design.md` §4b) is implemented in
`fix-op.mjs` and the class is measured. Reproduce:
`node scripts/kb-eval/classify-fix-ops.mjs --threshold 1` (`status_synonym` block).
| Population | STATUS_SYNONYM flags | Proven by §4b | Still aborting |
|---|---|---|---|
| Pilot (`not_grounded` ≥ 7) | 15 | **5** | 10 |
| Whole `not_grounded` corpus | 54 | **8** | 46 |
Why the other 46 abort, corpus-wide — this sub-distribution is the actionable
part, because the top-level `STATUS_SYNONYM` count alone says nothing:
| Reason | N | What it means |
|---|---|---|
| `NO_COMPLETE_FILE_LABEL` | 27 | the file writes the status inside a sentence — `(preview)` in a list item, `**DSPM (preview):**`, `"[Preview]: …"` in a JSON string. Constraint 2 refuses these, correctly. |
| `NO_SOURCE_STATUS` | 15 | the cited quote carries no listed lifecycle phrasing at all — the flag was never a status swap. |
| `SOURCE_STATUS_AMBIGUOUS` | 2 | the quote asserts two different rows (e.g. "…is now generally available. Partner solutions remain in preview."). |
| `FILE_ALREADY_MATCHES` | 2 | file and source agree; the mismatch was in the LLM-extracted claim, not in the corpus. |
**The class is REVIEW-grade, not apply-grade — 5 of 8 correct.** All eight were
hand-judged against the cited source (appendix B). Three defects, all one family:
§4b binds the table, the completeness of the file label and the written value, and
**nothing about whether the source phrasing refers to the row's own subject**.
That is the same provenance-without-referent defect that falsified §4 (§2), now
reproduced in the status class. `status` is therefore deliberately **absent from
`o1_recommended`**: the machine writes nothing, and every proposal reaches a human.
### Two candidate conditions, costed over the eight
Neither is implemented — extending a table the operator ratified as *closed* is an
operator decision, exactly as condition 5 was in §4a. Both are pure gain on this
population (they kill wrong proposals and no correct one), which is the number the
decision needs:
| Candidate | Kills | Correct proposals lost |
|---|---|---|
| **A** — count a bare `GA` in the quote as a GA-row source phrasing, so a quote saying both `GA` and "public preview" becomes ambiguous | 1 (`onelake:198`) | 0 |
| **B** — abort when the quote is a multi-entity enumeration (≥ 2 pipes, or a numbered list) | 2 (`onelake:198`, `owasp:79`) | 0 |
B subsumes A on these eight. Neither catches `security-copilot-integration.md:94`,
where the quote is prose and the `(Preview)` marker simply belongs to a different
agent. A stricter **referent-name** condition (the row's subject must appear in the
quote) would catch it — and would also kill two *correct* proposals (`:83`, `:93`),
where the source names the capability rather than the agent. That trade is real and
is why this is put to the operator rather than shipped.
## 9. §10 measurement #2 — how much of R8 resolves as O2, measured 2026-08-03
The last machine-answerable pilot measurement. §4.2 recorded it as *"not
answered, and not answerable by machine"* — true of a regex, not of prose
classification, which is what this ran. Reproduce the verification and the
tally: `node scripts/kb-eval/check-o2-returns.mjs`.
**Population.** The 46 pilot flags that are R8 ∧ `MULTI_PART_CLAIM` — the O2
candidate shape, re-derived from the ledger, not read from a plan
(`classify-fix-ops.mjs --threshold 7`). 17 distinct files.
**Method.** Eight subagents, six items each, classifying against §5's three
conditions. Constraints, all deliberate: read-only (no writes, no commits); **no
web or MCP lookups** — O2 is *defined* by requiring no new fact-finding, so the
`evidence_quote` is the only source evidence a classifier may use; and every
proposal must be obtainable from the file text by **deleting characters only**.
Because `claim` matches its file line verbatim in 0 of 202 cases (F4), each
classifier had to open the actual file and locate the real text rather than edit
the restatement. All 46 located it; `locator_failed` is 0.
| Verdict | N | Share |
|---|---|---|
| **O2 candidate** | **17** | 37 % |
| O3 | 29 | 63 % |
**The result that matters is not the split — it is what blocks the other 29.**
| Blocking condition(s) | N |
|---|---|
| condition 2 + condition 3 | 21 |
| condition 3 alone | 3 |
| all three | 5 |
**All 29 are foreclosed by condition 3: the source supplies a *corrected value*,
so the fix is a swap or a rewrite and subtraction would destroy true
information.** Condition 1 — "asserts strictly less", the one that sounds like
the hard one — blocks only 5, and never alone. This is F2 (§5) reproduced at
scale: `prebuilt-check.us`, `prebuilt-mortgage.us.closingDisclosure`,
`Set-DlpCompliancePolicy`, `jensen_shannon_distance`, F300 = 384 GB, `DurationMs`
/ `ResultSignature`, Claude 4.5 → 4.6. **R8's failing multi-part claims are
predominantly a wrong-value class, not a surplus-specificity class.** The design's
reading of R8 in §3 — *"where the grounded part stands on its own, the fix is O2"*
— holds for a minority of the class.
**Machine verification of the returns (V1/V2/V2b, §9.1).** 46 of 46 pass V1: the
quoted file text occurs verbatim in the named file, æ/ø/å and markup intact.
16 of the 17 O2 proposals are deletion-only; one is flagged
(`feedback-loops-continuous-improvement.md:555`, where `Automatically add`
`Add` recapitalises rather than merely deletes). That is a text change, not a
subtraction, and it goes to a human as such.
**The 17 are candidates, not admitted edits — and the split inside them is the
honest number:**
| | N |
|---|---|
| conditions 2 **and** 3 both affirmatively `yes` | **2** (idx 8, 17) |
| at least one condition marked `human_must_confirm` | 15 |
| classifier confidence `high` | 1 |
So the machine's own reading is that **2 of 46 (4 %) are clean subtractions on
the evidence available, and 15 more are worth a human's time.** This is a triage,
not a machine assist. It is **not** measurement #4 — review throughput still
requires human review sessions that have not happened (§4.4) — but it is the
input #4 needs: it says how many items enter review and in what state, which is
the half of throughput that does not require a stopwatch. The recurring reason for `human_must_confirm` on condition 3 is
structural and worth recording: the removed material is often **true of something
else** (Purview really does classify data; the Communication Compliance template
really exists; Redis really is in Norway West) — it is merely false *of the
subject the row names*. Deleting it is defensible; relocating it may be better.
That is a judgement about the corpus, not about the source, which is exactly why
§5 put conditions 2 and 3 in human hands.
**Extrapolation, flagged as such.** `MULTI_PART_CLAIM` is 161 corpus-wide under
R8. At the pilot's 37 % that is ~60 O2 candidates and ~6 clean ones. **This is an
extrapolation from one measured sample, not a measurement**, and the pilot was
deliberately drawn from the densest files.
### 9.1 What the machine checks, and what it deliberately does not
`scripts/kb-eval/lib/o2-return-check.mjs` (25 tests). The checks do **not** decide
O2 — conditions 2 and 3 stay human by ratified contract. They bound the two
failure modes a human reviewing 46 proposals cannot catch cheaply:
- **V1** — the quoted file text must occur verbatim in the file. Catches invented
text and silent æ/ø/å transliteration. Applied to **every** row, not just the
O2 ones: an O3 verdict resting on invented text is equally wrong, merely wrong
in the safe direction.
- **V2** — the remainder must be obtainable by deleting characters only.
- **V2b** — word-level and case-sensitive, because V2 alone is too weak: deleting
a leading word and recapitalising the next passes the character test, since the
capital already existed inside the deleted word. This check was added *after*
wave 1 produced exactly that case.
- **V3** — schema completeness and verdict/condition coherence.
The raw returns are committed under `scripts/kb-eval/data/r11-o2-returns/`
they are evidence, not regenerable output, same discipline as appendices A and B.
### 9.2 The two affirmative candidates, hand-verified — one of them fails
Before putting anything in front of a human ratifier, the two candidates the
classifier marked affirmative on **both** human conditions (idx 8, 17) were
checked by hand, 2026-08-03. The check was not a re-reading of the source: it
was a check of the classifier's *own* cond-2 reasoning, which cites other lines
in the same file as its justification. V1/V2/V2b never touch those citations —
they bound the quoted block and the remainder string, nothing else.
**Anchoring, verified for all 17.** Each `file_text_verbatim` occurs exactly
**once** in its file (17/17), so text-anchored editing is unambiguous. The line
numbers are not: `line` differs from `real_line` in **9 of 17** records. Any
edit must be anchored on the verbatim block, never on the line number.
**idx 17 — holds.** Both cross-references check out: line 164 does carry
`Start med 0.15`, and the policy sample does use `score-threshold="0.15"`
at line **238**, not 239 as the classifier wrote (an off-by-one in the citation;
the substance stands). The three bands appear nowhere else in the file
(`grep` for the band values and their labels returns only lines 254-256, the
block itself), so deleting them leaves nothing dangling internally. Two cosmetic
residues for the ratifier, both already visible in the record: the heading keeps
a now-trailing colon, and the `**Verified** (Microsoft Learn - Enable semantic
caching for LLM APIs)` stamp on line 258 afterwards stamps only the direction
statement. Neither makes the remainder false.
**idx 8 — not ratifiable as written, under either reading of cond 2 — but the
reason and the remedy differ, and choosing between them is an operator call.**
First, the fork, because everything below depends on it. Cond 2 says the
remainder must not be misleading. **Its scope was never fixed:**
- **Broad reading** — misleading *to a reader of the file*. Then the rest of the
file is in scope, and a remainder that contradicts a passage seventy lines
down fails.
- **Narrow reading** — misleading *as a statement of what the source grounds*.
Then only the edited passage is in scope, and a contradiction elsewhere in the
file is a **separate ungrounded claim**, to be flagged on its own, not a
defeater of this subtraction.
**The operator ratified the broad reading, 2026-08-03.** Cond 2 is measured
against the whole file: a remainder that contradicts the file it sits in is
misleading, whatever the source says. The stated ground is that these files are
publicly distributed and read as wholes — a self-contradicting file is a trust
defect regardless of which half is wrong. The narrow reading (cond 2 scoped to
the edited passage, with contradictions elsewhere handled as separate ungrounded
claims) was considered and rejected. Recorded here because the fork was real and
a later run must not silently re-open it.
The classifier justified cond 2 on one of the two sub-deletions and never
checked the other:
- Sub-deletion (b), the `Optional: Application Insights` bullet: **verified.**
Lines 202-203 do cover Application Insights on their own terms
(`**Azure Monitor + Application Insights** (Verified)` / "Drift metrics
emitteres til Application Insights"), so removing the prerequisite bullet
leaves no false implication.
- Sub-deletion (a), `eller managed compute cluster`: **fails.** The same file
documents
`**Managed Compute Cluster** (for store volumer)` as a real compute option at
lines 263-265, with pricing and a usage recommendation. Deleting the
alternative from the prerequisites leaves the file asserting serverless Spark
as the only compute requirement seventy lines above a cost section that prices
the alternative. That is a misleading remainder under the ratified reading.
The candidate therefore does not go in as written: **(a) must be dropped from
the subtraction**, leaving (b) alone.
There is a further limit the O2 envelope cannot resolve: cond 3 passes for (a)
only because the source does not *mention* managed compute — not-mentioned
passes cond 3 by construction. Whether Azure ML actually permits managed compute
for monitoring is a fact question, and O2 is defined by zero fact-gathering. It
is an operator call, not a gap to be read harder.
A reduced subtraction — sub-deletion (b) alone — would still be deletion-only
and is not defeated by anything found here. But it is a **different remainder
string** than the one V2b attested, so it is not machine-clean until
`check-o2-returns.mjs` is re-run against it. The same rule governs any operator
amendment, including dropping idx 17's trailing colon: **amended remainder →
re-run the check before it counts as verified.**
**The generalisable finding.** The classifier judged cond 2 against the *source*
and against citations it chose itself. It did not systematically judge it against
**the rest of the same file**. Under the ratified reading that is a defect:
internal consistency is a cond-2 dimension the wave prompts never assigned, and
idx 8 — the single `high` confidence record in the set — is the proof that it
bites. The remaining 15 have not had the check. It must be run per candidate
before any of them reaches a ratifier, and its output corrects the existing
cond-2 verdicts rather than merely adding to them; expect it to move some.
**Score after hand-verification: 1 of 46 clears both conditions, not 2.** Only
idx 17 survives. That is the number §10 measurement #2 should be read with — the
classifier's own "2 of 46" counted idx 8 on a cond-2 justification that was
half-unchecked.
### 9.3 The whole-file check run on the remaining 15
Run 2026-08-03, one candidate at a time, over the 15 O2 candidates other than
idx 8 and 17 — ten distinct files. No KB file was edited.
**Method, and its one calibration.** For each record, the deletion segments were
recovered by diffing `file_text_verbatim` against `proposed_remainder` (word-level
LCS), content phrases were extracted from each segment (markdown stripped,
Norwegian and English stopwords dropped, contiguous content runs of 2-3 words kept
as noun-phrase units), and each phrase was matched case-insensitively against every
line of the file *outside* the verbatim block. The extractor was calibrated on
idx 8 before the sweep: it must surface `Managed Compute Cluster` at line 263 from
the deleted `eller managed compute cluster`. It does, as the top-ranked multi-word
hit. Without that calibration a narrower extractor would have returned a clean
bill on idx 8 — and silently on others.
**The grep is necessary and not sufficient.** Two of the findings below have no
lexical overlap with the deleted tokens at all. Idx 14's surviving claim is the
Norwegian `automatisk` restating a deleted English `Automatically`; no token
search finds it. Idx 36's deleted `indiscriminate` appears nowhere else in its
file — the grep returns nothing at all — and the finding at line 310 came from
reading the section, not from a hit. Every candidate was therefore also read in place — the enclosing
section around `real_line`, plus every hit line with context — and asked the
second question the grep cannot: *does the surviving text now claim something
broader or narrower than before, and does anything else in the file depend on the
version that was there?*
**Result: 9 clean · 4 contradicted · 2 operator calls.**
| idx | file:line | outcome | the line that decides it |
|---|---|---|---|
| 7 | `document-intelligence-prebuilt-models.md:79` | clean | `prebuilt-document` occurs only in the deleted row; the one `General` hit (192) is `generalisering`; no count binds the table |
| 9 | `data-drift-monitoring-detection.md:218` | clean | the only other Foundry/RAG reference (320) asserts exactly `groundedness, relevance` — the remainder — and never claims drift detection over grounding data |
| 14 | `feedback-loops-continuous-improvement.md:555` | **contradicted (partial)** | 566: `Reviewed documents automatisk tilgjengelige i "Feedback loop" data source når modellen retraines` |
| 18 | `rag-caching-optimization.md:29` | **contradicted** | 303-318: a whole section `### Azure AI Search - Built-in Caching`, plus 510: `Azure AI Search caching \| **Verified**` |
| 19 | `rag-caching-optimization.md:297` | clean | the deleted bullet is the file's only indexing statement; the code sample's silence about explicit vector indexes predates the deletion |
| 26 | `transparency-documentation-standards.md:117` | **contradicted (partial)** | 300: `\| **Risk assessment** \| Responsible AI Scorecard: Error analysis, fairness assessment \|` |
| 27 | `transparency-documentation-standards.md:426` | **contradicted (partial)** | 216: `- **Copilot Studio**: "Powered by AI" disclosure i chat interface` |
| 28 | `transparency-documentation-standards.md:83` | clean | `Hugging` occurs only in the deleted bullet; no model-card template is claimed anywhere else |
| 31 | `ai-incident-response-procedures.md:139` | clean | `legalHold` occurs only in this block; `enabled` at 148 is Blob versioning; 82's "Legal hold på alle artifacts" is consistent with a tags-only hold |
| 33 | `ai-threat-modeling-stride.md:211` | operator call | 357 restates both deleted capabilities — but about the CAF *document*, not about Defender AISPM |
| 36 | `ai-threat-modeling-stride.md:38` | operator call | 310: `Backdoored models og data poisoning er Critical-severity trusler` — unqualified, while the remainder narrows the register to `(targeted)` |
| 38 | `data-leakage-prevention-ai.md:396` | clean | both deleted policy templates occur only here; the following one-click-policy block names neither |
| 40 | `supply-chain-security-ai-models.md:133` | clean | the CVSS band definition occurs only here; 158's `critical vulnerabilities` is container-image scanning, a different tool |
| 42 | `supply-chain-security-ai-models.md:200` | clean (strengthened) | every other HuggingFace reference (32, 240, 492, 513) treats it as an *unverified* source; the deleted bullet was the outlier |
| 45 | `semantic-caching-patterns.md:436` | clean (strengthened) | `Norway West` appears nowhere else; 303, 451, 488 and 628 are all Norway East |
**The four contradictions are one shape.** In each, an enumerating passage would
lose a member that the file continues to assert elsewhere — a section (18), a
mapping row (26), a bullet in a sibling section (27), a prose restatement (14).
That is idx 8's shape exactly: a list narrowed against a file that documents what
was removed. Idx 18 is the hardest of them, because the surviving claim is not a
stray sentence but a titled section *and* a row in the verification table stamping
it `**Verified**`. No deletion confined to line 29 can fix that file; the correct
edit is larger than the O2 envelope permits.
**Three of the four admit a reduced subtraction**, on the same terms as idx 8:
- **idx 14** — drop the `Automatically``Add` half, keep `/ SharePoint` (which is
the file's only occurrence). This also removes the V2b machine flag, since the
recapitalisation was the flagged part.
- **idx 26** — drop item 4 (Error analysis), keep item 5 (Counterfactual analysis);
324, 475 and 693 attach counterfactuals to the dashboard and to GDPR, never to
the scorecard. The renumbering artifact (`1,2,3,4,6,7`) survives either way.
- **idx 27** — drop the Chat-interface row, keep the Plugin-actions row; `grep -niE
"confirmation|plugin"` returns line 428 alone.
Every one of these is a **different remainder string** than the one V2b attested,
so none is machine-clean until `check-o2-returns.mjs` is re-run against it. Idx 18
has no reduction: it is a single deletion.
**The two operator calls are a distinct class, and are not being called
contradictions.** In both, the deleted content survives elsewhere in the file
without the remainder becoming false:
- **idx 33** — line 357 does say `AI asset inventory via Azure Resource Graph` and
`Microsoft Purview Insider Risk Management for prompt-basert data
exfiltration-deteksjon`, stamped `*(Verified MCP 2026-04)*`. But it says it about
what the *Cloud Adoption Framework document* now covers, while the deleted
bullets attributed those capabilities to *Defender for Cloud AISPM*. Different
subjects, so no contradiction — but the edit's benefit is smaller than it looks,
because the content it removes stays in the file under another attribution.
- **idx 36** — the remainder narrows the severity register to `Data Poisoning
(targeted)` while line 310 still justifies a recommendation with the unqualified
`data poisoning er Critical-severity`. A narrower statement does not contradict a
broader one; it is subsumed by it. What the edit produces is a file whose
severity table is more precise than the prose that cites it. Whether that is
acceptable, or whether 310 needs the same qualifier, is a judgement about the
file — and a companion edit at 310 exceeds the single-locator O2 envelope.
**What this does to the score — nothing, and that is deliberate.** The sweep
settles the whole-file dimension of cond 2 for all 15: affirmative for 9, negative
for 4, operator for 2. It settles nothing about cond 3, which stands at
`human_must_confirm` for ten of them.
**The verified score remains 1 of 46 (idx 17).** The temptation here is idx 19,
and it should be named rather than acted on. Its cond-2 doubt was itself
whole-file-shaped — the classifier worried that the preceding code sample's
silence about explicit vector indexes might read as "no setup needed" — and its
cond 3 the classifier already marked `yes`. The whole-file check finds no
contradiction: the silence predates the deletion, and the deletion removes an
affirmative false claim. That **narrows** the objection; it does not resolve it.
What is left is an omission judgement — does pre-existing silence mislead a reader
of this file? — and that is the ratifier's call, not the sweep's. Counting idx 19
would mean promoting a `human_must_confirm` to affirmative on the checker's own
authority, which is the one direction this gate exists to prevent. Note the
asymmetry: idx 14, 18 and 26 were also `cond2 = confirm`, and for those the sweep
*confirmed* the doubt. Moving a candidate the other way is not the sweep's to do.
**Idx 19 is therefore the strongest new candidate for ratification** — not a
member of the verified class.
**What the sweep says about the method.** The classifier's cond-2 column was
wrong in one direction only. Of the eleven candidates it marked `cond 2 = yes`,
the whole-file check overturns two (27, 36) and confirms nine. Of the four it
marked `human_must_confirm`, the check clears one (19) and confirms the doubt on
three (14, 18, 26) — in every one of those three the classifier had already
written the contradicting line number into its own evidence field without
treating it as a defeater. The information was in the returns; the contract just
never asked the classifier to act on it. That is a prompt gap, not a model
failure, and it is cheap to close in a later wave: name internal consistency as a
cond-2 dimension and require the citation.
### 9.4 The ratification packet — four edits prepared, none applied
Prepared 2026-08-03. **No KB file was edited.** This section exists so the
ratifier decides from attested strings rather than from prose.
**First, an ambiguity in §9.3 that had to be resolved before anything could be
written.** The three reduction bullets above use the form *"drop X, keep Y"*.
That reads two ways — drop X *from the subtraction*, or drop X *from the file* —
and the two readings produce **opposite edits**. The prose does not settle it.
What settles it is the contradiction each reduction exists to avoid: the member
the file continues to assert elsewhere must **survive in the file**, so it is the
*other* member that stays in the subtraction. Applied to each, and verified
against the live file this session:
| idx | the file asserts elsewhere | therefore survives | subtraction reduces to |
|---|---|---|---|
| 14 | 566 `Reviewed documents automatisk tilgjengelige …` | `Automatically` | `/ SharePoint` only |
| 26 | 300 `Responsible AI Scorecard: Error analysis, fairness assessment` | item 4, Error analysis | item 5, Counterfactual, only |
| 27 | 216 `**Copilot Studio**: "Powered by AI" disclosure i chat interface` | the Chat-interface row | the Plugin-actions row only |
§9.3's own parenthetical corroborates this independently: it records the surviving
numbering for idx 26 as `1,2,3,4,6,7`, which is what deleting item 5 produces.
The reading also matches STATE's summary. The `drop/keep` phrasing above should be
read in the subtraction sense throughout.
**The four prepared edits.** Each reduced remainder was derived by string surgery
on the original `file_text_verbatim` — never transcribed by hand — and re-run
through `checkRow` (V1/V2/V2b/V3):
| candidate | file:real_line | machine | cond 1 | cond 2 (whole file) | cond 3 |
|---|---|---|---|---|---|
| **17** as attested | `rag-caching-optimization.md:254` | clean | yes | yes (§9.2) | yes |
| **14** reduced | `feedback-loops-continuous-improvement.md:555` | clean | yes | **yes (verified here)** | yes |
| **26** reduced | `transparency-documentation-standards.md:117` | clean | yes | no contradiction, **but see below** | yes |
| **27** reduced | `transparency-documentation-standards.md:426` | clean | yes | **yes (verified here)** | `human_must_confirm` |
The cond-2 evidence gathered this session, since a reduced remainder is a new
remainder and inherits no clearance from the one V2b originally attested:
- **idx 14** — `sharepoint` occurs in the file **only** inside the verbatim block
(line 555). Nothing else asserts SharePoint as feedback storage; line 563's
"AI Builder feedback loop storage" is generic and consistent with Dataverse.
The classifier's cond-2 doubt was line 566's automaticity — and the reduction
deletes the automaticity change, so that defeater no longer applies to this
subtraction at all.
- **idx 26** — counterfactuals appear at 324, 475 and 693. Line 324 is a row in
the table *Azure Machine Learning — Built-in transparency tools*, attaching
`counterfactual what-if` to **Model interpretability**, not to the Scorecard;
475 and 693 are GDPR right-to-explanation and Azure ML explanations. None
asserts Counterfactual analysis as a scorecard segment, so the reduction
contradicts nothing.
- **idx 27** — `grep -niE "confirmation|plugin|sensitive action"` and the
Norwegian forms (`bekreft|godkjenn|samtykke`) return line 428 alone, i.e. only
the row being deleted. Nothing else in the file carries the claim.
**Nothing here promotes itself, and that is deliberate.** Two of the four still
carry an unresolved human condition, and the resolution is the ratifier's:
- **idx 26** — the classifier's cond-2 `human_must_confirm` was never about a
contradiction. It was the **renumbering artifact**: delete-only cannot renumber,
so the list reads `1,2,3,4,6,7`. The whole-file check does not touch that, and
the artifact survives the reduction. Accepting it is a judgement about the file.
- **idx 27** — cond 3 stands at `human_must_confirm` and the sweep settles only
cond 2.
**The verified score therefore remains 1 of 46 (idx 17)** until a ratifier acts.
Moving idx 14 to affirmative is defensible on the record — its only stated doubt
is deleted along with the automaticity change — but that is a ratification, and
§9.3's asymmetry rule applies: this pass may confirm doubt, never clear it on its
own authority.
**Idx 17: recommended as attested, both residues left standing.** The amended
variant with the trailing colon dropped was also run through `checkRow` and is
**also machine-clean**, so the choice is free on machine grounds — which means it
must be made on other grounds. Two argue for leaving it: the colon is not false,
and the attested string is the one the measurement was taken on. The `**Verified**`
stamp at line 258 is *outside* the verbatim block; editing it would be a second
locator, excluded by the same single-locator rule that put idx 36's companion edit
at 310 out of envelope. Both residues should be recorded as consciously left, not
overlooked.
**One coupling checked before any write, because it fails silently.** Applying
idx 17 makes its `file_text_verbatim` no longer occur in the file, so **V1 fails
for that row permanently** and "46/46 pass V1" stops being true. The test suite is
unaffected — `tests/kb-eval/test-o2-return-check.test.mjs` drives `checkRow` with
a synthetic `readFile` stub (`skills/x/references/y.md`) and never reads the live
corpus. But `check-o2-returns.mjs`, the CLI, *does* read live, and will report the
failure on every future run. The returns directory is evidence of a pre-edit
state and must be read as such; the CLI's V1 tally is only meaningful against an
unedited corpus. Recorded rather than worked around.
**A second-order note for whoever applies these.** Idx 26 and 27 are in the same
file. Applying either shifts the line numbers the other cites (216, 300, 324),
so re-derive references after the first write and anchor on
`file_text_verbatim` — never on the line number, which already differs from
`real_line` in 9 of 17 records.
### 9.5 Ratified and applied — the first corpus edit
Operator ratification 2026-08-03, applied the same session in `957ebef`.
**Four subtractions written to three publicly distributed KB files.**
| idx | file | what was deleted | form |
|---|---|---|---|
| **17** | `rag-caching-optimization.md` | the three score-threshold bands | as attested, trailing colon **kept** |
| **19** | `rag-caching-optimization.md` | `- Automatic indexing av vectors` | as attested |
| **33** | `ai-threat-modeling-stride.md` | `(via Azure Resource Graph)` + the whole Purview bullet | as attested |
| **14** | `feedback-loops-continuous-improvement.md` | `/ SharePoint` only | **reduced** — `Automatically` kept |
**Held back, and why** — each is a live item, not a rejection:
- **idx 26** — the renumbering artifact (`1,2,3,4,6,7`) is unresolved. Delete-only
cannot renumber; accepting the artifact is a judgement not yet made.
- **idx 27** — cond 3 still `human_must_confirm`; the sweep settled only cond 2.
- **idx 36** — applying it alone yields a severity table more precise than the
prose at 310 that cites it. The companion edit is out of envelope → **G7**.
- **idx 18** — no reduction exists; the correct fix spans a section and a
verification-table row → **G7**.
**idx 19 was promoted by the operator, not by the checker.** §9.3 deliberately
declined to count it, since its cond 2 stood at `human_must_confirm` and clearing
it would have meant a checker promoting its own doubt. The ratifier resolved the
omission question — pre-existing silence in the code sample at 275-292 is not
made worse by deleting an affirmatively false bullet — and that is the authority
the gate was waiting for. **Verified score: 1 → 4 of 46**, all four by
ratification.
**The driver.** `scripts/kb-eval/apply-o2-ratified.mjs` (+ 11 tests). Every string
comes from the tracked returns, never transcription; the single amendment
(idx 14) is a derivation that asserts its own effect and aborts on a drifted
record. Anchoring is on `file_text_verbatim` — which matters concretely here,
because idx 17 and 19 share a file and the first shifts the second's line
numbers. The run aborts, writing nothing, on an anchor that is not unique, a
remainder that is not deletion-only, or a novel word form. Writes are atomic.
**The V1 consequence, predicted in §9.4 and now measured.** `check-o2-returns.mjs`
reports `machine-clean O2 candidates: 13/17` (was 16/17) with four new V1
findings — idx 14, 17, 19, 33, each `file_text_verbatim NOT FOUND in the file`.
This is correct behaviour, not a regression: the returns directory is evidence of
a **pre-edit** corpus, and V1 asks whether the quoted text is still there. The
test suite is unaffected (1032/1032) because it drives `checkRow` with a
synthetic stub. **Any future reading of that CLI's V1 tally must subtract the
applied rows** — the number is only meaningful against an unedited corpus.
**Cross-corpus check — the dimension every check so far has missed.** Cond 2 is
ratified as *whole file*, and §9.3's sweep, §9.4's verification and the driver's
invariants are all **within-file** by construction. But the corpus is 389 files
that agents read together, so a citation of deleted content in *another* file
would be dangling in publicly distributed material — the same trust defect, one
scope out. Run over all of `skills/**/*.md` after the edit:
- **idx 17** — no file cites the deleted APIM bands. `semantic-caching-patterns.md:79`
carries its own `0.70-0.84: Liberal matching` rubric for a different threshold;
the transcription-confidence bands at `audio-video-transcription-workflow.md:420,547`
are unrelated.
- **idx 19** — `Automatic indexing` occurs nowhere else in the corpus.
- **idx 14** — every other `SharePoint` hit is SharePoint as a general M365 source,
channel or connector. None claims it as AI Builder feedback-loop storage.
- **idx 33** — ~40 corpus hits for `Azure Resource Graph` / `Insider Risk Management`,
and **not one attributes them to Defender for Cloud AISPM.** They are independent
claims in their own files, and the closest matches — `ai-incident-response-procedures.md:505,
575, 603`, `data-leakage-prevention-ai.md:780`, `norge-ai-strategy-government.md:150,152`
— attribute them to the **Cloud Adoption Framework Secure AI** document, the same
attribution that survives at line 357 of the edited file.
**No dangling reference, but note what the idx 33 result actually shows.** The
deleted capabilities are asserted in at least four other files under CAF
attribution, several stamped `Verified MCP 2026-04`. §9.5 records that the edit's
benefit is smaller than it looks because the content survives at line 357; the
corpus-wide measurement is that it survives in *five* places, not one. The
subtraction is still correct — the AISPM attribution was unsupported — but anyone
weighing whether this class of edit is worth the review cost should weigh it
against that number. This is the first cond-2 evidence gathered at corpus scope
rather than file scope, and whether corpus scope becomes a standing third reading
of cond 2 is unratified and deliberately left open.
### 9.6 The two held-back candidates resolved against first-party evidence
Run 2026-08-03, the session after the first corpus edit. §9.4 left idx 26 and 27
each carrying **one** human condition, and STATE framed both as operator calls.
Fetching the sources changed the answer for one of them and sharpened the other.
Neither was applied; §9.6 is evidence, not an edit.
**A framing correction that had to come first.** §9.4's reduction narrows idx 27's
subtraction to **the Plugin-actions row alone** — and that is precisely the row the
classifier's cond 3 said a human must confirm. So the reduction does not dilute the
open condition, it **concentrates it**: after reduction, 100 % of the edit is the
part nobody had cleared. The original two-row form had a clean half; the reduced
form has none. "One condition remains" understated it.
**idx 27 — leaves O2, on modality grounds. Cond 3 is NOT settled, and the scope
question underneath it is named rather than assumed.**
The file cites `microsoft-copilot-studio/responsible-ai-overview` for the whole
section. That page is a **hub**: it establishes nothing itself, it links the FAQ
set. One hop out, `faqs-generative-orchestration` says:
> "Makers can require user confirmation before executing tools that modify data."
**Two cautions, both against the stronger reading this section originally carried.**
*First, a scope decision, not a free move.* Grounding on a page the file does not
cite — reached one hop through the cited hub — is a **new reading of "the source"**,
structurally the same kind of expansion as the corpus-scope reading of cond 2 that
§9.5 deliberately left unratified. R11 has otherwise held grounding to the cited
URL throughout. **Whether a hub's linked children count as the cited source is
therefore UNRATIFIED**, and is recorded here as an open question rather than
exercised silently.
*Second, the quote is thinner than it looks.* "Makers **can require**" describes a
**configurable capability**, not a feature the source establishes as present. Cond 3
asks whether the subtraction destroys what the source establishes; an option a maker
may enable is weak evidence for that. So **cond 3 stands unresolved at
`human_must_confirm`** — this pass did not clear it and does not claim to have
failed it either.
**idx 27 leaves O2 anyway, and for a reason that does not depend on either point
above.** The defect is **modality**, not fabrication: the file asserts confirmation
prompts as a *built-in disclosure*, whereas the mechanism — on any reading of the
evidence — is maker-configured. Deleting the row would remove a claim whose core is
sound and whose framing is wrong. **The correct repair is a replacement, and
replacement is outside the delete-only envelope by definition.** That holds whether
cond 3 eventually fails or clears, which is why it is the load-bearing reason.
Same for the Chat-interface row the reduction kept: the FAQ documents a default
transparency **message** ("Just so you are aware, I sometimes use AI to answer your
questions."), not the "Powered by AI" **badge** the file claims. Not part of this
subtraction, but now on the record as imprecise.
**idx 26 — confirmed, and its held-back half is now positively false.** The canonical
scorecard segments, from `how-to-responsible-ai-scorecard`, are: summary/model
overview, data analysis, model performance, cohorts, top important factors, fairness
insights, causal insights. `concept-responsible-ai-dashboard` lists Error analysis
and Counterfactual analysis as **dashboard** components. The classifier was right
about **both** items 4 and 5.
That matters because §9.4's reduction deletes item 5 **and deliberately keeps item
4**, since line 300 asserts Error analysis as scorecard content too. Before this
session item 4 was merely *uncleared*; it is now **measured false in two places**
(117 and 300). So the operator question is not one part but two:
1. accept the renumbering artifact (`1,2,3,4,6,7` — delete-only cannot renumber), and
2. accept that a **known-false** claim stays in a publicly distributed file, with the
117+300 pair booked to G7.
Presenting only (1) would let the whole-file reading of cond 2 quietly convert a
defect into a permanent resident.
**A regression this session found in the edit already shipped.** §9.4 recorded idx
17's residue as a `**Verified**` stamp sitting outside the verbatim block. The live
file shows something worse. `957ebef` deleted the three bands that followed the lead-in,
leaving (`rag-caching-optimization.md:253`):
```
**Score Threshold Tuning** (APIM `score-threshold` er en DISTANSE: …likhet):
**Verified** (Microsoft Learn - Enable semantic caching for LLM APIs)
```
A lead-in ending in a colon, promising an enumeration that no longer exists, followed
by a verification stamp. **The subtraction was correct; the paragraph it left is not.**
The colon was kept as an operator choice on the grounds that both variants were
machine-clean — and they were. V1/V2/V2b/V3 are string invariants over the deleted
text; **none of them can see document coherence.** Dropping the colon would not have
saved it either, since the lead-in is empty in both variants. This is a genuine
defect introduced by our own edit into public material, and it belongs to the
**replacement** class, not the subtraction class. The file already carries the
correct guidance twice (164 and 429: "Start med 0.15, tune opp basert på metrics").
**G7 measured rather than extrapolated.** Of the four subtractions applied in
`957ebef`, **two left a residue** (17 — the dangling lead-in; 33 — the CAF-attributed
survival at 357 plus five cross-file assertions) and two were clean (19, 14). Adding
the held-back set, G7's membership is now **five**, and one is a live regression:
| idx | residue | class |
|---|---|---|
| 17 | dangling lead-in at 253 + `**Verified**` stamp | replacement — **shipped, live** |
| 33 | line 357 CAF attribution, stamped `Verified MCP 2026-04`, + 5 cross-file | replacement / corpus-scope |
| 26 | item 4 at 117 kept, asserted again at 300 — **measured false** | multi-locator |
| 36 | companion edit at 310 | multi-locator |
| 18 | whole section 303-318 + `**Verified**` row 510 | multi-locator |
**This is input the §9.4 write-up did not have, and it tilts the (a)/(b) choice.**
A 50 % residue rate on applied edits means residues are not an exception to be
queued; they are the **normal by-product** of a delete-only envelope. A named queue
into human review (b) absorbs a steady stream. An explicit multi-locator class with
its own return contract (a) would have to be built for the common case, not the
edge — and note that two of the five (17, 33) are not multi-locator at all but
**replacements**, which an O4 deletion-oriented class would not fix. On this
measurement (b) is the better fit, and (a) would be mis-sized against the evidence.
**Verified score is unchanged at 4 of 46.** idx 27 leaves O2 without becoming a
subtraction; idx 26 remains available to a ratifier as a two-part accept.
**Operator resolution, same session.** idx 26: the partial fix was **declined** —
a delete-only edit that knowingly leaves a measured-false claim standing while
introducing a renumbering artifact buys too little, so the whole 117+300 pair went
to G7. G7 form: **(b), the named queue**, on the measurement above. The idx 17
regression was fixed by changing the colon to a period, making the lead-in a
complete and independently true sentence that the `**Verified**` stamp correctly
covers; the corpus was swept for the same defect shape with no other occurrence.
### 9.7 idx 26 closed out of the queue — the first G7 entry repaired
**Both sources re-fetched live before writing, not read off §9.6.** The measurement
above was made in the same session that booked the entry, so it was treated as a
premise rather than a fact. `how-to-responsible-ai-scorecard` enumerates the
segments as summary/model overview, data analysis, model performance, cohorts, top
important factors, fairness insights, causal insights.
`concept-responsible-ai-dashboard` lists Error analysis and Counterfactual what-if
among the dashboard components. Both confirm §9.6 independently.
**Operator ratified form (b): relabel, not removal.** Items 4 and 5 leave the
numbered scorecard list, which renumbers cleanly to **15**. The `1,2,3,4,6,7`
artifact §9.4 and the queue entry both predicted was forced *only inside the
delete-only envelope*; an ordinary Edit renumbers for free. Carrying that
constraint forward would have been inheriting a stale cost. The two capabilities
survive in a blockquote explicitly marked as dashboard components, so
source-confirmed information is preserved and the reader is warned off exactly the
conflation that produced the defect. Locator 2 now reads `fairness insights` alone.
**A third defect sat inside neither anchor.** `**Confidence:** Verified (MCP:
microsoft-learn)`, twelve lines below locator 1, vouched for the false list. No
machine check could see it: V1/V2/V2b/V3 are string invariants over deleted text,
and `check-g7-queue.mjs` tests anchors only. The stamp was kept but dated
`2026-08-03` to record the re-verification. This is the same class as the idx 17
regression — **an edit can be anchor-correct and leave a false claim standing
somewhere the checks do not reach.** Read the neighbourhood, not the operation.
**Booked, not folded in: idx-26b.** The post-edit file sweep found
`| **Accuracy metrics** | Responsible AI Scorecard: Quantitative analyses |` in the
same table as locator 2 but outside both anchors. "Quantitative analyses" is not a
scorecard segment — the source calls it *model performance* — and it is a canonical
**Model Card** section, which this same file lists as one at line 77. So the defect
is a cross-attribution between two standards, not loose wording. It was entered as
its own queue member rather than repaired inside a ratified entry, per gap
discipline: folding unbooked work into a ratified entry launders it.
**And the repair itself opened one — idx-26c.** The same live fetch listed **seven**
canonical segments; the corrected list carries **five**. `model performance` and
`cohorts` are absent. That was a latent incompleteness under an undated stamp; dating
the stamp to `2026-08-03` converted it into a positive claim that *this* enumeration
was verified that day, over content the same day's verification showed to be short two
members. **The finding was already sitting inside idx 26's own `resolution` field,
which quotes all seven while the file lists five** — booked, not folded in, because
idx 26's ratified scope was the falsity of items 4 and 5, not the completeness of the
list. This is the third time in two sessions that the defect was next to the edit
rather than in it.
**One sentence is weaker than the rest, and is marked as such.** The blockquote's
second clause — that the two components have no segments of their own in the PDF — is
derived from *absence in an enumeration*, not from a positive statement in the source.
The "How to read your scorecard" walk-through is structurally exhaustive, so it is
near-certain, but it is an inference and is recorded here as one.
**Queue state: 6 open, 2 resolved.** Suite 1047/1047; both idx 26 anchors correctly
stopped matching, and the entry carries its `resolution`.
### 9.8 idx-26b and idx-26c closed together — and a stale constraint caught a second time
**The source was re-fetched live again, not read off §9.7.** §9.7 was written in the
session that booked both entries, so it is a premise. `how-to-responsible-ai-scorecard`
was fetched fresh and enumerates seven segments — summary/model overview, data
analysis, model performance, cohorts, top important factors, fairness insights, causal
insights — and names the accuracy segment *model performance*. Both measurements
confirmed independently of §9.7.
**A machine constraint asserted in STATE was false, and it was blocking the edit
form.** STATE said a resolved entry's anchor *must* stop matching or the check fails,
"by design". `lib/g7-queue.mjs:69-75` says otherwise: a `resolved` entry returns before
the anchor check, so it is fully exempt, and anchor drift only bites an entry left
open. The real mechanism is the inverse of the claim — the check catches *editing the
file without resolving the entry*, not *resolving without changing the text*. Had the
claim gone unchecked it would have forced insertion in source order for idx-26c on
machine grounds that do not exist. **This is §9.7's "stale constraints are inherited"
lesson recurring one session later, with the stale constraint now living in the
handover rather than in the queue.** Read the code, not the note about the code.
**idx-26b — operator ratified rename over removal.** The Accuracy metrics row now
reads `Responsible AI Scorecard: model performance`. Deletion was available and
rejected: the EU AI Act accuracy-metrics mapping is genuine and source-supported, so
dropping the row would have removed true information in order to repair a naming
defect. Line 77 keeps *Quantitative analyses* as a Model Card section — correct there,
and the reason the cross-attribution was visible at all.
**idx-26c — operator ratified adding the two over downgrading the list.** *Model
performance* and *Cohorts* were appended as items 6 and 7; the existing five were left
untouched, confining the change to the measured gap. Appending rather than inserting in
source order is defensible because the list carries no ordering claim and the original
five were already not in source order, so appending introduces no new falsity. Item 7
is worded *automatisk uttrukket av scorecard-en* to keep it distinct from the
`Cohort analysis` bullet in the Customization block, which is operator-defined and
enumerates no segment — the same distinction the queue entry warned would otherwise
look like a cure.
**Two trust markers, decided explicitly in opposite directions.** The
`Verified (MCP: microsoft-learn, 2026-08-03)` stamp at line 129 is **kept unchanged**;
its scope was checked rather than assumed, and it now vouches for a complete
seven-member enumeration re-verified against the live source on the date it already
carries. The `Verified (Baseline + MCP-inferred)` stamp under the compliance table is
**deliberately not upgraded**, even though the row beneath it was just measured against
first-party source: that stamp covers six rows and **four** of them are still unmeasured,
so strengthening it would extend a verification claim over unmeasured content. §9.7 taught
that renewing a marker commits you to everything it covers; the same rule read forward
says a marker may not be strengthened by a repair narrower than its scope.
**Correction to this section as first written (`527fb03`), caught on review.** The
sentence above said five of the six rows were unmeasured. That is wrong: idx 26's
locator 2 rewrote the Risk assessment row to *fairness insights*, which the same live
enumeration lists as a canonical segment, so **two** rows are source-measured and four
are not. The error is instructive, because the queue `resolution` for idx-26b states it
correctly — "only one was measured against the source **this session**" — and §9.8
restated it with the scoping clause dropped, turning a true scoped claim into a false
count. **That is the §9.7 defect class reproduced one paragraph after writing it up:**
a claim that was true inside its qualifier became false when the qualifier was left
behind. The conclusion the count supports is unaffected — four unmeasured rows forbid
strengthening the stamp exactly as five would.
**idx-26c's anchor still matches verbatim, and that is the first real exercise of the
exemption.** Item 5 was untouched, so `5. **Data quality**: …` still occurs in the
file. idx-26b's anchor did stop matching. Both validate, because `resolution` is what
the gate demands of a resolved entry. "First" was checked rather than assumed against
all four resolved entries: idx 26's two anchors both drifted, idx-26b's drifted, and
idx-17 carries an **empty** anchor array, so it could neither match nor drift and never
exercised the exemption. idx-26c is the first entry to close with a live anchor
standing.
**The sweep opened a fourth entry — idx-26d — but not in the same way §9.7 did.**
idx-26c was *created* by its repair: dating a stamp converted a latent gap into an
active claim. idx-26d was not created by this one. The paraphrase drift in items 15
predates both edits — item 5 *Data quality* is the source's *data analysis*, item 3
*Model interpretability* is *top important factors*, and item 1 describes the summary
segment as "Architecture, training data, intended use", which is Model Card content and
is the same cross-attribution class as idx-26b rather than mere imprecision. What the
repair changed is **visibility**: items 6 and 7 carry source names, so the list now
mixes two naming dialects and the drift is legible where it was not before. Full
canonicalisation was put to the operator as a third option for idx-26c and was not
chosen, so the ratified scope was completeness alone and this is booked, not folded in.
**Queue state: 5 open, 4 resolved.** Suite 1047/1047.
### 9.9 idx-26d and idx-27 closed — and what an unmeasured neighbour costs an edit
Run 2026-08-03, same day as §9.8, later pass. Both sources were re-fetched live
before writing; neither §9.7 nor §9.8 was read as fact.
**idx-26d closed by full canonicalisation, and the discriminator was not the one the
entry offered.** The entry framed an either/or: rename all five members, or rewrite
item 1 alone — item 1 being the only member whose drift produces a false attribution
rather than a recognisable paraphrase. Item 1 is genuinely the worst member, so the
narrow form is tempting. It is also the wrong form, for a reason the entry's own
summary contains: idx-26d is **booked as the list mixing two naming dialects** after
idx-26c added two source-named members. Rewriting item 1 alone removes a falsity and
leaves the booked defect standing — it would resolve an entry whose stated defect
survives the resolution. **A resolution has to close the defect the entry names, not
the worst defect the entry mentions.** The two are not the same thing, and the entry
text will not tell you which one you are looking at unless you re-read it against the
resolution you are about to write.
All five names were renamed to the source segments. Descriptions were rewritten for
items 1 and 3 only: item 1 carried Model Card content (this file lists Model details /
Intended use / Training data as Model Card sections at lines 7176), and item 3
carried "global/local explanations", RAI *dashboard* vocabulary standing inside a
scorecard enumeration that this same file's callout says is not a dashboard listing.
Source order was **not** imposed — §9.8 established the list carries no ordering claim.
**The descriptions that were deliberately left wrong.** Items 2 and 5 carry specifics
the source does not state: "(gender, ethnicity, age)" where the source says only "your
desired sensitive groups", and "missing values, outlier analysis" where it says only
that the segment "shows you characteristics of your data". Those were raised as
**idx-26f** in the same pass and left untouched in the file. The option presented to
the operator carried an illustrative sketch that *did* rewrite them; the option's own
label did not. Applying the sketch would have silently resolved an entry raised the
same session and left the queue incoherent — an entry pointing at text that no longer
exists, with nobody having ratified its removal. **When a decision is presented as
label plus illustration, the label is the ratified object.** The illustration is a
reading aid and may be wider than what was decided.
**idx-27 closed without ratifying the hub question.** Open question #8 — whether a hub
page's linked children count as the cited source — was answered **no**. R11's strict
reading stands. The entry closed anyway, by a move the question does not gate: the
child page was **added to the file as a cited source**. Grounding on an uncited page
is an expansion of "the source"; grounding on a page you then cite is not. This is
worth naming as a general move — *an unratified scope question can sometimes be
routed around by changing the artefact rather than the rule*, and routing around it
leaves the rule unweakened for every other entry that will meet it.
**The repair could not be made in place, and the reason is instructive.** The
Plugin-actions row asserts a confirmation prompt as a **built-in** disclosure; the
source makes it maker-configurable. The row sits in a four-row table under the heading
**Built-in disclosures**, and two of those four rows have never been measured. Three
repairs were available:
- Fix the row's text in place → the heading still asserts the false modality.
- Weaken the heading → silently restates the modality of the two unmeasured rows.
- Add a modality column → asserts "built-in" about the unmeasured rows outright.
The last two fix one unverified claim by minting two more. The row was **moved out**
into a separate maker-configured block instead. **Unmeasured neighbours constrain the
shape of a repair, not just its scope:** the cheapest true edit is the one that
restates nothing you have not checked, and that is frequently *not* the smallest diff.
This is the same neighbourhood discipline as §9.7/§9.8 read from the other side —
there, the neighbourhood held defects to find; here, it held claims not to touch.
**A false locator inside a tracked artefact, found by re-reading a resolution.**
idx-26c's resolution located the **Confidence:** stamp at "line 129". Line 129 is the
**Status:** line; the stamp is 131. The claim about the stamp was true, the locator
was not, and it was committed. The queue contract forbids line numbers as anchors
because `line ≠ real_line` in 9 of 17 R11 records — **that prohibition applies to
prose inside an entry too, and nothing checks it.** Corrected in place, with the
correction recorded rather than overwritten. STATE carried the same error.
**Three entries opened, none swept in.** idx-27b (the same "Powered by AI"
imprecision in Mønster 3's implementation list, outside idx-27's anchors), idx-26e (a
parallel five-member scorecard list in `stakeholder-communication-ai-decisions.md`, in
the dialect idx-26d just removed, under an *undated* Verified stamp, and additionally
framed as "configurable elements" which the source does not enumerate), and idx-26f
above. idx-26e was found by a **cross-file grep run before the edit** — the check that
asks whether a rename here creates an inconsistency there. It did not, but it found a
mirror the corpus was not known to contain.
**A scope defect caught in the resolution itself, before the closing line.** An
adversarial read of the idx-27 write-up asked a question the repair had not: the FAQ's
own scope line reads "the AI impact of **generative orchestration** for custom agents
built in Copilot Studio", and both repaired claims had been written into a section
headed *Microsoft Copilot Studio* with no qualifier. That is the widening class idx-27
was raised for, reproduced by its own repair — §9.7's pattern for a third time, and
§9.8's dropped-qualifier pattern for a second.
**Re-checked against the docs rather than reasoned about, and the two claims came
apart.** The maker-confirmation safeguard occurs *only* in the orchestration FAQ, in a
list about tool execution — orchestration-scoped, and the bullet now says so. The
default transparency message occurs there **and, verbatim, in `faqs-generative-answers`**
under the same protections question, framed there as a general best practice rather
than an orchestration feature. Two independent feature FAQs stating it without a
feature qualifier is why the Built-in disclosures row keeps none; the second FAQ was
added as a cited source so a reader can check that reasoning instead of trusting it.
Neither page is a Copilot-Studio-wide statement of record, so the standing is
**grounded-as-cited, not established-for-all-agents**.
The transferable part is not the fix. **A repair inherits the defect class it was
raised against unless something explicitly re-asks the question of the repair.** The
edit was correct on the axis the entry named (modality) and wrong on the axis it did
not (scope), and the resolution prose read as complete precisely because it answered
the named axis well.
**Queue state: 6 open, 6 resolved.** Suite 1047/1047.
### 9.10 idx-27b and idx-26f closed — the minimal diff was the unsound one
Both entries lived in `transparency-documentation-standards.md`, so they were taken
together and grep-checked as a pair, the same way idx-26b/26c and idx-26d/27 were.
**idx-27b was the near-mechanical one and behaved like it.** The Mønster 3
"Azure implementasjon" bullet asserted a Copilot Studio `"Powered by AI"` disclosure —
the same imprecision idx-27 had just removed from the Copilot Studio section. Because
idx-27 had already ratified a replacement wording a few hundred lines below, the repair
was a *copy of a ratified formulation*, not a fresh judgement: the bullet now carries
the standard transparency message verbatim, and the quote was grep-verified
byte-for-byte against that row **after** the edit. Hand-copying a verbatim string is
where quote-style drift enters, and nothing in the suite would have caught it.
What was deliberately not carried over matters more than what was. The bullet sits
under an *audience-layering* table; the ratified row sits under *Built-in disclosures*.
Restating the claim in layered-disclosure vocabulary would have added reach the source
does not state — which is exactly the scope class caught in `66fb567` one session
earlier. The repair for a scope defect is the first place that scope defect can recur.
**idx-26f carried a real choice, and the obvious form of it was wrong.** Items 2 and 5
of the scorecard enumeration carried specificity the live source does not state, under
a **Confidence:** Verified stamp. The apparent repair — delete the named specificity —
does not survive contact with the source:
| Item | Anchor | What deleting only the named specificity leaves | Why that fails |
|---|---|---|---|
| 2 | `... across sensitive groups (gender, ethnicity, age)` | `... across sensitive groups` | Drops the *you-choose-them* agency the source states ("target values **you set** for **your desired** sensitive groups"), and strands item 2 in the pre-idx-26d dialect while items 1/3/6/7 now carry "du har satt" |
| 5 | `Dataset statistics, missing values, outlier analysis` | `Dataset statistics` | The source says only "shows you characteristics of your data" — `Dataset statistics` is *also* unstated, and unlike item 4 this entry never adjudicated it checked-and-clear |
So the minimal diff would have re-created both defects it was raised against: unsourced
content under the stamp, and the mixed dialect idx-26d had just canonicalised — **in the
file idx-26d had just canonicalised it in.** Both descriptions were instead rewritten to
the source's own wording in the established dialect. Item 4 was left untouched, per the
entry's own adjudication that it is a recognisable paraphrase.
**Form (b) — narrowing the stamp's stated reach inside the file — was put to the
operator and declined.** The entry's own text supplies the reason: the stamp is honest
today *only because* idx-26d's resolution writes down an exception, and removing the
need for that written-down exception is what closing the entry means. Annotating the
file instead relocates the annotation from the queue into the corpus; it does not close
the defect as booked.
**A closure can falsify prose in a tracked artefact.** idx-26d's resolution ended
"Items 2, 4 and 5 descriptions are NOT covered by that widening" — true when written,
false the moment idx-26f closed, and checked by nothing. Same class as the false
"linje 129" locator caught last session. It was handled by appending a dated
supersession clause rather than rewriting the original, and idx-26f's resolution now
states positively what the stamp covers, so no exception has to be written down
anywhere for it to stay honest.
**Two neighbours booked, not swept: idx-27c and idx-27d.** The Foundry agent-transparency
bullet asserts a `"This chatbot uses AI"` embeddable component; the scenario 1 tooling
line names a "Copilot Studio disclosure widget" where this same file documents a
pre-built "AI disclosure" *topic*. Neither can be repaired by copying the ratified
Copilot Studio wording — one needs its own live fetch against Foundry docs, the other is
intra-file name drift inside a worked scenario rather than a source-attribution error.
Booking them was not optional bookkeeping: idx-27b's resolution *asserts* they exist, so
leaving them unwritten would have planted the same false-reference defect this section
documents catching.
Cross-file grep before editing put all three replaced strings at zero occurrences
elsewhere in the corpus.
**The repair failed on its own justifying axis, and the coherence check passed it.**
Item 2 first landed as `fairness-target values` — a hybrid compound that is neither
language, where item 1 renders the same source concept as `target-verdiene`. Dialect
consistency was the entire argument for choosing rewrite over deletion, so this is the
repair failing the axis it was justified by. The whole-list re-read after the edit
checked *content* alignment and passed; it did not check compound forms. **A coherence
check inherits the axis you had in mind when you wrote it** — the same blind spot as a
string invariant, one level up. Corrected to `fairness-målverdiene`.
Note where the ratification sat: the operator approved the label *"kildens ordlyd i
idx-26d-dialekten"*, while the illustrative preview carried the defective form. The
label is the ratified object, so correcting toward it needed no re-ratification — but
a preview that renders the label wrongly is a live way to smuggle an unratified form
past an operator who is reading the preview.
**The transferable part: the cheapest true edit is rarely the smallest diff.** A repair
scoped to exactly the words an entry names will silently adopt whatever unmeasured
content sits beside them — and adoption under a verification stamp is indistinguishable,
to a reader, from verification.
**One provenance seam left explicit rather than smoothed over.** The stamp's coverage
after this closure rests on two different acts: items 2 and 5 were re-verified against
the live source *this* session, items 1/3/6/7 under idx-26d earlier the same day. Same
page, same stamp date — but one is inherited, not re-run, and idx-26f's resolution now
says so in those terms rather than presenting all six as a single verification.
**Queue state: 6 open, 8 resolved.** Suite 1047/1047.
### 9.11 idx-26e closed — the reason it was booked separately was false
Run 2026-08-04. The entry was the last of the 26-family and the first to live in a
different file: a parallel five-member list in
`stakeholder-communication-ai-decisions.md`, in the paraphrase dialect §9.9/§9.10 had
just removed from `transparency-documentation-standards.md`.
**Ground truth killed the framing before it could drive the edit.** Both STATE's open
question #9 and the session brief carried it as *"should canonicalisation go as far in a
file that carries no Verified stamp over the list?"* — and the file does carry one. The
stamp `*Confidence: Verified (MCP microsoft-learn)*` at line 81 is **section-terminal**:
it closes §1 (lines 4681) and therefore reaches the list. That was the entire stated
reason the case was booked separately instead of folded into idx-26d, and with it gone
the honesty argument from §9.10 applied here with undiminished force. **Verifying an
entry's claim and verifying the reason it was booked are two different checks**, and
only the first one is habitual.
**Two source pages, and measuring only one of them would have manufactured findings.**
The seven segments are enumerated on `how-to-responsible-ai-scorecard`. But the page
*this file cites* (Kilder #1) is `concept-responsible-ai-scorecard`, which enumerates no
segments — and which turns out to ground the rest of the section almost verbatim:
| Section text | Measured against how-to alone | Measured against the page the file cites |
|---|---|---|
| "Muliggjøre multi-stakeholder alignment i ML-livssyklusen" | looks like paraphrase drift | "the need for effective multi-stakeholder alignment in an end-to-end machine learning lifecycle" |
| "Støtte auditability for risikoofficerer og regulatorer" | looks like invented specificity | "share model and data insights with auditors and risk officers for auditability purposes, as required by AI regulations" |
| the configurability framing | unsupported | "you can provide your desired model performance and fairness target values, such as target accuracy and target error rate" |
Two false findings and one wrongly-rejected framing, avoided only because the file's own
citation list was read before the stamp was dated. **The page a claim came from and the
page the file cites are different objects; a stamp covering a section has to be measured
against both.**
**The obvious repair was rejected on measurement, not taste.** Relabelling the five
existing members to "Komponenter i Scorecard" would have produced a five-member
components list — which is precisely the completeness defect idx-26c *closed* in the
sibling by adding the two missing segments. The repair would have imported a sibling
entry's already-closed defect. Same class as §9.10's finding, one file over: the repair
is the first place a defect recurs, including defects closed elsewhere.
What was written instead mirrors the sibling's *structure*: a seven-member
`**Komponenter i Scorecard**` list carrying the source's own segment names, plus a
separate one-line `**Du konfigurerer**` carrying only what the sources say the user sets.
The framing defect disappears because the members are no longer claimed to be
configurable; the name drift and the 26f-class unsourced specificity
(`Statistikk, distribusjoner, bias-indikatorer`, `Accuracy, error rates, fairness metrics`)
disappear with the paraphrases.
**Writing a parallel list in a second file creates a new diff surface, so the deviations
had to be written down.** Three were deliberate: source order and bullets rather than the
sibling's numbering (two numbered lists in different orders would make "item 2" denote
different members in the two files); a source-tight rendering of *causal insights* rather
than the sibling's looser `Causal vs correlational relationships i features`, which §9.10
adjudicated a recognisable paraphrase and which is **not** reopened — but importing a
looser paraphrase into *fresh* text under a stamp being dated today would be the class
this programme removes; and `target-verdiene` in every member, where the sibling leaves
one member in bare English. That last one is §9.10's own correction applied prospectively
rather than after the fact.
**The compound-form audit ran this time.** `80e17ec` exists because a coherence read
checked content and not compound forms. The new text was audited on that axis explicitly
before commit: `target-verdiene` ×3, `fairness-målverdiene` ×2, no hybrid, and the two
English metric names are the source's own examples already mirrored in the YAML block
below.
**Three neighbours booked, one of them in the file §9.9/§9.10 had already canonicalised.**
- **idx-26g** — §1 omits the public-preview banner both sources carry, while the file
header declares `Status: GA`; and the file's fifteen listed sources do not include the
page grounding the enumeration just written.
- **idx-26h** — in `transparency-documentation-standards.md`, the Status line recommends
what the source explicitly does not: *"anbefalt for production use"* against *"we don't
recommend it for production workloads."* The Customization block beneath it carries two
unsourced items. Both sit under the same stamp idx-26d dated and idx-26f widened.
**idx-26f's resolution states that no exception is written down anywhere for that stamp
to remain honest — true of the enumeration it measured, and not of the rest of the
block the stamp closes.** Two entries verified the list; neither read to the end of what
the stamp reaches.
- **idx-26i** — a finding about the *check*. idx-26f recorded a cross-file grep for
`(gender, ethnicity, age)` returning nothing; the same specificity exists in the corpus
as `(kjønn, etnisitet, alder)`. **An English-only grep cannot see a bilingual corpus's
Norwegian rendering of the same claim**, so every sweep in this queue that greped
English strings alone has an unmeasured Norwegian half. Booked as a candidate, not a
confirmed defect — the attribution there is to Foundry tooling, not a scorecard segment.
**Queue state: 8 open, 9 resolved.** Suite 1047/1047.
## Appendix A — the 15 admitted proposals, hand-verified
Every proposal the classifier (§4 + context condition) admitted over the whole
`not_grounded` population, with the verdict that produced §3's table. A later run
that admits a 16th can diff against this list; without it, "9 of 15" is an
unreproducible claim.
| # | File:line | Swap | Type | Verdict |
|---|---|---|---|---|
| 1 | `agent-orchestration/agent-evaluation-testing-frameworks.md:56` | `4.1` → `5` (`gpt-4.1-mini` → `gpt-5-mini`) | number | **unverified** — model identifier; the result is a real model name, but not checked against the source. Not applied. |
| 2 | `api-management/logging-analytics-ai-traffic.md:49` | `2023-09-01` → `2025-09-01` | iso_date | **correct** — ARM `loggers@` api-version bump |
| 3 | `azure-ai-services/translator-document-translation.md:162` | `40` → `10` (MB) | number | **unverified** — matched on the unit `MB`, but sync/async limits differ; metric-crossing risk. Not applied. |
| 4 | `monitoring-observability/log-analytics-kql-ai-queries.md:617` | `2025-09-01` → `2026-04-01` | iso_date | **correct** — `api-version=` inside a KQL string literal |
| 5 | `responsible-ai/responsible-ai-training-awareness.md:77` | `900` → `901` (`AI-900` → `AI-901`) | number | **wrong** — certification identifier mutilated |
| 6 | `bcdr/cost-analysis-dr-configurations.md:120` | `4` → `5.1` (`GPT-4o` → `GPT-5.1o`) | number | **wrong** — model identifier mutilated |
| 7 | `bcdr/multi-region-azure-openai-deployment.md:316` | `2024-06-01` → `2024-10-01` | iso_date | **correct** — `api-version=` in a management URL |
| 8 | `ai-security-engineering/ai-prompt-shield-network.md:309` | `2024-09-01` → `2024-09-15` | iso_date | **correct** — Content Safety api-version |
| 9 | `ai-security-engineering/content-safety-filter-calibration.md:277` | `2024-10-01` → `2024-10-21` | iso_date | **correct** — Azure OpenAI api-version in a curl sample |
| 10 | `ai-security-engineering/jailbreak-prevention-production.md:305` | `2024-09-01` → `2024-09-15` | iso_date | **correct** — Content Safety api-version in a curl sample |
| 11 | `cost-optimization/observability-cost-reduction.md:114` | `3.7.5` → `3.4.0` (Java Agent) | version | **wrong** — a downgrade; the quote's version is not the claim's referent |
| 12 | `cost-optimization/vector-storage-cost-optimization.md:266` | `2025-09-01` → `2026-04-01` | iso_date | **correct** — AI Search api-version |
| 13 | `cost-optimization/vector-storage-cost-optimization.md:318` | `2024-02-01` → `2024-10-21` | iso_date | **correct** — embeddings api-version |
| 14 | `performance-scalability/response-chunking-strategies.md:56` | `4` → `5.1` (`gpt-4o` → `gpt-5.1o`) | number | **wrong** — model identifier mutilated |
| 15 | `performance-scalability/token-per-second-optimization.md:295` | `2024-12-01` → `2025-01-01` | iso_date | **correct** — Azure OpenAI api-version |
**9 correct · 4 wrong · 2 unverified.** All nine correct are `iso_date`; every
wrong one is a digit inside a product, model or certification identifier, where a
matching prefix (`AI-`, `gpt-`, `Agent `) satisfies the context condition while
the digit is part of a name rather than a quantity. The two unverified are also
`number` and are excluded by the same class rule — verifying them costs a source
fetch each and would move the total to at most 11.
## Appendix B — the 8 §4b status proposals, hand-verified
Every status proposal the classifier admits over the whole `not_grounded`
population, judged against the cited source. Same discipline as appendix A: a
later run that admits a ninth can diff against this list, and "5 of 8" is
otherwise an unreproducible claim.
Reproduce: `node scripts/kb-eval/classify-fix-ops.mjs --threshold 1 --write`, then
read the `proposal.type === 'status'` items in
`scripts/kb-eval/data/r11-pilot-classification.json`.
| # | File:line | Swap | Verdict |
|---|---|---|---|
| 1 | `agent-orchestration/foundry-agent-service-ga.md:68` | `**Preview**` → `**GA**` | **correct** — quote: "hosted agents are generally available"; the row's subject is *Hosted agents* |
| 2 | `agent-orchestration/foundry-agent-service-ga.md:72` | `**GA**` → `**Preview**` | **correct** — quote: "Trigger an agent by using Logic Apps (preview)"; the row's subject is the Logic Apps trigger |
| 3 | `ai-security-engineering/security-copilot-integration.md:83` | `Public Preview` → `GA` | **correct** — quote: "Email and collaboration alert triage capabilities are already generally available (GA)"; the row is the phishing/email triage agent |
| 4 | `ai-security-engineering/security-copilot-integration.md:93` | `GA` → `Preview` | **correct** — quote is from the agent's own doc page: "This feature is in public preview" |
| 5 | `ai-security-engineering/entra-agent-id-zero-trust.md:439` | `Public Preview` → `GA` | **correct** — quote: "The Microsoft Entra Agent ID platform is now generally available"; the row's subject is Entra Agent ID (kjerne) |
| 6 | `ai-security-engineering/security-copilot-integration.md:94` | `GA` → `Preview` | **unproven** — the quote's `(Preview)` marker belongs to *Identity Risk Management Agent*, not to Access Review Agent. The judge's prose `reason` does support preview from a what's-new post, so the outcome is plausibly right; the cited evidence does not establish it. Not applied. |
| 7 | `data-engineering/onelake-data-strategy.md:198` | `GA` → `Preview` | **wrong** — the quote says `Lakehouse \| Yes \| GA`. "Public preview" in the same quote belongs to *Eventhouse*. Killed by candidate A and B. |
| 8 | `ai-security-engineering/owasp-llm-top10-azure-mitigations.md:79` | `GA` → `Preview` | **wrong** — the source marks only *Response Completeness* as preview; the row covers the groundedness/completeness pair, so the edit makes the groundedness half false. The correct fix is to split the row (O2/O3). Killed by candidate B. |
**5 correct · 1 unproven · 2 wrong.** All five correct ones carry the source
phrasing on the row's own subject; all three defects are the referent gap
described in §8. No proposal was applied — §4b output is a human review list.
## Appendix C — the 17 O2 candidates
Every proposal the prose classification admitted over the 46, with the two
human-judged conditions as the classifier left them. Same discipline as
appendices A and B: without this list, "17 of 46" is an unreproducible claim.
`confirm` = the classifier marked the condition `human_must_confirm`, i.e. it
could not settle it on the evidence available and is handing it over — not a
defect, it is the contract.
The full records, including each proposal's verbatim file text and the exact
remainder, are in `scripts/kb-eval/data/r11-o2-returns/`.
| # | File:line | Failing sub-assertion | Cond 2 | Cond 3 | Confidence |
|---|---|---|---|---|---|
| 7 | `ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md:79` | The third table row presenting `prebuilt-document` (General Document) as a current basic model — the source states the general document model is no l… | yes | confirm | medium |
| 8 | `ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md:191` | Two sub-assertions: (a) `eller managed compute cluster` as an alternative compute option — the how-to page and the monitor schema require a Spark poo… | yes | yes | high |
| 9 | `ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md:218` | The second sentence, `Støtter også drift detection for grounding data i RAG scenarios.` — the canonical observability page lists only Evaluation, Mon… | yes | confirm | medium |
| 14 | `ms-ai-engineering/mlops-genaiops/feedback-loops-continuous-improvement.md:555` | Two sub-assertions: 'SharePoint' as a feedback storage service, and the word 'Automatically' in 'Automatically add reviewed samples to training set' … | confirm | yes | medium |
| 17 | `ms-ai-engineering/rag-architecture/rag-caching-optimization.md:254` | The three-band rubric (0.1-0.2 strict / 0.3-0.5 balanced / 0.6-0.8 liberal) — undocumented, and the two upper bands contradict the source's warning t… | yes | yes | medium |
| 18 | `ms-ai-engineering/rag-architecture/rag-caching-optimization.md:29` | The list item 'Azure AI Search (built-in caching av search results)' — the source states each query operates on the current index view with no cachin… | confirm | yes | medium |
| 19 | `ms-ai-engineering/rag-architecture/rag-caching-optimization.md:297` | The bullet "Automatic indexing av vectors" — the judge states vector indexes must be declared explicitly in the indexing policy (only at container cr… | confirm | yes | medium |
| 26 | `ms-ai-governance/responsible-ai/transparency-documentation-standards.md:117` | Items 4 (Error analysis) and 5 (Counterfactual analysis) are listed as Responsible AI Scorecard components, but the canonical scorecard segment enume… | confirm | yes | medium |
| 27 | `ms-ai-governance/responsible-ai/transparency-documentation-standards.md:426` | The 'Chat interface' row (a "Powered by AI" badge in the chat window) and the 'Plugin actions' row (confirmation prompts before sensitive actions) ar… | yes | confirm | medium |
| 28 | `ms-ai-governance/responsible-ai/transparency-documentation-standards.md:83` | The second and third bullets — that Hugging Face model cards are synchronised automatically, and that a template exists for generating model cards fo… | yes | confirm | medium |
| 31 | `ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md:139` | The legalHold object is given a field named "enabled"; the Storage API's LegalHold model exposes tags and hasLegalHold, so the literal field name "en… | yes | confirm | medium |
| 33 | `ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md:211` | Two parts attributed to Defender for Cloud AI Security Posture Management that the AISPM page does not support: the discovery mechanism "(via Azure R… | yes | confirm | medium |
| 36 | `ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md:38` | The "/indiscriminate" qualifier, which extends the Tampering placement and the Critical severity to indiscriminate data poisoning; the source gives t… | yes | confirm | medium |
| 38 | `ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md:396` | The listing of 'DSPM for AI - Unethical behavior in AI apps' and 'DSPM for AI - Protect sensitive data from Copilot processing' as Insider Risk Manag… | yes | confirm | medium |
| 40 | `ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md:133` | The third bullet '**CVE severity mapping**' presented as a category of alert that dependency scanning generates; severity is a property of an alert, … | yes | confirm | medium |
| 42 | `ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md:200` | The second bullet presenting the HuggingFace Registry as a Microsoft channel for verified models with provenance tracking; the source calls it a comm… | yes | confirm | medium |
| 45 | `ms-ai-security/cost-optimization/semantic-caching-patterns.md:436` | The '/West' half of the region pair, i.e. the standing implication that Azure OpenAI can be deployed in Norway West. | yes | confirm | medium |
**2 affirmative on both conditions (8, 17) · 15 needing a human call.** Item 14
additionally carries a machine flag: its remainder recapitalises rather than
deletes (§9.1, V2b), so it is a text change and must be reviewed as one.
⚠️ **The `cond 2` column above is the classifier's claim, not a verified fact —
and it has now been corrected.** All 17 rows have had the whole-file check (idx 8
and 17 in §9.2, the other 15 in §9.3). Read the column together with §9.3's table,
which overrides it:
- **Contradicted, not ratifiable as written:** idx 8, 14, 18, 26, 27. All but 18
admit a reduced subtraction; every reduction is a new remainder string and must
be re-run through `check-o2-returns.mjs` before it counts as verified.
- **Operator call:** idx 33, 36 — deleted content survives elsewhere in the file
without the remainder becoming false.
- **Clean on the whole-file dimension:** idx 7, 9, 17, 19, 28, 31, 38, 40, 42, 45.
Cond 3 is still `human_must_confirm` for most of them; clean here means cond 2
only. **Idx 19 is a narrowing, not a resolution** — its cond 2 stands at
`human_must_confirm`, with the residual reduced to an omission question (§9.3).
**The verified score remains 1 of 46 — idx 17 alone.** The sweep removed no
member and added none; it corrected four rows and left idx 19 as the strongest
candidate for the next ratification.
Reproduce the tally and the machine checks:
`node scripts/kb-eval/check-o2-returns.mjs`.
### 9.12 idx-26h closed — the entry overstated its own coverage, and the stamp set the scope
Run 2026-08-09. Two locators under the `2026-08-03` stamp that §9.9 dated and §9.10
widened: a Status line reading *"anbefalt for production use med awareness om
SLA-limitations"*, and two Customization bullets the entry called unsourced.
**The Status finding held, and held harder than booked.** Three MS Learn pages —
`concept-responsible-ai-scorecard`, `how-to-responsible-ai-scorecard` and
`how-to-responsible-ai-insights-ui` — carry the identical banner: *"This preview version
is provided without a service-level agreement, and we don't recommend it for production
workloads."* The file recommended what the source explicitly does not. Replaced with the
source's own wording rather than deleted: preview status is the section's one
operationally decisive fact and nothing else in the section carries it.
**The second finding was false, and the entry was the one overstating.** It claims
*"neither source page states either"*. That holds for the two pages the entry measured
and fails against a third. `how-to-responsible-ai-insights-ui` documents scorecard
customisation directly — step 4: the Data analysis section *"enables cohort analysis"*
and you select *"features of interest to identify your model performance on their
underlying cohorts"*; step 1: *"an optional description about the model's functionality,
data it was trained and evaluated on, architecture type"*.
| Bullet as written | Entry's verdict | Measured verdict |
|---|---|---|
| "Cohort analysis: Disaggregated performance for identified risk groups" | unsourced specificity | grounded; only *"identified risk groups"* is not the source's — it says *features of interest*, and *sensitive features* belong to a different configuration step |
| "Narrative sections: Fritekst-forklaringer for decisions og mitigations" | unsourced specificity | the free-text field exists; *"decisions og mitigations"* is its wrong referent, and *"Narrative sections"* is not source vocabulary either |
So the repair was **reformulation, not deletion** — deleting would have removed
documented behaviour. §9.11's lesson recurred with the entry itself as the overstater:
*the page a claim came from is a different object from the pages an entry happened to
measure.* An entry is a hypothesis with an evidence list, and its evidence list is not a
census of the sources.
**The stamp, not preference, set the scope.** Re-dating `Verified` to 2026-08-09 asserts
the whole section is verified, so every measured-unsourced claim under it had to be
closed or the new date would knowingly reproduce the §9.10 class. Reading the section
end-to-end rather than to the entry's anchors surfaced two more: a third Customization
bullet (*"max error rate per subgroup"* — targets are set on the metric, and fairness
targets capture the difference or ratio **across** subgroups, never a per-subgroup
maximum) and a role-table row (*"Compliance officers | … (EU AI Act, sector-specific
regler)"* — neither the role name nor the parenthetical is in any source). **A stamp is a
scope-setting instrument: what it closes is what the edit owes, and an entry's anchors
are a lower bound on that.**
**The stamp was wrong when set, not stale.** The Status line was false on 2026-08-03, the
day idx-26d dated the stamp over it. Neither idx-26d nor idx-26f detected it because both
measured the seven-member enumeration and neither read to the end of the block the stamp
reaches. Staleness and falsity look identical in a dated stamp, and only re-measuring
tells them apart.
**The ratified preview was wider than the ratified label.** The option that closed the
role row carried a preview illustrating a row that *merged* `Compliance officers` into
`Auditors`; the label said *"rett raden"*, singular. The label was taken as the ratified
object and the `Auditors` row — grounded verbatim in the concept page — was left standing.
Following the illustration would have deleted correct content on the strength of a mockup.
**A ratification question can carry a false premise, and the operator ratifies it anyway.**
The source-list option presented in this session stated that the edit would close idx-26h
**and** idx-26g together. It does not: idx-26g is booked against a different file,
`stakeholder-communication-ai-decisions.md`, with its own source list, its own counting
claim and a `Status: GA` header finding untouched here. The renumbering was independently
correct for this file, so the edit stands — but the question was wrong, and the operator
had no way to see it. **The premise-verification duty applies to the questions you ask,
not only to the facts you act on.** Nothing in the queue, the suite or the CLI checks the
framing of a ratification prompt.
**Booked, not fixed:** `idx-26j` — the same file's footer provenance block, which no entry
has ever measured. `Total MCP calls: 5` enumerates components summing to 6; the line has
no defined referent (original generation run, or current state?); the `80% / 20%` split
**cannot be checked at all**, which is the defect rather than any particular mismatch —
read as list positions it fits neither 10/5 before nor 12/5 after, but that reading is an
assumption, and the sibling file counts the same kind of claim on a different partition
(*"15 (8 verified fra MCP, 7 baseline/code samples)"*), so no denominator can be asserted.
This edit also introduced a two-dialect compound pair (`model-` on line 101 against the
ratified `modell-` in the rewritten row) that no check can see, because the stamp claims
verification against source, not orthographic consistency.
**One locator adjudicated out of scope rather than missed.** The header carries
`**Status:** GA` at line 4 — the string idx-26g flags in the sibling as a false status
claim of this same class. It is not falsified by this edit's stamp: line 4 sits in the
bold-label header block above `## Innhold`, and the file's first `Confidence` stamp is
section-terminal at line 59, so no stamp reaches it. The real finding is that the field's
referent is undefined **corpus-wide**. Measured across `skills/*/references/`, `**Status:**`
mixes two vocabularies answering different questions — document maturity (`Established
Practice` 43, `Gjeldende` 18, `Reference` 7, `Komplett` 2) and product availability (`GA`
252, `Preview` 2, `GA — avvikles 31. mars 2029` 2) — and the corpus already carries the
precedent form for a file spanning both: `GA / Preview (varies by feature)`, 2 files. This
file spans Transparency Notes (GA), Model Cards (practice) and the scorecard (public
preview) at once. **So idx-26g's header question is not local to either file — it is a
corpus-wide field-semantics decision, and settling it one file at a time would set the
convention by accident.**
### 9.13 idx-26g closed — and §9.12's own corpus measurement was taken over the wrong population
**The entry was right about the sources this time.** Both pages it names carry the Important
banner verbatim — *"This feature is currently in public preview. This preview version is
provided without a service-level agreement, and we don't recommend it for production
workloads."* — and both titles read "(preview)". Fetched live 2026-08-09 rather than taken on
the entry's word, because §9.12's lesson was that an entry can overstate its own evidence.
**The two decisions were independent, and the entry bundled them.** The banner sits on
`concept-responsible-ai-scorecard`, which is *already* this file's Kilder #1. So the preview
fact needed **no new citation at all**; only the seven-segment enumeration needed
`how-to-responsible-ai-scorecard`. Separating them meant the ratification question could put
the a/b/c citation choice where it actually belonged instead of letting a citation decision
gate a status correction.
**Ratified:** header `**Status:** GA` → `**Status:** GA / Preview (Responsible AI scorecard)`,
and the how-to page added as Verified source #9 with 915 renumbered to 1016, `**Total
kilder**` 15 → 16. The parenthetical form was chosen over the terser `varies by feature`
after checking all six corpus files in the `X / Preview (Y)` family: in every one, the
parenthetical names *what is in preview*, never the GA surface. Renumbering was safe on
measurement, not assumption — grep over lines 1794 found **zero** prose cross-references to
source numbers. The chosen form also preserves the counting claim's pre-existing referent: it
still counts numbered entries (16 entries, 9 carrying URLs), whereas the "second URL under
Kilder #1" form would have left the file holding 16 URLs against 15 entries — the undefined-
referent class idx-26j had booked one commit earlier.
**The stamp was re-dated deliberately, because otherwise this edit would have authored the
idx-26h defect.** §1's stamp read `2026-08-04`. Adding a sentence written on 2026-08-09
underneath it would have produced text sitting under a verification date five days its senior
— *"the stamp was false when it was set, not stale."* Re-dating claims all of §1, so all of §1
was re-measured against both live pages: the seven segments map 1:1 to the how-to page's "How
to read your scorecard"; `Du konfigurerer` matches the concept page's "target accuracy and
target error rate"; the Formål bullets match its three unaddressed-needs bullets; the Verdi
bullets match "building trust and gaining their approval for deployment"; and the YAML
workflow block, the one piece that is synthesis rather than quotation, has every step grounded
in "Who should use a Responsible AI scorecard?".
**A third locator the entry did not name.** Kilder #1 read `Status: GA (public preview for
some features)` — two lines below its own title ending in "(preview)". The banner covers the
whole feature, so this was false rather than merely imprecise. Anchors are a lower bound,
tenth instance.
#### The correction: §9.12 measured `**Status:**` over the wrong population
§9.12 (and idx-26j's fifth locator, and the STATE hand-off built on both) reported the field's
value space as `GA` 252 · `Preview` 2 · `Komplett` 2 alongside `Established Practice` 43 ·
`Gjeldende` 18 · `Reference` 7. Those counts were taken over **every** `**Status:**` occurrence
under `skills/*/references/` — 434 of them — which silently merges the line-4 header field with
**47 section-level `**Status:**` lines inside file bodies**. Two different fields, two different
referents, one number.
Re-measured on the header field alone (`FNR<=10`): **387 occurrences over 389 files, 69
distinct values.** Corrected: `GA` **249** (not 252) · `Preview` **1** (not 2) · `Komplett`
**1** (not 2). `Established Practice` 43, `Gjeldende` 18 and `Reference` 7 hold as stated.
The two-vocabulary finding **survives the correction** — document maturity and product
availability genuinely are mixed in one field, and that observation was right. What does not
survive is the *consequence*: **"settling it one file at a time would set the convention by
accident" is false.** The header field carries 69 distinct values over 387 files, roughly 44 of
them appearing exactly once, and 35+ are compound GA/Preview forms. There is no single
convention available to be set by accident; an accurate compound value *joins* an established
practice. The claim that survives is narrower and does not need a corpus-wide decision to
license the edit: **this file had already chosen the product-availability axis for its header,
and the edit made its chosen axis truthful.** Choosing an axis for 389 files was never on the
table.
This is the §9.11 defect class turned on the *measurement* rather than the resolution — a
premise measured over a population wider than the field it was reporting on. The cheapest
ground-truth check (`FNR<=10`) was cheaper than the corpus-wide operator decision it would
have triggered.
#### Booked rather than repaired
- **idx-26k (new).** This file's footer reads `**MCP calls gjennomført**: 5 (3 docs_search, 2
docs_fetch, 1 code_sample_search)` — components summing to 6. Same class as idx-26j's first
locator in the sibling, different wording and different field name. That **confirms as
measured** what idx-26j could only call likely: the footer form is replicated. Not repaired
here, because the referent is undefined in the same way (historical generation run, or
current provenance?) and idx-26j owns the form decision. Changing 5 to 6 would pick a
referent by stealth.
- **idx-26j re-scoped.** Its fifth locator is no longer blocked on a corpus-wide decision: the
form is now precedent. Its header (`**Status:** GA`, spanning Transparency Notes GA, Model
Cards practice, scorecard preview) is false the same way, and a divergence from the ratified
form would now be *introduced* by that entry rather than inherited.
- **Untouched, on precedent:** `**Last updated:**` — verified that zero of the eleven prior
ratified corpus edits modified that field. `Verifisert: 2026-02` on sources 18 — those pages
were not re-fetched.
**Queue:** 8 open / 11 resolved. **Suite:** 1047/1047.
---
### 9.14 idx-26j and idx-26k closed — the footer's referent was already mixed, by a ratified edit
Two entries, one form decision. idx-26k exists only because idx-26j predicted the defect would
recur, and idx-26k's own text forbids closing it alone: *"idx-26j owns that decision and this
entry is its second measured instance."* So the unit of work was the form, not the file.
#### What made the referent decidable
The entry framed it as a choice between two readings of the footer provenance block — a record
of the original generation run, or a description of the file's current provenance — and said no
reader and no check could adjudicate it. Two measurements collapsed that.
**First: the block already carries a mixed referent, and a ratified G7 edit mixed it.** `99675d5`
(idx-26h, the same day) changed `**Unique sources:** 15 URLs` → `17 URLs` while leaving
`**Total MCP calls:** 5` and `**Confidence:** 80%` untouched. One field was maintained as current
provenance; two were left frozen. The precedent inside the block was therefore already set, and
set by the programme itself.
**Second: no generation log exists.** Grepped `scripts/` and `docs/` for any artifact recording
MCP calls per file — only the queue and this document mention them, both as analysis rather than
as a run record. That matters more than which reading wins: it removes `5 → 6` from the option
space under **both** branches, not just one. A number that cannot be checked under the historical
reading either is not a historical fact, it is an unsourced claim wearing a date's clothes.
#### The corpus measurement — and a count corrected before it was used
23 footer lines in 23 files state a total followed by a per-tool enumeration. A regex pass
returned 11 inconsistent. Hand-classification returned **8**:
| | count | |
|---|---|---|
| internally consistent | 10 | components sum to the stated total |
| **internally inconsistent** | **8** | stated total ≠ sum |
| ambiguous | 1 | `foundry-agent-service-ga.md:560` — reading *is* the question |
| not checkable | 4 | no per-tool numbers to sum |
The three regex artifacts were `translator-document-translation.md:402` and
`rag-cost-optimization.md:558`, which spell the arithmetic out (`4 (docs_search) + 3 (docs_fetch)
= **7**`) and are correct, and `foundry-agent-service-ga.md:560`, whose `4` inside a descriptive
clause the pattern read as a component. This is the §9.13 lesson applied one step earlier: the
number was going into an operator-facing ratification label, so it was hand-verified **before**
the question was asked rather than after the answer came back.
**All eight inconsistencies run the same direction** — stated total *below* the sum, never above.
A transcription error would scatter. A one-sided bias is evidence the total was never transcribed
from a run at all.
#### Ratified and applied
Current provenance, one referent for the whole block. `**Total MCP calls:**` **deleted** in both
files — unmeasurable, with no artifact that could ever adjudicate it. `**Unique sources:** 17
URLs` **kept**, verified true (17 numbered entries, 17 distinct URLs). The uncheckable
`80% Verified (MCP), 20% Baseline` became a structural count with a defined denominator. The
sibling needed only the deletion: its `**Total kilder**: 16 (9 verified fra MCP, 7
baseline/code samples)` was set by idx-26g and re-verified here by counting, and its
`**Confidence vurdering**` is a banded per-topic list, not a percentage split.
Two locators closed on other grounds. The **header** (`**Status:** GA` →
`GA / Preview (Responsible AI scorecard)`) was applied by precedent, not re-asked — §9.13 settled
the form. It needed no new citation: this file already states the preview fact in its own body at
line 129, under a stamp already dated 2026-08-09. The parenthetical's *completeness* was verified
rather than assumed — all seven `preview` occurrences grepped; six are the scorecard, and line 56
is content inside a Transparency Note example, not a surface this file makes an availability claim
about. The **dialect pair** (line 101 `model-` → `modell-`) aligned to the operator-ratified string
at line 109.
#### The replacement authored the defect, and was caught before commit
The first draft of the Confidence line read `12 av 17 kilder verifisert via MCP (nr. 112), 5
baseline (nr. 1317)`. It would have **asserted that source 7 is MCP-verified** — and source 7 is
itself labelled `(Status: Baseline — …)`. Writing it would have been the exact defect class the
edit was closing, in the fix. The shipped wording counts block membership only
(`17 kilder — 12 oppført under «Verified sources», 5 under «Baseline sources»`) and asserts nothing
about any individual source, which leaves the finding below genuinely open instead of quietly
asserted over.
#### Booked rather than repaired
- **idx-26s (new).** The two source blocks contradict themselves in exactly two places, in
*opposite* directions: #7 sits under **Verified sources** but is labelled `(Status: Baseline …)`;
#17 sits under **Baseline sources** but is labelled `(Status: Verified 2026-02 …)`. Not repaired,
because the direction of the fix is undecided and one direction renumbers the list. The shipped
count survives either way: by block membership 12/5, and swapping both mislabelled entries gives
verified {16, 812, 17} = 12 and baseline {7, 13, 14, 15, 16} = 5. **Eleventh instance of the
defect sitting next to the edit.**
- **idx-26l … idx-26r (new, 7).** One entry per remaining affected file — 6 inconsistent plus the
1 ambiguous — each with a verbatim anchor verified unique before being written. Operator ratified
per-file entries over a single class-wide entry precisely because the queue schema carries one
`file` per entry: a class-wide entry would have left 20 files invisible to
`check-g7-queue.mjs`, which is the `795d494` failure repeated at class scale.
- **`ai-red-team-operations-practical.md:87`** — the corpus's only other `model- og`. Left
untouched; the dialect fix was scoped to the one file.
- **Untouched, on precedent:** `**Last updated:**` — twelve *prior* ratified corpus edits, none of
which moved that field; this one makes thirteen.
#### No re-dating decision was required — and that was adjudicated before writing
§9.13 established that re-dating a stamp is a scope decision taken up front, so the absence of one
had to be *proved*, not assumed. Every stamp in the file sits at lines 59, 95, 131, 160, 312, 378 —
all section-terminal, all before 378. Nothing after 378 carries a stamp except the inline
`*(Verified MCP 2026-06-19)*` inside source 8, which is part of a source entry. Therefore: line 4
sits above the first stamp, so the header edit falsifies none; the Kilder section and footer sit
under no stamp at all; and line 101 sits under the section-3 stamp at 131, which is already dated
2026-08-09 **and** claims verification against source rather than orthographic consistency — the
entry says so itself. An orthographic alignment neither falsifies it nor requires re-dating it.
**Queue:** 14 open / 13 resolved. **Suite:** 1047/1047.
#### Scope addendum — the ground is unverifiability, not arithmetic
Written the same day, because the section above understated what the ratified ground covers. The
operator ratified deletion because the line is **unmeasurable** — no generation log exists to
adjudicate it. That property holds for all 23 lines, not only the 8 that also fail their own
arithmetic. Internal inconsistency was the *evidence* that the number was never transcribed from a
run; it was never the *reason* for deleting it.
But the ratification question's own option text defined "affected file" as the inconsistent set, so
booking stopped at 7. **No machine sees the framing of a question** — this is the same class as
§9.13's false premise, this time in a question written here rather than inherited.
So **14 files carry the condemned line with no queue entry**: 10 internally consistent, 4 not
checkable (a total with no per-tool numbers, so the arithmetic test cannot be applied — but the
total is just as unverifiable). They are not excluded on merit. They are unbooked because expanding
past the ratified 7 is a scope decision belonging to the operator, and it is raised as open question
**#17** rather than taken here. Both lists are enumerated in idx-26j's resolution so they are
greppable rather than prose.
#### Two counts this edit invalidated
- **The footer population is now 21 lines, not 23.** Two were deleted here, both from the
inconsistent bucket: **10 consistent · 6 inconsistent · 1 ambiguous · 4 not checkable.**
- **The binding `**Status:**` header counts were already stale before this session.** They read
`GA` 249 over 387 occurrences / 69 distinct — but that measurement was taken inside the idx-26g
session *before* `96639f7` was applied. Ground truth now: **387 occurrences, 70 distinct, `GA`
247**, and `GA / Preview (Responsible AI scorecard)` at 2. The chain is exact: 249 → 248 (26g) →
247 (this edit); distinct rose 69 → 70 when 26g first created that string, then held, because this
edit added a second instance of a value that already existed.
**Anchor uniqueness for idx-26s was verified after it was written, not before** — named here rather
than quietly fixed. The gate proves *presence* (`text.includes`), never uniqueness, so a non-unique
anchor passes silently. All three occur exactly once.
---
## §9.15 — idx-26l…idx-26r: den ratifiserte footer-formen anvendt på sju filer, og målingen som grunnla den viste seg å være halv
Andre batch av footer-klassen. Formen fra §9.14 ble anvendt uendret på de sju bokførte
filene, og alle sju er lukket. Det som gjør denne økten verdt å lese er ikke de sju
slettingene — de var mekaniske — men de to tingene som falt ut av å sjekke *hva som sto
igjen* og *hvor stor klassen faktisk er*.
### Anvendt
Sju filer, åtte slettede linjer (én fil hadde linja som eget avsnitt og mistet også
blanklinja under). Ingen erstatningstekst noe sted: formen sier *slett*, ikke *rett*.
| Entry | Fil | Slettet | Beholdt, og verifisert ved telling |
|-------|-----|---------|-------------------------------------|
| 26l | `feedback-loops-continuous-improvement.md` | `**MCP calls:** 6` | ingen tellinger igjen (kun en dato) |
| 26m | `rag-caching-optimization.md` | `**MCP calls:** 6` | `9 unike Microsoft Learn URLer` ✔ (9 entries, 9 URL-er) |
| 26n | `agent-to-agent-a2a-protocol.md` | `**MCP calls:** 4` | `10 unike URLer` ✘ — se under |
| 26o | `agent-to-agent-communication.md` | `**MCP calls**: 4` | `7 unique URLs` ✔ som tall, ✘ som proveniens |
| 26p | `agent-365-governance-and-deployment.md` | `**Total MCP calls:** 4` | `7 Microsoft Learn-artikler` ✔ |
| 26q | `azure-cost-management-ai.md` | `**MCP calls:** 4` | `8 unique Microsoft Learn URLs` ✔ |
| 26r | `foundry-agent-service-ga.md` | `**MCP calls:** 4` | `9 primærkilder` ✔ (9 entries, 9 URL-er) |
### idx-26r trengte aldri form-beslutningen den var bokført for
Entryen sa at linja «ikke kan adjudiseres uten å avgjøre om feltet teller kall eller
runder». Begge lesningene handler om **om tallet er riktig**. Under en form som sletter
i stedet for å rette, er riktighet ikke i mulighetsrommet: linja går fordi den er
uverifiserbar, og den er uverifiserbar under begge lesningene. Åpent spørsmål **#15** er
dermed **moot for slettingen** — samme figur som lukket 26j, der et *fravær* avgjorde et
valg ingen av grenene kunne avgjøre.
Hvorfor dette ikke var å gjenåpne en ratifisering, sjekket i stedet for påstått: `idx-26k`
bærer en eksplisitt sperre («Do not close this one in isolation»), `idx-26r` bærer ingen.
Scope-addendumet til 26j sier at grunnen er uverifiserbarhet, «a property that holds for
ALL 23 lines». STATEs advarsel om at fila «ikke er mekanisk» bar videre framingen fra
entry-teksten, skrevet 14:11 — ti minutter *før* addendumet (14:21) som oppløste den.
**Entryens framing overlevde addendumet som svarte på den.**
### Det formen ikke dekket: tre funn i linjene som ble BEHOLDT
Formen sier at sjekkbare tellinger «beholdes og verifiseres». Den forutsatte at de ville
passere, og sier ingenting om hva man gjør når én ikke gjør det. Tre gjorde ikke det, og
alle tre er **bokført, ikke reparert**:
- **`idx-26t`** — `agent-to-agent-a2a-protocol.md`: `10 unike URLer` er sann lest som
*kilder* og usann lest som *URL-er* (10 nummererte entries, 11 distinkte URL-er; entry 8
bærer to). Samme udefinerte referent, nå inne i linja formen ba oss beholde. **Avviket
går samme vei** som alle åtte MCP-avvikene.
- **`idx-26u`** — `azure-cost-management-ai.md`: `**File size:** ~14 KB` er **falsifiserbar**,
ikke bare uverifiserbar, fordi git *er* artefaktet som avgjør den. Målt ved hver commit
som har rørt fila: 17388 → 17388 → 17963 → 17965 → 18554 byte. Linja står ordrett i den
*eldste* versjonen, så påstanden var alt gal da fila kom inn i repoet og har **aldri**
vært sann på noe punkt repoet kan observere.
- **`idx-26v`** — samme fil: `80% Microsoft-verified, 20% domain-specific` er 26j-lokator-4-klassen.
Den ble **ikke** lukket i samme edit, og avviket fra 26j ble målt før det ble besluttet:
26j kunne bytte prosenten mot en strukturell telling fordi *den* fila hadde en partisjon
(12 Verified / 5 Baseline). Denne har ingen — alle 8 rader i URL-tabellen er `Verified`,
og seksjonstabellen er 3/3. Å re-uttrykke prosenten ville krevd å **finne på** en nevner,
som er nøyaktig «erstatningen forfatter defektklassen den lukker».
- **`idx-26w`** — `agent-to-agent-communication.md`: `7 unique URLs fra MCP-research`
stemmer som aritmetikk (7 entries, 7 URL-er) og feiler som **proveniens**: bare 5 ligger
på learn.microsoft.com; kilde 6 (`a2a-protocol.org`) og 7 (`jsonrpc.org`) kan ikke ha
kommet fra microsoft-learn-serveren. Kontrasten i samme batch gjør formen synlig — 26n
navngir «MCP-research + tavily-research» og dekker sine ikke-Learn-kilder.
### Korpusmålingen som grunnla ratifiseringen så under halvparten av klassen
Dette er øktens tyngste funn, og det gjelder ikke editene — det gjelder **tallet alle
entries i klassen siterer**.
Ratifiseringen hviler på «23 footerlinjer i 23 filer … fire feltnavn-dialekter». Etter de
ni slettingene (2 i `088cf06` + 7 her) skulle korpuset da hatt **14** igjen. En dialekt-bred
sveip finner **47 kandidatlinjer**. Håndsjekk av de tre som kunne vært kilde- snarere enn
kall-tellinger: to er falske positive (`Total MCP-kilder: 4 unique URLs`,
`MCP-kilder: 5 Microsoft Learn-dokumenter`), én er ekte under feil etikett
(`Totalt antall MCP-kilder: 3 docs_search calls + 2 docs_fetch calls = 5 MCP-kall`).
**Ground truth: 45 ekte linjer i 45 filer, i minst 16 distinkte feltnavn-skrivemåter** —
mot de fire målte. Den opprinnelige regexen var engelsk-språklig og så ikke `MCP-kall`,
`MCP-kall utført`, `Totalt antall MCP-kall`, `Total MCP-kall`, `MCP-kall totalt`,
`MCP-kall brukt`, `MCP-calls brukt`, `MCP call summary`, listeform (`- **MCP calls:**`)
eller overskriftsform (`### MCP Calls: 6`, `### Total MCP Calls: 4`).
**Dette ugyldiggjør ingen av de ni editene.** De sju her ble håndklassifisert som medlemmer,
og slettegrunnen — uverifiserbarhet — gjelder uansett populasjonsstørrelse. Det det
ugyldiggjør er **rekkevidden**: åpent spørsmål **#17** gjaldt «14 ubokførte filer». Den
reelle ubokførte populasjonen er **45**.
**Bøttene er bevisst IKKE oppgitt.** Et maskinelt forsøk i denne økten reproduserte nøyaktig
artefaktet hånden måtte rette sist gang: det leser `3 (search) + 2 (fetch) = 5 total` som
«oppgitt 3, sum 7». Regex sa 11 og hånden sa 8 forrige gang; en ny maskintelling har ikke
fortjent tillit her. Klassifiseringen av de 45 er **håndarbeid som ikke er gjort**.
### Gjenbrukbare funn
- **En entrys framing kan overleve addendumet som oppløste den.** 26r var bokført som
blokkert av et spørsmål et dokument skrevet ti minutter senere hadde gjort irrelevant.
Sjekk om sperren er *eksplisitt* (26k har en, 26r har ingen) før du behandler prosa som
en sperre.
- **«Behold og verifiser» forutsetter at verifiseringen passerer.** Tre av sju beholdte
linjer feilet. En form som ikke sier hva som skjer ved feil, har et hull der defekten
flytter inn.
- **Falsifiserbar er ikke det samme som uverifiserbar, og git kan være artefaktet.** `~14 KB`
kunne avgjøres; `MCP calls: 5` kunne ikke. Samme footer, to helt ulike bevissituasjoner.
- **En presedens som «lukk det i samme edit» må testes mot om forutsetningen finnes.**
26j kunne det fordi fila hadde en partisjon. Å kopiere handlingen uten forutsetningen
ville produsert en oppdiktet nevner.
- **Den dialekt-blinde regexen rammer korpus-målinger, ikke bare enkelt-editer.** Et tall
som er sitert ordrett i sju køoppføringer var målt over halve populasjonen. Ingen av de
sju kunne oppdage det; de siterte hverandre.
### §9.15 addendum, samme dag — «45» var et GULV, og min egen sveip bar defekten jeg nettopp navnga
Rettelse skrevet før økten ble avsluttet, ikke overlatt til neste. Avsnittet over
konkluderte med **45 ekte linjer i 45 filer** og kalte det ground truth. Det var
det ikke. Sveipen som produserte tallet krevde at **feltetiketten selv inneholder
`MCP`** — og en telling hvis etikett mangler ordet er usynlig for den. Det er
nøyaktig den dialekt-blindheten dette avsnittet har som hovedlærdom, i sveipen
som skulle rette den.
En etikett-fri sveip (linjer som nevner MCP-kall **og** bærer et tall, uansett
hva feltet heter) finner fem til. Håndklassifisert:
| Linje | Dom |
|-------|-----|
| `copilot-studio-nlp-configuration.md:553` | falsk positiv — metodikk-punkt, ingen telling |
| `m365-copilot-plugins-ecosystem.md:449` | falsk positiv — dato, ikke kalltelling |
| `chain-of-thought-prompting.md:500` | **EKTE** — `**Totalt:** 4 MCP-kall`, etikett `**Totalt:**` |
| `genaiops-llm-specific-practices.md:383` | **EKTE** — `**Totalt:** 18 kilder, 8 MCP-kall` (3+3+2 = 8 ✔) |
| `inferencing-optimization-caching.md:1029` | **EKTE, NY UNDERKLASSE** — se under |
I tillegg bar `ai-risk-taxonomy-classification.md:458` en overskrift uten siffer
(`### MCP Calls Summary`) og ble derfor aldri vist av noen linje-sveip. Under den
står klassen i **BLOKK-FORM**: `microsoft_docs_search: 3 calls` /
`microsoft_docs_fetch: 2 calls` / `microsoft_code_sample_search: 1 call`, uten
oppgitt total. **Ekte medlem, en strukturell form ingen linjebasert telling kan
enumerere.**
**Korrigert: ≥ 48 ekte linjer, og tallet er et GULV, ikke en måling.** De tre
formene som slapp unna er (a) telling under en etikett som ikke nevner MCP
(`**Totalt:**`), (b) telling i blokk under en sifferløs overskrift, (c) telling
gjentatt i prosa.
#### To funn som endrer hva neste økt må gjøre
**1. Slettingen er ikke selv-fullførende.** `inferencing-optimization-caching.md`
oppgir `12` i footeren (:1005) og **gjentar det i prosa** på :1029: «basert på 12
MCP-kall til offisiell Microsoft-dokumentasjon». Å slette footerlinja ville latt
tallet leve videre i samme fil. Kontrollert for denne økten: **ingen av de sju
filene har slik rest** — verifisert med et dialekt-bredt søk etter editen. Men
formen må heretter kreve et rest-søk per fil, ikke bare en linjesletting.
**2. «Alle avvik går samme vei» har nå en kandidat-motsigelse.** Det funnet —
åtte av åtte oppgitt < sum, null motsatt — er sitert som evidens for at tallet
aldri ble transkribert fra en kjøring. `chain-of-thought-prompting.md:500` oppgir
**4** MCP-kall over en nummerert liste med **3** verktøykall. Det er oppgitt >
enumerert, altså **motsatt vei**. Den er ikke ferdig adjudisert (listen kan ha
utelatt ett kall), men den er ikke trygt samme-vei, og den lå utenfor
populasjonen som ga «åtte av åtte». **Retningsargumentet må re-måles over den
korrigerte populasjonen før det siteres igjen.**
#### Lærdommen, skjerpet
- **Sveipen som retter en dialekt-blind måling må selv testes for dialekt-blindhet.**
Å binde regexen til feltetiketten er samme feilklasse som å binde den til
engelsk. Testen er billig: fjern etikett-kravet og se hva som dukker opp.
- **En sifferløs overskrift skjuler en hel strukturell form.** Linje-sveip finner
linjer; en klasse som også opptrer som blokk krever at du leser under hver
overskrift som nevner feltet.
- **En sletting kan etterlate tallet i live et annet sted i samme fil.** Verifiser
fravær i HELE fila etter editen, ikke bare at linja er borte.
## §9.16 — footer-klassen håndklassifisert: 58 medlemmer, en fjerde unnslippsform, og et retningsargument som ikke overlevde som absolutt
Leveransen denne økten er **en måling og en scope-anbefaling**, ikke editer. Ingen
korpusfil er rørt. `git status` viser kun `docs/` og måledatasettet.
`≥48` fra §9.15-addendumet var riktig **som gulv**. Håndklassifisert populasjon:
**58 ekte medlemmer i 58 distinkte filer.**
### Hva som ble sveipet, og hvorfor det er fire nett og ikke ett
Forrige økt lærte at et nett bundet til feltetiketten er blindt på samme måte som
et nett bundet til engelsk. Svaret er ikke et bredere ord — det er **flere nett
som er blinde på hver sin måte**, slik at residualen til hvert nett kan leses for
hånd:
| Nett | Definisjon | Blind for |
|------|------------|-----------|
| A | `MCP` + siffer + kall-ord (`kall`/`call`/`anrop`) | tellinger uten ordet «kall»; tellinger uten `MCP` i linja |
| B | verktøy-identifikator (`microsoft_docs_*`) + siffer | tellinger som ikke navngir verktøyet |
| C | proveniensblokkens **etikettvokabular** der verdien er et tall | ikke-fete etiketter; overskriftsform |
| D | forskningshandling-telling uten kall-ord OG uten verktøy-identifikator | brødtekst-gjentakelse uten begge deler |
A gav 59 linjer / 53 filer, B 72 / 48, D 35 kandidater etter innstramming. C er
completeness-argumentet og står omtalt under.
### Den FJERDE unnslippsformen — tellingen uttrykt som «søk», ikke «kall»
§9.15 navngav tre unnslippsformer. Det finnes en fjerde, og **mitt eget nett A var
blindt for den**:
> `- **MCP-søk** — 3 søk mot microsoft-learn (2026-02-04)`
> `**MCP-søk utført:** 3 søk (microsoft-learn)`
> `- **Total searches:** 3 (Azure ML registry, AI Foundry, MLOps lifecycle)`
> `**Antall dokumenter søkt:** 4 (search queries) + 2 (deep fetch)`
Tellingen bæres av handlingsordet (`søk`, `queries`, `hentinger`, `deep reads`,
`fetches`), ikke av `kall`. `domain-specific-prompt-optimization.md:589` ble bare
reddet av at NABOLINJA nevnte `microsoft_docs_fetch` — linja selv var usynlig for
både A og B. **`model-versioning-registry-management.md` var usynlig for begge i
sin helhet**: overskriften `### MCP research summary` bærer ingen siffer, og de tre
telle-linjene under den nevner verken `MCP` eller et verktøynavn. Den er et rent
produkt av nett D.
**Tre nye medlemmer kom kun fra nett D:** `model-versioning-registry-management.md`,
`ai-impact-assessment-framework.md`, `responsible-ai-framework-overview.md`.
### Completeness-argumentet: etikettvokabularet, ikke flere ordgjetninger
Å lete etter en femte form ved å gjette flere ord er den samme feilen en gang til.
I stedet ble **hele etikettvokabularet enumerert**: hver fete etikett inne i en
proveniensblokk hvis verdi *begynner med et heltall*. Det gir **68 distinkte
etiketter over 141 linjer**, og hver enkelt lar seg plassere i én av tre klasser:
- **kall-telling** (klassen under adjudisering) — `mcp calls`, `total mcp calls`,
`mcp-kall`, `mcp-kall utført`, `mcp call summary`, `mcp-søk utført`,
`total searches`, `document fetches`, `antall dokumenter søkt`, m.fl.
- **kilde-/URL-telling** (naboklassen — «behold og verifiser»-linjene) —
`unique sources`, `unike kilder`, `totalt antall kilder`, `total sources cited`, …
- **annet** — konfidensprosent, kodeeksempler, versjon, `verification status` (80/20-klassen).
Ingen etikett falt utenfor. **Ingen femte unnslippsform for kall-klassen dukket opp.**
**Residualen som gjensto, ble testet, ikke antatt.** Alle fire nett krever et
SIFFER. En telling skrevet med bokstaver («tre oppslag») ville sluppet gjennom alle
sammen. Sveip: linjer med tallord (`to`…`tolv`, `two`…`twelve`) i MCP-/kall-kontekst
**uten** siffer → 21 treff, alle brødtekst («To API-kall per dokument»,
`### Tre MCP-komponenter`). **Null medlemmer. Blindheten er tom, målt.**
### Målingen
58 medlemmer. Hvert anker er verifisert **ordrett og unikt** i fila (`str.count`
== 1) og på oppgitt linjenummer — 0 avvik. Datasett:
`scripts/kb-eval/data/r11-footer-class-2026-08-11.json`.
| Bøtte | Definisjon | Antall | Andel |
|-------|------------|--------|-------|
| konsistent | total + oppdeling oppgitt, og de stemmer under den eneste rimelige lesningen | **30** | 51,7 % |
| inkonsistent | total + oppdeling oppgitt, og de spriker under enhver rimelig lesning | **11** | 19,0 % |
| tvetydig | total + oppdeling oppgitt, men dommen snur mellom to forsvarlige lesninger | **1** | 1,7 % |
| ikke sjekkbar | ingen intern kryssjekk finnes (bar total, oppdeling uten total, eller et ledd uten tall) | **16** | 27,6 % |
Fordeling per skill: engineering 23 · governance 12 · security 12 · advisor 11.
⚠️ **BØTTENE MÅLER INTERN KOHERENS, IKKE SANNHET.** En «konsistent» linje er ikke
verifisert — den er bare ikke selvmotsigende. Ratifiseringsgrunnen er
**uverifiserbarhet**, og den gjelder alle 58 uansett bøtte. Bøttene er derfor
input til *hvordan* klassen bokføres, aldri til *om* den slettes.
### Retningsargumentet re-målt — falsifisert som absolutt, intakt som tendens
§9.15 flagget `chain-of-thought-prompting.md:500` som kandidat-motsigelse mot «alle
avvik går samme vei» (åtte av åtte, oppgitt < sum). Målt over den korrigerte
populasjonen, alle 11 inkonsistente:
| Δ | oppgitt | enumerert | fil |
|---|---------|-----------|-----|
| 3 | 3 | 6 | `response-quality-metrics-rag.md` |
| 2 | 3 | 5 | `prompt-testing-and-evaluation.md` |
| 2 | 4 | 6 | `model-monitoring-drift-detection.md` |
| 2 | 4 | 6 | `small-language-models-economics.md` |
| 2 | 4 | 6 | `token-counting-optimization.md` |
| 1 | 5 | 6 | `document-intelligence-prebuilt-models.md` |
| 1 | 7 | 8 | `speech-services-text-to-speech.md` |
| 1 | 4 | 5 | `responsible-ai-policy-development.md` |
| 1 | 4 | 5 | `output-validation-grounding-verification.md` |
| 1 | 4 | 5 | `azure-ai-foundry-cost-governance.md` |
| **+1** | **4** | **3** | **`chain-of-thought-prompting.md`** |
**10 av 11 samme vei, 1 motsatt.** I tillegg går det ene tvetydige medlemmet
(`microsoft-graph-api-copilot-integration.md`, oppgitt 7 mot 6 MCP-kall + 1
`ToolSearch`) motsatt vei under sin MCP-only-lesning.
**«Alle avvik går samme vei» er dermed FALSIFISERT som absolutt påstand og må
slutte å siteres i den formen.** Tendensen (10/11 ≈ 91 %) overlever og er fortsatt
forenlig med at tallet sjelden ble transkribert fra en faktisk kjøring — men
«null motsatt vei» var et artefakt av at populasjonen var målt til under halv
størrelse.
### Den ratifiserte formen dekker ikke hele klassen
Formen fra §9.14 sier **slett linja**. Målt fordeling:
- **46 av 58 (79,3 %) er LINJEFORM** — formen anvendes uendret.
- **12 av 58 (20,7 %) er BLOKKFORM** — tellingen bor i 27 linjer, ofte uten oppgitt
total, ofte under en egen overskrift (`### MCP Calls Summary`,
`### MCP research summary`, `**MCP-statistikk:**`, `**Research Coverage:**`).
For blokkformen er «slett linja» **udefinert**: å slette én linje av
`- **microsoft_docs_search calls:** 4` / `- ... fetch calls:** 3` / `- ... :** 1`
etterlater en halv telling. Og flere blokker blander klasser — i
`model-versioning-registry-management.md` står `- **Unique sources:** 7` (naboklassen,
skal BEHOLDES) i samme punktliste som de tre kall-linjene. **Sletteoperasjonen må
navngi hvilke linjer i blokken som går, ikke bare hvilken blokk.**
### Anbefaling for åpent spørsmål #17
Spørsmålet gjaldt opprinnelig «14 ubokførte filer», og beslutning **#14** sa **én
kø-entry per berørt fil**. Begge premissene er nå fire ganger for små. Anbefalingen
er å **splitte bokføringen etter om entryen bærer en åpen beslutning**:
**(1) ÉN klasse-entry for 42 medlemmer.** Disse er linjeform, uten flagg, med
ratifisert grunn (uverifiserbarhet) og ratifisert operasjon (slett). Det finnes
ingen åpen beslutning i dem — 42 individuelle entries ville vært 42 kopier av et
allerede besvart spørsmål. Fillisten bor i datasettet, ikke i 42 prosatekster.
**Presedensen taler for dette, ikke mot:** §9.15 viste at et tall sitert ordrett i
sju kø-entries var målt over halve populasjonen, og *ingen av de sju kunne oppdage
det — de siterte hverandre*. Duplisering er nettopp mekanismen som lot feilen leve.
**(2) 16 individuelle entries** — de som bærer en beslutning en klasse-entry ikke
kan bære:
| Fil | Bøtte | Hvorfor egen entry |
|-----|-------|--------------------|
| `microsoft-graph-api-copilot-integration.md` | tvetydig | `ToolSearch` talt som MCP-kall |
| `ai-act-annex-iii-checklist.md` | ikke sjekkbar | `WebSearch` + `tavily_extract` under overskrift «MCP-søk»; blokk |
| `ai-act-compliance-guide.md` | ikke sjekkbar | `WebSearch` under «MCP-søk utført»; blokk |
| `rag-document-preprocessing.md` | konsistent | etikett «MCP-kilder», verdi «MCP-kall» |
| `ai-impact-assessment-framework.md` | ikke sjekkbar | etikett «dokumenter søkt», verdi = handlinger |
| `inferencing-optimization-caching.md` | konsistent | tallet gjentatt i prosa (:1029) — krever rest-søk |
| `chain-of-thought-prompting.md` | inkonsistent | eneste motsatt-vei-avvik; blokk |
| `enterprise-governance-copilot-deployment.md` | ikke sjekkbar | blokk |
| `domain-specific-prompt-optimization.md` | ikke sjekkbar | blokk; ett ledd uten tall |
| `infrastructure-as-code-mlops.md` | ikke sjekkbar | blokk |
| `model-versioning-registry-management.md` | ikke sjekkbar | blokk; naboklasse i samme liste |
| `ai-risk-taxonomy-classification.md` | ikke sjekkbar | blokk |
| `algorithmic-accountability-auditability.md` | konsistent | blokk (overskrift bærer totalen) |
| `model-monitoring-drift-detection.md` | inkonsistent | blokk (overskrift bærer totalen) |
| `responsible-ai-framework-overview.md` | ikke sjekkbar | blokk over to etiketter |
| `prompt-injection-defense-patterns.md` | ikke sjekkbar | blokk |
De tre `provenance_mix`-tilfellene er **samme klasse som `idx-26w`** (åpent
spørsmål #21) og bør avgjøres sammen med det, ikke hver for seg.
**Kostnadsforskjellen er ikke hovedargumentet.** 58 entries mot 17 sparer arbeid,
men det som faktisk står på spill er at 42 nesten-identiske entries gjør køen til
et sted hvor en feilmåling kan gjemme seg bak sine egne kopier.
### Gjenbrukbare funn
- **Et nett kan ikke bevise sin egen fullstendighet. Flere nett med ULIKE blindsoner
kan** — når residualen til hvert nett leses for hånd. Ett bredere ord er ikke det
samme som ett nett til.
- **Enumerér vokabularet, ikke forekomstene.** Å telle hvor mange treff en gjetning
gir sier ingenting om hva gjetningen ikke ser. Å liste alle *etiketter* i klassens
strukturelle omgivelser er et argument som kan etterprøves.
- **Test den blindheten du VET at du har.** Alle fire nett krevde et siffer; det tok
én sveip å vise at tallord-formen er tom. Utestet blindhet er en påstand.
- **Et «alle X går samme vei»-argument dør av ett moteksempel, men tendensen kan
overleve.** Skill de to før du siterer noen av dem.
- **En form kan være ratifisert for én STRUKTUR og stille anta at hele klassen har
den.** «Slett linja» var aldri feil — den var udefinert for 20 % av klassen, og
ingenting i formen sa fra.
- **Duplisering i en kø er ikke bare kostnad — det er hvor en feilmåling gjemmer seg.**