fix(ms-ai-architect): G7 idx-26h lukket — entryen overdrev sin egen dekning, og stempelet satte scopet
Status-linja sa 'anbefalt for production use med awareness om SLA-limitations'. Tre MS Learn-sider sier ordrett det motsatte: 'provided without a service-level agreement, and we don't recommend it for production workloads'. Erstattet med kildens egen ordlyd, ikke slettet - preview-status er seksjonens eneste operativt avgjoerende faktum og staar ingen andre steder i den. ENTRYENS ANDRE FUNN VAR FALSKT. Den sier 'neither source page states either' om Customization-punktene. Det holder for de TO sidene entryen maalte, og faller mot en tredje: how-to-responsible-ai-insights-ui dokumenterer konfigurasjonen direkte (steg 4 'enables cohort analysis' + 'features of interest'; steg 1 'an optional description about the model's functionality'). Punktene var altsaa BELAGT INNHOLD MED UKILDET ORDLYD - reparasjonen er reformulering, ikke sletting. Bare 'identified risk groups' og 'decisions og mitigations' var ukildet. 9.11-lærdommen gjentok seg med entryen selv som den som overpaastod. STEMPELET SATTE SCOPET, ikke preferanse. Aa datere 'Verified' til 2026-08-09 paastaar at HELE seksjonen er verifisert, saa to locatorer entryen ikke navnga maatte ogsaa lukkes (operatoerratifisert hver for seg): 'max error rate per subgroup' (maalverdier settes paa metrikken; fairness-maal fanger differanse eller forhold PAA TVERS av undergrupper) og rad-en 'Compliance officers | ... (EU AI Act, sector-specific regler)' (verken rollenavnet eller parentesen finnes i kilden). Stempelet var FEIL DA DET BLE SATT, ikke foreldet: Status-linja var usann 2026-08-03, dagen idx-26d daterte stempelet over den. LABELEN, IKKE PREVIEWET: den ratifiserte previewet slo Auditors-raden inn i Risk officers-raden; labelen sa 'rett raden', entall. Auditors-raden er belagt ordrett i konsept-siden og staar urort. Kilder: begge how-to-sidene lagt inn som #11-12, gamle 11-15 renummerert 13-17, 'Unique sources: 15 URLs' -> 17. Ingen prosa i fila kryss-refererer kildenummer (verifisert med grep). RETTELSE: oektens egen ratifiseringsprompt paastod at editen ogsaa lukker idx-26g. Det gjoer den ikke - 26g er bokfoert mot en ANNEN fil (stakeholder-communication-ai-decisions.md), med egen kildeliste, egen telle-paastand og et 'Status: GA'-funn denne editen ikke roerer. Renummereringen var uansett riktig for DENNE fila, saa editen staar, men spoersmaalet var galt. Bokfoert som idx-26j: fottekstens proveniens-blokk (3+2+1 = 6, ikke 5; udefinert referent; 80/20 matcher hverken 10/5 foer eller 12/5 etter) + dialekt-paret 'model-' mot ratifisert 'modell-' som denne editen selv innfoerte. Suite 1047/1047. Koe: 8 aapne / 10 resolved.
This commit is contained in:
parent
4032dfcd43
commit
99675d576c
3 changed files with 111 additions and 14 deletions
|
|
@ -1402,3 +1402,77 @@ candidate for the next ratification.
|
|||
|
||||
Reproduce the tally and the machine checks:
|
||||
`node scripts/kb-eval/check-o2-returns.mjs`.
|
||||
|
||||
### 9.12 idx-26h closed — the entry overstated its own coverage, and the stamp set the scope
|
||||
|
||||
Run 2026-08-09. Two locators under the `2026-08-03` stamp that §9.9 dated and §9.10
|
||||
widened: a Status line reading *"anbefalt for production use med awareness om
|
||||
SLA-limitations"*, and two Customization bullets the entry called unsourced.
|
||||
|
||||
**The Status finding held, and held harder than booked.** Three MS Learn pages —
|
||||
`concept-responsible-ai-scorecard`, `how-to-responsible-ai-scorecard` and
|
||||
`how-to-responsible-ai-insights-ui` — carry the identical banner: *"This preview version
|
||||
is provided without a service-level agreement, and we don't recommend it for production
|
||||
workloads."* The file recommended what the source explicitly does not. Replaced with the
|
||||
source's own wording rather than deleted: preview status is the section's one
|
||||
operationally decisive fact and nothing else in the section carries it.
|
||||
|
||||
**The second finding was false, and the entry was the one overstating.** It claims
|
||||
*"neither source page states either"*. That holds for the two pages the entry measured
|
||||
and fails against a third. `how-to-responsible-ai-insights-ui` documents scorecard
|
||||
customisation directly — step 4: the Data analysis section *"enables cohort analysis"*
|
||||
and you select *"features of interest to identify your model performance on their
|
||||
underlying cohorts"*; step 1: *"an optional description about the model's functionality,
|
||||
data it was trained and evaluated on, architecture type"*.
|
||||
|
||||
| Bullet as written | Entry's verdict | Measured verdict |
|
||||
|---|---|---|
|
||||
| "Cohort analysis: Disaggregated performance for identified risk groups" | unsourced specificity | grounded; only *"identified risk groups"* is not the source's — it says *features of interest*, and *sensitive features* belong to a different configuration step |
|
||||
| "Narrative sections: Fritekst-forklaringer for decisions og mitigations" | unsourced specificity | the free-text field exists; *"decisions og mitigations"* is its wrong referent, and *"Narrative sections"* is not source vocabulary either |
|
||||
|
||||
So the repair was **reformulation, not deletion** — deleting would have removed
|
||||
documented behaviour. §9.11's lesson recurred with the entry itself as the overstater:
|
||||
*the page a claim came from is a different object from the pages an entry happened to
|
||||
measure.* An entry is a hypothesis with an evidence list, and its evidence list is not a
|
||||
census of the sources.
|
||||
|
||||
**The stamp, not preference, set the scope.** Re-dating `Verified` to 2026-08-09 asserts
|
||||
the whole section is verified, so every measured-unsourced claim under it had to be
|
||||
closed or the new date would knowingly reproduce the §9.10 class. Reading the section
|
||||
end-to-end rather than to the entry's anchors surfaced two more: a third Customization
|
||||
bullet (*"max error rate per subgroup"* — targets are set on the metric, and fairness
|
||||
targets capture the difference or ratio **across** subgroups, never a per-subgroup
|
||||
maximum) and a role-table row (*"Compliance officers | … (EU AI Act, sector-specific
|
||||
regler)"* — neither the role name nor the parenthetical is in any source). **A stamp is a
|
||||
scope-setting instrument: what it closes is what the edit owes, and an entry's anchors
|
||||
are a lower bound on that.**
|
||||
|
||||
**The stamp was wrong when set, not stale.** The Status line was false on 2026-08-03, the
|
||||
day idx-26d dated the stamp over it. Neither idx-26d nor idx-26f detected it because both
|
||||
measured the seven-member enumeration and neither read to the end of the block the stamp
|
||||
reaches. Staleness and falsity look identical in a dated stamp, and only re-measuring
|
||||
tells them apart.
|
||||
|
||||
**The ratified preview was wider than the ratified label.** The option that closed the
|
||||
role row carried a preview illustrating a row that *merged* `Compliance officers` into
|
||||
`Auditors`; the label said *"rett raden"*, singular. The label was taken as the ratified
|
||||
object and the `Auditors` row — grounded verbatim in the concept page — was left standing.
|
||||
Following the illustration would have deleted correct content on the strength of a mockup.
|
||||
|
||||
**A ratification question can carry a false premise, and the operator ratifies it anyway.**
|
||||
The source-list option presented in this session stated that the edit would close idx-26h
|
||||
**and** idx-26g together. It does not: idx-26g is booked against a different file,
|
||||
`stakeholder-communication-ai-decisions.md`, with its own source list, its own counting
|
||||
claim and a `Status: GA` header finding untouched here. The renumbering was independently
|
||||
correct for this file, so the edit stands — but the question was wrong, and the operator
|
||||
had no way to see it. **The premise-verification duty applies to the questions you ask,
|
||||
not only to the facts you act on.** Nothing in the queue, the suite or the CLI checks the
|
||||
framing of a ratification prompt.
|
||||
|
||||
**Booked, not fixed:** `idx-26j` — the same file's footer provenance block, which no entry
|
||||
has ever measured. `Total MCP calls: 5` enumerates components summing to 6; the line has
|
||||
no defined referent (original generation run, or current state?); the `80% / 20%` split
|
||||
matches neither the list's 10/5 before this edit nor its 12/5 after; and this edit itself
|
||||
introduced a two-dialect compound pair (`model-` on line 101 against the ratified
|
||||
`modell-` in the rewritten row) that no check can see, because the stamp claims
|
||||
verification against source, not orthographic consistency.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue