fix(ms-ai-architect): G7 idx-26h lukket — entryen overdrev sin egen dekning, og stempelet satte scopet

Status-linja sa 'anbefalt for production use med awareness om SLA-limitations'.
Tre MS Learn-sider sier ordrett det motsatte: 'provided without a service-level
agreement, and we don't recommend it for production workloads'. Erstattet med
kildens egen ordlyd, ikke slettet - preview-status er seksjonens eneste
operativt avgjoerende faktum og staar ingen andre steder i den.

ENTRYENS ANDRE FUNN VAR FALSKT. Den sier 'neither source page states either' om
Customization-punktene. Det holder for de TO sidene entryen maalte, og faller mot
en tredje: how-to-responsible-ai-insights-ui dokumenterer konfigurasjonen direkte
(steg 4 'enables cohort analysis' + 'features of interest'; steg 1 'an optional
description about the model's functionality'). Punktene var altsaa BELAGT INNHOLD
MED UKILDET ORDLYD - reparasjonen er reformulering, ikke sletting. Bare 'identified
risk groups' og 'decisions og mitigations' var ukildet. 9.11-lærdommen gjentok seg
med entryen selv som den som overpaastod.

STEMPELET SATTE SCOPET, ikke preferanse. Aa datere 'Verified' til 2026-08-09
paastaar at HELE seksjonen er verifisert, saa to locatorer entryen ikke navnga
maatte ogsaa lukkes (operatoerratifisert hver for seg): 'max error rate per
subgroup' (maalverdier settes paa metrikken; fairness-maal fanger differanse
eller forhold PAA TVERS av undergrupper) og rad-en 'Compliance officers | ...
(EU AI Act, sector-specific regler)' (verken rollenavnet eller parentesen finnes
i kilden).

Stempelet var FEIL DA DET BLE SATT, ikke foreldet: Status-linja var usann
2026-08-03, dagen idx-26d daterte stempelet over den.

LABELEN, IKKE PREVIEWET: den ratifiserte previewet slo Auditors-raden inn i
Risk officers-raden; labelen sa 'rett raden', entall. Auditors-raden er belagt
ordrett i konsept-siden og staar urort.

Kilder: begge how-to-sidene lagt inn som #11-12, gamle 11-15 renummerert 13-17,
'Unique sources: 15 URLs' -> 17. Ingen prosa i fila kryss-refererer kildenummer
(verifisert med grep).

RETTELSE: oektens egen ratifiseringsprompt paastod at editen ogsaa lukker
idx-26g. Det gjoer den ikke - 26g er bokfoert mot en ANNEN fil
(stakeholder-communication-ai-decisions.md), med egen kildeliste, egen
telle-paastand og et 'Status: GA'-funn denne editen ikke roerer. Renummereringen
var uansett riktig for DENNE fila, saa editen staar, men spoersmaalet var galt.

Bokfoert som idx-26j: fottekstens proveniens-blokk (3+2+1 = 6, ikke 5; udefinert
referent; 80/20 matcher hverken 10/5 foer eller 12/5 etter) + dialekt-paret
'model-' mot ratifisert 'modell-' som denne editen selv innfoerte.

Suite 1047/1047. Koe: 8 aapne / 10 resolved.
This commit is contained in:
Kjell Tore Guttormsen 2026-08-09 10:14:47 +02:00
commit 99675d576c
3 changed files with 111 additions and 14 deletions

View file

@ -1402,3 +1402,77 @@ candidate for the next ratification.
Reproduce the tally and the machine checks:
`node scripts/kb-eval/check-o2-returns.mjs`.
### 9.12 idx-26h closed — the entry overstated its own coverage, and the stamp set the scope
Run 2026-08-09. Two locators under the `2026-08-03` stamp that §9.9 dated and §9.10
widened: a Status line reading *"anbefalt for production use med awareness om
SLA-limitations"*, and two Customization bullets the entry called unsourced.
**The Status finding held, and held harder than booked.** Three MS Learn pages —
`concept-responsible-ai-scorecard`, `how-to-responsible-ai-scorecard` and
`how-to-responsible-ai-insights-ui` — carry the identical banner: *"This preview version
is provided without a service-level agreement, and we don't recommend it for production
workloads."* The file recommended what the source explicitly does not. Replaced with the
source's own wording rather than deleted: preview status is the section's one
operationally decisive fact and nothing else in the section carries it.
**The second finding was false, and the entry was the one overstating.** It claims
*"neither source page states either"*. That holds for the two pages the entry measured
and fails against a third. `how-to-responsible-ai-insights-ui` documents scorecard
customisation directly — step 4: the Data analysis section *"enables cohort analysis"*
and you select *"features of interest to identify your model performance on their
underlying cohorts"*; step 1: *"an optional description about the model's functionality,
data it was trained and evaluated on, architecture type"*.
| Bullet as written | Entry's verdict | Measured verdict |
|---|---|---|
| "Cohort analysis: Disaggregated performance for identified risk groups" | unsourced specificity | grounded; only *"identified risk groups"* is not the source's — it says *features of interest*, and *sensitive features* belong to a different configuration step |
| "Narrative sections: Fritekst-forklaringer for decisions og mitigations" | unsourced specificity | the free-text field exists; *"decisions og mitigations"* is its wrong referent, and *"Narrative sections"* is not source vocabulary either |
So the repair was **reformulation, not deletion** — deleting would have removed
documented behaviour. §9.11's lesson recurred with the entry itself as the overstater:
*the page a claim came from is a different object from the pages an entry happened to
measure.* An entry is a hypothesis with an evidence list, and its evidence list is not a
census of the sources.
**The stamp, not preference, set the scope.** Re-dating `Verified` to 2026-08-09 asserts
the whole section is verified, so every measured-unsourced claim under it had to be
closed or the new date would knowingly reproduce the §9.10 class. Reading the section
end-to-end rather than to the entry's anchors surfaced two more: a third Customization
bullet (*"max error rate per subgroup"* — targets are set on the metric, and fairness
targets capture the difference or ratio **across** subgroups, never a per-subgroup
maximum) and a role-table row (*"Compliance officers | … (EU AI Act, sector-specific
regler)"* — neither the role name nor the parenthetical is in any source). **A stamp is a
scope-setting instrument: what it closes is what the edit owes, and an entry's anchors
are a lower bound on that.**
**The stamp was wrong when set, not stale.** The Status line was false on 2026-08-03, the
day idx-26d dated the stamp over it. Neither idx-26d nor idx-26f detected it because both
measured the seven-member enumeration and neither read to the end of the block the stamp
reaches. Staleness and falsity look identical in a dated stamp, and only re-measuring
tells them apart.
**The ratified preview was wider than the ratified label.** The option that closed the
role row carried a preview illustrating a row that *merged* `Compliance officers` into
`Auditors`; the label said *"rett raden"*, singular. The label was taken as the ratified
object and the `Auditors` row — grounded verbatim in the concept page — was left standing.
Following the illustration would have deleted correct content on the strength of a mockup.
**A ratification question can carry a false premise, and the operator ratifies it anyway.**
The source-list option presented in this session stated that the edit would close idx-26h
**and** idx-26g together. It does not: idx-26g is booked against a different file,
`stakeholder-communication-ai-decisions.md`, with its own source list, its own counting
claim and a `Status: GA` header finding untouched here. The renumbering was independently
correct for this file, so the edit stands — but the question was wrong, and the operator
had no way to see it. **The premise-verification duty applies to the questions you ask,
not only to the facts you act on.** Nothing in the queue, the suite or the CLI checks the
framing of a ratification prompt.
**Booked, not fixed:** `idx-26j` — the same file's footer provenance block, which no entry
has ever measured. `Total MCP calls: 5` enumerates components summing to 6; the line has
no defined referent (original generation run, or current state?); the `80% / 20%` split
matches neither the list's 10/5 before this edit nor its 12/5 after; and this edit itself
introduced a two-dialect compound pair (`model-` on line 101 against the ratified
`modell-` in the rewritten row) that no check can see, because the stamp claims
verification against source, not orthographic consistency.

File diff suppressed because one or more lines are too long

View file

@ -106,7 +106,7 @@ PDF-rapport designet for å dele model- og data-innsikter mellom tekniske og ikk
|-------|-------------------|
| **Data scientists** | Ekstrahere insights fra Responsible AI dashboard for deployment approval |
| **Product managers** | Sette target performance/fairness metrics og verifisere at modellen møter dem |
| **Compliance officers** | Review for regulatory compliance (EU AI Act, sector-specific regler) |
| **Risk officers** | Dele modell- og data-innsikter for auditability, slik AI-regelverk krever |
| **Auditors** | Arkiverte scorecards i Azure ML Run History for retrospective review |
**Komponenter i Scorecard:**
@ -122,13 +122,13 @@ PDF-rapport designet for å dele model- og data-innsikter mellom tekniske og ikk
> **Ikke scorecard-segmenter:** Error analysis og Counterfactual analysis er komponenter i Responsible AI *dashboard*, ikke segmenter i scorecard-PDF-en. Scorecard-en eksporterer innsikter fra dashboard-et, men de to komponentene har ingen egne segmenter i PDF-rapporten.
**Customization:**
- Target values: Akseptabel accuracy, max error rate per subgroup
- Cohort analysis: Disaggregated performance for identified risk groups
- Narrative sections: Fritekst-forklaringer for decisions og mitigations
- Target values: Ytelsesmetrikkene du velger (inntil tre) med målverdier — f.eks. target accuracy og target error rate — pluss målverdi på fairness-metrikken du velger
- Cohort analysis: Modellytelsen på de underliggende kohortene for de features of interest du velger, og over-/underrepresentasjon i datasettet
- Scorecard summary: Tittel og en valgfri fritekstbeskrivelse av modellens funksjonalitet, dataene den er trent og evaluert på, og arkitekturtype
**Status:** Public preview (Azure ML) — anbefalt for production use med awareness om SLA-limitations.
**Status:** Public preview (Azure ML) — leveres uten SLA, og Microsoft anbefaler den ikke for production workloads.
**Confidence:** Verified (MCP: microsoft-learn, 2026-08-03)
**Confidence:** Verified (MCP: microsoft-learn, 2026-08-09)
---
@ -759,30 +759,38 @@ Return on investment: Transparency er billigere enn cleanup. Skal vi prioritere
https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai-agents/responsible-ai-across-organization
(Status: Verified 2026-02 — Cross-functional governance, auditing, transparency mechanisms)
11. **Generate Responsible AI insights in the studio UI**
https://learn.microsoft.com/en-us/azure/machine-learning/how-to-responsible-ai-insights-ui
(Status: Verified 2026-08-09 — Scorecard-konfigurasjon: målverdier, cohort analysis i Data analysis-seksjonen, fritekstbeskrivelse i scorecard summary)
12. **Use Responsible AI scorecard (preview) in Azure Machine Learning**
https://learn.microsoft.com/en-us/azure/machine-learning/how-to-responsible-ai-scorecard
(Status: Verified 2026-08-09 — De sju scorecard-segmentene; preview uten SLA, ikke anbefalt for production workloads)
**Baseline sources (model knowledge + MCP-inferred):**
11. **Model Cards for Model Reporting** (Mitchell et al., 2019)
13. **Model Cards for Model Reporting** (Mitchell et al., 2019)
https://arxiv.org/abs/1810.03993
(Academic origin of model card concept)
12. **Datasheets for Datasets** (Gebru et al., 2018)
14. **Datasheets for Datasets** (Gebru et al., 2018)
https://arxiv.org/abs/1803.09010
(Academic origin of datasheet concept)
13. **EU AI Act**
15. **EU AI Act**
https://artificialintelligenceact.eu/
(Status: Adopted 2024 — Transparency obligations for high-risk AI)
14. **NIST AI Risk Management Framework**
16. **NIST AI Risk Management Framework**
https://www.nist.gov/itl/ai-risk-management-framework
(US standard for AI governance)
15. **Developing Responsible Generative AI Applications (Windows)**
17. **Developing Responsible Generative AI Applications (Windows)**
https://learn.microsoft.com/en-us/windows/ai/rai
(Status: Verified 2026-02 — Model Cards reference, red teaming, governance processes)
**Total MCP calls:** 5 (microsoft_docs_search: 3, microsoft_docs_fetch: 2, microsoft_code_sample_search: 1)
**Unique sources:** 15 URLs
**Unique sources:** 17 URLs
**Confidence:** 80% Verified (MCP), 20% Baseline (established frameworks)
---