fix(ms-ai-architect): G7 idx-26e lukket — grunnen til at den ble bokfoert separat var falsk
Lista i stakeholder-communication-ai-decisions.md er kanonisert etter soesterfilas
struktur: sju kildenavngitte segmenter under 'Komponenter i Scorecard', pluss en egen
'Du konfigurerer'-linje som kun baerer det kildene faktisk sier brukeren setter.
Stempelet paa linje 81 datert til 2026-08-04 (operatoervalg, idx-26-presedens).
Premisset som begrunnet SEPARAT bokfoering var falskt. Baade STATE-spoersmaal #9 og
entryens ramme sa 'fil UTEN Verified-stempel over lista' - stempelet paa :81 er
seksjons-terminalt og lukker linje 46-81, altsaa ogsaa lista. AA verifisere en entrys
paastand og aa verifisere GRUNNEN til at den ble bokfoert er to ulike sjekker.
To kildesider, og aa maale mot bare den ene ville produsert to falske funn.
Segmentene staar paa how-to-siden; fila siterer selv konsept-siden, som viser seg aa
belegge 'multi-stakeholder alignment', 'risikoofficerer' OG konfigurerbarheten
naer-ordrett. Siden et krav kom fra og siden fila siterer er ulike objekter.
Aa bare re-merke de fem til 'Komponenter' ble forkastet paa maaling: en femmedlems
komponent-liste er nettopp completeness-defekten idx-26c lukket i soesterfila.
Reparasjonen ville importert en soester-entrys allerede lukkede defekt.
Kompositum-auditen ble kjoert denne gangen (80e17ec eksisterer fordi den ikke ble
det): target-verdiene 3x, fairness-maalverdiene 2x, ingen hybrid.
Tre naboer bokfoert, ikke feid inn:
- idx-26g: §1 utelater public-preview-banneret begge kilder baerer, mens fil-headeren
sier Status: GA; og how-to-siden staar ikke blant fila sine femten kilder.
- idx-26h: soesterfilas Status-linje ANBEFALER det kilden eksplisitt fraraader
('anbefalt for production use' vs 'we don't recommend it for production workloads'),
under samme stempel idx-26d daterte og idx-26f utvidet. idx-26f-resolutionen sier at
ingen unntak er skrevet ned for at stempelet skal vaere aerlig - sant om
enumerasjonen den maalte, ikke om resten av blokka stempelet lukker.
- idx-26i: funn om SJEKKEN. idx-26f grep etter '(gender, ethnicity, age)' og fikk null;
samme spesifisitet finnes som '(kjoenn, etnisitet, alder)'. En engelsk-bare grep ser
ikke et tospraaklig korpus' norske halvdel.
Koe: 8 aapne / 9 resolved. Suite 1047/1047.
This commit is contained in:
parent
80e17ec462
commit
08f31250c4
3 changed files with 139 additions and 9 deletions
|
|
@ -1200,6 +1200,93 @@ says so in those terms rather than presenting all six as a single verification.
|
|||
|
||||
**Queue state: 6 open, 8 resolved.** Suite 1047/1047.
|
||||
|
||||
### 9.11 idx-26e closed — the reason it was booked separately was false
|
||||
|
||||
Run 2026-08-04. The entry was the last of the 26-family and the first to live in a
|
||||
different file: a parallel five-member list in
|
||||
`stakeholder-communication-ai-decisions.md`, in the paraphrase dialect §9.9/§9.10 had
|
||||
just removed from `transparency-documentation-standards.md`.
|
||||
|
||||
**Ground truth killed the framing before it could drive the edit.** Both STATE's open
|
||||
question #9 and the session brief carried it as *"should canonicalisation go as far in a
|
||||
file that carries no Verified stamp over the list?"* — and the file does carry one. The
|
||||
stamp `*Confidence: Verified (MCP microsoft-learn)*` at line 81 is **section-terminal**:
|
||||
it closes §1 (lines 46–81) and therefore reaches the list. That was the entire stated
|
||||
reason the case was booked separately instead of folded into idx-26d, and with it gone
|
||||
the honesty argument from §9.10 applied here with undiminished force. **Verifying an
|
||||
entry's claim and verifying the reason it was booked are two different checks**, and
|
||||
only the first one is habitual.
|
||||
|
||||
**Two source pages, and measuring only one of them would have manufactured findings.**
|
||||
The seven segments are enumerated on `how-to-responsible-ai-scorecard`. But the page
|
||||
*this file cites* (Kilder #1) is `concept-responsible-ai-scorecard`, which enumerates no
|
||||
segments — and which turns out to ground the rest of the section almost verbatim:
|
||||
|
||||
| Section text | Measured against how-to alone | Measured against the page the file cites |
|
||||
|---|---|---|
|
||||
| "Muliggjøre multi-stakeholder alignment i ML-livssyklusen" | looks like paraphrase drift | "the need for effective multi-stakeholder alignment in an end-to-end machine learning lifecycle" |
|
||||
| "Støtte auditability for risikoofficerer og regulatorer" | looks like invented specificity | "share model and data insights with auditors and risk officers for auditability purposes, as required by AI regulations" |
|
||||
| the configurability framing | unsupported | "you can provide your desired model performance and fairness target values, such as target accuracy and target error rate" |
|
||||
|
||||
Two false findings and one wrongly-rejected framing, avoided only because the file's own
|
||||
citation list was read before the stamp was dated. **The page a claim came from and the
|
||||
page the file cites are different objects; a stamp covering a section has to be measured
|
||||
against both.**
|
||||
|
||||
**The obvious repair was rejected on measurement, not taste.** Relabelling the five
|
||||
existing members to "Komponenter i Scorecard" would have produced a five-member
|
||||
components list — which is precisely the completeness defect idx-26c *closed* in the
|
||||
sibling by adding the two missing segments. The repair would have imported a sibling
|
||||
entry's already-closed defect. Same class as §9.10's finding, one file over: the repair
|
||||
is the first place a defect recurs, including defects closed elsewhere.
|
||||
|
||||
What was written instead mirrors the sibling's *structure*: a seven-member
|
||||
`**Komponenter i Scorecard**` list carrying the source's own segment names, plus a
|
||||
separate one-line `**Du konfigurerer**` carrying only what the sources say the user sets.
|
||||
The framing defect disappears because the members are no longer claimed to be
|
||||
configurable; the name drift and the 26f-class unsourced specificity
|
||||
(`Statistikk, distribusjoner, bias-indikatorer`, `Accuracy, error rates, fairness metrics`)
|
||||
disappear with the paraphrases.
|
||||
|
||||
**Writing a parallel list in a second file creates a new diff surface, so the deviations
|
||||
had to be written down.** Three were deliberate: source order and bullets rather than the
|
||||
sibling's numbering (two numbered lists in different orders would make "item 2" denote
|
||||
different members in the two files); a source-tight rendering of *causal insights* rather
|
||||
than the sibling's looser `Causal vs correlational relationships i features`, which §9.10
|
||||
adjudicated a recognisable paraphrase and which is **not** reopened — but importing a
|
||||
looser paraphrase into *fresh* text under a stamp being dated today would be the class
|
||||
this programme removes; and `target-verdiene` in every member, where the sibling leaves
|
||||
one member in bare English. That last one is §9.10's own correction applied prospectively
|
||||
rather than after the fact.
|
||||
|
||||
**The compound-form audit ran this time.** `80e17ec` exists because a coherence read
|
||||
checked content and not compound forms. The new text was audited on that axis explicitly
|
||||
before commit: `target-verdiene` ×3, `fairness-målverdiene` ×2, no hybrid, and the two
|
||||
English metric names are the source's own examples already mirrored in the YAML block
|
||||
below.
|
||||
|
||||
**Three neighbours booked, one of them in the file §9.9/§9.10 had already canonicalised.**
|
||||
|
||||
- **idx-26g** — §1 omits the public-preview banner both sources carry, while the file
|
||||
header declares `Status: GA`; and the file's fifteen listed sources do not include the
|
||||
page grounding the enumeration just written.
|
||||
- **idx-26h** — in `transparency-documentation-standards.md`, the Status line recommends
|
||||
what the source explicitly does not: *"anbefalt for production use"* against *"we don't
|
||||
recommend it for production workloads."* The Customization block beneath it carries two
|
||||
unsourced items. Both sit under the same stamp idx-26d dated and idx-26f widened.
|
||||
**idx-26f's resolution states that no exception is written down anywhere for that stamp
|
||||
to remain honest — true of the enumeration it measured, and not of the rest of the
|
||||
block the stamp closes.** Two entries verified the list; neither read to the end of what
|
||||
the stamp reaches.
|
||||
- **idx-26i** — a finding about the *check*. idx-26f recorded a cross-file grep for
|
||||
`(gender, ethnicity, age)` returning nothing; the same specificity exists in the corpus
|
||||
as `(kjønn, etnisitet, alder)`. **An English-only grep cannot see a bilingual corpus's
|
||||
Norwegian rendering of the same claim**, so every sweep in this queue that greped
|
||||
English strings alone has an unmeasured Norwegian half. Booked as a candidate, not a
|
||||
confirmed defect — the attribution there is to Foundry tooling, not a scorecard segment.
|
||||
|
||||
**Queue state: 8 open, 9 resolved.** Suite 1047/1047.
|
||||
|
||||
## Appendix A — the 15 admitted proposals, hand-verified
|
||||
|
||||
Every proposal the classifier (§4 + context condition) admitted over the whole
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue