ms-ai-architect/scripts/kb-eval/data/g7-review-queue.json
Kjell Tore Guttormsen 08f31250c4 fix(ms-ai-architect): G7 idx-26e lukket — grunnen til at den ble bokfoert separat var falsk
Lista i stakeholder-communication-ai-decisions.md er kanonisert etter soesterfilas
struktur: sju kildenavngitte segmenter under 'Komponenter i Scorecard', pluss en egen
'Du konfigurerer'-linje som kun baerer det kildene faktisk sier brukeren setter.
Stempelet paa linje 81 datert til 2026-08-04 (operatoervalg, idx-26-presedens).

Premisset som begrunnet SEPARAT bokfoering var falskt. Baade STATE-spoersmaal #9 og
entryens ramme sa 'fil UTEN Verified-stempel over lista' - stempelet paa :81 er
seksjons-terminalt og lukker linje 46-81, altsaa ogsaa lista. AA verifisere en entrys
paastand og aa verifisere GRUNNEN til at den ble bokfoert er to ulike sjekker.

To kildesider, og aa maale mot bare den ene ville produsert to falske funn.
Segmentene staar paa how-to-siden; fila siterer selv konsept-siden, som viser seg aa
belegge 'multi-stakeholder alignment', 'risikoofficerer' OG konfigurerbarheten
naer-ordrett. Siden et krav kom fra og siden fila siterer er ulike objekter.

Aa bare re-merke de fem til 'Komponenter' ble forkastet paa maaling: en femmedlems
komponent-liste er nettopp completeness-defekten idx-26c lukket i soesterfila.
Reparasjonen ville importert en soester-entrys allerede lukkede defekt.

Kompositum-auditen ble kjoert denne gangen (80e17ec eksisterer fordi den ikke ble
det): target-verdiene 3x, fairness-maalverdiene 2x, ingen hybrid.

Tre naboer bokfoert, ikke feid inn:
- idx-26g: §1 utelater public-preview-banneret begge kilder baerer, mens fil-headeren
  sier Status: GA; og how-to-siden staar ikke blant fila sine femten kilder.
- idx-26h: soesterfilas Status-linje ANBEFALER det kilden eksplisitt fraraader
  ('anbefalt for production use' vs 'we don't recommend it for production workloads'),
  under samme stempel idx-26d daterte og idx-26f utvidet. idx-26f-resolutionen sier at
  ingen unntak er skrevet ned for at stempelet skal vaere aerlig - sant om
  enumerasjonen den maalte, ikke om resten av blokka stempelet lukker.
- idx-26i: funn om SJEKKEN. idx-26f grep etter '(gender, ethnicity, age)' og fikk null;
  samme spesifisitet finnes som '(kjoenn, etnisitet, alder)'. En engelsk-bare grep ser
  ikke et tospraaklig korpus' norske halvdel.

Koe: 8 aapne / 9 resolved. Suite 1047/1047.
2026-08-04 09:45:33 +02:00

230 lines
46 KiB
JSON

{
"_meta": {
"gap": "G7",
"form": "(b) named queue into the human review phase",
"ratified": "2026-08-03",
"rationale": "Measured in R11 §9.6: 2 of the 4 subtractions applied in 957ebef left a residue, so residues are the normal by-product of a delete-only envelope rather than an exception. Two of the members are replacements, not multi-locator cases, which a deletion-oriented O4 class would not have fixed. A queue absorbs both classes; an O4 return contract would have been mis-sized against the evidence.",
"contract": "Anchors are verbatim strings, never line numbers (line ≠ real_line in 9 of 17 R11 records). An open entry whose anchor no longer occurs in its file is drift, and check-g7-queue.mjs fails rather than passing it silently. Nothing in this queue is machine-appliable by definition — every entry is outside the O2 envelope. Resolution is a human review act.",
"evidence": "docs/r11-pilot-results.md §9.4, §9.5, §9.6; docs/ref-kb-correctness-program-2026-06.md §8 G7"
},
"entries": [
{
"id": "idx-17",
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "The idx 17 subtraction (957ebef) deleted the three score-threshold bands but left the lead-in ending in a colon, promising an enumeration that no longer existed, immediately followed by a **Verified** stamp. The subtraction was correct; the paragraph it left was not. V1/V2/V2b/V3 are string invariants over deleted text and cannot see document coherence, so the machine could not have caught it.",
"evidence": "docs/r11-pilot-results.md §9.6",
"anchors": [],
"resolution": "Operator-ratified 2026-08-03: colon changed to a period, making the lead-in a complete and independently true sentence that the **Verified** stamp correctly covers. Corpus swept for the same defect shape (bold lead-in ending in colon, blank line, **Verified**) — no other occurrence."
},
{
"id": "idx-33",
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "The idx 33 subtraction removed AI asset inventory via Azure Resource Graph and the Purview Insider Risk Management bullet from an unsupported Defender for Cloud AISPM attribution. Correct — but the same capabilities survive in this file under Cloud Adoption Framework Secure AI attribution, stamped 'Verified MCP 2026-04', and are asserted in four other corpus files. The edit's benefit is corpus-wide smaller than the single line suggested. Whether the CAF attribution is itself supported has not been checked.",
"evidence": "docs/r11-pilot-results.md §9.5 (cross-corpus check), §9.6",
"anchors": [
"Oppdatert 2026-04: inkluderer nå AI asset inventory via Azure Resource Graph"
]
},
{
"id": "idx-26",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "multi-locator",
"status": "resolved",
"raised": "2026-08-03",
"summary": "The Responsible AI Scorecard component list names Error analysis (item 4) and Counterfactual analysis (item 5). Both were measured false against first-party docs 2026-08-03: the canonical scorecard segments are summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights and causal insights; Error analysis and Counterfactual analysis are Responsible AI *dashboard* components. The delete-only reduction could only remove item 5, because line 300 asserts Error analysis as scorecard content too — so a partial fix would have left a known-false claim standing while introducing a renumbering artifact (1,2,3,4,6,7). Operator declined the partial fix 2026-08-03 and sent the whole case here. Correct repair spans both locators.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"4. **Error analysis**: Error rates per cohort, confusion matrices",
"| **Risk assessment** | Responsible AI Scorecard: Error analysis, fairness assessment |"
],
"resolution": "Operator-ratified 2026-08-03, form (b) — relabel rather than remove. Both source pages were re-fetched live this session and confirm the measurement independently of §9.6: how-to-responsible-ai-scorecard enumerates summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights and causal insights; concept-responsible-ai-dashboard lists Error analysis and Counterfactual what-if among the dashboard components. Locator 1: items 4 and 5 removed from the numbered scorecard list, which renumbers cleanly to 1-5 — the 1,2,3,4,6,7 artifact was forced only inside the delete-only envelope and does not apply to an ordinary Edit. The two capabilities are retained in a blockquote explicitly marked as dashboard components rather than scorecard segments, so genuine source-confirmed information survives and the reader is warned off precisely the conflation that produced the defect. Locator 2: the Risk assessment row now reads 'fairness insights' alone. The **Confidence:** Verified stamp twelve lines below vouched for the false list and was silently outside every machine check (V1/V2/V2b/V3 are string invariants; check-g7-queue only tests anchors); it is kept but dated to 2026-08-03 to record the re-verification. Document coherence around both locators was read after the edit, per the c569bdc lesson. A neighbouring defect surfaced by the file sweep is booked separately as idx-26b rather than folded in here."
},
{
"id": "idx-26c",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "The scorecard enumeration is incomplete, and idx 26's repair made that assertion active rather than latent. The live fetch performed for idx 26 lists seven canonical segments: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. The list after the idx 26 edit carries five — model overview, fairness assessment, model interpretability, causal inference, data quality — so **model performance** and **cohorts** are absent. Note for whoever repairs this: 'Cohort analysis' does appear a few lines below, but inside the *Customization* block, not as an enumerated segment, so it does not cure the omission. This was a latent incompleteness under an undated stamp before 2026-08-03; dating the stamp to 2026-08-03 as part of idx 26 turned it into a positive claim that this enumeration was verified that day, over content the same day's verification showed to be missing two members — the same class of defect §9.7 describes, where an anchor-correct edit leaves a false claim where no check reaches. Booked rather than folded into idx 26, whose ratified scope was the falsity of items 4 and 5, not the completeness of the list. The date on the stamp is deliberately NOT re-litigated here: the relabel it covers WAS verified 2026-08-03, and the operator may keep it once this entry closes the completeness gap. Adding the two missing segments is a content change requiring its own operator ratification.",
"evidence": "docs/r11-pilot-results.md §9.7; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"5. **Data quality**: Dataset statistics, missing values, outlier analysis"
],
"resolution": "Operator-ratified 2026-08-03, form (a) - add the two missing segments rather than downgrade the list to a selection. The source was re-fetched live this session and enumerates seven segments in order: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. Model performance and Cohorts were appended as items 6 and 7; the existing five were left untouched, so the change is confined to the measured gap. Appending rather than inserting in source order was a free choice, not a machine-forced one: STATE asserted that a resolved entry anchor must stop matching or the check fails, and that is wrong - lib/g7-queue.mjs:69-75 returns before the anchor check for resolved entries, so resolved is fully exempt and anchor drift only bites an entry left open. The list carries no ordering claim, and the existing five were already not in source order, so appending introduces no new falsity. Item 7 is worded \"automatisk uttrukket av scorecard-en\" to keep it distinct from the Cohort analysis bullet in the Customization block, which is operator-defined and does not enumerate a segment. The **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp closing §3 (Responsible AI Scorecard) is kept unchanged and its scope was checked explicitly rather than left implicit: it now vouches for a complete seven-member enumeration re-verified against the live source on the date it already carries. Post-edit sweep of both regions found the surrounding prose coherent; the paraphrase drift it did surface is booked as idx-26d, not folded in. Correction 2026-08-03 (same session, later pass): this resolution originally located that stamp at \"line 129\". It is not there - line 129 is the **Status:** Public preview line, and the **Confidence:** stamp is line 131. The claim about the stamp was true; the locator was false, and it was written into a tracked artefact. Same class as the row count corrected in a9d4724, and the reason the queue contract forbids line numbers as anchors - the prohibition applies to prose inside an entry too, not only to the anchors array."
},
{
"id": "idx-26b",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced by the post-edit file sweep for idx 26, in the same compliance-mapping table as idx 26's second locator, but outside both of its anchors — so booked separately rather than folded in (gap discipline; operator-ratified 2026-08-03). The Accuracy metrics row attributes 'Quantitative analyses' to the Responsible AI Scorecard. That is not a scorecard segment name: how-to-responsible-ai-scorecard calls the corresponding segment 'model performance'. The defect is a cross-attribution between two different standards rather than mere imprecision — 'Quantitative Analyses' is a canonical Model Card section, and this same file lists it as one at line 77. Repair is a replacement, so it is outside the delete-only envelope. Whether the right fix is to rename the segment or to drop the row is not yet decided.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"| **Accuracy metrics** | Responsible AI Scorecard: Quantitative analyses |"
],
"resolution": "Operator-ratified 2026-08-03, form (a) - rename rather than drop the row. The source page was re-fetched live this session (how-to-responsible-ai-scorecard) rather than taken from §9.6 as a premise, and it names the segment model performance: \"The model performance segment displays your model most important metrics and characteristics of your predictions and how well they satisfy your desired target values.\" The Accuracy metrics row now reads \"Responsible AI Scorecard: model performance\". Rename was chosen over deletion because the EU AI Act accuracy-metrics mapping is genuine and source-supported; dropping the row would have removed true information to repair a naming defect. Line 77 keeps Quantitative analyses as a Model Card section, which is correct there and is what made the cross-attribution visible. The **Confidence:** Verified (Baseline + MCP-inferred) stamp at the foot of the compliance table was deliberately NOT upgraded: it covers all six rows and only one was measured against the source this session, so re-dating or strengthening it would have extended a verification claim over unmeasured content - the §9.7 defect class."
},
{
"id": "idx-27",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Failed out of O2 on cond 3 (§9.6): the cited source DOES establish the mechanism — faqs-generative-orchestration states 'Makers can require user confirmation before executing tools that modify data' — so deleting the Plugin-actions row would destroy source-confirmed information. The defect is modality, not fabrication: the file presents confirmation prompts as a built-in disclosure, whereas the source makes them maker-configured. The Chat-interface row in the same table is imprecise for the same reason: the FAQ documents a default transparency message ('Just so you are aware, I sometimes use AI to answer your questions.'), not a 'Powered by AI' badge. Repair is a replacement, outside the delete-only envelope.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-orchestration; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-answers",
"anchors": [
"| **Plugin actions** | Confirmation prompts før sensitive actions (send email, delete file) |",
"| **Chat interface** | \"Powered by AI\" badge i chat window |"
],
"resolution": "Operator-ratified 2026-08-03. Open question #8 (does a hub page's linked children count as \"the source\") was answered NO - the strict reading R11 has held throughout stands, and no new standing grounding rule was introduced. idx-27 was instead resolved by making the child a CITED source: the Responsible AI FAQ block now lists both the hub (responsible-ai-overview) and faqs-generative-orchestration, naming the latter as the source for the two repaired claims. The FAQ was fetched live before writing; both quotes are verbatim from it. Chat-interface row: \"Powered by AI\" badge replaced with the default transparency message the source actually documents - \"Just so you are aware, I sometimes use AI to answer your questions.\" - which the source states agents include, so it belongs under Built-in disclosures. Plugin-actions row: the source says \"Makers can require user confirmation before executing tools that modify data\", a configurable safeguard, so the row was MOVED OUT of the Built-in disclosures table into a new \"Maker-konfigurerte kontroller (ikke innebygde disclosures)\" block. Repairing the text in place would have left the false modality asserted by the table heading; weakening the heading instead was rejected because it would have silently restated the modality of the two unmeasured rows (Generative answers, Data usage), and adding a modality column would have asserted \"built-in\" about them outright - creating new unverified claims to fix an old one. The deleted \"(send email, delete file)\" examples were not carried over: the source names neither. The same \"Powered by AI\" imprecision at line 218, in the Azure implementasjon list of Mønster 3, is outside this entry's anchors and is booked as idx-27b rather than swept in - the idx-26c/26d precedent. No **Confidence:** stamp covers the Copilot Studio section, so no marker obligation was triggered by this edit. Scope addendum, same session, after an adversarial read of the resolution above: both repaired claims were taken from a FAQ whose own scope line reads \"the AI impact of generative orchestration for custom agents built in Copilot Studio\", and they were written into a section headed Microsoft Copilot Studio with no qualifier - the same widening class this entry was raised for. Re-checked against the docs rather than reasoned about, and the two claims came apart. The maker-confirmation safeguard occurs ONLY in the generative-orchestration FAQ, in a safeguards list about tool execution, so it is orchestration-scoped and the bullet now says so explicitly (\"for agenter med generativ orkestrering\"). The default transparency message occurs in that FAQ AND, verbatim, in faqs-generative-answers under \"What protections are in place within Copilot Studio for responsible AI?\", framed there as a general best practice rather than an orchestration feature. Two independent feature FAQs stating it without a feature qualifier is why the Built-in disclosures row carries no scope qualifier; faqs-generative-answers was added as a third cited source so the reader can check that reasoning rather than take it on trust. Neither FAQ is a Copilot-Studio-wide statement of record, so this is grounded-as-cited, not established-for-all-agents - if a stronger claim is ever wanted, it needs a source that says so."
},
{
"id": "idx-36",
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-03",
"summary": "Applying idx 36 alone yields a severity table more precise than the prose that cites it, so a companion edit is required for the prose to match the narrowed table. Held back from 957ebef as out of envelope.",
"evidence": "docs/r11-pilot-results.md §9.4, §9.5",
"anchors": [
"Øker severity bar; krever mer robust adversarial defenses"
]
},
{
"id": "idx-18",
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-03",
"summary": "No reduction exists. The surviving claim is a whole titled section on Azure AI Search built-in caching plus a **Verified** row in the verification table, so no deletion confined to a single locator can repair the file. This is the member that most clearly motivated G7.",
"evidence": "docs/r11-pilot-results.md §9.3, §9.4",
"anchors": [
"**Automatic Caching Behavior:**",
"| Azure AI Search caching | **Verified** | Microsoft Learn docs (4, 6) |"
]
},
{
"id": "idx-26d",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced by the post-edit sweep for idx-26c. The five original members of the scorecard component list carry paraphrased names rather than the source segment names, and idx-26c added two members that DO carry source names, so the list is now mixed. Concretely: item 5 Data quality maps to the source segment data analysis; item 3 Model interpretability maps to top important factors; item 1 Model overview is described as \"Architecture, training data, intended use\", which is Model Card content - the source summary segment is a model overview plus the key target values the user set. That last one is the same cross-attribution class as idx-26b rather than mere imprecision, since Model details / Intended use / Training data are canonical Model Card sections listed in this same file at lines 71-76. Full canonicalisation was offered to the operator as part of the idx-26c decision and was NOT chosen; the ratified scope there was completeness only, so this is booked rather than folded in. Open choice: rename the five to the source segment names in source order, or rewrite item 1 alone (the only member where the drift produces a false attribution rather than a recognisable paraphrase) and leave the rest.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"1. **Model overview**: Architecture, training data, intended use"
],
"resolution": "Operator-ratified 2026-08-03, full canonicalisation. The source was re-fetched live before writing (how-to-responsible-ai-scorecard) and enumerates seven segments: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. All five original member NAMES were renamed to the source segment names, closing the defect this entry was actually booked under - that the list mixed two dialects after idx-26c added two source-named members. Rewriting item 1 alone was rejected for exactly that reason: it removes the false attribution but leaves the booked defect standing. DESCRIPTIONS were rewritten for two members only: item 1 (was \"Architecture, training data, intended use\", which is Model Card content - Model details / Intended use / Training data are canonical Model Card sections listed in this same file at lines 71-76 - now \"Modelloversikt og de target-verdiene du har satt\", matching the source summary segment) and item 3 (was \"Feature importance (global/local explanations)\", RAI *dashboard* vocabulary inside a scorecard enumeration, contradicting this file's own callout that dashboard components are not scorecard segments - now \"Faktorene som påvirker modellens prediksjoner mest\"). Descriptions of items 2, 4 and 5 were deliberately left UNTOUCHED even though they carry specifics the source does not state, because that is a distinct defect class (unsourced specificity under a verification marker) booked as idx-26f; rewriting them here would have silently resolved an entry raised the same session and left the queue incoherent. Source order was NOT imposed: idx-26c established the list carries no ordering claim and the five were already not in source order, so reordering would be a larger diff with no measured defect behind it. Cross-file check run before editing: the five names occur elsewhere in the corpus, but in dashboard/interpretability contexts, not as a mirror of this enumeration - except stakeholder-communication-ai-decisions.md:57-62, which is a parallel five-member list in the same paraphrase dialect and is booked as idx-26e rather than swept in. What the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp closing §3 now vouches for, stated explicitly rather than left implicit: a seven-member enumeration whose NAMES are all the source's own segment names and whose descriptions for items 1, 3, 6 and 7 were verified against the live source on the date the stamp already carries. Items 2, 4 and 5 descriptions are NOT covered by that widening - idx-26f is open against them. [Superseded 2026-08-03 by idx-26f's closure: items 2 and 5 were rewritten to source wording and item 4 adjudicated a recognisable paraphrase; see idx-26f's resolution for what the stamp covers now.]"
},
{
"id": "idx-27b",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced by the neighbourhood sweep for idx-27. The Mønster 3 \"Azure implementasjon\" list asserts a Copilot Studio \"Powered by AI\" disclosure in the chat interface - the same claim idx-27 removed from the Copilot Studio section, and false for the same reason: faqs-generative-orchestration documents a default transparency message (\"Just so you are aware, I sometimes use AI to answer your questions.\"), not a badge. Different section, outside idx-27's anchors, so booked separately rather than swept in. Repair is a replacement, outside the delete-only envelope. Note for whoever repairs this: after idx-27 the corrected wording already exists a few hundred lines below, so the repair is a copy of an already-ratified formulation rather than a fresh judgement.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-orchestration",
"anchors": [
"- **Copilot Studio**: \"Powered by AI\" disclosure i chat interface"
],
"resolution": "Operator-ratified 2026-08-03. The Moenster 3 \"Azure implementasjon\" bullet was replaced with exactly what the already-ratified row at the Copilot Studio section asserts: the standard transparency message \"Just so you are aware, I sometimes use AI to answer your questions.\" delivered in the chat interface. This is a copy of a formulation ratified earlier the same session (idx-27), not a fresh judgement, and the quote was grep-verified byte-for-byte against that row after the edit - hand-copying a verbatim string is where quote-style drift enters. Deliberately NOT carried over: any audience-tiering or layered-disclosure language from the surrounding Moenster 3 pattern. The bullet sits under an audience-layering table while the source row sits under \"Built-in disclosures\"; adding reach the source does not state is exactly the scope class caught in 66fb567, and this repair would have inherited it. Standing matches idx-27's addendum: the standard message is documented in both faqs-generative-orchestration and faqs-generative-answers (the latter with the broader \"best practice to communicate to users that the agent uses artificial intelligence\" framing), so it carries no scope qualifier - unlike the confirmation safeguard, which does. Cross-file grep before editing: \"Powered by AI\" occurred nowhere else in the corpus. Two neighbours were checked and deliberately NOT swept in - the Foundry agent-transparency bullet asserting a \"This chatbot uses AI\" embeddable component, and the scenario-1 tooling line naming a \"Copilot Studio disclosure widget\" where this file's own Copilot Studio section documents a pre-built \"AI disclosure\" topic in the generative AI toolkit, not a widget. Both are unverified product artefacts in other sections against other sources; they are booked as idx-27c and idx-27d rather than repaired here."
},
{
"id": "idx-26e",
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced by the cross-file grep run before the idx-26d edit. This file carries a parallel five-member list describing the Responsible AI Scorecard under the heading \"Konfigurerbare elementer\", in the same paraphrase dialect idx-26d just removed from transparency-documentation-standards.md: Dataset-helse, Modell-ytelse, Target values, Fortolkningsevne, Fairness assessment. Two problems, and they are distinguishable. First, the names are paraphrases of the source's scorecard SEGMENTS (data analysis, model performance, top important factors, fairness insights), so the list is the same drift class as idx-26d. Second, and separately, the heading claims these are CONFIGURABLE elements; the source page enumerates seven scorecard segments and does not enumerate a set of configurable elements, so the framing itself is unsupported. The list sits under a *Confidence: Verified (MCP microsoft-learn)* stamp carrying no date. Booked, not folded into idx-26d: idx-26d's ratified scope was one list in one file, and this file was not re-verified against the live source this session.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"**Konfigurerbare elementer**:",
"- **Fortolkningsevne**: Global/lokal feature importance"
],
"resolution": "Operator-ratified 2026-08-04: mirror the sibling file's structure. Both this entry's premise and STATE's framing of open question #9 were checked against ground truth before writing, and the framing was FALSE: the claim that this list sits in a file without a Verified stamp over it does not hold. The stamp *Confidence: Verified (MCP microsoft-learn)* at :81 is section-terminal and closes §1 (lines 46-81), so it reaches the list. That was the entire stated reason 26e was booked separately rather than folded into idx-26d/idx-26f, and with it gone the honesty argument from 26f applies here with the same force it had there. Two source pages were fetched live this session: how-to-responsible-ai-scorecard, which enumerates the seven segments, and concept-responsible-ai-scorecard, which is the page THIS file actually cites (Kilder #1) and which grounds the configurability claim. Relabelling the existing five members to 'Komponenter i Scorecard' without completing them was rejected on measurement rather than taste: idx-26c's ratified scope in the sibling was completeness - it ADDED the two missing segments - so a five-member components list here would have imported the exact defect idx-26c closed, the repair inheriting a sibling entry's defect class. What was written: the block is split in two, as in transparency-documentation-standards.md. 'Komponenter i Scorecard' now carries all seven source segment names, with descriptions verified against the live how-to page this session; 'Du konfigurerer' carries only what the sources state the user sets - target values for model performance, fairness target values for the sensitive groups you choose, with target accuracy and target error rate as the concept page's own examples. This closes all three defect classes the list carried, of which the entry had booked two. Name drift (26d class): the five names were paraphrases of source segments. Unsourced specificity under a verification marker (26f class, NOT booked here and found while measuring): 'Statistikk, distribusjoner, bias-indikatorer' and 'Accuracy, error rates, fairness metrics' are stated by neither source page, and 'på tvers av sensitive grupper' dropped the you-choose-them agency exactly as the sibling's item 2 did before 26f. Unsupported framing (the entry's second booked problem): neither source enumerates configurable elements; the concept page states only that you provide desired model performance and fairness target values. DELIBERATE DIVERGENCES FROM THE SIBLING, recorded so a future cross-file sweep does not book them as drift: (1) source order and bullets, not the sibling's numbering - idx-26c established the enumeration carries no ordering claim, this file's own lists are bulleted, and two numbered lists in different orders would make 'item 2' denote different members in the two files; (2) item 7 reads 'Om faktorene eller tiltakene du har identifisert har kausal effekt på utfallet i den virkelige verden' rather than the sibling's 'Causal vs correlational relationships i features' - the sibling's rendering was adjudicated a recognisable paraphrase under 26f and is NOT reopened, but 'vs correlational' and 'i features' are in neither source, and importing a looser paraphrase into FRESH text under a stamp being dated today would be the very class this programme exists to remove; (3) 'target values' is rendered 'target-verdiene' in every member, where the sibling leaves one member in bare English - 26f's ratified correction was precisely that one source concept must not carry two renderings inside one list. THE STAMP was dated to 2026-08-04 per operator choice, following idx-26's precedent of dating a stamp to record re-verification. What it vouches for, stated rather than left implicit: the seven-member enumeration and the configurability line, both verified against live sources this session; and the rest of §1 - 'Hva det er', 'Primært bruksområde', the three 'Formål' bullets and the four 'Verdi for ikke-tekniske' bullets - which were READ and checked against the concept page this session and are supported by it. Two of those bullets ('multi-stakeholder alignment i ML-livssyklusen', 'risikoofficerer') looked like paraphrase drift when measured against the how-to page alone and are near-verbatim on the concept page; measuring only the page a segment list came from would have produced two false findings. The YAML block is marked 'Typisk' and is illustrative. NOT covered by the stamp, measured and booked rather than folded in: §1 omits the public-preview banner both source pages carry while the file header declares Status: GA, and the file's fifteen listed sources do not include how-to-responsible-ai-scorecard, the page grounding the enumeration just written - booked as idx-26g. CROSS-FILE GREP before and after the edit: 'Konfigurerbare elementer', 'Dataset-helse' and 'Fortolkningsevne' occur nowhere else in the corpus; 'Modell-ytelse' and 'Fairness assessment' occur elsewhere only in dashboard, monitoring and Foundry-tooling contexts, never as a mirror of this enumeration. One of those hits is booked as idx-26i. The whole section was re-read for coherence after the edit, and - per the failure that forced 80e17ec - the new text was separately audited on the axis this repair was justified by rather than on content alone: compound forms count target-verdiene 3x and fairness-målverdiene 2x with no hybrid, and the two English metric names are the source's own examples, already mirrored in the YAML block immediately below."
},
{
"id": "idx-26f",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced while writing idx-26d, and deliberately excluded from it. Items 2 and 5 of the scorecard enumeration carry specifics the live source does not state: the fairness insights segment is described with \"(gender, ethnicity, age)\" where the source says only \"your desired sensitive groups\", and the data analysis segment with \"missing values, outlier analysis\" where the source says only that it \"shows you characteristics of your data\". Neither is false on its face - both are plausible instances - but both sit under the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp, which now vouches for an enumeration re-verified against the live source. This is a distinct class from idx-26d's name drift and idx-26b's cross-attribution: unsourced SPECIFICITY under a verification marker, where the defect is the marker's reach rather than the sentence. idx-26d's resolution states explicitly that its widening of the stamp does not cover items 2, 4 and 5, so the stamp is currently honest only because that exclusion is written down - closing this entry is what would make it honest on its own terms. Item 4's description was checked in the same pass and is a recognisable paraphrase of the source, not added specificity, so it is named here as checked-and-clear rather than left ambiguous.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"2. **Fairness insights**: Performance disparities across sensitive groups (gender, ethnicity, age)",
"5. **Data analysis**: Dataset statistics, missing values, outlier analysis"
],
"resolution": "Operator-ratified 2026-08-03, form (a): descriptions rewritten to the source's own wording, NOT surgical deletion of the named parentheticals. The source was re-fetched live before writing (how-to-responsible-ai-scorecard) and states only \"how well your model is satisfying the fairness target values you set for your desired sensitive groups\" and \"The data analysis segment shows you characteristics of your data\". Item 2 is now \"Hvor godt modellen moeter fairness-maalverdiene du har satt for de sensitive gruppene du velger\" and item 5 \"Karakteristikker ved dataene dine\" (both written with correct Norwegian diacritics in the corpus file). Deleting only the specificity this entry NAMED was rejected on two counts, each of which would have made the repair inherit the defect class it was raised against: it leaves item 5 as \"Dataset statistics\", which the source does not state either and which this entry did NOT adjudicate as checked-and-clear the way it did item 4; and it leaves item 2 as \"across sensitive groups\", dropping the you-choose-them agency and stranding item 2 in the pre-idx-26d dialect while items 1, 3, 6 and 7 carry \"du har satt\" - reopening the mixed-dialect defect idx-26d was booked under, in the file idx-26d had just canonicalised. Form (b) - narrowing the stamp's stated reach in the file instead - was put to the operator and declined: it is honest-by-annotation relocated from this queue's resolution field into the corpus, i.e. the same mechanism this entry exists to remove the need for. Item 4 was left untouched per this entry's own adjudication. Cross-file grep before editing: \"gender, ethnicity, age\", \"outlier analysis\" and \"Dataset statistics\" occurred nowhere else in the corpus. The enumeration was re-read whole after the edit for coherence, not just for the two edited strings. SUPERSEDES the closing clause of idx-26d's resolution: what the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp vouches for is now a seven-member enumeration whose names are all the source's own segment names, whose descriptions for items 2 and 5 were re-verified against the live source THIS session and whose descriptions for items 1, 3, 6 and 7 were verified against the live source under idx-26d earlier the same day - the same stamp date and the same page, but an inherited verification rather than one re-run here, stated separately so the artefact is honest about its own provenance, and whose item 4 description was adjudicated a recognisable paraphrase of the source's causal-insights passage. No exception is written down anywhere for the stamp to remain honest. POST-EDIT CORRECTION, same session: item 2 first landed as \"fairness-target values\", a hybrid compound that is neither language and that broke the very dialect consistency this repair was justified by - item 1 renders the same source concept as \"target-verdiene\". Corrected to \"fairness-maalverdiene\". The operator-ratified LABEL was \"kildens ordlyd i idx-26d-dialekten\"; the illustrative preview carried the defective form, and the label is the ratified object, so the correction needed no re-ratification. The coherence re-read after the first edit checked the enumeration for CONTENT alignment and passed it - it did not check compound forms, which is the axis that failed."
},
{
"id": "idx-27c",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced by the neighbourhood sweep for idx-27b and deliberately excluded from it. The Microsoft Foundry \"Agent transparency\" list asserts \"Disclosure widgets: \\\"This chatbot uses AI\\\" embeddable component\" - a named product artefact carrying a verbatim-looking user-facing string. Same class as idx-27 and idx-27b (a plausible disclosure phrasing asserted as a built-in surface), but a different platform section against a different source, so it cannot be repaired by copying the ratified Copilot Studio formulation. Needs its own live fetch against Foundry documentation to establish whether an embeddable disclosure widget exists and, if so, what string it actually carries. Repair is a replacement, outside the delete-only envelope.",
"evidence": "docs/r11-pilot-results.md §9.10",
"anchors": [
"- Disclosure widgets: \"This chatbot uses AI\" embeddable component"
]
},
{
"id": "idx-27d",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced by the neighbourhood sweep for idx-27b and deliberately excluded from it. The scenario 1 tooling line names a \"Copilot Studio disclosure widget\", but this same file's Copilot Studio section documents a pre-built \"AI disclosure\" TOPIC in the generative AI toolkit - not a widget. This is intra-file name drift of the same shape as idx-26d, not a source-attribution error: the file contradicts itself about what the artefact is called. Note the context differs from idx-27/idx-27b - this line sits inside a worked scenario recommendation rather than a product-capability claim, so the repair should decide whether to name the toolkit topic or drop the artefact name, not merely restate a source. Repair is a replacement, outside the delete-only envelope.",
"evidence": "docs/r11-pilot-results.md §9.10",
"anchors": [
"**Tooling:** Azure OpenAI Transparency Note + Copilot Studio disclosure widget"
]
},
{
"id": "idx-26g",
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-04",
"summary": "Surfaced while closing idx-26e and named in its resolution rather than folded in. Two provenance and status facts about §1 of this file, both now under the *Confidence: Verified (MCP microsoft-learn, 2026-08-04)* stamp idx-26e dated. First: both source pages carry an Important banner stating the Responsible AI scorecard is in public preview, provided without a service-level agreement and not recommended for production workloads. §1 says nothing about it, and the file header declares Status: GA. Second: the seven-member enumeration idx-26e wrote is grounded in how-to-responsible-ai-scorecard, which is not among the fifteen sources this file lists - Kilder #1 is the concept page, concept-responsible-ai-scorecard, which does not enumerate the segments. The Total kilder line was counted against the list before booking and is accurate as it stands (15 = 8 verified + 4 baseline + 3 code samples), so adding a source is a renumbering edit across entries 9-15 plus a counting claim - which is why it was not swept into idx-26e's ratified scope. Open choice for the reviewer: add the how-to page as a numbered source and renumber, add it as a second URL under Kilder #1 (no renumber, but 'Total kilder' then counts entries rather than URLs), or leave the citation as is and narrow what the file claims.",
"evidence": "https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard ; https://learn.microsoft.com/en-us/azure/machine-learning/concept-responsible-ai-scorecard?view=azureml-api-2 ; idx-26e resolution; docs/r11-pilot-results.md §9.11",
"anchors": [
"**Status:** GA",
"**Total kilder**: 15 (8 verified fra MCP, 7 baseline/code samples)"
]
},
{
"id": "idx-26h",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-04",
"summary": "Surfaced while measuring idx-26e against the live sources, in the file idx-26d and idx-26f canonicalised. Two locators under the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp that idx-26 dated and idx-26d/26f widened. First, and the stronger of the two: the Status line reads 'Public preview (Azure ML) - anbefalt for production use med awareness om SLA-limitations', which contradicts both source pages rather than paraphrasing them. Both state: 'This preview version is provided without a service-level agreement, and we don't recommend it for production workloads.' The file recommends what the source explicitly does not recommend. Second: the Customization block lists 'Cohort analysis: Disaggregated performance for identified risk groups' and 'Narrative sections: Fritekst-forklaringer for decisions og mitigations'; neither source page states either, making both 26f-class unsourced specificity under the same stamp. This matters for an artefact claim, not only for the corpus: idx-26f's resolution states that no exception is written down anywhere for that stamp to remain honest. That is true of the seven-member enumeration 26f measured, and not of the Customization and Status lines the same stamp also closes, which 26f did not measure. Both entries verified the enumeration; neither read down to the end of the block the stamp reaches.",
"evidence": "https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard ; https://learn.microsoft.com/en-us/azure/machine-learning/concept-responsible-ai-scorecard?view=azureml-api-2 ; idx-26f resolution; docs/r11-pilot-results.md §9.11",
"anchors": [
"**Status:** Public preview (Azure ML) — anbefalt for production use med awareness om SLA-limitations.",
"- Cohort analysis: Disaggregated performance for identified risk groups"
]
},
{
"id": "idx-26i",
"file": "skills/ms-ai-governance/references/norwegian-public-sector-governance/public-sector-ai-ethics-framework.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-04",
"summary": "Surfaced by the cross-file grep run while closing idx-26e, and it is as much a finding about the CHECK as about the corpus. idx-26f removed '(gender, ethnicity, age)' from the sibling's fairness insights description and recorded a cross-file grep for that string returning nothing elsewhere. That grep was dialect-blind: this file renders the same specificity in Norwegian - '**Fairness assessment** - evaluerer modellrettferdighet på tvers av sensitive grupper (kjønn, etnisitet, alder)' - under the heading '### Microsoft Foundry RAI-tools:'. Booked as an unmeasured CANDIDATE, not a confirmed defect: the attribution here is to Foundry RAI tooling rather than to a scorecard segment, and neither this file's stamp coverage nor its source page has been checked. What IS confirmed is the check gap - an English-only grep cannot see a bilingual corpus's Norwegian rendering of the same claim, so every cross-file sweep recorded in this queue that greped English strings alone has an unmeasured Norwegian half. Resolving this entry should decide the corpus question AND whether the sweep convention changes.",
"evidence": "idx-26f resolution (cross-file grep clause); idx-26e cross-file grep; docs/r11-pilot-results.md §9.11",
"anchors": [
"- **Fairness assessment** — evaluerer modellrettferdighet på tvers av sensitive grupper (kjønn, etnisitet, alder)"
]
}
]
}