fix(ms-ai-architect): G7 idx-26d + idx-27 lukket — scorecard-lista kanonisert, modalitet rettet, FAQ-barnet sitert

idx-26d: alle fem scorecard-navn omdøpt til kildens segmentnavn; beskrivelsene
på punkt 1 (Model Card-innhold) og 3 (dashboard-vokabular) skrevet om. Punkt
2/4/5 bevisst urørt — deres ubelagte spesifisitet er egen klasse (idx-26f).

idx-27: åpent spørsmål #8 besvart NEI (hub-barn teller ikke som kilden).
Løst ved å SITERE barnet i stedet: faqs-generative-orchestration lagt til.
Chat interface-raden bærer nå kildens faktiske standardmelding; Plugin
actions FLYTTET ut av Built-in-tabellen fordi den er maker-konfigurerbar —
å svekke overskriften eller legge til en modalitets-kolonne ville laget nye
ubekreftede påstander om to umålte rader.

idx-26c: falsk locator («linje 129») rettet i køa; stempelet er linje 131.

Tre nye entries: idx-27b, idx-26e (funnet av kryss-fil-grep FØR editen),
idx-26f. Kø: 6 åpne / 6 resolved. Suite 1047/1047.

[skip-docs]
This commit is contained in:
Kjell Tore Guttormsen 2026-08-03 22:09:17 +02:00
commit beddba8dd3
3 changed files with 137 additions and 12 deletions

View file

@ -1001,6 +1001,88 @@ chosen, so the ratified scope was completeness alone and this is booked, not fol
**Queue state: 5 open, 4 resolved.** Suite 1047/1047.
### 9.9 idx-26d and idx-27 closed — and what an unmeasured neighbour costs an edit
Run 2026-08-03, same day as §9.8, later pass. Both sources were re-fetched live
before writing; neither §9.7 nor §9.8 was read as fact.
**idx-26d closed by full canonicalisation, and the discriminator was not the one the
entry offered.** The entry framed an either/or: rename all five members, or rewrite
item 1 alone — item 1 being the only member whose drift produces a false attribution
rather than a recognisable paraphrase. Item 1 is genuinely the worst member, so the
narrow form is tempting. It is also the wrong form, for a reason the entry's own
summary contains: idx-26d is **booked as the list mixing two naming dialects** after
idx-26c added two source-named members. Rewriting item 1 alone removes a falsity and
leaves the booked defect standing — it would resolve an entry whose stated defect
survives the resolution. **A resolution has to close the defect the entry names, not
the worst defect the entry mentions.** The two are not the same thing, and the entry
text will not tell you which one you are looking at unless you re-read it against the
resolution you are about to write.
All five names were renamed to the source segments. Descriptions were rewritten for
items 1 and 3 only: item 1 carried Model Card content (this file lists Model details /
Intended use / Training data as Model Card sections at lines 7176), and item 3
carried "global/local explanations", RAI *dashboard* vocabulary standing inside a
scorecard enumeration that this same file's callout says is not a dashboard listing.
Source order was **not** imposed — §9.8 established the list carries no ordering claim.
**The descriptions that were deliberately left wrong.** Items 2 and 5 carry specifics
the source does not state: "(gender, ethnicity, age)" where the source says only "your
desired sensitive groups", and "missing values, outlier analysis" where it says only
that the segment "shows you characteristics of your data". Those were raised as
**idx-26f** in the same pass and left untouched in the file. The option presented to
the operator carried an illustrative sketch that *did* rewrite them; the option's own
label did not. Applying the sketch would have silently resolved an entry raised the
same session and left the queue incoherent — an entry pointing at text that no longer
exists, with nobody having ratified its removal. **When a decision is presented as
label plus illustration, the label is the ratified object.** The illustration is a
reading aid and may be wider than what was decided.
**idx-27 closed without ratifying the hub question.** Open question #8 — whether a hub
page's linked children count as the cited source — was answered **no**. R11's strict
reading stands. The entry closed anyway, by a move the question does not gate: the
child page was **added to the file as a cited source**. Grounding on an uncited page
is an expansion of "the source"; grounding on a page you then cite is not. This is
worth naming as a general move — *an unratified scope question can sometimes be
routed around by changing the artefact rather than the rule*, and routing around it
leaves the rule unweakened for every other entry that will meet it.
**The repair could not be made in place, and the reason is instructive.** The
Plugin-actions row asserts a confirmation prompt as a **built-in** disclosure; the
source makes it maker-configurable. The row sits in a four-row table under the heading
**Built-in disclosures**, and two of those four rows have never been measured. Three
repairs were available:
- Fix the row's text in place → the heading still asserts the false modality.
- Weaken the heading → silently restates the modality of the two unmeasured rows.
- Add a modality column → asserts "built-in" about the unmeasured rows outright.
The last two fix one unverified claim by minting two more. The row was **moved out**
into a separate maker-configured block instead. **Unmeasured neighbours constrain the
shape of a repair, not just its scope:** the cheapest true edit is the one that
restates nothing you have not checked, and that is frequently *not* the smallest diff.
This is the same neighbourhood discipline as §9.7/§9.8 read from the other side —
there, the neighbourhood held defects to find; here, it held claims not to touch.
**A false locator inside a tracked artefact, found by re-reading a resolution.**
idx-26c's resolution located the **Confidence:** stamp at "line 129". Line 129 is the
**Status:** line; the stamp is 131. The claim about the stamp was true, the locator
was not, and it was committed. The queue contract forbids line numbers as anchors
because `line ≠ real_line` in 9 of 17 R11 records — **that prohibition applies to
prose inside an entry too, and nothing checks it.** Corrected in place, with the
correction recorded rather than overwritten. STATE carried the same error.
**Three entries opened, none swept in.** idx-27b (the same "Powered by AI"
imprecision in Mønster 3's implementation list, outside idx-27's anchors), idx-26e (a
parallel five-member scorecard list in `stakeholder-communication-ai-decisions.md`, in
the dialect idx-26d just removed, under an *undated* Verified stamp, and additionally
framed as "configurable elements" which the source does not enumerate), and idx-26f
above. idx-26e was found by a **cross-file grep run before the edit** — the check that
asks whether a rename here creates an inconsistency there. It did not, but it found a
mirror the corpus was not known to contain.
**Queue state: 6 open, 6 resolved.** Suite 1047/1047.
## Appendix A — the 15 admitted proposals, hand-verified
Every proposal the classifier (§4 + context condition) admitted over the whole

View file

@ -56,7 +56,7 @@
"anchors": [
"5. **Data quality**: Dataset statistics, missing values, outlier analysis"
],
"resolution": "Operator-ratified 2026-08-03, form (a) - add the two missing segments rather than downgrade the list to a selection. The source was re-fetched live this session and enumerates seven segments in order: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. Model performance and Cohorts were appended as items 6 and 7; the existing five were left untouched, so the change is confined to the measured gap. Appending rather than inserting in source order was a free choice, not a machine-forced one: STATE asserted that a resolved entry anchor must stop matching or the check fails, and that is wrong - lib/g7-queue.mjs:69-75 returns before the anchor check for resolved entries, so resolved is fully exempt and anchor drift only bites an entry left open. The list carries no ordering claim, and the existing five were already not in source order, so appending introduces no new falsity. Item 7 is worded \"automatisk uttrukket av scorecard-en\" to keep it distinct from the Cohort analysis bullet in the Customization block, which is operator-defined and does not enumerate a segment. The **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp at line 129 is kept unchanged and its scope was checked explicitly rather than left implicit: it now vouches for a complete seven-member enumeration re-verified against the live source on the date it already carries. Post-edit sweep of both regions found the surrounding prose coherent; the paraphrase drift it did surface is booked as idx-26d, not folded in."
"resolution": "Operator-ratified 2026-08-03, form (a) - add the two missing segments rather than downgrade the list to a selection. The source was re-fetched live this session and enumerates seven segments in order: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. Model performance and Cohorts were appended as items 6 and 7; the existing five were left untouched, so the change is confined to the measured gap. Appending rather than inserting in source order was a free choice, not a machine-forced one: STATE asserted that a resolved entry anchor must stop matching or the check fails, and that is wrong - lib/g7-queue.mjs:69-75 returns before the anchor check for resolved entries, so resolved is fully exempt and anchor drift only bites an entry left open. The list carries no ordering claim, and the existing five were already not in source order, so appending introduces no new falsity. Item 7 is worded \"automatisk uttrukket av scorecard-en\" to keep it distinct from the Cohort analysis bullet in the Customization block, which is operator-defined and does not enumerate a segment. The **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp closing §3 (Responsible AI Scorecard) is kept unchanged and its scope was checked explicitly rather than left implicit: it now vouches for a complete seven-member enumeration re-verified against the live source on the date it already carries. Post-edit sweep of both regions found the surrounding prose coherent; the paraphrase drift it did surface is booked as idx-26d, not folded in. Correction 2026-08-03 (same session, later pass): this resolution originally located that stamp at \"line 129\". It is not there - line 129 is the **Status:** Public preview line, and the **Confidence:** stamp is line 131. The claim about the stamp was true; the locator was false, and it was written into a tracked artefact. Same class as the row count corrected in a9d4724, and the reason the queue contract forbids line numbers as anchors - the prohibition applies to prose inside an entry too, not only to the anchors array."
},
{
"id": "idx-26b",
@ -75,14 +75,15 @@
"id": "idx-27",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Failed out of O2 on cond 3 (§9.6): the cited source DOES establish the mechanism — faqs-generative-orchestration states 'Makers can require user confirmation before executing tools that modify data' — so deleting the Plugin-actions row would destroy source-confirmed information. The defect is modality, not fabrication: the file presents confirmation prompts as a built-in disclosure, whereas the source makes them maker-configured. The Chat-interface row in the same table is imprecise for the same reason: the FAQ documents a default transparency message ('Just so you are aware, I sometimes use AI to answer your questions.'), not a 'Powered by AI' badge. Repair is a replacement, outside the delete-only envelope.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-orchestration",
"anchors": [
"| **Plugin actions** | Confirmation prompts før sensitive actions (send email, delete file) |",
"| **Chat interface** | \"Powered by AI\" badge i chat window |"
]
],
"resolution": "Operator-ratified 2026-08-03. Open question #8 (does a hub page's linked children count as \"the source\") was answered NO - the strict reading R11 has held throughout stands, and no new standing grounding rule was introduced. idx-27 was instead resolved by making the child a CITED source: the Responsible AI FAQ block now lists both the hub (responsible-ai-overview) and faqs-generative-orchestration, naming the latter as the source for the two repaired claims. The FAQ was fetched live before writing; both quotes are verbatim from it. Chat-interface row: \"Powered by AI\" badge replaced with the default transparency message the source actually documents - \"Just so you are aware, I sometimes use AI to answer your questions.\" - which the source states agents include, so it belongs under Built-in disclosures. Plugin-actions row: the source says \"Makers can require user confirmation before executing tools that modify data\", a configurable safeguard, so the row was MOVED OUT of the Built-in disclosures table into a new \"Maker-konfigurerte kontroller (ikke innebygde disclosures)\" block. Repairing the text in place would have left the false modality asserted by the table heading; weakening the heading instead was rejected because it would have silently restated the modality of the two unmeasured rows (Generative answers, Data usage), and adding a modality column would have asserted \"built-in\" about them outright - creating new unverified claims to fix an old one. The deleted \"(send email, delete file)\" examples were not carried over: the source names neither. The same \"Powered by AI\" imprecision at line 218, in the Azure implementasjon list of Mønster 3, is outside this entry's anchors and is booked as idx-27b rather than swept in - the idx-26c/26d precedent. No **Confidence:** stamp covers the Copilot Studio section, so no marker obligation was triggered by this edit."
},
{
"id": "idx-36",
@ -113,12 +114,51 @@
"id": "idx-26d",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"status": "resolved",
"raised": "2026-08-03",
"summary": "Surfaced by the post-edit sweep for idx-26c. The five original members of the scorecard component list carry paraphrased names rather than the source segment names, and idx-26c added two members that DO carry source names, so the list is now mixed. Concretely: item 5 Data quality maps to the source segment data analysis; item 3 Model interpretability maps to top important factors; item 1 Model overview is described as \"Architecture, training data, intended use\", which is Model Card content - the source summary segment is a model overview plus the key target values the user set. That last one is the same cross-attribution class as idx-26b rather than mere imprecision, since Model details / Intended use / Training data are canonical Model Card sections listed in this same file at lines 71-76. Full canonicalisation was offered to the operator as part of the idx-26c decision and was NOT chosen; the ratified scope there was completeness only, so this is booked rather than folded in. Open choice: rename the five to the source segment names in source order, or rewrite item 1 alone (the only member where the drift produces a false attribution rather than a recognisable paraphrase) and leave the rest.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"1. **Model overview**: Architecture, training data, intended use"
],
"resolution": "Operator-ratified 2026-08-03, full canonicalisation. The source was re-fetched live before writing (how-to-responsible-ai-scorecard) and enumerates seven segments: summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights, causal insights. All five original member NAMES were renamed to the source segment names, closing the defect this entry was actually booked under - that the list mixed two dialects after idx-26c added two source-named members. Rewriting item 1 alone was rejected for exactly that reason: it removes the false attribution but leaves the booked defect standing. DESCRIPTIONS were rewritten for two members only: item 1 (was \"Architecture, training data, intended use\", which is Model Card content - Model details / Intended use / Training data are canonical Model Card sections listed in this same file at lines 71-76 - now \"Modelloversikt og de target-verdiene du har satt\", matching the source summary segment) and item 3 (was \"Feature importance (global/local explanations)\", RAI *dashboard* vocabulary inside a scorecard enumeration, contradicting this file's own callout that dashboard components are not scorecard segments - now \"Faktorene som påvirker modellens prediksjoner mest\"). Descriptions of items 2, 4 and 5 were deliberately left UNTOUCHED even though they carry specifics the source does not state, because that is a distinct defect class (unsourced specificity under a verification marker) booked as idx-26f; rewriting them here would have silently resolved an entry raised the same session and left the queue incoherent. Source order was NOT imposed: idx-26c established the list carries no ordering claim and the five were already not in source order, so reordering would be a larger diff with no measured defect behind it. Cross-file check run before editing: the five names occur elsewhere in the corpus, but in dashboard/interpretability contexts, not as a mirror of this enumeration - except stakeholder-communication-ai-decisions.md:57-62, which is a parallel five-member list in the same paraphrase dialect and is booked as idx-26e rather than swept in. What the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp closing §3 now vouches for, stated explicitly rather than left implicit: a seven-member enumeration whose NAMES are all the source's own segment names and whose descriptions for items 1, 3, 6 and 7 were verified against the live source on the date the stamp already carries. Items 2, 4 and 5 descriptions are NOT covered by that widening - idx-26f is open against them."
},
{
"id": "idx-27b",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced by the neighbourhood sweep for idx-27. The Mønster 3 \"Azure implementasjon\" list asserts a Copilot Studio \"Powered by AI\" disclosure in the chat interface - the same claim idx-27 removed from the Copilot Studio section, and false for the same reason: faqs-generative-orchestration documents a default transparency message (\"Just so you are aware, I sometimes use AI to answer your questions.\"), not a badge. Different section, outside idx-27's anchors, so booked separately rather than swept in. Repair is a replacement, outside the delete-only envelope. Note for whoever repairs this: after idx-27 the corrected wording already exists a few hundred lines below, so the repair is a copy of an already-ratified formulation rather than a fresh judgement.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-orchestration",
"anchors": [
"- **Copilot Studio**: \"Powered by AI\" disclosure i chat interface"
]
},
{
"id": "idx-26e",
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced by the cross-file grep run before the idx-26d edit. This file carries a parallel five-member list describing the Responsible AI Scorecard under the heading \"Konfigurerbare elementer\", in the same paraphrase dialect idx-26d just removed from transparency-documentation-standards.md: Dataset-helse, Modell-ytelse, Target values, Fortolkningsevne, Fairness assessment. Two problems, and they are distinguishable. First, the names are paraphrases of the source's scorecard SEGMENTS (data analysis, model performance, top important factors, fairness insights), so the list is the same drift class as idx-26d. Second, and separately, the heading claims these are CONFIGURABLE elements; the source page enumerates seven scorecard segments and does not enumerate a set of configurable elements, so the framing itself is unsupported. The list sits under a *Confidence: Verified (MCP microsoft-learn)* stamp carrying no date. Booked, not folded into idx-26d: idx-26d's ratified scope was one list in one file, and this file was not re-verified against the live source this session.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"**Konfigurerbare elementer**:",
"- **Fortolkningsevne**: Global/lokal feature importance"
]
},
{
"id": "idx-26f",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced while writing idx-26d, and deliberately excluded from it. Items 2 and 5 of the scorecard enumeration carry specifics the live source does not state: the fairness insights segment is described with \"(gender, ethnicity, age)\" where the source says only \"your desired sensitive groups\", and the data analysis segment with \"missing values, outlier analysis\" where the source says only that it \"shows you characteristics of your data\". Neither is false on its face - both are plausible instances - but both sit under the **Confidence:** Verified (MCP: microsoft-learn, 2026-08-03) stamp, which now vouches for an enumeration re-verified against the live source. This is a distinct class from idx-26d's name drift and idx-26b's cross-attribution: unsourced SPECIFICITY under a verification marker, where the defect is the marker's reach rather than the sentence. idx-26d's resolution states explicitly that its widening of the stamp does not cover items 2, 4 and 5, so the stamp is currently honest only because that exclusion is written down - closing this entry is what would make it honest on its own terms. Item 4's description was checked in the same pass and is a recognisable paraphrase of the source, not added specificity, so it is named here as checked-and-clear rather than left ambiguous.",
"evidence": "docs/r11-pilot-results.md §9.8; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"2. **Fairness insights**: Performance disparities across sensitive groups (gender, ethnicity, age)",
"5. **Data analysis**: Dataset statistics, missing values, outlier analysis"
]
}
]

View file

@ -111,11 +111,11 @@ PDF-rapport designet for å dele model- og data-innsikter mellom tekniske og ikk
**Komponenter i Scorecard:**
1. **Model overview**: Architecture, training data, intended use
2. **Fairness assessment**: Performance disparities across sensitive groups (gender, ethnicity, age)
3. **Model interpretability**: Feature importance (global/local explanations)
4. **Causal inference**: Causal vs correlational relationships i features
5. **Data quality**: Dataset statistics, missing values, outlier analysis
1. **Summary / model overview**: Modelloversikt og de target-verdiene du har satt
2. **Fairness insights**: Performance disparities across sensitive groups (gender, ethnicity, age)
3. **Top important factors**: Faktorene som påvirker modellens prediksjoner mest
4. **Causal insights**: Causal vs correlational relationships i features
5. **Data analysis**: Dataset statistics, missing values, outlier analysis
6. **Model performance**: Modellens viktigste metrikker og prediksjonskarakteristikker, målt mot target values du har satt
7. **Cohorts**: Best og dårligst presterende datakohorter og subgrupper, automatisk uttrukket av scorecard-en
@ -425,17 +425,20 @@ https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/transparen
| Component | Disclosure |
|-----------|------------|
| **Chat interface** | "Powered by AI" badge i chat window |
| **Chat interface** | Standard transparensmelding: "Just so you are aware, I sometimes use AI to answer your questions." |
| **Generative answers** | Attribution links til source documents |
| **Plugin actions** | Confirmation prompts før sensitive actions (send email, delete file) |
| **Data usage** | Privacy statement link i bot settings |
**Maker-konfigurerte kontroller (ikke innebygde disclosures):**
- Bekreftelse før dataendrende verktøy: makers *kan* kreve brukerbekreftelse før verktøy som endrer data kjøres. Kilden oppgir dette blant sikringene makeren konfigurerer, og sier ikke at bekreftelse gis uten slik konfigurasjon.
**Customization:**
- Copilot Studio generative AI toolkit: Pre-built "AI disclosure" topic
- Adaptive cards: Template for transparency notices
**Responsible AI FAQ:**
https://learn.microsoft.com/en-us/microsoft-copilot-studio/responsible-ai-overview
- Oversikt (hub): https://learn.microsoft.com/en-us/microsoft-copilot-studio/responsible-ai-overview
- Generativ orkestrering — kilden for transparensmeldingen og bekreftelses-sikringen over: https://learn.microsoft.com/en-us/microsoft-copilot-studio/faqs-generative-orchestration
---