ms-ai-architect/scripts/kb-eval/data/g7-review-queue.json
Kjell Tore Guttormsen 8cde10f09a fix(ms-ai-architect): G7 idx 26 lukket — scorecard-segmenter rettet, dashboard-komponenter merket
Error analysis og Counterfactual analysis sto som item 4/5 i 'Komponenter i
Scorecard'. Begge kilder hentet live denne okten og bekrefter malingen
uavhengig av 9.6: how-to-responsible-ai-scorecard enumererer summary/model
overview, data analysis, model performance, cohorts, top important factors,
fairness insights og causal insights; concept-responsible-ai-dashboard lister
Error analysis og Counterfactual what-if som *dashboard*-komponenter.

Form (b), operatorratifisert: relabel framfor fjerning. Lista renummererer
rent til 1-5 — 1,2,3,4,6,7-artefakten var tvungen kun inne i delete-only-
konvolutten, ikke for en ordinaer Edit. De to kapabilitetene beholdes i et
sitatblokk eksplisitt merket som dashboard-komponenter, sa kildebekreftet
informasjon overlever og leseren advares mot nettopp den forvekslingen som
skapte defekten. Lokator 2: Risk assessment-raden leser na 'fairness insights'.

Confidence-stempelet tolv linjer under vouchet for den falske lista og la
utenfor enhver maskinsjekk (V1/V2/V2b/V3 er strenginvarianter; check-g7-queue
tester kun ankere). Beholdt, men datert 2026-08-03 for a fore re-verifiseringen.
Dokument-koherens lest etter editen (c569bdc-laerdommen).

Nabodefekt funnet av filsveipet bokfort separat som idx-26b, ikke foldet inn:
'Quantitative analyses' attribueres til Scorecard, men er en Model Card-seksjon
(samme fil, linje 77) — kryss-attribuering mellom to standarder.

Suite 1047/1047.

[skip-docs]
2026-08-03 21:23:35 +02:00

99 lines
9.6 KiB
JSON

{
"_meta": {
"gap": "G7",
"form": "(b) named queue into the human review phase",
"ratified": "2026-08-03",
"rationale": "Measured in R11 §9.6: 2 of the 4 subtractions applied in 957ebef left a residue, so residues are the normal by-product of a delete-only envelope rather than an exception. Two of the members are replacements, not multi-locator cases, which a deletion-oriented O4 class would not have fixed. A queue absorbs both classes; an O4 return contract would have been mis-sized against the evidence.",
"contract": "Anchors are verbatim strings, never line numbers (line ≠ real_line in 9 of 17 R11 records). An open entry whose anchor no longer occurs in its file is drift, and check-g7-queue.mjs fails rather than passing it silently. Nothing in this queue is machine-appliable by definition — every entry is outside the O2 envelope. Resolution is a human review act.",
"evidence": "docs/r11-pilot-results.md §9.4, §9.5, §9.6; docs/ref-kb-correctness-program-2026-06.md §8 G7"
},
"entries": [
{
"id": "idx-17",
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
"class": "replacement",
"status": "resolved",
"raised": "2026-08-03",
"summary": "The idx 17 subtraction (957ebef) deleted the three score-threshold bands but left the lead-in ending in a colon, promising an enumeration that no longer existed, immediately followed by a **Verified** stamp. The subtraction was correct; the paragraph it left was not. V1/V2/V2b/V3 are string invariants over deleted text and cannot see document coherence, so the machine could not have caught it.",
"evidence": "docs/r11-pilot-results.md §9.6",
"anchors": [],
"resolution": "Operator-ratified 2026-08-03: colon changed to a period, making the lead-in a complete and independently true sentence that the **Verified** stamp correctly covers. Corpus swept for the same defect shape (bold lead-in ending in colon, blank line, **Verified**) — no other occurrence."
},
{
"id": "idx-33",
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "The idx 33 subtraction removed AI asset inventory via Azure Resource Graph and the Purview Insider Risk Management bullet from an unsupported Defender for Cloud AISPM attribution. Correct — but the same capabilities survive in this file under Cloud Adoption Framework Secure AI attribution, stamped 'Verified MCP 2026-04', and are asserted in four other corpus files. The edit's benefit is corpus-wide smaller than the single line suggested. Whether the CAF attribution is itself supported has not been checked.",
"evidence": "docs/r11-pilot-results.md §9.5 (cross-corpus check), §9.6",
"anchors": [
"Oppdatert 2026-04: inkluderer nå AI asset inventory via Azure Resource Graph"
]
},
{
"id": "idx-26",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "multi-locator",
"status": "resolved",
"raised": "2026-08-03",
"summary": "The Responsible AI Scorecard component list names Error analysis (item 4) and Counterfactual analysis (item 5). Both were measured false against first-party docs 2026-08-03: the canonical scorecard segments are summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights and causal insights; Error analysis and Counterfactual analysis are Responsible AI *dashboard* components. The delete-only reduction could only remove item 5, because line 300 asserts Error analysis as scorecard content too — so a partial fix would have left a known-false claim standing while introducing a renumbering artifact (1,2,3,4,6,7). Operator declined the partial fix 2026-08-03 and sent the whole case here. Correct repair spans both locators.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"4. **Error analysis**: Error rates per cohort, confusion matrices",
"| **Risk assessment** | Responsible AI Scorecard: Error analysis, fairness assessment |"
],
"resolution": "Operator-ratified 2026-08-03, form (b) — relabel rather than remove. Both source pages were re-fetched live this session and confirm the measurement independently of §9.6: how-to-responsible-ai-scorecard enumerates summary/model overview, data analysis, model performance, cohorts, top important factors, fairness insights and causal insights; concept-responsible-ai-dashboard lists Error analysis and Counterfactual what-if among the dashboard components. Locator 1: items 4 and 5 removed from the numbered scorecard list, which renumbers cleanly to 1-5 — the 1,2,3,4,6,7 artifact was forced only inside the delete-only envelope and does not apply to an ordinary Edit. The two capabilities are retained in a blockquote explicitly marked as dashboard components rather than scorecard segments, so genuine source-confirmed information survives and the reader is warned off precisely the conflation that produced the defect. Locator 2: the Risk assessment row now reads 'fairness insights' alone. The **Confidence:** Verified stamp twelve lines below vouched for the false list and was silently outside every machine check (V1/V2/V2b/V3 are string invariants; check-g7-queue only tests anchors); it is kept but dated to 2026-08-03 to record the re-verification. Document coherence around both locators was read after the edit, per the c569bdc lesson. A neighbouring defect surfaced by the file sweep is booked separately as idx-26b rather than folded in here."
},
{
"id": "idx-26b",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Surfaced by the post-edit file sweep for idx 26, in the same compliance-mapping table as idx 26's second locator, but outside both of its anchors — so booked separately rather than folded in (gap discipline; operator-ratified 2026-08-03). The Accuracy metrics row attributes 'Quantitative analyses' to the Responsible AI Scorecard. That is not a scorecard segment name: how-to-responsible-ai-scorecard calls the corresponding segment 'model performance'. The defect is a cross-attribution between two different standards rather than mere imprecision — 'Quantitative Analyses' is a canonical Model Card section, and this same file lists it as one at line 77. Repair is a replacement, so it is outside the delete-only envelope. Whether the right fix is to rename the segment or to drop the row is not yet decided.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/azure/machine-learning/how-to-responsible-ai-scorecard",
"anchors": [
"| **Accuracy metrics** | Responsible AI Scorecard: Quantitative analyses |"
]
},
{
"id": "idx-27",
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
"class": "replacement",
"status": "open",
"raised": "2026-08-03",
"summary": "Failed out of O2 on cond 3 (§9.6): the cited source DOES establish the mechanism — faqs-generative-orchestration states 'Makers can require user confirmation before executing tools that modify data' — so deleting the Plugin-actions row would destroy source-confirmed information. The defect is modality, not fabrication: the file presents confirmation prompts as a built-in disclosure, whereas the source makes them maker-configured. The Chat-interface row in the same table is imprecise for the same reason: the FAQ documents a default transparency message ('Just so you are aware, I sometimes use AI to answer your questions.'), not a 'Powered by AI' badge. Repair is a replacement, outside the delete-only envelope.",
"evidence": "docs/r11-pilot-results.md §9.6; https://learn.microsoft.com/microsoft-copilot-studio/faqs-generative-orchestration",
"anchors": [
"| **Plugin actions** | Confirmation prompts før sensitive actions (send email, delete file) |",
"| **Chat interface** | \"Powered by AI\" badge i chat window |"
]
},
{
"id": "idx-36",
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-03",
"summary": "Applying idx 36 alone yields a severity table more precise than the prose that cites it, so a companion edit is required for the prose to match the narrowed table. Held back from 957ebef as out of envelope.",
"evidence": "docs/r11-pilot-results.md §9.4, §9.5",
"anchors": [
"Øker severity bar; krever mer robust adversarial defenses"
]
},
{
"id": "idx-18",
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
"class": "multi-locator",
"status": "open",
"raised": "2026-08-03",
"summary": "No reduction exists. The surviving claim is a whole titled section on Azure AI Search built-in caching plus a **Verified** row in the verification table, so no deletion confined to a single locator can repair the file. This is the member that most clearly motivated G7.",
"evidence": "docs/r11-pilot-results.md §9.3, §9.4",
"anchors": [
"**Automatic Caching Behavior:**",
"| Azure AI Search caching | **Verified** | Microsoft Learn docs (4, 6) |"
]
}
]
}