fix(ms-ai-architect): Sesjon 9 - K4+K9 body-cleanup security + governance
Tynner duplisert/volatil body-info i de to skills som feilet begge kriterier, til pekere mot kanoniske ref-filer. Fersk operatør-gated LLM-judge: security K4 5/5 + K9 PASS; governance K4 5/5 + K9 PASS. Security (K9): Defender GA/preview/Azure-Gov-status, §3-ytelsestall, GPT-4o-par og PTU break-even ut av body. Uverifiserte tall (20-50ms, 5-10x, 80% latens — fantes i ingen ref) droppet, ikke flyttet. Security (K4): risikoklassifiserings-tabellen (divergerte fra rubrikk: 6 band vs 5, manglet Uakseptabel) + P10/P50/P90-tabellen (flat x0.6/x1.8 motsa kanonisk per-komponent-modell, agentens OBLIGATORISK-kilde) -> peker. Antakelse-test bestaatt: ingenting asserter body-tallene. Governance (K9): EU Data Boundary §2.3 - droppet volatile regioner (Sweden Central/West Europe; body motsa refs som sier Norway East/West Europe), repek kryss-ref til gdpr-compliance-ai-systems.md. Governance (K4): §6.2 AI Act-tre - kollapset oppramsede (og ufullstendige) Art.5 + Annex III-lister til pekere mot ai-act-classification-methodology.md; §2.1 eneste oversiktstabell. judge-prompt.md K9 strammet: stabile identifikatorer (forordningsaar, OWASP 2025, MADR v3.0, saksnr) eksplisitt ute av scope. judge-results.json oppdatert med ferske sec+gov-verdikter. Gates: validate 239 · kb-eval 15 · kb-update 122 · kb-integrity 192/192. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01REiKFhP4w6xGXXqWKpPCJJ
This commit is contained in:
parent
2356d211ee
commit
c4230d2e88
4 changed files with 40 additions and 56 deletions
|
|
@ -20,11 +20,12 @@
|
|||
"K9_noTimeSensitive": { "pass": false, "findings": ["Line 119: 'Foundry Agent Service GA' — explicit GA-status claim in body", "GPT-4o/Whisper/text-embedding-3/Florence — version-pinned product names in body", "Line 48: 'Agent Framework (erstatter Semantic Kernel Agents)' — lifecycle/transition claim", "A2A/CUA/Foundry Workflows — era-bound feature names"] }
|
||||
},
|
||||
"ms-ai-governance": {
|
||||
"K1_triggerPrecision": { "provisional": true, "precision": 0.95, "notes": "19/20. Miss: Schrems II / data-transfer — body section 2.3 covers it but description has NO Schrems/dataoverføring keyword. Add trigger phrase to description." },
|
||||
"K4_noDuplication": { "score": 3, "pass": false, "evidence": "Moderate duplication: §6.1 DPIA risk-factor tree (incl. >=2-faktorer threshold) duplicates dpia-norwegian-methodology-ai.md; §6.2 + §2.1 AI Act taxonomy duplicates ai-act-classification-methodology.md; §1.2 Digdir 7-principle table restates per-principle files. Regulatory update must be applied twice." },
|
||||
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 imperative." },
|
||||
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers + Status + inline source citations (Lovdata/NSM/europalov). BONUS DEFECT: SKILL.md line 191 references 'drift-detection-automated-retraining.md' which does NOT exist (actual: model-performance-drift-detection.md) — broken ref path." },
|
||||
"K9_noTimeSensitive": { "pass": false, "findings": ["Line 121: 'Microsoft EU Data Boundary ... Azure OpenAI (Sweden Central, West Europe)' — volatile region/availability claim in body (most genuine finding)", "Lines 81/121: 'Regulation 2024/1689' + 'Schrems II (C-311/18)' — legal identifiers (stable, borderline)", "§2.1/§6.2 AI Act obligation status presented without Digital Omnibus caveat — may date"] }
|
||||
"_updated": "S9 (2026-06-20) — K4+K9 body-cleanup re-judge",
|
||||
"K1_triggerPrecision": { "provisional": true, "precision": 0.85, "notes": "17/20 (7/10 in-domain hits, 10/10 out-of-domain correct). RECALL GAP: description has no Schrems II/dataoverføring/TIA trigger though §2.3 covers it; GDPR + 'consult Datatilsynet' also under-represented vs body scope. Operator must curate + add Schrems II trigger (S11)." },
|
||||
"K4_noDuplication": { "score": 5, "pass": true, "evidence": "S9 FIX: §6.2 now a compact decision-flow with explicit pointers ('[full forbudsliste i ai-act-classification-methodology.md]', '[åtte kategorier ...]') — Art.5 + Annex III lists no longer enumerated in body; they live only in references/responsible-ai/ai-act-classification-methodology.md. §2.1 is the single 4-level overview table ('ikke gjenta dem her'). §6.1 (DPIA tree), §1.2 (Digdir table), §6.3/§6.4 are routing/orientation, not verbatim copies of ref files." },
|
||||
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative/infinitive." },
|
||||
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 sampled refs carry Last updated + Status + Category headers." },
|
||||
"K9_noTimeSensitive": { "pass": true, "findings": [] }
|
||||
},
|
||||
"ms-ai-infrastructure": {
|
||||
"K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20. Tightly scoped + named trigger phrases, low false-positive risk. Borderline (hybrid RAG, multi-region cost) need real-corpus validation. Operator must curate." },
|
||||
|
|
@ -34,10 +35,11 @@
|
|||
"K9_noTimeSensitive": { "pass": false, "findings": ["Lines 93-98: SLA table w/ hardcoded percentages (99.9/99.999/99.95%) + 'Standard v2'", "Lines 149/158/184/186/249: Phi-3/Phi-4 model versions + param counts (3.8B/14B) in body", "Line 141: 'Azure Local (tidl. Azure Stack HCI)' rename note", "Line 87: hardcoded 'peak + 30% buffer'"] }
|
||||
},
|
||||
"ms-ai-security": {
|
||||
"K1_triggerPrecision": { "provisional": true, "precision": 0.95, "notes": "19/20. Clean sibling separation (AI Act/DPIA/platform/RAG/BCDR not triggered). One ambiguity: model feature-comparison could be pulled by broad 'performance optimization for AI'. Operator must curate." },
|
||||
"K4_noDuplication": { "score": 3, "pass": false, "evidence": "DUPLICATION-WITH-CONTRADICTION: body weighting table (L52-59 Standard: Identity 20/Network 15/Data 20/Content 20/Compliance 15/Monitoring 10) CONTRADICTS canonical rubric security-scoring-rubrics-6x5.md (Compliance 25/Data 20/Identity 20/Content 15/Network 10/Monitoring 10). Body scoring rule (weighted sum) also diverges from rubric (Ja-checkpoint count). Real correctness bug." },
|
||||
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 imperative." },
|
||||
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers; mostly no header source-URL (URLs in body cost-register rows)." },
|
||||
"K9_noTimeSensitive": { "pass": false, "findings": ["Line 73: 'OWASP LLM Top 10 (2025)' — dated standard version", "Line 92: 'Defender ... GA for AI applications, Preview for AI agents' + 'ikke i Azure Government' — GA/preview + region status in body", "Line 138: 'GPT-4o mini vs GPT-4o' — model versions", "Lines 149-153: hardcoded perf/price figures (20-50ms, 5-10x, 50%/80%, Batch 50% @ 24h SLA)"] }
|
||||
"_updated": "S9 (2026-06-20) — K4+K9 body-cleanup; K4 re-judged after S8 weighting fix + S9 risk-table/P10P50P90 fixes",
|
||||
"K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20 from description; clean sibling separation (governance/engineering/advisor/infrastructure/license not triggered). Borderline: generic latency/perf prompts overlap engineering, but description explicitly claims 'performance optimization for AI'. Operator must curate final 20." },
|
||||
"K4_noDuplication": { "score": 5, "pass": true, "evidence": "S9 re-judge (cold, post-fix): (a) 6x5 weights — body L50 routes to security-scoring-rubrics-6x5.md ('Ikke dupliser vekttallene her'), no body numbers; (b) risk-classification thresholds — body L54 routes to same rubric ('Ikke dupliser terskeltallene her'), canonical mapping incl. 1.00-1.49 Uakseptabel lives only in rubric; (c) P10/P50/P90 — body L94 affirms 'per komponent (ikke flat multiplikator)' owned by deterministic-cost-calculation-model.md §3, concrete factors only in cost model — body affirms, does not contradict; (d) OWASP table + §3 perf are routing/orientation with explicit volatile-numbers-live-in-refs note. No duplication/contradiction." },
|
||||
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative/infinitive." },
|
||||
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers + Status; 3/5 also carry Verified: MCP <date>." },
|
||||
"K9_noTimeSensitive": { "pass": true, "findings": [] }
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -24,8 +24,12 @@ Vær streng og adversariell — ikke ros, ikke pynt på tall.
|
|||
imperativ/infinitiv. pass = ratio ≥ 0.80.
|
||||
- **K8 kildehenvisning i ref-filer:** for 5 samplede ref-filer, har header `Last updated` /
|
||||
`Verified` / kilde-URL? Rapporter andel. pass = ratio ≥ 0.80.
|
||||
- **K9 ingen tid-sensitiv info i SKILL.md body:** finnes datoer/versjoner/GA/preview-status DIREKTE
|
||||
i body (ikke i ref-filer)? List funn. pass = ingen funn.
|
||||
- **K9 ingen VOLATIL tid-sensitiv info i SKILL.md body:** finnes **volatile** påstander DIREKTE i body
|
||||
(ikke i ref-filer) — GA/preview-release-status, modellversjoner, SLA-/ytelses-/pris-tall, regions-
|
||||
tilgjengelighet? List funn. pass = ingen volatile funn. **Utenfor scope (ikke funn):** stabile
|
||||
identifikatorer som forordningsår (2024/1689), lovsaksnr (C-311/18), standard-versjonsnavn
|
||||
(OWASP LLM Top 10 2025, MADR v3.0), filnavn og generisk forklarende «preview»/«GA» uten konkret
|
||||
produkt-status. Disse er navngitte identifikatorer, ikke ferskhets-sensitive påstander.
|
||||
|
||||
RETURNER KUN dette JSON-objektet (ingen annen tekst, ingen markdown-fence):
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue