fix(ms-ai-architect): Sesjon 10 — K9-destale eng+infra (volatil status → ref)

Begge S10-kriterier møtt (operatør-gated kald LLM-judge): engineering K9 PASS,
infrastructure K9 PASS. Deterministisk ikke regredert: validate 239 · kb-eval 15
· kb-update 122 · kb-integrity 192/192. K4 hevet 4→5 begge skills.

Engineering: modellversjoner (GPT-4o/Whisper/Florence/text-embedding-3) →
generisk kapabilitets-framing; MAF-«erstatter SK»-transisjon + Foundry Agent
Service GA → pekere; «Start med GPT-4o» (ABSENT i ref) droppet.

Infrastructure: SLA-tabell (ABSENT i ref) → relativ-veiledning + MCP-peker;
RTO/RPO-tabell → sammendrag + peker (ref bærer hele tier-modellen); «peak+30%»
(ref sier 20%) → generisk; Phi-3/Phi-4 + parametertall → «Phi-familien»
(fjernet faktafeil «Phi-4 14B»).

Metode: read-only verifiserings-agenter kartla ref-dekning FØR sletting.
judge-results.json oppdatert med ferske kalde eng+infra-verdikter (_updated: S10).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REiKFhP4w6xGXXqWKpPCJJ
This commit is contained in:
Kjell Tore Guttormsen 2026-06-20 09:31:21 +02:00
commit 6b3c44adc3
4 changed files with 48 additions and 39 deletions

View file

@ -13,11 +13,12 @@
"K9_noTimeSensitive": { "pass": true, "findings": ["Only meta-instructions (preview/GA as dynamic-to-verify) + stable identifiers (M365 SKUs, MADR v3.0). No stale-able product claim in body."] }
},
"ms-ai-engineering": {
"K1_triggerPrecision": { "provisional": true, "precision": 0.95, "notes": "19/20. Miss: 'Semantic Kernel agent governance' overlaps security/governance because description lists 'governance' under agent-orchestration. Operator must curate." },
"K4_noDuplication": { "score": 4, "pass": true, "evidence": "Mostly router. Two inline tables (RAG-vs-finetuning; MLOps test-types w/ Ragas/red-teaming) restate comparison content also in ref files — summary-level, not verbatim." },
"K7_imperativeStyle": { "ratio": 0.9, "pass": true, "notes": "9/10; section intros are descriptive prose by design." },
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers; schema drift ('Dato:' vs 'Last updated'); no source-URL in headers." },
"K9_noTimeSensitive": { "pass": false, "findings": ["Line 119: 'Foundry Agent Service GA' — explicit GA-status claim in body", "GPT-4o/Whisper/text-embedding-3/Florence — version-pinned product names in body", "Line 48: 'Agent Framework (erstatter Semantic Kernel Agents)' — lifecycle/transition claim", "A2A/CUA/Foundry Workflows — era-bound feature names"] }
"_updated": "S10 (2026-06-20) — K9-destale re-judge (cold)",
"K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20 from description (cold re-judge). Clean sibling separation (governance/security/cost/advisor/infra not triggered). Operator must curate final 20." },
"K4_noDuplication": { "score": 5, "pass": true, "evidence": "S10 re-judge: 7 section intros are orientation prose routing to references/<domain>/ + named kjernefiler; the two body tables (RAG-vs-finetuning, MLOps test-types) have no verbatim row-match in refs. No duplication." },
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative." },
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers across 5 domains; format inconsistent (EN/NO, month vs day granularity)." },
"K9_noTimeSensitive": { "pass": true, "findings": [] }
},
"ms-ai-governance": {
"_updated": "S9 (2026-06-20) — K4+K9 body-cleanup re-judge",
@ -28,11 +29,12 @@
"K9_noTimeSensitive": { "pass": true, "findings": [] }
},
"ms-ai-infrastructure": {
"K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20. Tightly scoped + named trigger phrases, low false-positive risk. Borderline (hybrid RAG, multi-region cost) need real-corpus validation. Operator must curate." },
"K4_noDuplication": { "score": 4, "pass": true, "evidence": "Largely summary/routing w/ '> Ref:' deferral. Minor: RTO/RPO table + SLA table put specific values in body that also live in ref files." },
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 imperative (concentrated in per-section directives)." },
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 'Last updated: 2026-02' + Status. NO source-URL in headers; identical batch date reads as generation timestamp, not per-file verification." },
"K9_noTimeSensitive": { "pass": false, "findings": ["Lines 93-98: SLA table w/ hardcoded percentages (99.9/99.999/99.95%) + 'Standard v2'", "Lines 149/158/184/186/249: Phi-3/Phi-4 model versions + param counts (3.8B/14B) in body", "Line 141: 'Azure Local (tidl. Azure Stack HCI)' rename note", "Line 87: hardcoded 'peak + 30% buffer'"] }
"_updated": "S10 (2026-06-20) — K9-destale re-judge (cold)",
"K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20 from description (cold). Sharply scoped via explicit 'Triggers on:' phrases; borderline negatives (plain RAG vs hybrid-RAG; MLOps vs infra) resolve correctly. Operator must curate final 20." },
"K4_noDuplication": { "score": 5, "pass": true, "evidence": "S10 re-judge: consistent summary+pointer pattern; §1.2 RTO/RPO now ~2 lines delegating to bcdr/rto-rpo-planning-ai-services.md (265 lines). SLA table replaced by relative-guidance prose. No procedural duplication." },
"K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative/infinitive." },
"K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 'Last updated: 2026-02' + Status; no source-URL on header line." },
"K9_noTimeSensitive": { "pass": true, "findings": [] }
},
"ms-ai-security": {
"_updated": "S9 (2026-06-20) — K4+K9 body-cleanup; K4 re-judged after S8 weighting fix + S9 risk-table/P10P50P90 fixes",