{ "_meta": { "rubric": "scripts/kb-eval/judge-prompt.md", "judge_model": "opus", "method": "5 parallel adversarial LLM-judges, one per skill", "note": "K1 precision is PROVISIONAL — operator must curate the final 20 trigger-prompts per skill before K1 is authoritative." }, "ms-ai-advisor": { "K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20 from description. BORDERLINE: cost/diagram/DPIA/AI-Act prompts are MS-AI-adjacent; broad 'Microsoft AI architecture' phrase risks over-triggering vs sibling commands. Operator must curate + stress-test sibling overlap." }, "K4_noDuplication": { "score": 5, "pass": true, "evidence": "Body = persona + 7-phase workflow + ref-index; no ref detail reproduced. Only internal MCP-table redundancy (SKILL-internal, not SKILL<->ref)." }, "K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative." }, "K8_sourceCitation": { "ratio": 0.8, "pass": true, "notes": "AT THRESHOLD: architecture/decision-trees.md lacks dated header (footer source only). Add dated header to harden margin." }, "K9_noTimeSensitive": { "pass": true, "findings": ["Only meta-instructions (preview/GA as dynamic-to-verify) + stable identifiers (M365 SKUs, MADR v3.0). No stale-able product claim in body."] } }, "ms-ai-engineering": { "K1_triggerPrecision": { "provisional": true, "precision": 0.95, "notes": "19/20. Miss: 'Semantic Kernel agent governance' overlaps security/governance because description lists 'governance' under agent-orchestration. Operator must curate." }, "K4_noDuplication": { "score": 4, "pass": true, "evidence": "Mostly router. Two inline tables (RAG-vs-finetuning; MLOps test-types w/ Ragas/red-teaming) restate comparison content also in ref files — summary-level, not verbatim." }, "K7_imperativeStyle": { "ratio": 0.9, "pass": true, "notes": "9/10; section intros are descriptive prose by design." }, "K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers; schema drift ('Dato:' vs 'Last updated'); no source-URL in headers." }, "K9_noTimeSensitive": { "pass": false, "findings": ["Line 119: 'Foundry Agent Service GA' — explicit GA-status claim in body", "GPT-4o/Whisper/text-embedding-3/Florence — version-pinned product names in body", "Line 48: 'Agent Framework (erstatter Semantic Kernel Agents)' — lifecycle/transition claim", "A2A/CUA/Foundry Workflows — era-bound feature names"] } }, "ms-ai-governance": { "_updated": "S9 (2026-06-20) — K4+K9 body-cleanup re-judge", "K1_triggerPrecision": { "provisional": true, "precision": 0.85, "notes": "17/20 (7/10 in-domain hits, 10/10 out-of-domain correct). RECALL GAP: description has no Schrems II/dataoverføring/TIA trigger though §2.3 covers it; GDPR + 'consult Datatilsynet' also under-represented vs body scope. Operator must curate + add Schrems II trigger (S11)." }, "K4_noDuplication": { "score": 5, "pass": true, "evidence": "S9 FIX: §6.2 now a compact decision-flow with explicit pointers ('[full forbudsliste i ai-act-classification-methodology.md]', '[åtte kategorier ...]') — Art.5 + Annex III lists no longer enumerated in body; they live only in references/responsible-ai/ai-act-classification-methodology.md. §2.1 is the single 4-level overview table ('ikke gjenta dem her'). §6.1 (DPIA tree), §1.2 (Digdir table), §6.3/§6.4 are routing/orientation, not verbatim copies of ref files." }, "K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative/infinitive." }, "K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 sampled refs carry Last updated + Status + Category headers." }, "K9_noTimeSensitive": { "pass": true, "findings": [] } }, "ms-ai-infrastructure": { "K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20. Tightly scoped + named trigger phrases, low false-positive risk. Borderline (hybrid RAG, multi-region cost) need real-corpus validation. Operator must curate." }, "K4_noDuplication": { "score": 4, "pass": true, "evidence": "Largely summary/routing w/ '> Ref:' deferral. Minor: RTO/RPO table + SLA table put specific values in body that also live in ref files." }, "K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 imperative (concentrated in per-section directives)." }, "K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 'Last updated: 2026-02' + Status. NO source-URL in headers; identical batch date reads as generation timestamp, not per-file verification." }, "K9_noTimeSensitive": { "pass": false, "findings": ["Lines 93-98: SLA table w/ hardcoded percentages (99.9/99.999/99.95%) + 'Standard v2'", "Lines 149/158/184/186/249: Phi-3/Phi-4 model versions + param counts (3.8B/14B) in body", "Line 141: 'Azure Local (tidl. Azure Stack HCI)' rename note", "Line 87: hardcoded 'peak + 30% buffer'"] } }, "ms-ai-security": { "_updated": "S9 (2026-06-20) — K4+K9 body-cleanup; K4 re-judged after S8 weighting fix + S9 risk-table/P10P50P90 fixes", "K1_triggerPrecision": { "provisional": true, "precision": 1.0, "notes": "20/20 from description; clean sibling separation (governance/engineering/advisor/infrastructure/license not triggered). Borderline: generic latency/perf prompts overlap engineering, but description explicitly claims 'performance optimization for AI'. Operator must curate final 20." }, "K4_noDuplication": { "score": 5, "pass": true, "evidence": "S9 re-judge (cold, post-fix): (a) 6x5 weights — body L50 routes to security-scoring-rubrics-6x5.md ('Ikke dupliser vekttallene her'), no body numbers; (b) risk-classification thresholds — body L54 routes to same rubric ('Ikke dupliser terskeltallene her'), canonical mapping incl. 1.00-1.49 Uakseptabel lives only in rubric; (c) P10/P50/P90 — body L94 affirms 'per komponent (ikke flat multiplikator)' owned by deterministic-cost-calculation-model.md §3, concrete factors only in cost model — body affirms, does not contradict; (d) OWASP table + §3 perf are routing/orientation with explicit volatile-numbers-live-in-refs note. No duplication/contradiction." }, "K7_imperativeStyle": { "ratio": 1.0, "pass": true, "notes": "10/10 sampled instruction sentences imperative/infinitive." }, "K8_sourceCitation": { "ratio": 1.0, "pass": true, "notes": "5/5 dated headers + Status; 3/5 also carry Verified: MCP ." }, "K9_noTimeSensitive": { "pass": true, "findings": [] } } }