docs(architect): weekly KB update — 52 files refreshed (2026-04)

Key content changes:
- MLOps: MLflow 3 scorers expanded (RetrievalRelevance, Fluency, multi-turn judges)
- MLflow 3 A/B eval: mirror_traffic GA confirmed, new scorer catalog
- CI/CD: OIDC auth replaces deprecated --sdk-auth (Azure ML GitHub Actions)
- Agent framework A2A: updated SDK patterns (A2ACardResolver, BearerAuth)
- AG-UI backend tool rendering: accurate TOOL_CALL_* event shapes
- Computer Use agents: US region requirement, credentials patterns
- Purview governance: bulk term edit, expire/delete workflows
- CAF AI Secure: 3-phase structure confirmed current
- Copilot Studio: Claude Sonnet 4.5/4.6 GA, new orchestration controls
- M365 manifest: v1.26 GA (April 2026), copilotAgents node
- Power Platform: agent flow capacity enforcement corrected
- Azure Monitor: Simple Log Alerts GA, AMBA for policy-based alerting
- Security Copilot: SCU capacity model (400 SCU/1000 users)
- EU Data Boundary: all EU + EFTA countries confirmed
- gateway-multi-backend: added 4th topology, subscription-level quota note

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-04-10 11:31:11 +02:00
commit be4925a8ff
40 changed files with 398 additions and 239 deletions

View file

@ -1,6 +1,6 @@
# Model Evaluation Frameworks and Metrics
**Last updated:** 2026-02
**Last updated:** 2026-04
**Verified:** MCP 2026-04
**Status:** GA
**Category:** MLOps & GenAIOps
@ -119,10 +119,11 @@ MLflow 3 provides the evaluation framework for both traditional ML and GenAI app
| Type | Customization | Use Case |
|------|--------------|---------|
| Built-in judges | Minimal | Quick evaluation: `Correctness`, `RetrievalGroundedness`, `Safety` |
| Guidelines judges | Moderate | Custom natural-language rules (pass/fail) |
| Built-in judges | Minimal | Quick evaluation: `Correctness`, `RetrievalGroundedness`, `Safety`, `RelevanceToQuery`, `Fluency`, `Equivalence` — Verified (MCP 2026-04) |
| Guidelines judges | Moderate | Custom natural-language rules (pass/fail): `Guidelines`, `ExpectationsGuidelines` |
| Custom LLM judges | Full | Domain-specific criteria, detailed scoring |
| Code-based scorers | Full | Deterministic: exact match, format validation, business logic |
| Multi-turn judges | Minimal | Conversation-level: `ConversationCompleteness`, `UserFrustration`, `KnowledgeRetention`, `ConversationalSafety` — Verified (MCP 2026-04) |
**Key evaluation functions**:
```python