docs(architect): weekly KB update — 52 files refreshed (2026-04)
Key content changes: - MLOps: MLflow 3 scorers expanded (RetrievalRelevance, Fluency, multi-turn judges) - MLflow 3 A/B eval: mirror_traffic GA confirmed, new scorer catalog - CI/CD: OIDC auth replaces deprecated --sdk-auth (Azure ML GitHub Actions) - Agent framework A2A: updated SDK patterns (A2ACardResolver, BearerAuth) - AG-UI backend tool rendering: accurate TOOL_CALL_* event shapes - Computer Use agents: US region requirement, credentials patterns - Purview governance: bulk term edit, expire/delete workflows - CAF AI Secure: 3-phase structure confirmed current - Copilot Studio: Claude Sonnet 4.5/4.6 GA, new orchestration controls - M365 manifest: v1.26 GA (April 2026), copilotAgents node - Power Platform: agent flow capacity enforcement corrected - Azure Monitor: Simple Log Alerts GA, AMBA for policy-based alerting - Security Copilot: SCU capacity model (400 SCU/1000 users) - EU Data Boundary: all EU + EFTA countries confirmed - gateway-multi-backend: added 4th topology, subscription-level quota note Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
parent
6645e93205
commit
be4925a8ff
40 changed files with 398 additions and 239 deletions
|
|
@ -1,6 +1,6 @@
|
|||
# Model Evaluation Frameworks and Metrics
|
||||
|
||||
**Last updated:** 2026-02
|
||||
**Last updated:** 2026-04
|
||||
**Verified:** MCP 2026-04
|
||||
**Status:** GA
|
||||
**Category:** MLOps & GenAIOps
|
||||
|
|
@ -119,10 +119,11 @@ MLflow 3 provides the evaluation framework for both traditional ML and GenAI app
|
|||
|
||||
| Type | Customization | Use Case |
|
||||
|------|--------------|---------|
|
||||
| Built-in judges | Minimal | Quick evaluation: `Correctness`, `RetrievalGroundedness`, `Safety` |
|
||||
| Guidelines judges | Moderate | Custom natural-language rules (pass/fail) |
|
||||
| Built-in judges | Minimal | Quick evaluation: `Correctness`, `RetrievalGroundedness`, `Safety`, `RelevanceToQuery`, `Fluency`, `Equivalence` — Verified (MCP 2026-04) |
|
||||
| Guidelines judges | Moderate | Custom natural-language rules (pass/fail): `Guidelines`, `ExpectationsGuidelines` |
|
||||
| Custom LLM judges | Full | Domain-specific criteria, detailed scoring |
|
||||
| Code-based scorers | Full | Deterministic: exact match, format validation, business logic |
|
||||
| Multi-turn judges | Minimal | Conversation-level: `ConversationCompleteness`, `UserFrustration`, `KnowledgeRetention`, `ConversationalSafety` — Verified (MCP 2026-04) |
|
||||
|
||||
**Key evaluation functions**:
|
||||
```python
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue