docs(ms-ai-architect): KB-refresh tema-b — Foundry-navnesveip «Azure AI Foundry»→«Microsoft Foundry» (233 filer)
Verifisert mot offisiell MS-doc (juni 2026): «Microsoft Foundry» er det gjeldende produkt-/portalnavnet; «Foundry (classic)» = gamle «Azure AI Foundry» (/azure/foundry/ vs /azure/foundry-classic/). Premiss bekreftet før sveip. Multi-regel, IKKE naiv s/Azure AI Foundry/Microsoft Foundry/ — MS dropper «Azure AI» (legger IKKE til «Microsoft») for to produktvarianter: - «Azure AI Foundry Agent[ Service|s]» → «Foundry Agent Service/Agents» (MS-form) - «Azure AI Foundry Models» → «Foundry Models» (i «Azure OpenAI in Foundry Models») - «Azure AI Foundry SDK» → «Microsoft Foundry SDK» (operatør-valg) - «Azure AI Foundry portal/project» + generisk → «Microsoft Foundry» - Pre-eksisterende «Microsoft Foundry Models» (4) normalisert → «Foundry Models» Bevart: «Azure OpenAI», «Azure AI Inference SDK», «Azure AI Search», «Azure AI Services», kode-IDer. Historisk ref «(tidligere Azure AI Foundry)» i model-catalog-2026.md beskyttet via lookbehind. URL /azure/ai-foundry/→ /azure/foundry/ kun i owasp-llm-top10 (KB-ref); docs/-filer deferred. Scope: skills (inkl. 3 SKILL.md) + commands + agents + README + CLAUDE. Ekskludert: docs/ (interne), playground/+tests/ fixtures (testdata), CHANGELOG.md (historisk logg), STATE.md (gitignored). 3 SKILL.md endret (advisor/engineering/security) → judge-cache teknisk invalidert for disse, men scorer uendret: advisor 91, eng/gov/infra/sec 96 (alle ≥90). validate 239/0. 0 «Azure AI Foundry» igjen (utenom bevart ref). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
20b522ab10
commit
03d596e4ec
233 changed files with 810 additions and 810 deletions
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
**Kategori:** MLOps & GenAIOps
|
||||
**Sist oppdatert:** 2026-06-19
|
||||
**Confidence:** High (basert på offisiell Microsoft dokumentasjon, Azure AI Foundry SDK, og MLflow 3)
|
||||
**Confidence:** High (basert på offisiell Microsoft dokumentasjon, Microsoft Foundry SDK, og MLflow 3)
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -36,7 +36,7 @@ Production evaluation i Microsoft AI-stakken består av fem hovedkomponenter som
|
|||
|
||||
### 1. Tracing Infrastructure
|
||||
|
||||
**Azure AI Foundry Tracing** og **MLflow Tracing** gir den datainfrastrukturen som all evaluering bygger på. Tracing logger automatisk:
|
||||
**Microsoft Foundry Tracing** og **MLflow Tracing** gir den datainfrastrukturen som all evaluering bygger på. Tracing logger automatisk:
|
||||
|
||||
- Input prompts og kontekst
|
||||
- Mellomsteg (retrieval-resultater, tool calls, reasoning)
|
||||
|
|
@ -121,7 +121,7 @@ Matematisk-baserte metrikker for tekstlikhet (krever ground truth):
|
|||
|
||||
Kontinuerlig evaluering kjører evaluatorer automatisk på production traffic med konfigurerbar sampling rate.
|
||||
|
||||
**Azure AI Foundry Continuous Evaluation (for Agents):**
|
||||
**Microsoft Foundry Continuous Evaluation (for Agents):**
|
||||
|
||||
```python
|
||||
from azure.ai.projects.models import (
|
||||
|
|
@ -228,7 +228,7 @@ Visualisering og alerting er kritisk for actionable insights.
|
|||
|
||||
Production evaluation er ikke komplett uten human-in-the-loop validering.
|
||||
|
||||
**Azure AI Foundry Review App:**
|
||||
**Microsoft Foundry Review App:**
|
||||
|
||||
- Domain experts kan review AI-genererte svar direkte fra dashboard
|
||||
- Thumbs up/down feedback lagres som evaluation data for future training
|
||||
|
|
@ -266,7 +266,7 @@ Dashboard + Alerts
|
|||
**Implementering:**
|
||||
|
||||
```python
|
||||
# Azure AI Foundry: sampling via max_hourly_runs
|
||||
# Microsoft Foundry: sampling via max_hourly_runs
|
||||
action=ContinuousEvaluationRuleAction(
|
||||
eval_id=eval_object.id,
|
||||
max_hourly_runs=100 # Hvis traffic er 1000/hour → 10% sampling
|
||||
|
|
@ -475,7 +475,7 @@ vulnerability_report = scan_results.get_vulnerability_summary()
|
|||
**A. Async evaluation** (anbefalt):
|
||||
|
||||
```python
|
||||
# Azure AI Foundry: Evaluation kjører async etter response er returnert
|
||||
# Microsoft Foundry: Evaluation kjører async etter response er returnert
|
||||
# Ingen user-facing latency impact
|
||||
event_type=EvaluationRuleEventType.RESPONSE_COMPLETED # Trigger AFTER response
|
||||
```
|
||||
|
|
@ -513,7 +513,7 @@ with mlflow.start_run():
|
|||
|
||||
## Integrasjon med Microsoft-stakken
|
||||
|
||||
### Azure AI Foundry + Application Insights
|
||||
### Microsoft Foundry + Application Insights
|
||||
|
||||
**Full stack monitoring:**
|
||||
|
||||
|
|
@ -648,7 +648,7 @@ traces = spark.read.table("catalog.schema.agent_traces")
|
|||
|
||||
Power Platform har begrenset native support for LLM evaluation i production. Anbefalt mønster:
|
||||
|
||||
1. **Lag custom connector til Azure AI Foundry Evaluation API**
|
||||
1. **Lag custom connector til Microsoft Foundry Evaluation API**
|
||||
2. **Lagre evaluation results i Dataverse**
|
||||
3. **Bygg Power BI dashboard for visualisering**
|
||||
|
||||
|
|
@ -674,7 +674,7 @@ Results → Dataverse custom table
|
|||
Power BI report
|
||||
```
|
||||
|
||||
**Gap:** Ingen out-of-the-box production evaluation. Microsoft roadmap (Q2 2026) inkluderer native integration med Azure AI Foundry evaluation. *(Medium confidence – based on public roadmap)*
|
||||
**Gap:** Ingen out-of-the-box production evaluation. Microsoft roadmap (Q2 2026) inkluderer native integration med Microsoft Foundry evaluation. *(Medium confidence – based on public roadmap)*
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -720,7 +720,7 @@ For AI-systemer klassifisert som high-risk (helse, lov, kritisk infrastruktur):
|
|||
- **Article 15:** Logging av input/output data – **tracing er lovpålagt**
|
||||
- **Article 61:** Post-market monitoring plan – **production evaluation er compliance requirement**
|
||||
|
||||
**Anbefaling:** Bruk Azure AI Foundry continuous evaluation med 100% sampling for high-risk AI. Lagre evaluation logs i minimum 5 år for audit purposes. *(High confidence – based on AI Act legal text)*
|
||||
**Anbefaling:** Bruk Microsoft Foundry continuous evaluation med 100% sampling for high-risk AI. Lagre evaluation logs i minimum 5 år for audit purposes. *(High confidence – based on AI Act legal text)*
|
||||
|
||||
### GDPR & Privacy i Production Evaluation
|
||||
|
||||
|
|
@ -799,7 +799,7 @@ if eval_results["metrics"]["violence_violations"] > 0:
|
|||
|
||||
## Kostnad og lisensiering
|
||||
|
||||
### Prismodell for Azure AI Foundry Evaluation
|
||||
### Prismodell for Microsoft Foundry Evaluation
|
||||
|
||||
**Komponenter:**
|
||||
|
||||
|
|
@ -845,7 +845,7 @@ if eval_results["metrics"]["violence_violations"] > 0:
|
|||
- 10k traces/day × 5 KB × 30 days = 1.5 GB = **$0.03/month**
|
||||
|
||||
- **LLM Judge API calls:**
|
||||
- Same as Azure AI Foundry (charged by OpenAI/Azure OpenAI)
|
||||
- Same as Microsoft Foundry (charged by OpenAI/Azure OpenAI)
|
||||
|
||||
**Total monthly cost (10k req/day, daily batch eval):**
|
||||
|
||||
|
|
@ -854,11 +854,11 @@ if eval_results["metrics"]["violence_violations"] > 0:
|
|||
- LLM calls: $1000 (assume 3 evaluators, 100% sampling)
|
||||
- **Total:** ~$1007.50/month
|
||||
|
||||
**vs. Azure AI Foundry (continuous):** MLflow batch er billigere for compute ($7.50 vs. $0 for serverless continuous), men krever samme LLM judge cost. **Break-even:** Hvis du kan leve med daily batch i stedet for real-time, spar ~$400/month på Application Insights og serverless overhead. *(Medium confidence – varies by implementation)*
|
||||
**vs. Microsoft Foundry (continuous):** MLflow batch er billigere for compute ($7.50 vs. $0 for serverless continuous), men krever samme LLM judge cost. **Break-even:** Hvis du kan leve med daily batch i stedet for real-time, spar ~$400/month på Application Insights og serverless overhead. *(Medium confidence – varies by implementation)*
|
||||
|
||||
### Lisenskrav
|
||||
|
||||
**Azure AI Foundry SDK:**
|
||||
**Microsoft Foundry SDK:**
|
||||
|
||||
- Open source (MIT license)
|
||||
- Krever Azure subscription med:
|
||||
|
|
@ -873,7 +873,7 @@ if eval_results["metrics"]["violence_violations"] > 0:
|
|||
- Databricks: Requires **Premium** or **Enterprise** workspace tier for Unity Catalog governance
|
||||
- Self-hosted MLflow: Gratis, men krever infrastruktur og vedlikehold
|
||||
|
||||
**Recommendation for offentlig sektor:** Azure AI Foundry for compliance-ready, managed service. MLflow for kostnadskontroll og data sovereignty (kan kjøres on-prem/Azure Gov Cloud). *(High confidence)*
|
||||
**Recommendation for offentlig sektor:** Microsoft Foundry for compliance-ready, managed service. MLflow for kostnadskontroll og data sovereignty (kan kjøres on-prem/Azure Gov Cloud). *(High confidence)*
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -936,7 +936,7 @@ Production evaluation er ikke komplett uten human review loop. Anbefal:
|
|||
- **Monthly calibration** av LLM judges mot human-labeled golden dataset
|
||||
- **Quarterly retrospective** – oppdater evaluators basert på learnings
|
||||
|
||||
**Tooling:** Azure AI Foundry Review App eller custom Power Apps interface til Dataverse.
|
||||
**Tooling:** Microsoft Foundry Review App eller custom Power Apps interface til Dataverse.
|
||||
|
||||
### Red flags å se etter
|
||||
|
||||
|
|
@ -1022,7 +1022,7 @@ Production evaluation er ikke komplett uten human review loop. Anbefal:
|
|||
|
||||
**Rekommandasjon (standard scenario):**
|
||||
|
||||
> "Jeg anbefaler å starte med Azure AI Foundry continuous evaluation for safety metrics (Violence, Self-harm) ved 100% sampling, kombinert med scheduled daily batch evaluation for quality metrics (Groundedness, Relevance) ved 30% sampling. Dette gir dere incident detection innen 1 time for safety issues, mens dere holder evalueringskostnaden under $500/måned for en app med 5000 requests/dag. Vi integrerer med Application Insights dere allerede bruker, og setter opp Azure Monitor alerts for automatisk varsling når metrics faller under acceptable thresholds."
|
||||
> "Jeg anbefaler å starte med Microsoft Foundry continuous evaluation for safety metrics (Violence, Self-harm) ved 100% sampling, kombinert med scheduled daily batch evaluation for quality metrics (Groundedness, Relevance) ved 30% sampling. Dette gir dere incident detection innen 1 time for safety issues, mens dere holder evalueringskostnaden under $500/måned for en app med 5000 requests/dag. Vi integrerer med Application Insights dere allerede bruker, og setter opp Azure Monitor alerts for automatisk varsling når metrics faller under acceptable thresholds."
|
||||
|
||||
**Trade-off diskusjon:**
|
||||
|
||||
|
|
@ -1052,7 +1052,7 @@ Production evaluation er ikke komplett uten human review loop. Anbefal:
|
|||
|
||||
### Primærkilder (Official Microsoft Documentation)
|
||||
|
||||
1. **Azure AI Foundry Evaluation SDK:**
|
||||
1. **Microsoft Foundry Evaluation SDK:**
|
||||
[Evaluate your generative AI application locally with the Azure AI Evaluation SDK](https://learn.microsoft.com/en-us/azure/foundry-classic/how-to/develop/evaluate-sdk) – Comprehensive guide til local og cloud evaluation
|
||||
|
||||
2. **Continuous Evaluation for Agents:**
|
||||
|
|
@ -1062,7 +1062,7 @@ Production evaluation er ikke komplett uten human review loop. Anbefal:
|
|||
[Evaluate and monitor AI agents - Azure Databricks](https://learn.microsoft.com/en-us/azure/databricks/mlflow3/genai/eval-monitor/) – MLflow 3 evaluation harness og production scorers
|
||||
|
||||
4. **Observability Overview:**
|
||||
[Observability in generative AI - Azure AI Foundry](https://learn.microsoft.com/en-us/azure/foundry/concepts/observability) – High-level GenAIOps lifecycle og evaluator taxonomy
|
||||
[Observability in generative AI - Microsoft Foundry](https://learn.microsoft.com/en-us/azure/foundry/concepts/observability) – High-level GenAIOps lifecycle og evaluator taxonomy
|
||||
|
||||
5. **Model Monitoring for Generative AI:**
|
||||
[Model monitoring for generative AI applications (preview)](https://learn.microsoft.com/en-us/azure/machine-learning/prompt-flow/how-to-monitor-generative-ai-applications) – Azure ML Prompt Flow monitoring approach. **NB (MCP 2026-06-19):** Prompt Flow pensjoneres 2027-04-20 (migrer til Microsoft Agent Framework); monitoring-tilnærmingen er fortsatt gyldig for eksisterende flows frem til fristen.
|
||||
|
|
@ -1085,7 +1085,7 @@ Production evaluation er ikke komplett uten human review loop. Anbefal:
|
|||
|
||||
**High confidence areas (basert på offisiell dokumentasjon og code samples):**
|
||||
|
||||
- Azure AI Foundry SDK API usage og evaluator configuration
|
||||
- Microsoft Foundry SDK API usage og evaluator configuration
|
||||
- MLflow 3 production monitoring patterns
|
||||
- Cost estimation for LLM judges (basert på Azure OpenAI pricing)
|
||||
- Compliance requirements (AI Act, GDPR) – basert på legal text
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue