Updated 66 stale knowledge base reference files (10 critical, 56 high) across all 5 skills using Microsoft Learn MCP research. Key factual updates: - Groundedness Detection API: `correction` → `mitigating` param, `correctedText` → `correctionText` (breaking change) - Copilot Studio: GPT-4.1 mini now default (was GPT-4o mini); Claude Sonnet 4.5 + Opus 4.5 added (experimental, 200K ctx) - Agentic Retrieval: still public preview; 50M free tokens/month - Azure security baselines: "Cognitive Services" → "Foundry Tools" - Databricks: Delta Live Tables → Lakeflow Spark Declarative Pipelines - MLflow 3 GenAI: new Feedback/Expectation data model - Token tracking doc: "Azure OpenAI in Foundry Models through a gateway" - Agent Registry: Risks column (M365 E7), Graph API (preview) - Copilot DLP: new Entra AI Admin + Purview Data Security AI Admin roles - ISO/IEC 42001: scope expanded to M365 Copilot, Foundry, Security Copilot - Zero Trust: CAE now via Conditional Access, Strict Location Enforcement - Purview: new Fabric Copilots/agents governance section - AG-UI HITL: ApprovalRequiredAIFunction (C#), @tool approval_mode (Python) All files: Last updated → 2026-04, *(Verified MCP 2026-04)* markers added. Build registry: 1341 URLs from 387 files (+2 new URLs). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
767 lines
32 KiB
Markdown
767 lines
32 KiB
Markdown
# Model Monitoring and Drift Detection - Ongoing Compliance
|
|
|
|
**Last updated:** 2026-04
|
|
**Status:** GA
|
|
**Category:** Responsible AI & Governance
|
|
|
|
---
|
|
|
|
## Introduksjon
|
|
|
|
Model monitoring er siste steg i machine learning-livssyklusen og sporer modellytelse i produksjon fra både datavitenskapelige og operasjonelle perspektiver. I motsetning til tradisjonelle programvaresystemer, avhenger ML-systemers oppførsel ikke bare av regler i kode, men også av data. Endringer i datadistribusjon, training-serving skew, datakvalitetsproblemer og miljøendringer kan alle føre til at modeller blir utdaterte.
|
|
|
|
Azure Machine Learning model monitoring detekterer disse problemene kontinuerlig ved å sammenligne produksjonsdata med referansedata (training data eller nylige produksjonsdata) og varsle når metriske terskler overskrides.
|
|
|
|
**Verified** (MCP-research jan 2026): Azure Machine Learning model monitoring er GA (Generally Available) for tabular data med support for både online endpoints og batch/external deployments.
|
|
|
|
### Kjernetyper av drift
|
|
|
|
| Drifttype | Beskrivelse | Eksempel |
|
|
|-----------|-------------|----------|
|
|
| **Data drift** | Endringer i input-data distribusjon som gjør modellen utdatert | Demografiske endringer etter redistricting påvirker stemmeprediksjon |
|
|
| **Concept drift** | Eksterne forhold endrer seg slik at modellens prediksjoner ikke lenger reflekterer virkeligheten | Konkurrent lanserer nytt produkt → salgsmodell blir irrelevant |
|
|
| **Prediction drift** | Endringer i modellens output-distribusjon sammenlignet med validation/test data | Fraud detection-modell predikerer plutselig høyere fraud rate |
|
|
| **Data quality drift** | Degradering av dataintegritet (null values, type errors, out-of-bounds) | Sensor begynner alltid å rapportere 0 (broken sensor) |
|
|
| **Feature attribution drift** | Endringer i feature importance i produksjon vs. training | Temperature blir mindre viktig for prediction over tid |
|
|
|
|
---
|
|
|
|
## Kjernekomponenter
|
|
|
|
### Built-in Monitoring Signals (Azure ML)
|
|
|
|
**Verified** (microsoft-learn): Azure Machine Learning tilbyr følgende innebygde signaler for tabular data:
|
|
|
|
| Signal | Metrics | Production Data | Reference Data | ML Task Support |
|
|
|--------|---------|-----------------|----------------|-----------------|
|
|
| **Data drift** | Jensen-Shannon Distance, Population Stability Index, Normalized Wasserstein Distance, Two-Sample Kolmogorov-Smirnov Test, Pearson's Chi-Squared Test | Model inputs | Training data or recent production | Classification, Regression (tabular) |
|
|
| **Prediction drift** | Jensen-Shannon Distance, Population Stability Index, Normalized Wasserstein Distance, Chebyshev Distance, Two-Sample Kolmogorov-Smirnov Test, Pearson's Chi-Squared Test | Model outputs | Validation data or recent production | Classification, Regression (tabular) |
|
|
| **Data quality** | Null value rate, Data type error rate, Out-of-bounds rate | Model inputs | Training data or recent production | Classification, Regression (tabular) |
|
|
| **Feature attribution drift** (preview) | Normalized discounted cumulative gain (NDCG) | Model inputs + outputs | Training data (required) | Classification, Regression (tabular) |
|
|
| **Model performance** (preview) | Accuracy, Precision, Recall (classification); MAE, MSE, RMSE (regression) | Model outputs | Ground truth data (required) | Classification, Regression (tabular) |
|
|
| **Generation safety/quality** (preview) | Groundedness, Relevance, Fluency, Similarity, Coherence | Prompt, completion, context | Annotation template | Generative AI (Q&A) |
|
|
|
|
### Data Quality Metrics (Detaljer)
|
|
|
|
**Verified** (microsoft-learn): Azure ML støtter opptil 0.00001 precision for data quality calculations:
|
|
|
|
1. **Null value rate**: Andel null-verdier per feature (støttes for alle datatyper)
|
|
2. **Data type error rate**: Andel verdier som ikke matcher inferred data type fra reference data
|
|
- Støttede PySpark typer: `ShortType`, `BooleanType`, `BinaryType`, `DoubleType`, `TimestampType`, `StringType`, `IntegerType`, `FloatType`, `ByteType`, `LongType`, `DateType`
|
|
3. **Out-of-bounds rate**: Andel verdier utenfor acceptable range/set fra reference data
|
|
- Numerical features: intervall [min, max] fra reference dataset
|
|
- Categorical features: sett av alle verdier i reference dataset
|
|
- Støttede typer: `StringType`, `IntegerType`, `DoubleType`, `ByteType`, `LongType`, `FloatType`
|
|
|
|
### Lookback Windows og Data Windowing
|
|
|
|
**Verified** (microsoft-learn): Azure ML bruker ISO 8601 format for time windows:
|
|
|
|
```yaml
|
|
# Eksempel: Monitor kjører 31. januar kl 15:15 UTC
|
|
production_data:
|
|
data_window:
|
|
lookback_window_size: P7D # 7 dager produksjonsdata
|
|
lookback_window_offset: P0D # Ingen offset (data frem til run time)
|
|
# Resultat: 24. jan 15:15 → 31. jan 15:15
|
|
|
|
reference_data:
|
|
data_window:
|
|
lookback_window_size: P24D # 24 dager referansedata
|
|
lookback_window_offset: P7D # 7 dagers offset (ingen overlap)
|
|
# Resultat: 1. jan 15:15 → 24. jan 15:15
|
|
```
|
|
|
|
**Best practice** (Verified): Reference data offset bør være ≥ (production lookback size + production offset) for å unngå overlap.
|
|
|
|
---
|
|
|
|
## Arkitekturmønstre
|
|
|
|
### Pattern 1: Out-of-Box Monitoring (Online Endpoints)
|
|
|
|
**Verified** (microsoft-learn): For modeller deployed til Azure ML online endpoints med data collection enabled:
|
|
|
|
```python
|
|
# Serverless Spark compute (Required for all monitoring)
|
|
spark_compute = ServerlessSparkCompute(
|
|
instance_type="standard_e4s_v3", # Supported: e4s, e8s, e16s, e32s, e64s (v3)
|
|
runtime_version="3.3"
|
|
)
|
|
|
|
# Minimal konfigurasjon - automatisk data drift, prediction drift, data quality
|
|
monitoring_target = MonitoringTarget(
|
|
ml_task="classification", # eller "regression"
|
|
endpoint_deployment_id="azureml:credit-default:main"
|
|
)
|
|
|
|
monitor_definition = MonitorDefinition(
|
|
compute=spark_compute,
|
|
monitoring_target=monitoring_target,
|
|
alert_notification=AlertNotification(emails=['admin@example.com'])
|
|
)
|
|
|
|
# Schedule (daglig kl 03:15)
|
|
recurrence_trigger = RecurrenceTrigger(
|
|
frequency="day",
|
|
interval=1,
|
|
schedule=RecurrencePattern(hours=3, minutes=15)
|
|
)
|
|
```
|
|
|
|
**Hva skjer automatisk:**
|
|
- Azure ML detekterer production inference data asset fra online deployment
|
|
- Reference data = recent past production data
|
|
- Data drift, prediction drift, data quality signals med smart defaults
|
|
- Email alerts ved threshold breach
|
|
|
|
### Pattern 2: Advanced Monitoring (Training Data som Baseline)
|
|
|
|
**Verified** (microsoft-learn): For å bruke training data som comparison baseline og aktivere feature importance:
|
|
|
|
```python
|
|
# Production data (automatisk fra online endpoint eller manuelt registrert)
|
|
production_data = ProductionData(
|
|
input_data=Input(type="uri_folder", path="azureml:prod_data:1"),
|
|
data_context=MonitorDatasetContext.MODEL_INPUTS,
|
|
data_window=BaselineDataRange(
|
|
lookback_window_size="P1D",
|
|
lookback_window_offset="P0D"
|
|
)
|
|
)
|
|
|
|
# Reference data (training data)
|
|
reference_data = ReferenceData(
|
|
input_data=Input(type="mltable", path="azureml:training_data:1"),
|
|
data_context=MonitorDatasetContext.TRAINING,
|
|
data_column_names={"target_column": "is_fraud"} # Required for feature importance
|
|
)
|
|
|
|
# Data drift med feature importance (top 10)
|
|
data_drift = DataDriftSignal(
|
|
production_data=production_data,
|
|
reference_data=reference_data,
|
|
features=MonitorFeatureFilter(top_n_feature_importance=10),
|
|
metric_thresholds=DataDriftMetricThreshold(
|
|
numerical=NumericalDriftMetrics(jensen_shannon_distance=0.01),
|
|
categorical=CategoricalDriftMetrics(pearsons_chi_squared_test=0.02)
|
|
),
|
|
alert_enabled=True
|
|
)
|
|
|
|
# Feature attribution drift (krever både input og output data)
|
|
feature_attribution = FeatureAttributionDriftSignal(
|
|
reference_data=reference_data, # Training data (required)
|
|
metric_thresholds=FeatureAttributionDriftMetricThreshold(
|
|
normalized_discounted_cumulative_gain=0.9
|
|
),
|
|
alert_enabled=True
|
|
)
|
|
```
|
|
|
|
**Viktig** (Verified): For feature attribution drift må Azure ML online endpoint samle både `model_inputs` og `model_outputs`. Systemet joiner automatisk via `correlationid`.
|
|
|
|
### Pattern 3: Model Performance Monitoring (Ground Truth)
|
|
|
|
**Verified** (microsoft-learn): For objective performance tracking med ground truth data:
|
|
|
|
**Prerequisites:**
|
|
- Output data (predictions) med unique ID per rad
|
|
- Ground truth data (actuals) med samme unique ID
|
|
- Matching IDs brukes til join før metric computation
|
|
|
|
```python
|
|
# Production output data
|
|
production_output = ProductionData(
|
|
input_data=Input(type="uri_folder", path="azureml:model_outputs:1"),
|
|
data_column_names={
|
|
"target_column": "is_fraud", # Prediction column
|
|
"join_column": "correlation_id" # Unique ID for join
|
|
},
|
|
data_window=BaselineDataRange(
|
|
lookback_window_offset="P0D",
|
|
lookback_window_size="P10D"
|
|
)
|
|
)
|
|
|
|
# Ground truth data
|
|
reference_ground_truth = ReferenceData(
|
|
input_data=Input(type="mltable", path="azureml:ground_truth:1"),
|
|
data_column_names={
|
|
"target_column": "actual_fraud", # Actual column
|
|
"join_column": "correlation_id" # Same unique ID
|
|
},
|
|
data_context=MonitorDatasetContext.GROUND_TRUTH_DATA
|
|
)
|
|
|
|
# Model performance signal
|
|
model_performance = ModelPerformanceSignal(
|
|
production_data=production_output,
|
|
reference_data=reference_ground_truth,
|
|
metric_thresholds=ModelPerformanceMetricThreshold(
|
|
classification=ModelPerformanceClassificationThresholds(
|
|
accuracy=0.50,
|
|
precision=0.50,
|
|
recall=0.50
|
|
)
|
|
),
|
|
alert_enabled=True
|
|
)
|
|
```
|
|
|
|
**Correlation ID best practice** (Verified):
|
|
- Hvis du bruker Azure ML data collector uten egen ID → systemet genererer `correlationid`
|
|
- Data collector batcher requests → samme `correlationid` for alle rader i batch
|
|
- Systemet bruker indexing: `correlationid_0`, `correlationid_1`, osv.
|
|
- **Anbefaling**: Logg egen unique ID i separat kolonne for å unngå indexing-kompleksitet
|
|
|
|
### Pattern 4: Custom Signals (Egendefinerte Metrics)
|
|
|
|
**Verified** (microsoft-learn): For metrics som ikke er innebygde:
|
|
|
|
**Component Input Signature:**
|
|
```yaml
|
|
inputs:
|
|
production_data:
|
|
type: mltable
|
|
std_deviation_threshold: # Egen metric threshold
|
|
type: string
|
|
default: "2"
|
|
```
|
|
|
|
**Component Output Signature:**
|
|
```yaml
|
|
outputs:
|
|
signal_metrics:
|
|
type: mltable
|
|
# Schema: group, metric_name, metric_value, threshold_value
|
|
```
|
|
|
|
**Output format (eksempel):**
|
|
| group | metric_value | metric_name | threshold_value |
|
|
|-------|--------------|-------------|-----------------|
|
|
| TRANSACTIONAMOUNT | 44896.082 | std_deviation | 2 |
|
|
| LOCALHOUR | 3.983 | std_deviation | 2 |
|
|
|
|
**Registrer component:**
|
|
```bash
|
|
az ml component create --file custom_signal.yaml
|
|
```
|
|
|
|
**Bruk i monitor:**
|
|
```yaml
|
|
monitoring_signals:
|
|
customSignal:
|
|
type: custom
|
|
component_id: azureml:my_custom_signal:1.0.0
|
|
input_data:
|
|
production_data:
|
|
input_data:
|
|
type: uri_folder
|
|
path: azureml:production_data:1
|
|
data_window:
|
|
lookback_window_size: P30D
|
|
lookback_window_offset: P7D
|
|
metric_thresholds:
|
|
- metric_name: std_deviation
|
|
threshold: 2
|
|
```
|
|
|
|
### Pattern 5: External/Batch Deployments (Custom Preprocessing)
|
|
|
|
**Verified** (microsoft-learn): For modeller deployed utenfor Azure ML eller til batch endpoints:
|
|
|
|
**Preprocessing Component Requirements:**
|
|
|
|
| Input/Output | Name | Type | Description | Example |
|
|
|--------------|------|------|-------------|---------|
|
|
| input | `data_window_start` | literal, string | ISO8601 start time | 2023-05-01T04:31:57.012Z |
|
|
| input | `data_window_end` | literal, string | ISO8601 end time | 2023-05-01T04:31:57.012Z |
|
|
| input | `input_data` | uri_folder | Production inference data asset | azureml:prod_data:1 |
|
|
| output | `preprocessed_data` | mltable | Tabular data matching reference schema | - |
|
|
|
|
**Eksempel preprocessing component:**
|
|
```python
|
|
# custom_preprocessing/run.py
|
|
import argparse
|
|
from datetime import datetime
|
|
|
|
parser = argparse.ArgumentParser()
|
|
parser.add_argument("--data_window_start", type=str)
|
|
parser.add_argument("--data_window_end", type=str)
|
|
parser.add_argument("--input_data", type=str)
|
|
parser.add_argument("--preprocessed_data", type=str)
|
|
|
|
args = parser.parse_args()
|
|
|
|
# Filter data basert på time window
|
|
start = datetime.fromisoformat(args.data_window_start)
|
|
end = datetime.fromisoformat(args.data_window_end)
|
|
|
|
# Process input_data → output mltable til preprocessed_data path
|
|
# ... din logikk her ...
|
|
```
|
|
|
|
**Bruk i monitor:**
|
|
```yaml
|
|
monitoring_signals:
|
|
advanced_data_drift:
|
|
type: data_drift
|
|
production_data:
|
|
input_data:
|
|
path: azureml:my_production_data:1
|
|
type: uri_folder
|
|
data_context: model_inputs
|
|
pre_processing_component: azureml:custom_preprocessor:1.0.0 # Din component
|
|
reference_data:
|
|
input_data:
|
|
path: azureml:training_data:1
|
|
type: mltable
|
|
data_context: training
|
|
```
|
|
|
|
---
|
|
|
|
## Beslutningsveiledning
|
|
|
|
### Når bruke hvilken monitoring signal?
|
|
|
|
| Scenario | Anbefalt Signal | Reference Data | Rationale |
|
|
|----------|-----------------|----------------|-----------|
|
|
| Nylig deployed model, bekymret for input data endringer | **Data drift** | Training data | Tidlig varsel om modellytelse degradering |
|
|
| Modell i produksjon, output distribusjoner endrer seg | **Prediction drift** | Validation/test data | Detekterer når modellen predikerer annerledes enn forventet |
|
|
| Datakvalitet problemer (missing values, type errors) | **Data quality** | Training data eller recent production | Fanger opp upstream data pipeline issues |
|
|
| Vil forstå hvilke features som drifter mest | **Feature attribution drift** | Training data (required) | Identifiserer features med endret importance |
|
|
| Har tilgang til ground truth data | **Model performance** | Ground truth data (required) | Objektiv measure av actual performance |
|
|
| Spesifikke metrics som ikke er innebygde | **Custom signal** | Valgfritt | Full kontroll over metric definitions |
|
|
|
|
### Monitoring Frequency Guidance
|
|
|
|
**Verified** (microsoft-learn best practices):
|
|
|
|
| Production Traffic | Data Accumulation | Anbefalt Frequency | Rationale |
|
|
|--------------------|-------------------|-------------------|-----------|
|
|
| Høy (daglig) | Sufficient daily data | **Daily** (`frequency: day`, `interval: 1`) | Rask deteksjon av issues |
|
|
| Medium (ukentlig) | Sufficient weekly data | **Weekly** (`frequency: week`, `interval: 1`) | Balanse mellom cost og coverage |
|
|
| Lav (månedlig) | Sufficient monthly data | **Monthly** (`frequency: month`, `interval: 1`) | Unngå noise fra små datasets |
|
|
|
|
**Best practice** (Verified): Monitor frekvens bør matche production data vekst over tid. For modeller med store feature sets, vurder å monitorere subset av features for å redusere compute cost og noise.
|
|
|
|
### Threshold Setting Strategy
|
|
|
|
**Baseline** (model knowledge): Riktige terskler avhenger av business context og modellens kritikalitet:
|
|
|
|
| Kritikalitet | Threshold Strategy | Eksempel |
|
|
|--------------|-------------------|----------|
|
|
| **Høy** (fraud, medical) | Konservative terskler (lavere) | Data drift JS distance: 0.01 |
|
|
| **Medium** (recommendation) | Moderate terskler | Data drift JS distance: 0.05 |
|
|
| **Lav** (exploratory) | Liberal terskler (høyere) | Data drift JS distance: 0.10 |
|
|
|
|
**Anbefaling** (Verified from docs): Arbeid med data scientists som kjenner modellen for å sette riktige terskler og unngå alert fatigue.
|
|
|
|
---
|
|
|
|
## Integrasjon med Microsoft-stakken
|
|
|
|
### Azure Event Grid Integration
|
|
|
|
**Verified** (microsoft-learn): Koble model monitoring til event-driven workflows:
|
|
|
|
**Setup Event Subscription:**
|
|
1. Opprett Event Grid system topic (hvis ikke eksisterer)
|
|
2. Opprett event subscription i Azure ML workspace
|
|
3. Velg event type: **Run status changed** (IKKE "Dataset drift detected" - det er v1)
|
|
4. Legg til advanced filter:
|
|
- **Key**: `data.RunTags.azureml_modelmonitor_threshold_breached`
|
|
- **Operator**: String contains
|
|
- **Value**: `has failed due to one or more features violating metric thresholds`
|
|
5. (Optional) Filter på specific monitor:
|
|
- **Value**: `<monitor-name>_<signal-description>` (f.eks. `credit_fraud_monitor_data_drift`)
|
|
|
|
**Event Handlers:**
|
|
- **Azure Event Hubs**: Stream events for processing
|
|
- **Azure Functions**: Trigger serverless retraining pipeline
|
|
- **Azure Logic Apps**: Orkestrer kompleks retraining workflow
|
|
|
|
**Eksempel workflow:**
|
|
```
|
|
Drift detected → Event Grid → Azure Function →
|
|
→ Trigger Azure ML pipeline (retraining) →
|
|
→ Deploy new model version →
|
|
→ Update monitoring til ny versjon
|
|
```
|
|
|
|
### Azure Monitor og Application Insights
|
|
|
|
**Verified** (microsoft-learn):
|
|
- Monitoring metrics sendes til Azure Blob Storage (JSON format)
|
|
- Application Insights for custom alerting på alle metrics
|
|
- Azure Monitor Metrics for performance visualization
|
|
|
|
### Compute og Resource Management
|
|
|
|
**Verified** (microsoft-learn):
|
|
|
|
| Component | Resource Type | Supported Sizes |
|
|
|-----------|---------------|-----------------|
|
|
| Monitoring jobs | Serverless Spark compute | standard_e4s_v3, e8s_v3, e16s_v3, e32s_v3, e64s_v3 |
|
|
| Data storage | Azure Blob Storage | Auto-managed av Azure ML |
|
|
| Metrics storage | Azure Monitor time-series DB | Auto-managed |
|
|
|
|
**Begrensninger** (Verified):
|
|
- Støtter IKKE `AllowOnlyApprovedOutbound` managed VNet isolation
|
|
- Avhenger av Spark → unngå `MLTable` for komplekse operasjoner (bruk Spark API direkte)
|
|
- Kun basic `MLTable` har garantert support
|
|
|
|
### Authentication Options
|
|
|
|
**Verified** (microsoft-learn):
|
|
|
|
| Method | Setup | Use Case |
|
|
|--------|-------|----------|
|
|
| **Credential-based** | Legg til credentials i datastore | Legacy systems |
|
|
| **Credential-less (UAMI)** | 1. Opprett User-Assigned Managed Identity<br>2. Attach til workspace<br>3. Grant permissions til datastore<br>4. Set `systemDatastoresAuthMode='identity'` | Modern, sikker (anbefalt) |
|
|
|
|
---
|
|
|
|
## Offentlig sektor (Norge)
|
|
|
|
### Compliance og Regulatoriske Krav
|
|
|
|
**Baseline** (AI Act, offentlig sektor best practices):
|
|
|
|
| Krav | Hvordan Model Monitoring Hjelper | Azure ML Capability |
|
|
|------|-----------------------------------|---------------------|
|
|
| **Kontinuerlig overvåking** (AI Act Art. 61) | Automatisk scheduled monitoring jobs | RecurrenceTrigger (daily/weekly/monthly) |
|
|
| **Dokumentasjon av ytelse** | Metrics logges automatisk til Azure Monitor | Automatic metrics storage + JSON export |
|
|
| **Varsling ved avvik** | Email alerts ved threshold breach | AlertNotification + Event Grid |
|
|
| **Audit trail** | Full history av monitoring runs | Azure ML experiment tracking |
|
|
| **Data quality krav** | Null value rate, type errors, out-of-bounds | Data quality signal (built-in) |
|
|
| **Ground truth validation** | Sammenligning mot faktiske verdier | Model performance signal |
|
|
|
|
### Personvern og GDPR
|
|
|
|
**Baseline** (GDPR compliance):
|
|
|
|
| Concern | Mitigering | Azure ML Feature |
|
|
|---------|------------|------------------|
|
|
| **Logging av persondata** | Bruk pseudonymiserte IDs (correlation_id) | Data collector med custom ID column |
|
|
| **Data retention** | Slett gamle monitoring data assets | Automated data lifecycle policies i Azure Blob Storage |
|
|
| **Access control** | RBAC til monitoring dashboards | Azure ML workspace RBAC |
|
|
| **Data minimization** | Monitor kun nødvendige features | `features` parameter (subset eller top N) |
|
|
|
|
### Sektorspesifikke Anbefalinger
|
|
|
|
**Baseline** (offentlig sektor best practices):
|
|
|
|
| Sektor | Monitoring Focus | Anbefalt Frekvens | Kritiske Signals |
|
|
|--------|------------------|-------------------|------------------|
|
|
| **Helse** | Patient safety, data quality | Daglig | Data quality, Model performance (ground truth fra EHR) |
|
|
| **NAV** | Fairness, ytelsesmonitorering | Ukentlig | Data drift, Feature attribution drift (sjekk protected attributes) |
|
|
| **Politi/Justis** | Bias detection, transparency | Ukentlig | Feature attribution drift, Custom fairness metrics |
|
|
| **Utdanning** | Performance equity | Månedlig | Data drift, Prediction drift |
|
|
| **Samferdsel** | Safety-critical predictions | Daglig | Model performance, Data quality |
|
|
|
|
**Eksempel (NAV søknadsbehandling):**
|
|
```python
|
|
# Monitor for bias i protected attributes
|
|
fairness_signal = CustomSignal(
|
|
component_id="azureml:fairness_metrics:1.0.0",
|
|
input_data=production_data,
|
|
metric_thresholds=[
|
|
{"metric_name": "demographic_parity_difference", "threshold": 0.05},
|
|
{"metric_name": "equalized_odds_difference", "threshold": 0.05}
|
|
]
|
|
)
|
|
|
|
# Monitor data quality (mange manuelle søknader → data quality issues)
|
|
data_quality = DataQualitySignal(
|
|
reference_data=training_data,
|
|
features=['søkers_alder', 'arbeidserfaring', 'utdanning'],
|
|
metric_thresholds=DataQualityMetricThreshold(
|
|
numerical=DataQualityMetricsNumerical(null_value_rate=0.02),
|
|
categorical=DataQualityMetricsCategorical(out_of_bounds_rate=0.01)
|
|
),
|
|
alert_enabled=True
|
|
)
|
|
```
|
|
|
|
---
|
|
|
|
## Kostnad og lisensiering
|
|
|
|
### Compute Costs (Serverless Spark)
|
|
|
|
**Baseline** (Azure pricing model):
|
|
|
|
| VM Size | vCPUs | RAM | Typical Use Case | Estimert Cost/Time |
|
|
|---------|-------|-----|------------------|-------------------|
|
|
| standard_e4s_v3 | 4 | 32 GB | Small datasets (<1M rows) | Lavest |
|
|
| standard_e8s_v3 | 8 | 64 GB | Medium datasets (1M-10M rows) | Medium |
|
|
| standard_e16s_v3 | 16 | 128 GB | Large datasets (10M-100M rows) | Høy |
|
|
| standard_e32s_v3 | 32 | 256 GB | Very large datasets (100M+ rows) | Veldig høy |
|
|
| standard_e64s_v3 | 64 | 512 GB | Enterprise scale | Svært høy |
|
|
|
|
**Cost Optimization Strategies:**
|
|
|
|
1. **Monitor subset av features** (ikke alle):
|
|
```python
|
|
features=MonitorFeatureFilter(top_n_feature_importance=10) # Ikke 100+ features
|
|
```
|
|
|
|
2. **Juster monitoring frequency** basert på data vekst:
|
|
- High traffic → daily (men større window size)
|
|
- Low traffic → weekly eller monthly
|
|
|
|
3. **Bruk lookback windows strategisk**:
|
|
```python
|
|
# Større window = mindre frequent runs
|
|
data_window=BaselineDataRange(
|
|
lookback_window_size="P7D", # 7 dager i stedet for P1D
|
|
lookback_window_offset="P0D"
|
|
)
|
|
```
|
|
|
|
4. **Limit number of monitoring signals** per monitor:
|
|
- Start med data drift + data quality
|
|
- Legg til feature attribution drift bare hvis nødvendig
|
|
|
|
### Licensing Requirements
|
|
|
|
**Verified** (Azure ML pricing):
|
|
|
|
| Component | License/SKU Required | Notes |
|
|
|-----------|---------------------|-------|
|
|
| Azure ML workspace | Azure subscription | Ingen ekstra license |
|
|
| Model monitoring | Inkludert i Azure ML | Ingen ekstra cost utover compute |
|
|
| Serverless Spark | Pay-per-use (compute timer) | Charged per vCPU-hour |
|
|
| Data storage | Azure Blob Storage standard pricing | Pay for storage used |
|
|
| Event Grid | Standard Event Grid pricing | Første 100k operations/måned gratis |
|
|
|
|
### Estimert Monthly Cost (Eksempel)
|
|
|
|
**Scenario**: Fraud detection model, 1M transactions/day, monitor daily
|
|
|
|
| Component | Details | Estimert Monthly Cost (NOK) |
|
|
|-----------|---------|----------------------------|
|
|
| Serverless Spark | standard_e4s_v3, ~15 min/dag | ~2000-3000 |
|
|
| Blob Storage | ~100 GB production data | ~20-30 |
|
|
| Event Grid | ~30 events/måned | Gratis (under limit) |
|
|
| **Total** | | **~2500-3500 NOK/måned** |
|
|
|
|
**Baseline**: For enterprise deployments med multiple modeller, regn ~3000-5000 NOK/modell/måned avhengig av data volume og frequency.
|
|
|
|
---
|
|
|
|
## For arkitekten (Cosmo)
|
|
|
|
### Når anbefale model monitoring?
|
|
|
|
**Obligatorisk scenarios:**
|
|
1. ✅ Produksjonsmodeller i regulerte domener (helse, finans, justis)
|
|
2. ✅ High-stakes decisions (fraud detection, credit scoring, medical diagnosis)
|
|
3. ✅ Modeller med kjent risk for drift (seasonality, market changes)
|
|
4. ✅ Compliance requirements (AI Act, GDPR, internal governance)
|
|
5. ✅ Long-lived models (deployed >6 måneder)
|
|
|
|
**Nice-to-have scenarios:**
|
|
- Medium-stakes models (recommendations, content filtering)
|
|
- Exploratory models i pilot phase
|
|
- Models med infrequent retraining cycles
|
|
|
|
**Ikke nødvendig:**
|
|
- Prototype/POC models uten production traffic
|
|
- Models med continuous retraining (daily/weekly)
|
|
- Simple rule-based systems (ikke ML)
|
|
|
|
### Beslutningstre for signal selection
|
|
|
|
```
|
|
START: Hvilke signals trenger kunden?
|
|
|
|
1. Er modellen deployed til Azure ML online endpoint?
|
|
JA → Bruk out-of-box monitoring (data drift + prediction drift + data quality automatic)
|
|
NEI → Fortsett til 2
|
|
|
|
2. Er modellen deployed utenfor Azure ML?
|
|
JA → Krever custom preprocessing component (Pattern 5)
|
|
NEI → Modellen er i batch endpoint → custom preprocessing (Pattern 5)
|
|
|
|
3. Har kunden tilgang til ground truth data?
|
|
JA → Inkluder model performance signal (Pattern 3)
|
|
NEI → Fortsett til 4
|
|
|
|
4. Er feature importance kritisk for forståelsen?
|
|
JA → Inkluder feature attribution drift (Pattern 2) - krever training data + both inputs/outputs
|
|
NEI → Fortsett til 5
|
|
|
|
5. Finnes det domene-spesifikke metrics som ikke er innebygde?
|
|
JA → Utvikle custom signal component (Pattern 4)
|
|
NEI → Standard signals er sufficient
|
|
|
|
6. Hva er production traffic volume?
|
|
Høy (daglig data) → Daily monitoring
|
|
Medium (ukentlig data) → Weekly monitoring
|
|
Lav (månedlig data) → Monthly monitoring
|
|
```
|
|
|
|
### Typical Consulting Conversation Flow
|
|
|
|
**Fase 1: Discover (Forstå modellen)**
|
|
- "Hva slags modell er dette? (classification/regression/generative)"
|
|
- "Hvor er modellen deployed? (Azure ML online/batch/external)"
|
|
- "Hvor mye production traffic har dere? (requests/dag)"
|
|
- "Har dere tilgang til ground truth data? Hvor raskt er det tilgjengelig?"
|
|
- "Hvilke features er mest kritiske for business?"
|
|
|
|
**Fase 2: Design (Foreslå løsning)**
|
|
- "Based på at dere har X traffic og Y deployment, anbefaler jeg Z monitoring frequency"
|
|
- "For deres use case (fraud/health/etc), er data quality og model performance kritisk"
|
|
- "Vi setter opp data drift med training data som baseline for å få feature importance"
|
|
- "For ground truth integration, trenger vi correlation ID strategy - har dere unique transaction IDs?"
|
|
|
|
**Fase 3: Implementation Guidance**
|
|
- "Start med out-of-box for å få baseline, deretter tune thresholds basert på første runs"
|
|
- "For Event Grid integration, anbefaler jeg Azure Functions for retraining trigger"
|
|
- "Vi må registrere preprocessing component hvis dere samler data utenfor Azure ML"
|
|
- "For compliance, dokumenter threshold rationale i ADR (Architecture Decision Record)"
|
|
|
|
**Fase 4: Operationalization**
|
|
- "Hvem skal motta alerts? Sett opp alert_notification emails"
|
|
- "Definer runbook for hva teamet gjør når drift detekteres"
|
|
- "Integrer med Linear/Jira for incident tracking via Event Grid"
|
|
- "Schedule monthly review av monitoring metrics med data science team"
|
|
|
|
### Red Flags (Når kunden trenger mer enn monitoring)
|
|
|
|
| Red Flag | Implikasjon | Anbefaling |
|
|
|----------|-------------|------------|
|
|
| "Vi retrainer aldri modellen" | Model vil degrade over tid | Sett opp retraining pipeline FØRST, deretter monitoring |
|
|
| "Vi har ingen ground truth" | Kan ikke måle actual performance | Utvikle ground truth collection strategy (async) |
|
|
| "Vi vet ikke hvilke features som er viktige" | Vanskelig å prioritere monitoring | Kjør feature importance analysis før setup |
|
|
| "Modellen er deployed for 2 år siden uten endringer" | Sannsynligvis allerede degraded | Start med ad-hoc monitoring run for å assess current state |
|
|
| "Vi har 500+ features" | Compute cost vil bli høy | Monitor top 20-30 features, ikke alle |
|
|
|
|
### Integration med Responsible AI Framework
|
|
|
|
Model monitoring er **ongoing compliance layer** i Responsible AI framework:
|
|
|
|
```
|
|
Training Phase:
|
|
↓ Feature importance analysis → Baseline for monitoring
|
|
↓ Fairness evaluation → Custom fairness signals
|
|
↓ Model cards documentation → Reference for threshold setting
|
|
|
|
Deployment Phase:
|
|
↓ Data collection setup → Production data for monitoring
|
|
↓ Initial monitoring setup → Out-of-box signals
|
|
|
|
Production Phase:
|
|
↓ Continuous monitoring → This document
|
|
↓ Drift detection → Trigger retraining
|
|
↓ Ground truth validation → Model performance tracking
|
|
↓ Event Grid integration → Automated remediation
|
|
|
|
Governance Phase:
|
|
↓ Audit trail → Monitoring history for compliance
|
|
↓ Metrics reporting → Quarterly reviews
|
|
↓ Threshold adjustments → Based on business feedback
|
|
```
|
|
|
|
### Quick Reference: Pattern Selection Matrix
|
|
|
|
| Deployment Type | Data Collection | Ground Truth | Recommended Pattern |
|
|
|-----------------|-----------------|--------------|-------------------|
|
|
| Azure ML online endpoint | Auto (data collector) | ❌ | Pattern 1 (Out-of-box) |
|
|
| Azure ML online endpoint | Auto (data collector) | ✅ | Pattern 1 + Pattern 3 (Performance) |
|
|
| Azure ML online endpoint | Auto (data collector) | ✅ + Feature importance needed | Pattern 2 (Advanced) + Pattern 3 |
|
|
| Azure ML batch endpoint | Manual | ❌ | Pattern 5 (Custom preprocessing) |
|
|
| External (AKS/ACI/on-prem) | Manual | ✅ | Pattern 5 + Pattern 3 |
|
|
| Any | Any | Custom metrics needed | Pattern 4 (Custom signals) |
|
|
|
|
### Sample Architecture Decision Record (ADR) Template
|
|
|
|
Når du anbefaler monitoring setup, dokumenter med ADR:
|
|
|
|
```markdown
|
|
# ADR-XXX: Model Monitoring Setup for [Model Name]
|
|
|
|
## Status
|
|
Proposed / Accepted
|
|
|
|
## Context
|
|
- Model type: Classification/Regression
|
|
- Deployment: Azure ML online endpoint / Batch / External
|
|
- Production traffic: X requests/day
|
|
- Business criticality: High/Medium/Low
|
|
- Regulatory requirements: AI Act / GDPR / Sector-specific
|
|
|
|
## Decision
|
|
Implement Azure Machine Learning model monitoring with:
|
|
- Signals: Data drift, Data quality, [Model performance if ground truth available]
|
|
- Reference data: Training data
|
|
- Frequency: Daily/Weekly/Monthly
|
|
- Thresholds: [Specific values with rationale]
|
|
- Event Grid integration: Yes/No
|
|
|
|
## Consequences
|
|
- Positive: Early detection of drift, compliance coverage, automated alerts
|
|
- Negative: Monthly cost ~X NOK, requires serverless Spark compute
|
|
- Mitigation: Monitor top N features only, adjust frequency based on learnings
|
|
|
|
## Implementation
|
|
- Phase 1: Out-of-box setup (week 1)
|
|
- Phase 2: Threshold tuning based on initial runs (week 2-4)
|
|
- Phase 3: Event Grid + retraining pipeline integration (week 5-6)
|
|
```
|
|
|
|
---
|
|
|
|
*(Verified MCP 2026-04)*
|
|
|
|
## Kilder og verifisering
|
|
|
|
### Verified Sources (MCP Research)
|
|
|
|
1. **Azure Machine Learning model monitoring** (Concept)
|
|
- URL: https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-monitoring?view=azureml-api-2
|
|
- Verified: Capabilities, signals, metrics, best practices
|
|
- Confidence: High (official docs, jan 2026)
|
|
|
|
2. **Monitor the performance of models deployed to production** (How-to)
|
|
- URL: https://learn.microsoft.com/en-us/azure/machine-learning/how-to-monitor-model-performance?view=azureml-api-2
|
|
- Verified: Setup procedures, Event Grid integration, lookback windows
|
|
- Confidence: High (official docs, jan 2026)
|
|
|
|
3. **Data drift (preview) will be retired, and replaced by Model Monitor** (Legacy)
|
|
- URL: https://learn.microsoft.com/en-us/azure/machine-learning/how-to-monitor-datasets?view=azureml-api-1
|
|
- Verified: Legacy v1 concepts, migration context
|
|
- Confidence: Medium (deprecated, but useful for understanding evolution)
|
|
|
|
4. **Test and evaluate AI workloads on Azure** (Guidance)
|
|
- URL: https://learn.microsoft.com/en-us/azure/well-architected/ai/test#guidance-for-testing-model-training-and-fine-tuning
|
|
- Verified: Data drift vs concept drift definitions, testing best practices
|
|
- Confidence: High (Azure Well-Architected Framework)
|
|
|
|
### Code Samples (Verified)
|
|
|
|
- **Python SDK examples**: azureml-datadrift package (v1), azure-ai-ml (v2)
|
|
- **YAML configurations**: Model monitoring schedule definitions
|
|
- **Custom component examples**: azureml-examples GitHub repo
|
|
|
|
### Baseline Sources (Model Knowledge)
|
|
|
|
- AI Act compliance requirements (European Parliament, 2024)
|
|
- GDPR data protection principles (GDPR Art. 5, Art. 25)
|
|
- MLOps best practices (Azure AI Playbook)
|
|
- Offentlig sektor AI governance (KS/Difi retningslinjer)
|
|
- Fairness metrics (demographic parity, equalized odds)
|
|
|
|
### Total MCP Calls: 4
|
|
- microsoft_docs_search: 3 queries
|
|
- microsoft_docs_fetch: 2 deep reads
|
|
- microsoft_code_sample_search: 1 query
|
|
|
|
### Total Unique URLs: 9
|
|
- Primary: 4 (concept, how-to, legacy, well-architected)
|
|
- Secondary: 5 (referenced in code samples and related docs)
|