feat(ms-ai-architect): R11 §10 måling #2 bølge 1 — 24 av 46 R8-enumerasjoner prosa-klassifisert [skip-docs]
Fire subagenter (Opus, read-only, ingen writes/commits, ingen web-oppslag) klassifiserte idx 1-24 mot §5s tre O2-betingelser. Returene er evidens, ikke regenererbare, og lagres derfor tracked etter samme mønster som fase0-returns. Maskin-verifikasjon av returene (V1/V2/V2b): - V1: alle 24 file_text_verbatim finnes ordrett i de faktiske filene (æøå og markup intakt) — oppfunnet tekst ville ryket her. - V2: alle O2-forslag er deletion-only mot den teksten. - V2b: ett flagg (idx 14) — 'Automatically add' -> 'Add' er en rekapitalisering, altså en tekstendring og ikke ren fjerning. Går til menneskelig review. Bølge 1: 7 O2-kandidater, 17 O3, 0 locator-bom.
This commit is contained in:
parent
e4925c6b28
commit
b174db0936
4 changed files with 416 additions and 0 deletions
104
scripts/kb-eval/data/r11-o2-returns/batch-01.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-01.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 1,
|
||||
"id": "ms-ai-engineering/agent-orchestration/foundry-agent-service-ga.md#13",
|
||||
"file": "skills/ms-ai-engineering/references/agent-orchestration/foundry-agent-service-ga.md",
|
||||
"line": 187,
|
||||
"real_line": 198,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Verktøy | Type | Formål | Tilgjengelighet |\n|---------|------|--------|-----------------|\n| **Code Interpreter** | Action | Kjøre Python-kode i sandkasse, generere filer og visualiseringer | GA |\n| **File Search** | Knowledge | RAG over opplastede filer via Azure AI Search | GA (ikke tilgjengelig i Italy North, Brazil South) |\n| **Grounding with Bing Search** | Knowledge | Webgrunnlag via Bing | GA |\n| **Bing Custom Search** | Knowledge | Webgrunnlag begrenset til definerte domener | GA |\n| **SharePoint** | Knowledge | Tilgang til interne dokumenter via SharePoint | Preview |\n| **Azure Functions** | Action | Kalle serverless-funksjoner (synkron via MCP eller asynkron via Queue) | GA |\n| **Azure Logic Apps** | Action/Trigger | Over 1400 forhåndsbygde koblinger, event-trigget invokasjon | GA |\n| **OpenAPI tool** | Action | Kalle HTTP-endepunkter beskrevet med OpenAPI 3.0-spec | GA |\n| **MCP tool** | Action/Knowledge | Koble til MCP-servere (remote) | GA (juni 2025) |\n| **Deep Research tool** | Knowledge | Flerstegs research via o3-deep-research + Bing | GA (juni 2025) |\n| **Fabric Data Agent** | Knowledge | Chat med strukturert data i Microsoft Fabric | GA |\n| **Morningstar tool** | Knowledge | Finansdata fra Morningstar | GA |",
|
||||
"failing_part": "The table presents Deep Research (line 198, stated as 'GA (juni 2025)') and Morningstar (line 200, stated as 'GA') as current built-in tools; the source's migration table shows Deep Research as classic-only Public Preview with no equivalent in new Foundry, and Morningstar is said to be absent from the current catalog.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the Deep Research and Morningstar rows leaves a table that asserts a strict subset of the original twelve tool claims."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The judge's reason is that the table as a whole describes the superseded classic tool set, so the surviving ten rows would still stand under the heading 'Innebygde verktøy' as the current new-Foundry catalog with their unqualified GA labels — the same standing-implication failure the contract names for this exact file."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote 'Deep Research | Yes (Public Preview) | No (Recommendation: Deep Research model with Web Search tool)' confirms the tool does exist in classic and names its new-Foundry replacement, so the source supports a corrected value rather than deletion; and the payload carries no evidence at all about Morningstar, so its removal cannot be justified without new fact-finding."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: the source confirms Deep Research's existence and its replacement (a value/qualifier fix, not subtraction), Morningstar's absence is unevidenced in the payload, and the classic-vs-new framing of the whole table cannot be fixed by deleting rows.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 2,
|
||||
"id": "ms-ai-engineering/agent-orchestration/foundry-agent-service-ga.md#22",
|
||||
"file": "skills/ms-ai-engineering/references/agent-orchestration/foundry-agent-service-ga.md",
|
||||
"line": 340,
|
||||
"real_line": 352,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Foundry Agent Service er tilgjengelig i følgende Azure-regioner (per februar 2026):\n\n| Region | Status |\n|--------|--------|\n| **Norway East** | **Tilgjengelig** |\n| Sweden Central | Tilgjengelig |\n| West Europe | Tilgjengelig |\n| Germany West Central | Tilgjengelig |\n| France Central | Tilgjengelig |\n| Switzerland North | Tilgjengelig |\n| UK South | Tilgjengelig |\n| East US / East US 2 | Tilgjengelig |\n| ... (19 regioner totalt) | Se docs for full liste |",
|
||||
"failing_part": "The total region count '19 regioner totalt' on line 352; the eight individually named regions are all confirmed available.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Replacing '... (19 regioner totalt)' with a countless '... (flere regioner)' would drop the numeric assertion while keeping every surviving region claim intact."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The eight listed regions are confirmed present in the Agents column, and the retained 'Se docs for full liste' pointer prevents a reader from taking the shown rows as the complete set."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The judge states the source's Agents column is Yes for 30 regions, so the source supports a corrected value (30) for exactly the failing part — which the contract routes to O1 value swap, never subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails: the source supports a corrected total (30), making this a value swap (O1) rather than a subtraction; deciding whether to swap the count, re-date the 'per februar 2026' qualifier, or generalise is a decision about what to assert.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 3,
|
||||
"id": "ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md#10",
|
||||
"file": "skills/ms-ai-engineering/references/azure-ai-services/document-intelligence-prebuilt-models.md",
|
||||
"line": 391,
|
||||
"real_line": 391,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Tier | Pris per side (USD) | Inkludert |\n|------|---------------------|-----------|\n| **Free (F0)** | $0 | 500 sider/måned, 2 sider per dokument, 20 calls/min |\n| **Standard (S0)** | $1.50 per 1000 sider (prebuilt models) | 2,000 sider per dokument, 15 TPS |",
|
||||
"failing_part": "The F0 rate limit '20 calls/min'; F0's 2 pages/document, S0's 2,000 pages and S0's 15 TPS are all confirmed, and '500 sider/måned' is neither confirmed nor contradicted by the fetched source.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting ', 20 calls/min' from the F0 row would leave the remaining page-allowance assertions untouched and assert strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "In a two-row tier table where the S0 row still states '15 TPS', an F0 row with no throughput figure can read as 'F0 has no throughput cap' — the opposite of the source's 1 transaction/second — and the unverified '500 sider/måned' would remain standing."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote 'Analyze transactions Per Second limit | 1 | 15 (default value)' confirms the correct F0 value (1 TPS), so the source supports a corrected value for the failing part and the contract routes it to O1, not subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails (source confirms F0 = 1 analyze transaction/second, a swap target), and condition 2 is at best unresolved because dropping the F0 throughput next to S0's stated 15 TPS implies an absent limit.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 4,
|
||||
"id": "ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md#4",
|
||||
"file": "skills/ms-ai-engineering/references/azure-ai-services/document-intelligence-prebuilt-models.md",
|
||||
"line": 40,
|
||||
"real_line": 46,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Check** | `prebuilt-check` | Sjekkbehandling | Check number, amount, payee, date |",
|
||||
"failing_part": "The model ID string `prebuilt-check`; the source's model ID is `prebuilt-check.us`. The other six IDs in the table are confirmed.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the Check row (or the ID cell) from the Financial Services table would assert strictly less than the current seven-model list."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A Financial Services table without a check model implies to a reader that Document Intelligence has no prebuilt bank-check model, which the source contradicts."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote 'prebuilt-check.us | ✓ | ✓' confirms the model exists under a corrected ID, and this is the exact failure case the ratified contract names for subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: the source confirms a corrected copy-into-code SKU string (`prebuilt-check.us`), so the fix is a value swap (O1), never subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 5,
|
||||
"id": "ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md#5",
|
||||
"file": "skills/ms-ai-engineering/references/azure-ai-services/document-intelligence-prebuilt-models.md",
|
||||
"line": 52,
|
||||
"real_line": 56,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Marriage Certificate** | `prebuilt-marriageCertificate` | Vigselattester |",
|
||||
"failing_part": "The model ID string `prebuilt-marriageCertificate`; the source's model ID is `prebuilt-marriageCertificate.us`. The ID and tax model IDs in the same table are confirmed.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the Marriage Certificate row from the Identity & Tax table would assert strictly less than the current eight-model list."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Removing the row leaves the Identity & Tax table implying no prebuilt marriage-certificate model exists, which the source contradicts."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote 'prebuilt-marriageCertificate.us | ✓ | ✓' confirms the model under a corrected ID, so a supported value exists and subtraction would destroy true information."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: the source supports a corrected SKU string (`prebuilt-marriageCertificate.us`), making this an O1 value swap.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 6,
|
||||
"id": "ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md#6",
|
||||
"file": "skills/ms-ai-engineering/references/azure-ai-services/document-intelligence-prebuilt-models.md",
|
||||
"line": 65,
|
||||
"real_line": 71,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Disclosure** | `prebuilt-mortgage.us.disclosure` | Endelige lånevilkår |",
|
||||
"failing_part": "The model ID string `prebuilt-mortgage.us.disclosure`; the source's model ID is `prebuilt-mortgage.us.closingDisclosure`. The four other mortgage IDs are confirmed.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the Disclosure row from the US Mortgage table would assert strictly less than the current five-model list."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Removing the row leaves the US Mortgage table implying no closing-disclosure model exists, which the source contradicts."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote '| Closing Disclosure | Extract closing, transaction costs, and loan details. | prebuilt-mortgage.us.closingDisclosure |' confirms the model under a corrected ID, so the supported fix is a swap, not deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: the source supports a corrected SKU string (`prebuilt-mortgage.us.closingDisclosure`), making this an O1 value swap.",
|
||||
"confidence": "high"
|
||||
}
|
||||
]
|
||||
104
scripts/kb-eval/data/r11-o2-returns/batch-02.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-02.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 7,
|
||||
"id": "ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md#7",
|
||||
"file": "skills/ms-ai-engineering/references/azure-ai-services/document-intelligence-prebuilt-models.md",
|
||||
"line": 77,
|
||||
"real_line": 79,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "### Grunnleggende modeller\n\n| Modell | Model ID | Formål |\n|--------|----------|--------|\n| **Read** | `prebuilt-read` | OCR: tekst, linjer, ord, språkdeteksjon |\n| **Layout** | `prebuilt-layout` | Struktur: tabeller, selection marks, seksjoner, key-value pairs (valgfritt) |\n| **General Document** | `prebuilt-document` | Key-value pairs, tabeller, selection marks fra generiske dokumenter |",
|
||||
"failing_part": "The third table row presenting `prebuilt-document` (General Document) as a current basic model — the source states the general document model is no longer supported and its capabilities live in the layout model.",
|
||||
"proposed_remainder": "### Grunnleggende modeller\n\n| Modell | Model ID | Formål |\n|--------|----------|--------|\n| **Read** | `prebuilt-read` | OCR: tekst, linjer, ord, språkdeteksjon |\n| **Layout** | `prebuilt-layout` | Struktur: tabeller, selection marks, seksjoner, key-value pairs (valgfritt) |",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "One whole table row is deleted and nothing is added, so the table asserts the existence of two basic models instead of three."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The remaining two rows (`prebuilt-read`, `prebuilt-layout`) are both judged grounded and the deleted row is the only mention of `prebuilt-document` in the file (grep: line 79 only), so no dangling reference or implied-availability trap is left behind."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The source confirms no replacement value for this cell — it says the model is no longer supported and its capabilities are in the layout model — so subtraction destroys no confirmed current fact, but a human may prefer a rewrite that explicitly records the deprecation and the redirect to layout."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 8,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md#14",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/data-drift-monitoring-detection.md",
|
||||
"line": 191,
|
||||
"real_line": 191,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Azure Machine Learning Workspace** (Verified)\nData drift monitoring krever:\n- Azure ML workspace (v2 API)\n- Compute resources (serverless Spark eller managed compute cluster)\n- Datastore for production inference data (Azure Blob Storage eller ADLS Gen2)\n- Optional: Application Insights for custom metrics logging",
|
||||
"failing_part": "Two sub-assertions: (a) `eller managed compute cluster` as an alternative compute option — the how-to page and the monitor schema require a Spark pool; and (b) the `Optional: Application Insights for custom metrics logging` prerequisite line, which is not among the v2 model-monitoring prerequisites.",
|
||||
"proposed_remainder": "**Azure Machine Learning Workspace** (Verified)\nData drift monitoring krever:\n- Azure ML workspace (v2 API)\n- Compute resources (serverless Spark)\n- Datastore for production inference data (Azure Blob Storage eller ADLS Gen2)",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "One alternative is struck from a disjunction and one whole bullet is deleted, with no word added, so the prerequisite list asserts a strict subset of what it asserted before."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The remainder states serverless Spark as the compute requirement, which is exactly what the evidence quote supports (\"Schedule model monitoring jobs to run on serverless Spark compute pools\"), and dropping the Application Insights bullet leaves no false implication because Application Insights is still covered on its own terms at lines 202-203."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The judge states the source does not support `managed compute cluster` as an option and does not list Application Insights in the prerequisites, so neither deletion removes anything the source confirms."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 9,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md#18",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/data-drift-monitoring-detection.md",
|
||||
"line": 218,
|
||||
"real_line": 218,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Microsoft Foundry (tidligere Azure AI Studio)** (Baseline + Verified)\nFor generative AI workloads: Microsoft Foundry har egen monitoring med observability features og generation quality metrics (groundedness, relevance, fluency). Støtter også drift detection for grounding data i RAG scenarios.",
|
||||
"failing_part": "The second sentence, `Støtter også drift detection for grounding data i RAG scenarios.` — the canonical observability page lists only Evaluation, Monitoring and Tracing, and covers `system drift` via scheduled evaluation on test datasets, not drift detection over RAG grounding data.",
|
||||
"proposed_remainder": "**Microsoft Foundry (tidligere Azure AI Studio)** (Baseline + Verified)\nFor generative AI workloads: Microsoft Foundry har egen monitoring med observability features og generation quality metrics (groundedness, relevance, fluency).",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "A complete sentence is deleted and the surviving sentence is untouched, so the paragraph asserts strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The judge states the first sentence holds on its own (Foundry has its own gen-AI monitoring with groundedness, relevance and fluency), and it makes no claim about drift that the deletion would leave half-standing."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The deleted sentence is not confirmed by the source, so nothing confirmed is lost; however the source does confirm an adjacent true fact (scheduled evaluation detects system drift) that a human may prefer to assert instead, which would make the fix a rewrite rather than a subtraction."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 10,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md#21",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/data-drift-monitoring-detection.md",
|
||||
"line": 395,
|
||||
"real_line": 396,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Setup options**:\n- **Out-of-box**: Automatically configured for Azure ML online endpoints (no configuration required)\n- **Advanced**: Custom monitoring for models deployed outside Azure ML (batch endpoints, external)\n- **Azure Event Grid integration**: Route monitoring alerts for automated response",
|
||||
"failing_part": "The `Advanced` bullet's category mapping: the source defines advanced setup as more signals, training/validation data as the reference dataset and top-N features, while models deployed outside Azure ML and batch endpoints belong to a separate setup path (\"Set up model monitoring for production data\").",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the `Advanced` bullet outright would assert strictly less, so condition 1 is not what blocks this item."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The heading `Setup options` frames the list as the option space, so a remainder of only `Out-of-box` plus an Event Grid bullet would imply that out-of-box is the only real setup path, which is exactly the standing-implication failure the contract warns about; and trimming only the parenthetical `(batch endpoints, external)` would leave the false `Advanced = models deployed outside Azure ML` mapping intact."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The source confirms both that an advanced setup exists (with a different meaning) and that models outside Azure ML / batch endpoints can be monitored, so removing the bullet destroys confirmed information rather than merely dropping unsupported specificity."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: fixing this requires deciding what `Advanced` denotes and re-splitting the option space (advanced signal configuration vs. the separate production-data setup), which is a rewrite, not a subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 11,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md#22",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/data-drift-monitoring-detection.md",
|
||||
"line": 400,
|
||||
"real_line": 400,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Statistical methods used**:\n- Jensen-Shannon divergence for categorical features\n- Wasserstein distance (Earth Mover's Distance) for numerical features\n- Population Stability Index (PSI) for feature stability",
|
||||
"failing_part": "Both the metric names and the feature-type mapping: the allowed metric is `jensen_shannon_distance` (Jensen-Shannon Distance, valid for numerical and categorical features, not categorical alone) and `normalized_wasserstein_distance` (Normalized Wasserstein Distance), and the per-feature-type mapping the claim constructs does not exist in the reference.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Stripping the `for ... features` qualifiers would assert strictly less, so condition 1 alone would not block the item."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A remainder of bare names would still read `Jensen-Shannon divergence` and `Wasserstein distance (Earth Mover's Distance)`, which are the wrong metric names, so the misleading part survives the subtraction."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence quote confirms the corrected values verbatim (`jensen_shannon_distance`, `normalized_wasserstein_distance`, `population_stability_index`), and the contract says that where the source supports a corrected value the fix is a swap, never subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails outright — the source supplies corrected metric names, making this a value-swap/rewrite; condition 2 also fails because subtraction leaves the wrong names standing.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 12,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/feedback-loops-continuous-improvement.md#8",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/feedback-loops-continuous-improvement.md",
|
||||
"line": 443,
|
||||
"real_line": 443,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Komponent | Azure-tjeneste | Formål |\n|-----------|----------------|--------|\n| **Data collection** | Inference tables (managed endpoints) | Capture production inputs/outputs |\n| **Monitoring** | Model Monitor (Azure ML) | Data drift, prediction drift, performance |\n| **Alerting** | Azure Monitor Alerts | Email/webhook ved threshold breach |\n| **Retraining** | Azure ML Pipelines | Triggered retraining workflow |\n| **A/B testing** | Staging endpoints | Champion vs challenger validation |\n| **Deployment** | Managed Online Endpoints | Blue-green deployment |",
|
||||
"failing_part": "The `Azure-tjeneste` cell of the `Data collection` row: `Inference tables (managed endpoints)` — Azure ML has no inference tables (a Databricks concept); collection is done by the Azure Machine Learning Data collector, which logs to Azure Blob Storage.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": false, "evidence": "There is no subtraction that repairs this: the cell must name a service, so emptying it breaks the table and deleting the row removes a component the source confirms exists."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Deleting the row would leave an Azure ML feedback-loop table with monitoring, alerting and retraining but no data-collection step, implying production data reaches Model Monitor by itself; the same wrong term also appears in the architecture diagram at line 296, which the subtraction would not touch."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence quote confirms the corrected value directly (\"Azure Machine Learning Data collector provides real-time logging of input and output data from models that are deployed to managed online endpoints ... stores the logged inference data in Azure blob storage\"), so the supported fix is a swap, not subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 (and 1) fails — the source names the correct service, so this is a value swap (Inference tables -> Data collector / Azure Blob Storage), and the wrong term recurs at line 296.",
|
||||
"confidence": "high"
|
||||
}
|
||||
]
|
||||
104
scripts/kb-eval/data/r11-o2-returns/batch-03.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-03.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 13,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/feedback-loops-continuous-improvement.md#9",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/feedback-loops-continuous-improvement.md",
|
||||
"line": 471,
|
||||
"real_line": 473,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "### Microsoft Foundry (GenAI)\n\n**Feedback loop-komponenter:**\n\n| Komponent | Azure-tjeneste | Formål |\n|-----------|----------------|--------|\n| **Production tracing** | MLflow Tracing (Databricks) | Span-level telemetry |\n| **User feedback** | Review App | Thumbs up/down, textual feedback |\n| **LLM judges** | Agent Evaluation | Automated quality scoring |\n| **Monitoring dashboard** | Microsoft Foundry Observability | Quality trends, latency, errors |\n| **Eval datasets** | MLflow Datasets (Unity Catalog) | Versioned test sets |\n| **Red teaming** | AI Red Teaming Agent | Adversarial testing for safety |",
|
||||
"failing_part": "Three of the six rows name Databricks-MLflow components (Review App, Agent Evaluation, MLflow Datasets in Unity Catalog) as the Azure services of a table headed '### Microsoft Foundry (GenAI)'; the row 'MLflow Tracing (Databricks)' is of the same family though the judge does not name it.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the three misattributed rows would leave a table asserting three service mappings instead of six, which is strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The surviving table would still head a section titled 'Microsoft Foundry (GenAI)' while carrying the row 'Production tracing | MLflow Tracing (Databricks)', so a Databricks component keeps standing as a Foundry component — the exact misleading-remainder pattern condition 2 forbids."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence quote confirms Foundry itself provides Evaluation, Monitoring and Tracing as core capabilities, so deleting the 'LLM judges' and 'Production tracing' rows destroys true information whose correct fix is to name Foundry's own evaluation/tracing services (a swap), not subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: the remainder still attributes a Databricks component to Foundry, and the source confirms Foundry has evaluation/tracing capabilities, so correcting the table means deciding which Foundry services to name.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 14,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/feedback-loops-continuous-improvement.md#11",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/feedback-loops-continuous-improvement.md",
|
||||
"line": 552,
|
||||
"real_line": 555,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Komponent | Power Platform-tjeneste | Formål |\n|-----------|-------------------------|--------|\n| **Automated feedback collection** | Power Automate | Route low-confidence predictions til human review |\n| **Storage** | Dataverse / SharePoint | Lagre feedback data |\n| **Model improvement** | AI Builder Feedback Loop | Automatically add reviewed samples to training set |\n| **Retraining** | AI Builder | Manual/scheduled retraining |",
|
||||
"failing_part": "Two sub-assertions: 'SharePoint' as a feedback storage service, and the word 'Automatically' in 'Automatically add reviewed samples to training set' (the source requires the documents to be selected, tagged and the model retrained).",
|
||||
"proposed_remainder": "| Komponent | Power Platform-tjeneste | Formål |\n|-----------|-------------------------|--------|\n| **Automated feedback collection** | Power Automate | Route low-confidence predictions til human review |\n| **Storage** | Dataverse | Lagre feedback data |\n| **Model improvement** | AI Builder Feedback Loop | Add reviewed samples to training set |\n| **Retraining** | AI Builder | Manual/scheduled retraining |",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Dropping the '/ SharePoint' alternative and the adverb 'Automatically' removes two assertions and adds none, so the edited rows assert strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "The edited rows themselves read correctly against the source (Dataverse storage; the feedback loop is the mechanism for adding reviewed samples before retraining), but line 566 of the same file still states 'Reviewed documents automatisk tilgjengelige i \"Feedback loop\" data source når modellen retraines', so the automaticity claim keeps standing a few lines below unless it receives the same subtraction."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The source names only the Dataverse 'AI Builder Feedback Loop' table and describes a select-tag-retrain flow, so neither 'SharePoint' nor the automaticity is confirmed, while the confirmed parts (Dataverse, adding reviewed samples, manual/scheduled retraining) are all retained."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 15,
|
||||
"id": "ms-ai-engineering/mlops-genaiops/model-evaluation-frameworks.md#3",
|
||||
"file": "skills/ms-ai-engineering/references/mlops-genaiops/model-evaluation-frameworks.md",
|
||||
"line": 39,
|
||||
"real_line": 39,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Risk & Safety** | Self-harm, Hateful content, Violence, Sexual content, Protected material, Indirect attack | Nei | Nei (Foundry-hosted GPT-4) | Content moderation og sikkerhetsvurdering |",
|
||||
"failing_part": "The parenthetical 'Foundry-hosted GPT-4' in the 'Krever judge model?' column — the source says these evaluators run against Microsoft's hosted safety models and contrasts them explicitly with GPT-based LLM-as-judge evaluators.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the parenthetical would leave a bare 'Nei' in the judge-model column, asserting strictly less than before."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "A bare 'Nei' is exactly what the source supports, since risk & safety evaluators need no user-supplied judge model."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The source supports a corrected value for the very slot being emptied — Microsoft's hosted safety models — so per the contract the fix is a value swap ('Foundry-hosted GPT-4' to 'Microsoft-hostede sikkerhetsmodeller'), not subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3: the source states the replacement fact (hosted safety models), which makes this an O1-style swap rather than a subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 16,
|
||||
"id": "ms-ai-engineering/rag-architecture/rag-caching-optimization.md#7",
|
||||
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
|
||||
"line": 186,
|
||||
"real_line": 188,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Tiers:**\n- **Premium tier** — 99.9% SLA, up to 120GB per shard\n- **Enterprise tier** — 99.99% SLA, active-active geo-replication, Flash storage support\n- **Enterprise Flash tier** — Up to 13TB cache size, 20% RAM + 80% NVMe Flash",
|
||||
"failing_part": "'Up to 13TB cache size' for the Enterprise Flash tier — the source states 300 GB – 4.5 TB.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the capacity clause would leave the Enterprise Flash bullet asserting only the 20% RAM / 80% NVMe split, which is strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "A tier bullet describing only the RAM/Flash composition carries no false standing implication, since the two sibling bullets keep their own accurate SLA and capacity statements."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence quote gives the documented Enterprise Flash range (300 GB – 4.5 TB), so the source supports a corrected value and the contract routes this to a swap rather than subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3: an explicit corrected capacity range exists in the cited source, making this an O1-style value swap.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 17,
|
||||
"id": "ms-ai-engineering/rag-architecture/rag-caching-optimization.md#12",
|
||||
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
|
||||
"line": 253,
|
||||
"real_line": 254,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Score Threshold Tuning** (APIM `score-threshold` er en DISTANSE: lavere = strengere, krever høyere semantisk likhet):\n- 0.1-0.2 → Strict matching, lavere hit rate, høy relevance\n- 0.3-0.5 → Balanced, medium hit rate, god relevance\n- 0.6-0.8 → Liberal matching, høyere hit rate, noe lavere relevance",
|
||||
"failing_part": "The three-band rubric (0.1-0.2 strict / 0.3-0.5 balanced / 0.6-0.8 liberal) — undocumented, and the two upper bands contradict the source's warning that a threshold above 0.2 may lead to cache mismatch.",
|
||||
"proposed_remainder": "**Score Threshold Tuning** (APIM `score-threshold` er en DISTANSE: lavere = strengere, krever høyere semantisk likhet):",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "The three bullet lines are deleted and the surviving heading line is kept byte-for-byte, so the passage asserts only the direction of the threshold and nothing about band values."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The retained direction statement is exactly what the source grounds ('lower values require higher semantic similarity for a match'), and the doc's own operative recommendations elsewhere (line 164 'Start med 0.15' and the policy sample at line 239 using score-threshold=\"0.15\") stay below the source's 0.2 mismatch warning; the only cosmetic residue is the now-dangling colon on the heading line, which a human may drop."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The source documents only a recommended starting point (0.05) and a mismatch warning above 0.2 — it confirms none of the three bands, so no confirmed information is lost, and it offers no replacement rubric that a swap could install."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 18,
|
||||
"id": "ms-ai-engineering/rag-architecture/rag-caching-optimization.md#2",
|
||||
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
|
||||
"line": 29,
|
||||
"real_line": 29,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Microsoft-stakken tilbyr flere tjenester optimalisert for AI-workloads: Azure Cache for Redis (traditional og semantic caching), Azure Cosmos DB (semantic cache med vektorsøk), Azure AI Search (built-in caching av search results), og Azure API Management (semantic caching for LLM APIs). Valget av løsning avhenger av cache-type, scale-requirements, og compliance-krav.",
|
||||
"failing_part": "The list item 'Azure AI Search (built-in caching av search results)' — the source states each query operates on the current index view with no caching or snapshot of results.",
|
||||
"proposed_remainder": "Microsoft-stakken tilbyr flere tjenester optimalisert for AI-workloads: Azure Cache for Redis (traditional og semantic caching), Azure Cosmos DB (semantic cache med vektorsøk), og Azure API Management (semantic caching for LLM APIs). Valget av løsning avhenger av cache-type, scale-requirements, og compliance-krav.",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "One of four enumerated services is dropped and the remaining sentence is otherwise untouched, so it asserts strictly less."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "The edited sentence itself is clean — the three surviving services are all genuine caching services — but the same file still carries the section '### Azure AI Search - Built-in Caching' (lines 303-318, 'Azure AI Search cacher automatisk content etter første query') and the verification row 'Azure AI Search caching | **Verified**' at line 510, so removing only the intro mention leaves the contradicted claim standing further down."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The cited source denies result caching outright rather than supplying a corrected value, and AI Search's only documented cache (the preview enrichment cache for skillset output in Azure Storage) is not query-result caching, so no confirmed fact is lost by removing the item from a list of RAG response caches."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
}
|
||||
]
|
||||
104
scripts/kb-eval/data/r11-o2-returns/batch-04.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-04.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 19,
|
||||
"id": "ms-ai-engineering/rag-architecture/rag-caching-optimization.md#14",
|
||||
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
|
||||
"line": 296,
|
||||
"real_line": 297,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Fordeler:**\n- Globally distributed, multi-region writes\n- Automatic indexing av vectors\n- 99.999% SLA med multi-region setup\n- Built-in TTL support",
|
||||
"failing_part": "The bullet \"Automatic indexing av vectors\" — the judge states vector indexes must be declared explicitly in the indexing policy (only at container creation) and the vector path is placed in excludedPaths, i.e. vectors are deliberately kept out of automatic indexing.",
|
||||
"proposed_remainder": "**Fordeler:**\n- Globally distributed, multi-region writes\n- 99.999% SLA med multi-region setup\n- Built-in TTL support",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "One of four bullets is deleted and the remaining three are untouched, so the passage asserts a strict subset of what it asserted before."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "The three surviving bullets (multi-region writes, 99.999% SLA, TTL) are each grounded per the judge's reason and none of them implies anything about vector indexing, but the preceding code sample (lines 275-292) issues VectorDistance queries without showing index creation, so a human should confirm that silence about the required explicit vector index is acceptable rather than an implicit \"no setup needed\"."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The evidence_quote contradicts rather than corrects the deleted bullet — there is no confirmed \"advantage\" value to swap in, since the source's fact (explicit vector index in the indexing policy, vector path in excludedPaths) is the negation of the deleted assertion, not a corrected form of it."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 20,
|
||||
"id": "ms-ai-engineering/rag-architecture/rag-caching-optimization.md#19",
|
||||
"file": "skills/ms-ai-engineering/references/rag-architecture/rag-caching-optimization.md",
|
||||
"line": 373,
|
||||
"real_line": 377,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Tier | Size | Kapasitet | Månedskostnad (NOK) | Best For |\n|------|------|-----------|---------------------|----------|\n| Basic C0 | 250 MB | N/A (no SLA) | ~400 | Dev/Test |\n| Standard C1 | 1 GB | 2 replicas, 99.9% SLA | ~1,200 | Small production |\n| Premium P1 | 6 GB | Clustering, geo-replication | ~7,000 | Enterprise |\n| Enterprise E10 | 12 GB | Active-active, 99.99% SLA | ~25,000 | Mission-critical |\n| Enterprise Flash F300 | 345 GB | 20% RAM + 80% Flash | ~60,000 | Large-scale AI |",
|
||||
"failing_part": "The Size cell of the Enterprise Flash F300 row: \"345 GB\" (source table gives F300 = 384 GB); secondarily the Kapasitet cell of the Standard C1 row: \"2 replicas\" (Learn describes Standard as two VMs in a replicated configuration, i.e. primary + one replica).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Emptying the Size cell of the F300 row would technically assert less, but that is not the operative test here because condition 3 already forecloses subtraction."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A pricing table whose Size column is populated for every tier except Enterprise Flash reads as an omission the reader must fill in, and the ~60,000 NOK cost line would stand with no capacity to justify it."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote positively confirms the corrected value (\"| F300 | 384 GB |\"), so the supported fix is a value swap 345 GB -> 384 GB (and correspondingly 2 replicas -> 1 replica for Standard), never deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails: the source confirms the corrected figure (384 GB), which makes this an O1 value swap — subtraction would destroy true information. A second independent failing sub-assertion (Standard C1 \"2 replicas\") is likewise source-corrected, not source-negated.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 21,
|
||||
"id": "ms-ai-governance/responsible-ai/content-safety-implementation.md#1",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/content-safety-implementation.md",
|
||||
"line": 37,
|
||||
"real_line": 42,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Protected Material (Code)** | Oppdager kopiert kode fra public repos | LLM-generert kode | Match med source citation URL | GA |",
|
||||
"failing_part": "The Status cell \"GA\" on the Protected Material (Code) row — What's new and the current quickstart both title the feature \"(preview)\", and the Aug 2024 GA covered only Prompt Shields and Protected Material for text.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Blanking the Status cell would assert strictly less than asserting \"GA\", so condition 1 alone is not what disqualifies this item."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "This is exactly the contract's documented failure case: a row with an empty Status cell sitting in a table where every other row is labelled GA or Preview leaves Protected Material (Code) standing among generally available features, which is the misleading standing implication the judge flagged."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote \"Protected material detection for code (preview)\" confirms the corrected status outright, so the supported fix is the swap GA -> Preview."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: the source confirms the replacement status (Preview), making this an O1 swap, and a blanked Status cell would still read as \"available\" beside the GA rows.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 22,
|
||||
"id": "ms-ai-governance/responsible-ai/content-safety-implementation.md#2",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/content-safety-implementation.md",
|
||||
"line": 37,
|
||||
"real_line": 40,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Groundedness Detection** | Verifiserer at LLM-svar er grunnlagt i kildemateriale | Query + grounding sources (maks 55K tegn) | Grounded/ungrounded score | Preview |",
|
||||
"failing_part": "The Input cell's scoping of the limit: \"Query + grounding sources (maks 55K tegn)\" — the live overview caps grounding sources at 55,000 characters and text/query separately at 7,500 characters, so the 55K ceiling is misattributed to the combined input.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Dropping the parenthetical to leave \"Query + grounding sources\" would assert strictly less, since the quantitative ceiling would simply be gone."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "A blank limit in a column where every neighbouring row states a hard limit (10K tegn, 4MB, min 110 tegn) invites the reader to assume none applies, though it asserts nothing false on its own."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote confirms \"Maximum length for grounding sources: 55,000 characters (per API call)\" — the 55,000 figure is true information about this feature, and the source additionally supplies the missing 7,500-character query cap, so the supported fix restates the two scoped limits rather than deleting the number."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails in the contract's documented pattern (the prebuilt-check case): the number is real but mis-scoped, and the source hands over both corrected values (55,000 for sources, 7,500 for text/query), so correcting requires asserting the split, not subtracting.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 23,
|
||||
"id": "ms-ai-governance/responsible-ai/stakeholder-communication-ai-decisions.md#6",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
|
||||
"line": 374,
|
||||
"real_line": 379,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Causal Inference** | \"What if\" analysis for counterfactuals | Business: Inform strategy. End users: \"What can I change to get different outcome?\" |",
|
||||
"failing_part": "The Funksjon gloss on the Causal Inference row: \"'What if' analysis for counterfactuals\" — counterfactual what-if is a separate dashboard component (DiCE) while causal inference (EconML) concerns causal treatment effects; the row merges the two and the table omits counterfactual what-if as a tool in its own right.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": false, "evidence": "The Stakeholder-verdi cell of the same row carries the counterfactual framing independently (\"What can I change to get different outcome?\"), so deleting only the Funksjon gloss leaves the same conflation asserted, and deleting the whole row would remove Causal Inference, a component the source confirms exists."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Any subtraction leaves a Responsible AI tool table that still presents six tools with counterfactual what-if absent, preserving the very omission the judge named."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote confirms both components and what each does (causal inference = causal treatment effects from historical data; counterfactual what-if = what to change for a different outcome), so the source supports a corrected gloss plus an added row, not a deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "All three conditions fail: fixing this requires deciding what to assert — a corrected causal-inference gloss and a separate counterfactual what-if entry — which is a rewrite, not a subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 24,
|
||||
"id": "ms-ai-governance/responsible-ai/stakeholder-communication-ai-decisions.md#9",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
|
||||
"line": 437,
|
||||
"real_line": 437,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Stakeholder communication features**:\n\n1. **Agent observability**: Alle agenter har unik identitet (owner, version, lifecycle status)\n - **Verdi**: Governance team kan tracke hvem som er ansvarlig for hvilke agenter\n\n2. **Centralized logging**: Key events logges til Azure Log Analytics\n - **Verdi**: Audit trail for compliance\n\n3. **Cost tracking**: Token consumption og compute usage per agent\n - **Verdi**: CFO/Finance kan allokere kostnader til avdelinger\n\n4. **User disclosure**: Agents identifiserer seg som AI (ikke menneske)\n - **Verdi**: Etisk transparency overfor sluttbrukere",
|
||||
"failing_part": "Items 1-3 under the \"### Copilot Studio\" heading: the attribution of unique agent identity to Copilot Studio (the source attributes it to Microsoft Entra Agent ID, and the inventory/registry to Agent 365), the identity field list (source tracks ownership, purpose, platform, access scope — not owner, version, lifecycle status), and the telemetry/cost mechanisms (Copilot Studio's own facilities are Application Insights telemetry and Copilot Credits analytics, not an Azure Log Analytics key-event log with token/compute per agent).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting items 1-3 and leaving only item 4 (user disclosure) would assert strictly less, but the remaining conditions foreclose that route."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The heading \"### Copilot Studio\" plus \"**Stakeholder communication features**\" would survive, and the governance workflow immediately below still instructs \"Assign agent identity (owner, cost center, compliance tags)\" at line 451, so the mis-attributed identity capability would keep standing in the section even after the bullets go."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote confirms that agent identity, lifecycle controls and an organizational registry genuinely exist (via Entra Agent ID and Agent 365), and the judge names Copilot Studio's real equivalents (Application Insights, Copilot Credits), so the supported fix is re-attribution and substitution, not deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: correcting requires deciding what to assert (re-attributing identity/registry to Entra Agent ID and Agent 365, and replacing Log Analytics/token-compute with Application Insights and Copilot Credits), and the surrounding heading and workflow keep the wrong attribution alive after any subtraction.",
|
||||
"confidence": "high"
|
||||
}
|
||||
]
|
||||
Loading…
Add table
Add a link
Reference in a new issue