feat(ms-ai-architect): R11 §10 måling #2 bølge 2 — alle 46 R8-enumerasjoner klassifisert [skip-docs]
Fire subagenter til (idx 25-46), samme rammer: read-only, ingen writes/commits, ingen web-oppslag, evidence_quote som eneste kilde. To presiseringer i prompten etter bølge 1: forslaget må være oppnåelig ved kun sletting (behov for omskriving er i seg selv O3-bevis), og eksplisitt leveringsplikt. Totalt 46/46: 17 O2-kandidater, 29 O3, 0 locator-bom. Maskin-verifikasjon: 46/46 passerer V1 (ordrett filtekst), 16 av 17 O2-forslag passerer V2+V2b; ett flagget (idx 14, rekapitalisering). Alle 29 O3-er felles av betingelse 3 — kilden leverer en korrigert verdi.
This commit is contained in:
parent
b174db0936
commit
94c99c46dd
4 changed files with 382 additions and 0 deletions
104
scripts/kb-eval/data/r11-o2-returns/batch-05.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-05.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 25,
|
||||
"id": "ms-ai-governance/responsible-ai/stakeholder-communication-ai-decisions.md#2",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/stakeholder-communication-ai-decisions.md",
|
||||
"line": 91,
|
||||
"real_line": 94,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Nivå | Målgruppe | Eksempel | Microsoft-verktøy |\n|------|-----------|----------|-------------------|\n| **Global explanations** | Business ledere, produkteiere | \"Hvilke faktorer påvirker lånegodkjenning generelt?\" | Azure ML Interpretability component |\n| **Local explanations** | Sluttbrukere, saksbehandlere | \"Hvorfor ble *min* lånesøknad avslått?\" | Counterfactual What-If |\n| **Cohort explanations** | Compliance, fairness officers | \"Påvirker modellen lavlønnede søkere ulikt?\" | Responsible AI Dashboard |",
|
||||
"failing_part": "The tool assignment in the 'Microsoft-verktøy' column: local explanations are attributed to 'Counterfactual What-If', whereas the source attributes both local and cohort explanations to the interpretability component itself (counterfactual what-if is a separate DiCE-based component for feature perturbations).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the 'Microsoft-verktøy' column (or emptying the offending cell) is pure character deletion and would leave the row asserting only the level/audience/example mapping."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Emptying only the offending cell leaves a blank tool cell in a tool column, which stands as the false implication that no Microsoft tool produces local explanations, while deleting the whole column also erases the two mappings the evidence supports."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote confirms that the interpretability views (global, local, and cohort) belong to the same dashboard/interpretability component, so the source supports a corrected value for the local row rather than its removal."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails: the source supports a corrected tool value for the local-explanations row (the interpretability component / Responsible AI dashboard), which makes this a value swap or rewrite, never a subtraction; condition 2 also fails for the narrow variant.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 26,
|
||||
"id": "ms-ai-governance/responsible-ai/transparency-documentation-standards.md#4",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
|
||||
"line": 114,
|
||||
"real_line": 117,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Komponenter i Scorecard:**\n\n1. **Model overview**: Architecture, training data, intended use\n2. **Fairness assessment**: Performance disparities across sensitive groups (gender, ethnicity, age)\n3. **Model interpretability**: Feature importance (global/local explanations)\n4. **Error analysis**: Error rates per cohort, confusion matrices\n5. **Counterfactual analysis**: What-if scenarios (e.g., \"loan approved if income +10k\")\n6. **Causal inference**: Causal vs correlational relationships i features\n7. **Data quality**: Dataset statistics, missing values, outlier analysis",
|
||||
"failing_part": "Items 4 (Error analysis) and 5 (Counterfactual analysis) are listed as Responsible AI Scorecard components, but the canonical scorecard segment enumeration does not contain them (they are dashboard components, not scorecard segments).",
|
||||
"proposed_remainder": "**Komponenter i Scorecard:**\n\n1. **Model overview**: Architecture, training data, intended use\n2. **Fairness assessment**: Performance disparities across sensitive groups (gender, ethnicity, age)\n3. **Model interpretability**: Feature importance (global/local explanations)\n6. **Causal inference**: Causal vs correlational relationships i features\n7. **Data quality**: Dataset statistics, missing values, outlier analysis",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Two list items are deleted whole and nothing else changes, so the passage asserts five scorecard components instead of seven."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "The five surviving items all map onto canonical segments named in the judge's reason (model overview, fairness insights, top important factors, causal insights, data analysis) so no false implication stands, but the raw numbering becomes 1,2,3,6,7 (renumbering would exceed delete-only) and the human must accept that artifact."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "yes", "evidence": "The judge states plainly that neither Error analysis nor Counterfactual analysis is a scorecard component, and the evidence_quote confirms nothing about them, so no source-confirmed information is destroyed and neither is a mangled form of a canonical segment name that a swap could repair."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 27,
|
||||
"id": "ms-ai-governance/responsible-ai/transparency-documentation-standards.md#12",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
|
||||
"line": 426,
|
||||
"real_line": 426,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Built-in disclosures:**\n\n| Component | Disclosure |\n|-----------|------------|\n| **Chat interface** | \"Powered by AI\" badge i chat window |\n| **Generative answers** | Attribution links til source documents |\n| **Plugin actions** | Confirmation prompts før sensitive actions (send email, delete file) |\n| **Data usage** | Privacy statement link i bot settings |",
|
||||
"failing_part": "The 'Chat interface' row (a \"Powered by AI\" badge in the chat window) and the 'Plugin actions' row (confirmation prompts before sensitive actions) are asserted as built-in Copilot Studio disclosures but are absent from the canonical enumeration of built-in safety components.",
|
||||
"proposed_remainder": "**Built-in disclosures:**\n\n| Component | Disclosure |\n|-----------|------------|\n| **Generative answers** | Attribution links til source documents |\n| **Data usage** | Privacy statement link i bot settings |",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Two full table rows are deleted and the header, separator and remaining rows are untouched, so the table asserts two built-in disclosures instead of four."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The judge confirms both surviving rows (citations for generative answers and the privacy statement link), and a shorter list of built-in disclosures carries no standing implication that the removed features are unavailable or deprecated."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The evidence_quote documents only human-oversight guidance ('review AI-generated outputs and automated actions before applying them'), which is adjacent to but does not confirm a confirmation-prompt feature, so the human should confirm that dropping the 'Plugin actions' row destroys nothing the source actually establishes."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 28,
|
||||
"id": "ms-ai-governance/responsible-ai/transparency-documentation-standards.md#3",
|
||||
"file": "skills/ms-ai-governance/references/responsible-ai/transparency-documentation-standards.md",
|
||||
"line": 82,
|
||||
"real_line": 83,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Microsoft implementasjon:**\n- Microsoft Foundry: Model catalog med built-in model cards for pretrained models\n- Hugging Face integration: Model cards synces automatisk\n- Custom models: Template for å generere egne model cards",
|
||||
"failing_part": "The second and third bullets — that Hugging Face model cards are synchronised automatically, and that a template exists for generating model cards for custom models — are not covered by the canonical model card enumeration.",
|
||||
"proposed_remainder": "**Microsoft implementasjon:**\n- Microsoft Foundry: Model catalog med built-in model cards for pretrained models",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Two whole bullets are deleted with no other change, leaving only the model-catalog assertion the judge says holds."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The single surviving bullet is exactly the part the judge confirms, and a one-line 'Microsoft implementasjon' list states less without implying anything false about Hugging Face or custom models."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The evidence_quote only covers what a model card contains, so it confirms neither removed bullet, but deleting the whole Hugging Face bullet also removes the bare notion of a Hugging Face integration, and the human must confirm that the integration itself is not separately source-confirmed (which would make it a narrower edit)."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 29,
|
||||
"id": "ms-ai-infrastructure/bcdr/monitoring-alerting-failover-detection.md#5",
|
||||
"file": "skills/ms-ai-infrastructure/references/bcdr/monitoring-alerting-failover-detection.md",
|
||||
"line": 180,
|
||||
"real_line": 185,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "AzureDiagnostics\n| where ResourceProvider == \"MICROSOFT.COGNITIVESERVICES\"\n| where Category == \"RequestResponse\"\n| where TimeGenerated > ago(1h)\n| extend\n deploymentName = tostring(properties_s.modelDeploymentName),\n latencyMs = duration_s * 1000,\n statusCode = resultCode_d\n| summarize\n P50 = percentile(latencyMs, 50),\n P95 = percentile(latencyMs, 95),\n P99 = percentile(latencyMs, 99),\n SuccessRate = round(countif(statusCode < 400) * 100.0 / count(), 2),\n TotalRequests = count()\n by bin(TimeGenerated, 5m), deploymentName",
|
||||
"failing_part": "The column names duration_s and resultCode_d in the extend clause: the documented AzureDiagnostics columns for Azure OpenAI are DurationMs and ResultSignature.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": false, "evidence": "Removing the two extend assignments would also force deletion of the summarize lines that consume latencyMs and statusCode, gutting the query rather than weakening one assertion, and the surviving KQL would no longer be a runnable statement without rewording."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A code block presented as a working KQL query that references undefined or removed identifiers stands as a broken artifact readers would copy and run, which is a worse standing implication than the original."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote explicitly projects DurationMs and ResultSignature, so the source supplies the corrected values and the contract routes this to O1/O3, never subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails outright (the source supports corrected column names DurationMs and ResultSignature, making this a value swap), and condition 1 fails because no delete-only edit leaves a coherent query.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 30,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#6",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md",
|
||||
"line": 103,
|
||||
"real_line": 103,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "workflow Isolate-CompromisedVM {\n param([string]$VMResourceId, [string]$IncidentId)\n\n $nsg = Get-AzNetworkSecurityGroup -ResourceId $VMResourceId\n Add-AzNetworkSecurityRuleConfig -NetworkSecurityGroup $nsg `\n -Name \"Block-All-Incident-$IncidentId\" `\n -Priority 100 -Access Deny -Protocol * -Direction Inbound `\n -SourceAddressPrefix * -DestinationAddressPrefix *\n Set-AzNetworkSecurityGroup -NetworkSecurityGroup $nsg",
|
||||
"failing_part": "The -ResourceId parameter on Get-AzNetworkSecurityGroup, which exists in no parameter set of that cmdlet (only -Name, -ResourceGroupName, -ExpandResource, -DefaultProfile).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the trailing ' -ResourceId $VMResourceId' is pure character deletion and leaves the line asserting only that Get-AzNetworkSecurityGroup is called."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A bare Get-AzNetworkSecurityGroup returns every NSG in the subscription, so the runbook would read as isolating the named VM while actually piping a collection into a deny-all rule addition, and the now-unused $VMResourceId parameter leaves a visibly broken script."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote gives the real parameter set including -Name and -ResourceGroupName, so the source supports a corrected targeting mechanism and the contract routes this to O1/O3 rather than subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: the delete-only remainder is an actively dangerous script that targets all NSGs, and the source supports a corrected parameterisation, so the fix requires deciding what to assert.",
|
||||
"confidence": "high"
|
||||
}
|
||||
]
|
||||
104
scripts/kb-eval/data/r11-o2-returns/batch-06.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-06.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 31,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#8",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md",
|
||||
"line": 132,
|
||||
"real_line": 139,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "{\n \"storageAccount\": \"forensicstorage\",\n \"immutabilityPolicy\": {\n \"immutabilityPeriodSinceCreationInDays\": 2190,\n \"allowProtectedAppendWrites\": false,\n \"state\": \"Locked\"\n },\n \"legalHold\": {\n \"tags\": [\"incident-2026-02-001\", \"model-theft-investigation\"],\n \"enabled\": true\n }\n}",
|
||||
"failing_part": "The legalHold object is given a field named \"enabled\"; the Storage API's LegalHold model exposes tags and hasLegalHold, so the literal field name \"enabled\" does not exist.",
|
||||
"proposed_remainder": "{\n \"storageAccount\": \"forensicstorage\",\n \"immutabilityPolicy\": {\n \"immutabilityPeriodSinceCreationInDays\": 2190,\n \"allowProtectedAppendWrites\": false,\n \"state\": \"Locked\"\n },\n \"legalHold\": {\n \"tags\": [\"incident-2026-02-001\", \"model-theft-investigation\"]\n }\n}",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Only characters are deleted (the trailing comma on the tags line plus the whole \"enabled\": true line), so the block asserts the same immutabilityPolicy fields but no longer asserts any boolean field on legalHold."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "A legalHold object carrying only tags is exactly what the evidence describes as the operative state, since the quote says hasLegalHold is set to true by SRP whenever at least one tag exists, so nothing in the remainder implies the hold is inactive."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The source confirms a neighbouring field name (hasLegalHold) which could argue for a swap rather than a deletion, but the quote also states hasLegalHold is set by SRP rather than by the caller, so writing it into a desired-state config payload would assert something the API does not accept — a human should confirm that deletion, not swap, is the right call here."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 32,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#17",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md",
|
||||
"line": 474,
|
||||
"real_line": 474,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Defender XDR** | M365 E5 Security or E5 | Includes Defender for Endpoint, Identity, M365 |",
|
||||
"failing_part": "Two parts: the Required License cell asserting that M365 E5 Security or E5 is what Defender XDR requires, and the product name \"M365\" in the component list (the real component is Microsoft Defender for Office 365).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": false, "evidence": "The failing text sits in a mandatory table cell under the column header Required License, so deleting it leaves an empty cell that a reader parses as a positive claim (no licence needed / unknown), not as a narrower claim."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Even the survivable deletion of \", M365\" from the third cell would leave the licence cell standing with the contradicted E5-only requirement, which the evidence directly refutes by listing M365 E3 with the Defender Suite add-on and M365 E3 with EMS E5 as qualifying licences."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence confirms a corrected licence set and the judge names the correct product name (Microsoft Defender for Office 365), so both failing parts have supported replacement values and must be swapped, not subtracted."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 1 and 3 both fail: a mandatory table cell cannot be emptied without asserting something new, and the source supplies corrected values for both the licence list and the product name.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 33,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#11",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||||
"line": 211,
|
||||
"real_line": 211,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Capabilities:** *(Verified MCP 2026-04)*\n- Automated detection of AI workloads across Azure subscriptions (via Azure Resource Graph)\n- AI security posture management: automate detection and remediation of generative AI risks\n- Security recommendations for AI models, data stores, network isolation\n- Integration with Purview for data classification, DLP og Insider Risk Management for prompt-based data exfiltration",
|
||||
"failing_part": "Two parts attributed to Defender for Cloud AI Security Posture Management that the AISPM page does not support: the discovery mechanism \"(via Azure Resource Graph)\" and the entire Purview-integration bullet.",
|
||||
"proposed_remainder": "**Capabilities:** *(Verified MCP 2026-04)*\n- Automated detection of AI workloads across Azure subscriptions\n- AI security posture management: automate detection and remediation of generative AI risks\n- Security recommendations for AI models, data stores, network isolation",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "The parenthetical and the fourth bullet are removed by deleting characters only, leaving the surviving bullets byte-identical, so the section attributes strictly fewer capabilities to AISPM."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The surviving first bullet matches the evidence quote, which states that Defender for Cloud automatically and continuously discovers deployed AI workloads, and the remainder makes no claim at all about how discovery is implemented or about Purview."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "Azure Resource Graph and Purview are real tools the CAF page describes as separate from Defender for Cloud, so deleting the bullet drops content that is true of Purview itself even though it is false of AISPM — a human should confirm that relocating rather than deleting it is not required."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 34,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#14",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||||
"line": 281,
|
||||
"real_line": 281,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Microsoft Defender for Cloud (AI)** | ~$15/server/month (standard tier) | AI workload discovery, security posture management, threat detection |",
|
||||
"failing_part": "The License/Cost cell: both the plan name \"standard tier\" and the per-server unit price, since the AI capabilities come from the Defender CSPM and Defender for AI Services plans and are billed per resource and per scanned tokens respectively.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": false, "evidence": "The whole cell is the failing assertion, and blanking a cell under the column header License/Cost reads to a table user as a positive statement about price rather than as silence."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "A row that names an AI capability set in a cost table but shows no licence or cost implies the capability is free or licence-free, which the evidence contradicts by naming the Defender CSPM plan as the securing plan."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The judge's reason cites supported replacement values (Defender CSPM plus Defender for AI Services, resource-based and token-based billing capped at 75 billion tokens scanned), so the correct fix supplies a value rather than removing one."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "All three conditions fail; the cell cannot be emptied without asserting something new, and the source supports corrected plan names and billing units.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 35,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#2",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||||
"line": 37,
|
||||
"real_line": 37,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Spoofing** | Neural Net Reprogramming, Malicious ML Providers | Important-Critical | Strong API authentication, access control, client-server mutual auth |",
|
||||
"failing_part": "The placement of Malicious ML Providers under Spoofing, and the Important-Critical severity band applied to it, since the source treats it as information disclosure with severity Important if data is PII and Moderate otherwise.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting \", Malicious ML Providers\" is a pure character deletion and removes one of the two threats the Spoofing row asserts."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "After the deletion the Important-Critical band stands alone against Neural Net Reprogramming, whose severity and Spoofing placement the evidence never establishes and whose treatment in the source as \"an abuse scenario\" is what the judge calls into question, so the subtraction leaves an unsupported severity assertion looking newly precise."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The source confirms Malicious ML Providers is a real threat and supplies its correct home (information disclosure) and its correct severity, and the file already has an Information Disclosure row at line 40 to receive it, so the supported fix is a relocation with a corrected severity, not a deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: the surviving severity band becomes an unsupported claim about the one remaining threat, and the source confirms both the threat and its corrected categorisation, which makes this a move/rewrite rather than a subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 36,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#3",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||||
"line": 38,
|
||||
"real_line": 38,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Tampering** | Data Poisoning (targeted/indiscriminate), Backdoored Models | Critical | Training data validation, anomaly detection, RONI defense, bagging |",
|
||||
"failing_part": "The \"/indiscriminate\" qualifier, which extends the Tampering placement and the Critical severity to indiscriminate data poisoning; the source gives that variant severity Important and the traditional parallel authenticated denial of service.",
|
||||
"proposed_remainder": "| **Tampering** | Data Poisoning (targeted), Backdoored Models | Critical | Training data validation, anomaly detection, RONI defense, bagging |",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the eleven characters \"/indiscriminate\" narrows the row from both poisoning variants to the targeted variant only, with every other character untouched."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The judge states explicitly that targeted poisoning and Backdoored Models are in fact Critical, so the surviving row is fully supported and it makes no claim whatsoever about the indiscriminate variant."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The source does confirm indiscriminate data poisoning exists with severity Important, but the row's single shared Severity cell already reads Critical for the two remaining threats, so the confirmed value cannot be swapped in place and retaining the variant would require adding a new row — a human should confirm that dropping the coverage is acceptable rather than mandating that rewrite."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
}
|
||||
]
|
||||
104
scripts/kb-eval/data/r11-o2-returns/batch-07.json
Normal file
104
scripts/kb-eval/data/r11-o2-returns/batch-07.json
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
[
|
||||
{
|
||||
"idx": 37,
|
||||
"id": "ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#5",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||||
"line": 40,
|
||||
"real_line": 40,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Information Disclosure** | Model Inversion, Membership Inference, Model Stealing | Important-Critical | Rate limiting, access control, output obfuscation, differential privacy |",
|
||||
"failing_part": "The severity band Important-Critical asserted for all three threats, plus the placement of Membership Inference under Information Disclosure (source files it as a Data Privacy issue with no security severity, and Model Stealing is Important only in security-sensitive models, Moderate otherwise).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting 'Membership Inference, ' from the threat cell would leave the row asserting a strict subset of the original threat list."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Even after removing Membership Inference the remainder still asserts the severity band Important-Critical for Model Stealing, which the judge documents as Moderate outside security-sensitive models, so the surviving cell keeps an unsupported severity floor."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The evidence_quote confirms Membership Inference is a real, source-documented AI threat (filed under Data Privacy), so deleting the term drops information the source does carry, merely under a different heading."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 2 fails: no deletion-only edit repairs the severity cell, and correcting 'Important-Critical' requires deciding what severity to assert for the remaining threats — a rewrite, not a subtraction.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 38,
|
||||
"id": "ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#18",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md",
|
||||
"line": 396,
|
||||
"real_line": 396,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Policy templates:**\n- \"DSPM for AI - Detect risky AI usage\"\n- \"DSPM for AI - Unethical behavior in AI apps\"\n- \"DSPM for AI - Protect sensitive data from Copilot processing\"",
|
||||
"failing_part": "The listing of 'DSPM for AI - Unethical behavior in AI apps' and 'DSPM for AI - Protect sensitive data from Copilot processing' as Insider Risk Management policy templates; the source assigns the former to Communication Compliance and the judge assigns the latter to DLP.",
|
||||
"proposed_remainder": "**Policy templates:**\n- \"DSPM for AI - Detect risky AI usage\"",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the two mis-assigned bullet lines leaves a strict subset of the original list under the same heading, asserting one template instead of three."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The surviving bullet is the one template the judge confirms is genuinely an Insider Risk Management policy, and the heading '**Policy templates:**' under section 4.3 does not claim exhaustiveness, so nothing false is left standing."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The source confirms 'DSPM for AI - Unethical behavior in AI apps' exists as a Communication Compliance template, so an operator may prefer relocating both bullets to correctly-labelled sections rather than deleting them — though nothing the source confirms *about Insider Risk Management* is lost."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 39,
|
||||
"id": "ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#25",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md",
|
||||
"line": 606,
|
||||
"real_line": 607,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "**Viktige cmdlets:**\n- `New-DlpCompliancePolicy`: Create DLP policy\n- `New-DlpComplianceRule`: Add rule til policy\n- `Get-DlpCompliancePolicy`: List policies\n- `Set-DlpPolicy`: Update existing policy\n- `Get-Label`: List sensitivity labels med GUIDs",
|
||||
"failing_part": "The bullet '`Set-DlpPolicy`: Update existing policy' — the cmdlet is retired from the cloud-based service and functional only in on-premises Exchange.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the `Set-DlpPolicy` bullet would leave four cmdlets, a strict subset of the original five."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "A four-cmdlet list with New- and Get- verbs but no update verb could leave a reader to infer no supported update cmdlet exists, in a section explicitly titled 'Viktige cmdlets'."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote states outright 'Use the Set-DlpCompliancePolicy and Set-DlpComplianceRule cmdlets instead', so the source supports a corrected value and the contract forbids fixing this by subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails — this is the documented prebuilt-check pattern: the source names the replacement cmdlet, so the correct fix is an O1 value swap (`Set-DlpPolicy` -> `Set-DlpCompliancePolicy`), never deletion.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 40,
|
||||
"id": "ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#5",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md",
|
||||
"line": 134,
|
||||
"real_line": 133,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Dependency scanning genererer alerts for:\n- **Direct vulnerabilities**: Pakker i `requirements.txt`\n- **Transitive vulnerabilities**: Pakker som direkte dependencies bruker\n- **CVE severity mapping**: Critical (CVSS ≥9.0), High (7.0-9.0), Medium (4.0-7.0), Low (1.0-4.0)",
|
||||
"failing_part": "The third bullet '**CVE severity mapping**' presented as a category of alert that dependency scanning generates; severity is a property of an alert, not an alert category.",
|
||||
"proposed_remainder": "Dependency scanning genererer alerts for:\n- **Direct vulnerabilities**: Pakker i `requirements.txt`\n- **Transitive vulnerabilities**: Pakker som direkte dependencies bruker",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the third bullet reduces the enumeration from three alert categories to two without touching any other assertion."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The two surviving bullets map exactly onto the evidence_quote's 'any open-source component, direct or transitive, found to be vulnerable', so the remainder states precisely what the source states."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The evidence_quote says nothing about CVSS bands, so no source-confirmed fact is lost, but an operator may prefer an O3 rewrite that keeps the severity bands re-framed as an alert property rather than deleting them."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 41,
|
||||
"id": "ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#8",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md",
|
||||
"line": 157,
|
||||
"real_line": 156,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Defender for Containers:\n- Genererer vulnerability assessments automatisk når image pushes til Azure Container Registry\n- Blokkerer deployment av images med critical vulnerabilities (konfigurerbart via Azure Policy)\n- Integrerer med Azure Monitor for alerting",
|
||||
"failing_part": "The parenthetical '(konfigurerbart via Azure Policy)' — blocking is configured through Defender for Containers' gated-deployment security rules in Defender for Cloud, and the relevant Azure Policy definition offers only AuditIfNotExists and Disabled.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting ' (konfigurerbart via Azure Policy)' removes the mechanism attribution while leaving the blocking assertion untouched, so strictly less is asserted."},
|
||||
"cond2_remainder_not_misleading": {"holds": "human_must_confirm", "evidence": "The bare remainder 'Blokkerer deployment av images med critical vulnerabilities' reads as out-of-the-box behaviour, whereas the evidence_quote describes an admission-controller capability that must be configured with Deny rules and applies at Kubernetes admission, not at ACR push."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote supplies the corrected mechanism — gated deployment via an admission controller with Deny rules — so the source supports a replacement value and the contract routes this to O1/O3 rather than subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails (and condition 2 is doubtful): the source names the correct configuration mechanism, so the parenthetical should be swapped, not deleted.",
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 42,
|
||||
"id": "ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#9",
|
||||
"file": "skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md",
|
||||
"line": 202,
|
||||
"real_line": 200,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Microsoft tilbyr verifiserte modeller via:\n\n- **Azure Machine Learning Model Catalog**: Curated models med security attestation\n- **HuggingFace Registry i Azure**: Integrert med Azure ML, med provenance tracking",
|
||||
"failing_part": "The second bullet presenting the HuggingFace Registry as a Microsoft channel for verified models with provenance tracking; the source calls it a community registry Microsoft support doesn't cover, with weights not hosted on Azure.",
|
||||
"proposed_remainder": "Microsoft tilbyr verifiserte modeller via:\n\n- **Azure Machine Learning Model Catalog**: Curated models med security attestation",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the HuggingFace bullet leaves one channel where two were asserted, with no other text altered."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "The surviving bullet is the channel the judge leaves unchallenged, and the heading 'Microsoft tilbyr verifiserte modeller via:' remains true of the Model Catalog alone — the disputed entry is removed rather than left standing in a list of vetted options."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The source does confirm HuggingFace models can be deployed from Azure ML, so the 'Integrert med Azure ML' fragment is true; an operator may prefer an O3 rewrite that keeps the entry with a community-registry/no-Microsoft-support caveat instead of deleting it."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
}
|
||||
]
|
||||
70
scripts/kb-eval/data/r11-o2-returns/batch-08.json
Normal file
70
scripts/kb-eval/data/r11-o2-returns/batch-08.json
Normal file
|
|
@ -0,0 +1,70 @@
|
|||
[
|
||||
{
|
||||
"idx": 43,
|
||||
"id": "ms-ai-security/cost-optimization/gpt5-gpt41-pricing-models.md#9",
|
||||
"file": "skills/ms-ai-security/references/cost-optimization/gpt5-gpt41-pricing-models.md",
|
||||
"line": 203,
|
||||
"real_line": 202,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Modell | Takst-nivå | Copilot Credits | Power Platform Credits |\n|--------|-----------|----------------|----------------------|\n| `gpt-4.1-mini` | **Basic** | Laveste forbruk | Laveste forbruk |\n| `gpt-4.1` | **Standard** | Moderat forbruk | Moderat forbruk |\n| `gpt-5-chat` (preview) | **Standard** | Moderat forbruk | Moderat forbruk |\n| `gpt-5-reasoning` (preview) | **Premium** | Høyeste forbruk | Høyeste forbruk |\n| `o3` | **Premium** | Høyeste forbruk | Høyeste forbruk |\n| `Claude Sonnet 4.5` (experimental) | **Standard** | Moderat forbruk | Moderat forbruk |\n| `Claude Opus 4.5` (experimental) | **Premium** | Høyeste forbruk | Høyeste forbruk |",
|
||||
"failing_part": "The `Claude Sonnet 4.5` / `Claude Opus 4.5` rows (source lists 4.6 versions) and the `o3` row (absent from the source rate table).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the o3 row and the two Claude 4.5 rows would leave a table asserting rate tiers only for the four gpt-* models, which is strictly fewer assertions."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The surrounding prose is not deletable in the same stroke — line 214 still names o3, line 216 still names 'Claude Opus 4.5' and line 217 still states 'Claude Sonnet 4.5 og Opus 4.5 er nå tilgjengelig i Copilot Studio', so a table without those rows reads as a contradiction of its own section."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote positively confirms 'Claude Sonnet 4.6 | Standard rate', so the source supports a corrected value (4.5 -> 4.6) and deleting the row destroys true information — by contract that is O1/O3, never subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 both fail: the source confirms corrected Claude version/tier values (value swap territory), and any table-only subtraction leaves the section's prose asserting the deleted models.",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 44,
|
||||
"id": "ms-ai-security/cost-optimization/gpt5-gpt41-pricing-models.md#16",
|
||||
"file": "skills/ms-ai-security/references/cost-optimization/gpt5-gpt41-pricing-models.md",
|
||||
"line": 468,
|
||||
"real_line": 468,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| Modell | Tilgjengelighet | Registrering |\n|--------|----------------|-------------|\n| `gpt-5` | GA (begrenset) | Krever godkjenning (aka.ms/oai/gpt5access) |\n| `gpt-5-mini` | GA | Ikke nødvendig |\n| `gpt-5-nano` | GA | Ikke nødvendig |\n| `gpt-5-chat` | Preview (2 versjoner) | Ikke nødvendig |\n| `gpt-5-codex` | GA (begrenset) | Krever godkjenning |\n| `gpt-5-pro` | GA (begrenset) | Kun MCA-E/Default-abonnementer |",
|
||||
"failing_part": "The '(begrenset)' availability qualifier plus the Registrering cells 'Krever godkjenning (aka.ms/oai/gpt5access)', 'Krever godkjenning' and 'Kun MCA-E/Default-abonnementer' for gpt-5, gpt-5-codex and gpt-5-pro.",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting ' (begrenset)' and the three Registrering cell contents leaves the table asserting only that the models are GA, which is strictly less than before."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "Three empty Registrering cells sitting beside rows that explicitly say 'Ikke nødvendig' reads as 'registration status unknown/omitted' rather than 'no registration required', which is precisely the fact the source establishes."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "'Access is no longer restricted for this model.' is an affirmative source statement supporting the corrected values 'GA' and 'Ikke nødvendig', so the correct fix is a value swap, not a deletion."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Condition 3 fails (source supports a corrected value for both columns, making this O1-shaped) and condition 2 fails (blank cells in a column whose other rows are filled in are misleading).",
|
||||
"confidence": "high"
|
||||
},
|
||||
{
|
||||
"idx": 45,
|
||||
"id": "ms-ai-security/cost-optimization/semantic-caching-patterns.md#14",
|
||||
"file": "skills/ms-ai-security/references/cost-optimization/semantic-caching-patterns.md",
|
||||
"line": 436,
|
||||
"real_line": 436,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "| **Data residency** | Bruk Norway East/West for Redis og OpenAI for å sikre data forblir i Norge/EU. |",
|
||||
"failing_part": "The '/West' half of the region pair, i.e. the standing implication that Azure OpenAI can be deployed in Norway West.",
|
||||
"proposed_remainder": "| **Data residency** | Bruk Norway East for Redis og OpenAI for å sikre data forblir i Norge/EU. |",
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting the five characters '/West' narrows a two-region recommendation to a one-region recommendation and asserts nothing new."},
|
||||
"cond2_remainder_not_misleading": {"holds": "yes", "evidence": "'Bruk Norway East for Redis og OpenAI' is true for both products (Azure Managed Redis is available in Norway East and the source's regional table has a norwayeast column), and it is consistent with line 451 of the same file, 'Schrems II: Azure OpenAI i EU-region (Norway East)'."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "human_must_confirm", "evidence": "The deletion incidentally drops the true fact that Redis is also available in Norway West, but the line is a single bundled recommendation for both products — keeping '/West' is impossible without re-asserting the false OpenAI half, and no researched value is needed to obtain the remainder."},
|
||||
"verdict": "O2_CANDIDATE",
|
||||
"o3_reason": null,
|
||||
"confidence": "medium"
|
||||
},
|
||||
{
|
||||
"idx": 46,
|
||||
"id": "ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#11",
|
||||
"file": "skills/ms-ai-security/references/cost-optimization/vector-storage-cost-optimization.md",
|
||||
"line": 73,
|
||||
"real_line": 73,
|
||||
"locator_failed": false,
|
||||
"file_text_verbatim": "Azure AI Search lagrer vektorer i to kopier:\n1. **Index copy** (i minne, brukes til query execution)\n2. **Stored copy** (på disk, brukes til retrieval i query response)\n\nVed å sette `stored: false` kan man spare opptil 50 % disklagring, men man mister muligheten til å returnere vektorer i query-responser. Dette er akseptabelt i de fleste RAG-scenarier der kun tekst/metadata returneres.",
|
||||
"failing_part": "The count 'to kopier' and the implied exhaustiveness of the two-item enumeration — the source lists a third stored instance (original full-precision vectors retained for rescoring).",
|
||||
"proposed_remainder": null,
|
||||
"cond1_strictly_less": {"holds": true, "evidence": "Deleting 'to ' would drop the numeric assertion, so strictly less would be asserted — but no deletion-only edit yields grammatical Norwegian ('lagrer vektorer i kopier:')."},
|
||||
"cond2_remainder_not_misleading": {"holds": "no", "evidence": "The numbered list 1./2. survives any deletion of the count and still reads as a complete inventory, and the dependent sentence at line 77 ('spare opptil 50 % disklagring') derives its arithmetic from exactly two copies."},
|
||||
"cond3_nothing_confirmed_removed": {"holds": "no", "evidence": "The evidence_quote enumerates three storage instances, so the source supports a corrected count and a third list item — supplying them is a rewrite, not a subtraction."},
|
||||
"verdict": "O3",
|
||||
"o3_reason": "Conditions 2 and 3 fail: no character-only deletion produces a grammatical, non-exhaustive remainder, the downstream 50 % cost claim depends on the wrong count, and the source confirms the corrected three-copy value.",
|
||||
"confidence": "high"
|
||||
}
|
||||
]
|
||||
Loading…
Add table
Add a link
Reference in a new issue