De tre siste medlemmene (idx-26ab, idx-26ae, idx-26ah) baerte hver sitt flagg som holdt dem utenfor klasse-entryen. Ingen av flaggene overlevde maalingen som den saken de var bokfoert som. MAALT, IKKE ARVET: 58 av 58 medlemmer i r11-footer-class-2026-08-11.json har naa ingen ordrett ankerlinje igjen i korpus. Klassen er lukket. idx-26ab (prose_repeat). Entryens egen note sa at andre lokator «is a rewrite, not a line deletion». Det ble falsifisert ved aa skrive ut resultatstrengen: sletting av de 16 sammenhengende tegnene «12 MCP-kall til » gir en grammatisk norsk setning som fila selv baerer (8 nummererte kilder, 20 kodeblokker), og forfatter null ord. Argumentet sa rewrite; strengen sa sletting. => RATIFISERT #25: en delete-klasse-telling gjentatt INNE i en linje (hodefelt eller loepende prosa) lukkes ved aa slette det sammenhengende fragmentet. Tre vilkaar som ALLE maales: resten er grammatisk, resten er grunnet i fila, null ord forfattes. Holder ett av dem ikke, kommer saken tilbake som spoersmaal. Widening av #24 fra den ene parentesen den ble skrevet for. idx-26ae (label_value_mismatch). Spoersmaalet — lukkes en feilmerket linje av aa slettes, naar det er etiketten og ikke tallet som villeder — trengte ingen ny form. Linja har INGEN keep-verdi (3 docs_search + 2 docs_fetch, og dens eget resultat sier «= 5 MCP-kall»), saa den er ikke en #23-blandet linje; slettingen tar etiketten med seg; og den eneste alternative reparasjonen, aa doepe om «MCP-kilder» til «MCP-kall», er noeyaktig det #22 forbyr. idx-26ah (label_value_mismatch). Den smale editen var tilgjengelig og ble forkastet med grunn: :653 er eneste medlem innenfor de 58, men sletting av den alene ville latt :651 (naar genereringen kjoerte) og :652 (hvilket verktoey) staa rett over editen med samme referent. #22s tidsstempel-klausul sier INSIDE THE BLOCK, og her finnes ingen merkelapp — den rekker ikke, og kunne ikke strekkes uten den stille utvidelsen #22 selv nektet. => RATIFISERT #26, SCOPET TIL ÉN FIL: en avsluttende umerket rekke der HVER linje feiler referent-testen slettes hel. Den generelle klassen staar fortsatt aapen paa idx-26ar. NY DEFEKT FUNNET VED AA MAALE NABOEN FOER DEN BLE STOLT PAA (#21/#23-plikten): rag-document-preprocessings :792 «8 Microsoft Learn-artikler + 4 GitHub-repos = 12 kilder» er keep-klasse og to av tre deler er usanne. Maalt over to populasjoner som er enige (Kilder-seksjonen og hele fila): 8 distinkte learn.microsoft.com-artikler stemmer, men det er 2 distinkte github.com-URLer, ikke 4, og dermed 10 navngitte kilde-URLer, ikke 12 — eller 13 om de tre pris-URLene teller som kilder, en lesning linja ikke oppgir. Ingen lesning gir 12. Aritmetikken stemmer bare fordi GitHub-halvdelen er feil. BOKFOERT SOM idx-26at, IKKE REPARERT: aa rette 4 til 2 forfatter tall inn i korpus (idx-26v-faren), og ingen ratifisert form dekker en keep-telling som bare er gal — #24 slettet en slik telling kun fordi BEGGE halvdeler feilet. Beslutningen operatoeren skylder entryen er om dette programmet faar skrive et tall det har maalt. Koe: 54 entries (13 aapne, 41 resolved). Suite 1052/1052. Alle _meta-pekere sjekket for haand mot noekkelsettet — ingenting validerer _meta. De 5 utrackede maaledatafilene i scripts/kb-eval/data/ er tracket etter operatoerbeslutning (spurt 2026-08-11, avgjort 2026-08-12), paa linje med r11-footer-class-2026-08-11.json.
257 lines
1.1 MiB
257 lines
1.1 MiB
[
|
||
{
|
||
"file": "skills/ms-ai-infrastructure/references/hybrid-edge/disconnected-ai-scenarios.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-infrastructure/references/hybrid-edge/disconnected-ai-scenarios.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-infrastructure/references/hybrid-edge/disconnected-ai-scenarios.md`)\n\n[\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#1\",\n \"claim\": \"Foundry Local (on-device) er tilgjengelig som lokal modellruntime i klientapper på Windows, macOS og Linux, og krever ikke Azure-abonnement.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-local/what-is-foundry-local\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#2\",\n \"claim\": \"Foundry Local (on-device) tilbyr SDK for C# | JavaScript | Rust | Python.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-local/what-is-foundry-local\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#3\",\n \"claim\": \"Foundry Local on Azure Local (enterprise-skala inferens på on-prem Arc-enabled Kubernetes, tidligere Azure Stack HCI) er i Preview med tilgang på søknad og krever Azure-abonnement.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#4\",\n \"claim\": \"Foundry Local (on-device) sin modellkatalog omfatter Phi | Qwen | DeepSeek | Mistral | GPT OSS (chat) | Whisper (audio), og kjører offline etter at modeller er lastet ned og caches lokalt uten per-token-kostnad.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-local/what-is-foundry-local\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#5\",\n \"claim\": \"Air-gapped/disconnected deployment av Foundry Local on Azure Local er offisielt støttet fra juni 2026, i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#6\",\n \"claim\": \"I disconnected (air-gapped) modus hentes modeller fra et lokalt `edgeartifacts` container registry fylt fra expansion packs, og Arc-extension installeres ved at expansion pack lastes ned og importeres manuelt — mot Foundry cloud-katalog og standard Arc-extension i connected modus.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#7\",\n \"claim\": \"Sertifikathåndtering i disconnected Foundry Local on Azure Local bruker `cert-manager` + `trust-manager` (levert i expansion pack), mot `azure-cert-manager` i connected modus.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#8\",\n \"claim\": \"I disconnected modus sendes ingen telemetri til Microsoft, og autentisering er integrert med lokal Active Directory i stedet for offentlige Entra ID-endepunkter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#9\",\n \"claim\": \"Expansion packs importeres i disconnected Foundry Local on Azure Local med PowerShell-cmdletene `Start-AldoExpansionPackUpload` | `Start-AldoExpansionPackInstallation`, og publiseres til `edgeartifacts`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#10\",\n \"claim\": \"Inferens-runtimes i Foundry Local on Azure Local er ONNX-GenAI (CPU/GPU) | vLLM (GPU).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#11\",\n \"claim\": \"Multi-node Kubernetes-støtte for Foundry Local on Azure Local kom i juni 2026 og gir concurrent inferens og støtte for større modeller.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#12\",\n \"claim\": \"Foundry Local on Azure Local er del av Microsoft Sovereign Private Cloud, og eligibility for disconnected-modus krever dokumentert forretnings- eller regulatorisk begrunnelse.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#13\",\n \"claim\": \"Følgende Microsoft Foundry Tools-containere er GA i disconnected modus: Speech to Text | Custom Speech to Text | Neural Text to Speech | Translator (text-translation) | Language Detection | Key Phrase Extraction | Named Entity Recognition | PII Detection | Sentiment Analysis | CLU | Read OCR (vision-read) | Document Intelligence.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#14\",\n \"claim\": \"Summarization-containeren (text-summarization) er tilgjengelig disconnected, men i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#15\",\n \"claim\": \"Content Safety (Text) | Content Safety (Image) | Prompt Shields er tilgjengelige som disconnected containere, alle i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#16\",\n \"claim\": \"Azure Stack Edge i disconnected modus mangler Azure Portal-administrasjon (kun lokal UI) | Azure Arc-integrasjon | Azure Monitor (erstattes av lokalt Kubernetes dashboard) | Azure Container Registry (erstattes av Edge Container Registry) | Arc-enabled VM-styring (erstattes av lokalt PowerShell/UI), mens GPU-arbeidsbelastninger har full støtte.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#17\",\n \"claim\": \"Disconnected containers prises via commitment tier med årlig forpliktelse og krever Enterprise Agreement, samt godkjent søknad for betaling per container-tjeneste.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#18\",\n \"claim\": \"Commitment tier for disconnected containers har kalenderårs-binding med 12 måneders minimum og automatisk fornyelse.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n },\n {\n \"id\": \"ms-ai-infrastructure/hybrid-edge/disconnected-ai-scenarios.md#19\",\n \"claim\": \"Disconnected containers er ikke tilgjengelige i sovereign clouds — opprettelse skjer kun i public cloud.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/azure-sovereign-clouds/private/foundry-local/disconnected-operations/concept-overview\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-infrastructure/references/hybrid-edge/disconnected-ai-scenarios.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/adversarial-input-robustness-testing.md",
|
||
"claim_count": 20,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/adversarial-input-robustness-testing.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/adversarial-input-robustness-testing.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#1\",\n \"claim\": \"Microsofts Adversarial Machine Learning Threat Taxonomy omfatter perturbasjonsbaserte angrep: Targeted misclassification | Source/Target misclassification | Random misclassification | Confidence reduction.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#2\",\n \"claim\": \"Innholdsbaserte angrep i taksonomien består av: Prompt injection | Jailbreaking | Indirect prompt injection (XPIA).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#3\",\n \"claim\": \"Agentspesifikke angrepskategorier består av: Prohibited actions | Sensitive data leakage | Task adherence violations.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#4\",\n \"claim\": \"Microsoft Foundry tilbyr AI Red Teaming Agent som automatiserer adversarial testing av modeller og agenter.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#5\",\n \"claim\": \"AI Red Teaming Agent integrerer PyRIT (Python Risk Identification Tool) og Azure AI Risk and Safety Evaluations.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#6\",\n \"claim\": \"Risikokategoriene AI Red Teaming Agent støtter er: Hateful and Unfair Content | Sexual Content | Violent Content | Self-Harm-Related Content | Protected Materials (copyright) | Code Vulnerability | Ungrounded Attributes | Prohibited Actions (kun agenter) | Sensitive Data Leakage (kun agenter) | Task Adherence (kun agenter).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#7\",\n \"claim\": \"Testfasene for AI red teaming er: Design (velg sikreste foundation model) | Development (test modelloppgraderinger og fine-tuning) | Pre-deployment (valider før produksjonsutrulling) | Post-deployment (kontinuerlig testing på syntetiske adversarial data).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#8\",\n \"claim\": \"PyRIT tilbyr 20+ attack strategies for generering av testtilfeller.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#9\",\n \"claim\": \"Encoding-baserte attack strategies i PyRIT: Base64 | Binary | ASCII Art | Morse | ROT13 | Atbash | Caesar cipher | URL encoding | Unicode substitution | Unicode confusables.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#10\",\n \"claim\": \"Obfuskeringsbaserte attack strategies: Leetspeak | Diacritic marks | Character spacing | CharSwap | Flip (mirroring) | AsciiSmuggler | ANSI escape sequences.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#11\",\n \"claim\": \"Jailbreak-baserte attack strategies: User Prompt Injected Attacks (UPIA) | Indirect Prompt Injection Attacks | SuffixAppend (adversarial suffix) | Multi-turn attacks (context accumulation) | Crescendo (gradvis eskalering).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#12\",\n \"claim\": \"AdversarialSimulator og AdversarialScenario importeres fra modulen azure.ai.evaluation.simulator i Azure AI Evaluation SDK.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#13\",\n \"claim\": \"PyRIT er et open-source rammeverk fra Microsoft for AI red teaming.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#14\",\n \"claim\": \"PyRITs arkitektur består av komponentene: Orchestrator | Target | Scorers | Attack Strategy | Memory.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#15\",\n \"claim\": \"Attack-scenarioene som kan velges i en PyRIT-basert skanning inkluderer ADVERSARIAL_QA | UPIA | XPIA.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#16\",\n \"claim\": \"Mock tools i AI Red Teaming Agent støtter kun data retrieval, ikke komplekse handlinger (complex behaviors).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#17\",\n \"claim\": \"Attack Success Rate defineres per risikokategori slik: Hateful/Sexual/Violent Content = modell genererer harmful content over severity-terskel | Jailbreak = safety guardrails omgås | Prohibited Actions = agent utfører forbudt handling uten human-in-the-loop | Sensitive Data Leakage = format-nivå lekkasje detektert via pattern matching | Task Adherence = agent feiler i goal/rule/procedure compliance.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#18\",\n \"claim\": \"Alvorlighetsnivåene er: Critical (remote EOP, modellkontroll, dataeksfiltrering) | Important (targeted misclassification, model stealing, personvernlekkasjer) | Moderate (random misclassification, confidence reduction).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/failure-modes-in-machine-learning\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#19\",\n \"claim\": \"AdversarialSimulator kalles med parameterne scenario, max_conversation_turns, max_simulation_results, target og language (SupportedLanguages.English).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/ai-red-teaming-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/adversarial-input-robustness-testing.md#20\",\n \"claim\": \"Foundry Control Plane tilbyr sentralisert governance for agenters høyrisiko-/forbudte handlinger.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/adversarial-input-robustness-testing.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md",
|
||
"claim_count": 21,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#1\",\n \"claim\": \"Azure-verktøyene som knyttes til AI-spesifikke deteksjonstriggere er Azure AI Anomaly Detector | Microsoft Purview | Azure API Management analytics | Microsoft Sentinel | Azure AI Content Safety | Azure Monitor Log Analytics.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#2\",\n \"claim\": \"Microsoft Defender for AI Services / AI Security Posture Management er del av Microsoft Defender for Cloud og gir automatisk deteksjon og remediation av generative AI-risikoer på tvers av Azure-miljøet.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-cloud-introduction\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#3\",\n \"claim\": \"Microsoft Purview Insider Risk Management integrerer med andre security-suiter og identifiserer risikofylte AI-atferdsmønstre og prompt-basert data exfiltration.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#4\",\n \"claim\": \"Azure API Management støtter sikring av Model Context Protocol (MCP) server-endepunkter som del av AI communication channel security.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#5\",\n \"claim\": \"Azure AI Content Safety har en «strict mode» som kan slås på for forsterket input-filtrering ved prompt injection-hendelser.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#6\",\n \"claim\": \"Azure PowerShell-cmdleten Get-AzNetworkSecurityGroup støtter parameteren -ResourceId, og Add-AzNetworkSecurityRuleConfig støtter parameterne -Name | -Priority | -Access | -Protocol | -Direction | -SourceAddressPrefix | -DestinationAddressPrefix.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/backup/backup-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#7\",\n \"claim\": \"Azure PowerShell-cmdleten New-AzSnapshot støtter parameterne -SnapshotName og -Disk for å ta et forensisk øyeblikksbilde av en disk.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/virtual-machines/snapshot-copy-managed-disk\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#8\",\n \"claim\": \"Azure Storage sin immutabilityPolicy for blob-lagring består av feltene immutabilityPeriodSinceCreationInDays | allowProtectedAppendWrites | state (f.eks. «Locked»), med legalHold som eget objekt med tags og enabled.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/storage/blobs/immutable-storage-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#9\",\n \"claim\": \"Den Azure-native incident response-stacken består av Microsoft Defender for AI Services (threat detection) | Microsoft Sentinel (SIEM/SOAR) | Microsoft Defender XDR (XDR) | Azure Monitor + Log Analytics (forensics) | Azure Blob Immutable Storage (bevisbevaring) | Microsoft Entra ID + PIM (identitetsrespons) | Azure Firewall + NSG (nettverksisolering) | Azure ML + Purview (modellstyring).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#10\",\n \"claim\": \"Azure Monitor / Log Analytics har 30 dagers «hot» retention for KQL-basert etterforskning.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#11\",\n \"claim\": \"Azure Blob Immutable Storage støtter legal hold | tidsbaserte retention-policyer for bevisbevaring.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/storage/blobs/immutable-storage-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#12\",\n \"claim\": \"En Sentinel-playbook for AI-modellforgiftning bruker Logic Apps-konnektorene Azure Monitor | Azure ML | Microsoft Sentinel | Microsoft Teams | Azure Resource Manager | Azure DevOps.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/sentinel/tutorial-respond-threats-playbook\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#13\",\n \"claim\": \"Azure PowerShell-cmdleten Set-AzSecurityContact støtter parameterne -Name | -Email | -Phone | -AlertAdmin | -NotifyOnAlert for å konfigurere sikkerhetskontakter i Microsoft Defender for Cloud.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-cloud-introduction\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#14\",\n \"claim\": \"Microsoft Defender for Cloud tilbyr planen «Defender for Servers Plan 2», som prises per server per måned.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-cloud-introduction\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#15\",\n \"claim\": \"Azure Monitor Log Analytics har de første 5 GB per dag gratis, deretter betaling per GB.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#16\",\n \"claim\": \"Microsoft Sentinel lisensieres enten frittstående (standalone) eller via Microsoft 365 E5 Security.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/sentinel/tutorial-respond-threats-playbook\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#17\",\n \"claim\": \"Microsoft Defender XDR krever Microsoft 365 E5 Security eller E5, og inkluderer Defender for Endpoint | Defender for Identity | Defender for Microsoft 365.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/backup/backup-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#18\",\n \"claim\": \"Microsoft Entra ID P2 kreves for PIM og risikobaserte Conditional Access-policyer.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/backup/backup-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#19\",\n \"claim\": \"Azure Automation er gratis for de første 500 minuttene per måned, deretter betaling per minutt.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/backup/backup-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#20\",\n \"claim\": \"Microsoft Sentinel har commitment tiers på 100 | 200 | 300 GB per dag, med 15-50 % rabatt sammenlignet med pay-as-you-go.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/sentinel/tutorial-respond-threats-playbook\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md#21\",\n \"claim\": \"Microsoft CAF-dokumentet «Secure AI» dekker AI asset inventory (Azure Resource Graph) | AI communication channel security (Managed Identities, Virtual Networks, APIM for MCP) | data boundary definition (Microsoft Purview) | DLP (Purview DLP + innholdsfiltrering) | AI-spesifikk incident response (Defender for Cloud AI posture management).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/ai-incident-response-procedures.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-red-team-operations-practical.md",
|
||
"claim_count": 21,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/ai-red-team-operations-practical.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/ai-red-team-operations-practical.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#1\",\n \"claim\": \"Azure AI Red Teaming Agent er merket preview, er integrert i Microsoft Foundry og er basert på PyRIT.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#2\",\n \"claim\": \"AI Red Teaming Agent støtter disse målene (targets): Azure OpenAI-modeller via AzureOpenAIModelConfiguration | Foundry-hostede agenter (prompt agents, container agents) | Simple callbacks (custom Python-funksjoner) | PyRIT PromptChatTarget.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#3\",\n \"claim\": \"AI Red Teaming Agent støtter disse risikokategoriene: Hateful and Unfair Content | Sexual Content | Violent Content | Self-Harm Content | Protected Materials (lyrics, oppskrifter) | Code Vulnerability (SQL injection, tar-slip) | Ungrounded Attributes (demographics, emotional state).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#4\",\n \"claim\": \"De agent-spesifikke risikokategoriene er kun tilgjengelige i cloud og består av: Prohibited Actions | Sensitive Data Leakage | Task Adherence.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#5\",\n \"claim\": \"Støttede encoding-angrepsstrategier i AI Red Teaming Agent: Base64 | ROT13 | Caesar | Binary | Morse | URL | Atbash.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#6\",\n \"claim\": \"Støttede obfuscation-angrepsstrategier i AI Red Teaming Agent: Leetspeak | AsciiArt | Diacritic | CharacterSpace | UnicodeConfusable.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#7\",\n \"claim\": \"Støttede injection-angrepsstrategier i AI Red Teaming Agent: Jailbreak (UPIA) | Indirect Jailbreak (XPIA) | SuffixAppend.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#8\",\n \"claim\": \"Støttede multi-turn-angrepsstrategier i AI Red Teaming Agent: Crescendo (gradvis eskalering) | Multi turn (context accumulation).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#9\",\n \"claim\": \"AI Red Teaming Agent installeres med pakken azure-ai-evaluation med ekstra-avhengigheten [redteam] (uv pip install \\\"azure-ai-evaluation[redteam]\\\").\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#10\",\n \"claim\": \"For lokal scan importeres RedTeam og RiskCategory fra modulen azure.ai.evaluation.red_team.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#11\",\n \"claim\": \"For cloud-scan ligger RedTeam, AzureOpenAIModelConfiguration, AttackStrategy og RiskCategory i azure.ai.projects.models, og scannen opprettes med project_client.red_teams.create.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#12\",\n \"claim\": \"AI Red Teaming Agent er kun tilgjengelig i regionene East US2 | Sweden Central | France Central | Switzerland West.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#13\",\n \"claim\": \"PyRIT (Python Risk Identification Tool) er et open-source rammeverk fra Microsoft for adversarial testing.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#14\",\n \"claim\": \"PyRITs nøkkelkonsepter er: Prompt Targets (systemet du tester) | Attack Strategies (conversion methods) | Scorers (evaluering av om angrepet lyktes).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#15\",\n \"claim\": \"Begrensninger i AI Red Teaming Agent: mock tools henter kun syntetiske data (ikke real-world distributions) | ingen behavior mocking, kun data mocking | adversarial nature er kontrollert for å unngå real-world impact.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#16\",\n \"claim\": \"Transient agents brukes ved red teaming fordi chat completions ikke lagres i Foundry Agent Service.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#17\",\n \"claim\": \"ASR rapporteres per attack complexity med nivåene Easy | Moderate | Difficult.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#18\",\n \"claim\": \"Resultater vises i Foundry på Evaluation-siden under fanen AI red teaming, med Report view per risikokategori | Report view per attack complexity | Data-side med attack-response-par (full samtalehistorikk, brukt attack strategy, success/failure-status, human feedback).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#19\",\n \"claim\": \"Copilot Studio har ikke native integrasjon med AI Red Teaming Agent; man må bruke PyRIT eller custom scripting.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#20\",\n \"claim\": \"For M365 Copilot red teamer Microsoft selv plattformen, mens kundene tester custom plugins og declarative agents.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/ai-red-team/training\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-red-team-operations-practical.md#21\",\n \"claim\": \"Microsoft AI Red Team Training Series består av 10 episoder fordelt på: episode 1-2 Fundamentals | episode 3-6 Attack Techniques | episode 7 Defense | episode 8-10 Automation.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/training\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/ai-red-team-operations-practical.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-security-scoring-framework.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/ai-security-scoring-framework.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/ai-security-scoring-framework.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#1\",\n \"claim\": \"Microsofts tilnærming til AI-sikkerhetsscoring er basert på AI Risk Assessment Framework versjon 4.1.4.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#2\",\n \"claim\": \"Severity-nivåene i rammeverket er: Critical | High | Medium | Low | Informational.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#3\",\n \"claim\": \"Likelihood har to hovedkomponenter: Attack Surface Availability | Attack Technique Availability.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#4\",\n \"claim\": \"Microsoft bruker en 5x3 severity matrix for ML-spesifikke angrepstyper: Extraction | Evasion | Inference | Inversion | Poisoning.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#5\",\n \"claim\": \"Responsible AI-prinsippene brukt som scoring-dimensjoner er: Privacy & Security | Reliability & Safety | Fairness | Inclusiveness | Transparency | Accountability.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/govern\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#6\",\n \"claim\": \"Risikokategoriene som måles med Risk Category ASR er: hate | violence | self-harm | sexual.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/how-to/evaluate-generative-ai-app\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#7\",\n \"claim\": \"Verktøy for adversarial testing er: PyRIT (Python Risk Identification Tool for Generative AI) | Microsoft Foundry safety evaluations | egne jailbreak-testsuiter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#8\",\n \"claim\": \"Secure Score for AI-ressurser i Microsoft Defender for Cloud dekker: Azure OpenAI-endepunkter (network isolation checks) | Azure ML-workspaces (validering av RBAC-konfigurasjon) | lagringskontoer med treningsdata (verifisering av kryptering i hvile).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/resource-graph-samples\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#9\",\n \"claim\": \"Innebygde Azure Policy-policyer for AI-styring er: «Azure AI services should use private endpoints» | «Azure Machine Learning workspaces should disable public network access» | «Diagnostic logs in Azure AI services should be enabled».\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#10\",\n \"claim\": \"Azure Monitor er inkludert i Azure-abonnementet og faktureres per GB dataingestion.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#11\",\n \"claim\": \"Microsoft Defender for Cloud har en Standard tier som lisensieres per beskyttet ressurs per måned.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#12\",\n \"claim\": \"Microsoft Purview Compliance lisensieres per datakilde.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#13\",\n \"claim\": \"Safety evaluations for Azure OpenAI er inkludert i Azure OpenAI og faktureres token-basert.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/how-to/evaluate-generative-ai-app\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#14\",\n \"claim\": \"PyRIT er open source og gratis (kun compute-kostnader).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/ai-red-team/ai-risk-assessment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#15\",\n \"claim\": \"Power BI Pro lisensieres per bruker per måned.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#16\",\n \"claim\": \"Microsoft Defender for Cloud har en gratis tier med begrenset dekning.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#17\",\n \"claim\": \"Microsoft har omdøpt «Cognitive Services» til «Foundry Tools» i sikkerhetsbaselines, og URL-en for cognitive-services-security-baseline er fortsatt aktiv men omdirigerer til «Azure security baseline for Foundry Tools».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#18\",\n \"claim\": \"MCSB v2 (Microsoft Cloud Security Benchmark versjon 2) har en egen kontrollside for Artificial Intelligence Security med sikkerhetskontroller for AI-workloads (innholdsfiltrering, meta-prompts, modellgodkjenning).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-security-scoring-framework.md#19\",\n \"claim\": \"Evaluering av generative AI-modeller i Microsoft Foundry dekker AI-kvalitetsmetrikker (NLP-baserte | AI-assisterte) og risiko- og sikkerhetsmetrikker (content harm | ASR).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/how-to/evaluate-generative-ai-app\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/ai-security-scoring-framework.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#1\",\n \"claim\": \"Microsoft har utvidet STRIDE-rammeverket — som består av Spoofing | Tampering | Repudiation | Information Disclosure | Denial of Service | Elevation of Privilege — til å dekke AI-spesifikke trusler som datapoisoning, adversarial attacks, model inversion og prompt injection.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#2\",\n \"claim\": \"STRIDE-kategorien Spoofing dekker AI-truslene Neural Net Reprogramming | Malicious ML Providers, med alvorlighetsgrad Important-Critical.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#3\",\n \"claim\": \"STRIDE-kategorien Tampering dekker AI-truslene Data Poisoning (målrettet/vilkårlig) | Backdoored Models, med alvorlighetsgrad Critical.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#4\",\n \"claim\": \"STRIDE-kategorien Repudiation dekker AI-truslene manipulasjon av modelloutput | tap av lineage for treningsdata, med alvorlighetsgrad Moderate.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#5\",\n \"claim\": \"STRIDE-kategorien Information Disclosure dekker AI-truslene Model Inversion | Membership Inference | Model Stealing, med alvorlighetsgrad Important-Critical.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#6\",\n \"claim\": \"STRIDE-kategorien Denial of Service dekker AI-truslene Confidence Reduction | Random Misclassification, med alvorlighetsgrad Important.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#7\",\n \"claim\": \"STRIDE-kategorien Elevation of Privilege dekker AI-truslene Adversarial Perturbation | Excessive Agency | Physical Domain Attacks, med alvorlighetsgrad Critical.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#8\",\n \"claim\": \"Azure AI Content Safety tilbyr Prompt Shields for jailbreak-deteksjon | innholdsfiltre for usikker output-håndtering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/ai/playbook/technology-guidance/generative-ai/mlops-in-openai/security/security-plan-llm-application\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#9\",\n \"claim\": \"Azure OpenAI Service leverer personvernforpliktelser der kundedata ikke brukes til trening | innholdsfiltrering | abuse monitoring.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/ai/playbook/technology-guidance/generative-ai/mlops-in-openai/security/security-plan-llm-application\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#10\",\n \"claim\": \"Microsoft Foundry tilbyr sikre MLOps-pipelines | managed identities | private endpoints | modellregister med versjonering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/ai/playbook/technology-guidance/generative-ai/mlops-in-openai/security/security-plan-llm-application\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#11\",\n \"claim\": \"Microsoft Defender for Cloud AI Security Posture Management omfatter automatisk oppdagelse av AI-arbeidsbelastninger på tvers av Azure-abonnementer via Azure Resource Graph | automatisert deteksjon og utbedring av risiko i generativ AI | sikkerhetsanbefalinger for AI-modeller, datalagre og nettverksisolasjon | integrasjon med Purview for dataklassifisering, DLP og Insider Risk Management for prompt-basert dataeksfiltrering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#12\",\n \"claim\": \"Microsoft Threat Modeling Tool har AI-spesifikke maler: ML Training Pipeline | Model API | LLM Agent.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/security/develop/threat-modeling-tool\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#13\",\n \"claim\": \"Microsoft Threat Modeling Tool er gratis nedlasting (ingen lisenskostnad) og gir STRIDE-automatisering, AI-spesifikke maler og trusselrapporter.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/security/develop/threat-modeling-tool\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#14\",\n \"claim\": \"AI-kapabilitetene i Microsoft Defender for Cloud (oppdagelse av AI-arbeidsbelastninger, posture management, trusseldeteksjon) krever standard tier, lisensiert per server.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#15\",\n \"claim\": \"«Exploit software dependencies» er trussel nummer 11 i Microsofts AI/ML-trusselliste.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#16\",\n \"claim\": \"Microsoft Learn-dokumentet «Threat Modeling AI/ML Systems and Dependencies» inneholder 11 trusselkategorier med tilhørende tiltak.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#17\",\n \"claim\": \"«Secure AI» i Cloud Adoption Framework ble oppdatert 2026-04 og inkluderer nå AI asset inventory via Azure Resource Graph | sikring av AI-kommunikasjonskanaler med Managed Identities og Virtual Networks | Azure API Management for sikring av MCP-server-endepunkter | Microsoft Purview Insider Risk Management for deteksjon av prompt-basert dataeksfiltrering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/scenarios/ai/secure\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#18\",\n \"claim\": \"Microsoft Learn-dokumentet «Security Planning for LLM-based Applications» beskriver 11 LLM-spesifikke trusler kartlagt mot STRIDE, med mitigeringsmønstre for Azure OpenAI.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai/playbook/technology-guidance/generative-ai/mlops-in-openai/security/security-plan-llm-application\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md#19\",\n \"claim\": \"Microsoft Learn-dokumentet «Securing the Future of AI and ML at Microsoft» introduserer de AI-spesifikke sikkerhetspivotene Resilience | Discretion.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/securing-artificial-intelligence-machine-learning\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/ai-threat-modeling-stride.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/content-safety-filter-calibration.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/content-safety-filter-calibration.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/content-safety-filter-calibration.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#1\",\n \"claim\": \"Content Safety-filtrering i Microsoft AI-stakken har status GA (referansen oppdatert 2026-06-19).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#2\",\n \"claim\": \"Azure AI Content Safety tilbyr fire alvorlighetsgrader — safe | low | medium | high — for fire skadekategorier: hate | sexual | violence | self-harm.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#3\",\n \"claim\": \"Alvorlighetsgradene har numeriske scoreverdier: Safe = 0 | Low = 2 | Medium = 4 | High = 6.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#4\",\n \"claim\": \"Standard (default) terskel er medium: Low filtreres IKKE som default, mens Medium filtreres og High filtreres alltid som default.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#5\",\n \"claim\": \"Konfigurerbare parametere for innholdsfilter er: Severity threshold (per kategori hate/sexual/violence/self-harm, separat for prompts og completions) | Annotate-only mode | Blocklists (custom termlister for text og image) | Custom categories (basert på RAI-policy, text og image) | No filters.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#6\",\n \"claim\": \"Annotate-only mode og \\\"No filters\\\" krever godkjenning via Limited Access; \\\"No filters\\\" er kun tilgjengelig for managed customers, mens justering av severity threshold til low/medium/high ikke krever godkjenning.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#7\",\n \"claim\": \"Azure AI Content Safety støtter 100+ språk, inkludert norsk bokmål og nynorsk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#8\",\n \"claim\": \"Content Safety er default aktivert for alle Azure OpenAI-deployments, med unntak av Whisper.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/content-filters\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#9\",\n \"claim\": \"Innholdsfilter kan overstyres på request-nivå ved å sende headeren x-policy-id per API-kall mot Azure OpenAI.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/content-filters\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#10\",\n \"claim\": \"Azure OpenAI chat completions-endepunktet kalles med api-version=2024-10-01 på stien /openai/deployments/<model>/chat/completions.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/content-filters\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#11\",\n \"claim\": \"I Microsoft Foundry konfigureres innholdsfilter under Guardrails + controls → Content filters, med separate Input filters (user prompts) og Output filters (completions), severity threshold per kategori (low/medium/high), samt Prompt Shields og Protected Material detection som kan aktiveres.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/content-filters\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#12\",\n \"claim\": \"Innholdsfiltrering støtter streaming mode som filtrerer i near-real-time mens output genereres, og dermed reduserer latens.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/content-filters\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#13\",\n \"claim\": \"Python SDK-en azure.ai.contentsafety tilbyr ContentSafetyClient (med AzureKeyCredential), AnalyzeTextOptions, TextCategory og metoden analyze_text som returnerer categories_analysis med severity per kategori.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/python/api/overview/azure/ai-contentsafety-readme\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#14\",\n \"claim\": \"Det finnes en Content Safety-connector tilgjengelig i Power Automate.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#15\",\n \"claim\": \"Azure AI Content Safety prosesserer data i regionen ressursen opprettes i, for eksempel Norway East eller West Europe.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#16\",\n \"claim\": \"Både Text API og Image API i Azure AI Content Safety har en Free tier med 5000 transaksjoner per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#17\",\n \"claim\": \"Content Safety krever kun et Azure-abonnement (alle tiers, inkludert Free Trial) og ingen spesifikk Azure OpenAI-lisens — tjenesten fungerer standalone.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/content-safety-filter-calibration.md#18\",\n \"claim\": \"Content Safety er inkludert i Azure OpenAI-deployments (default aktivert) og Microsoft Foundry-prosjekter, men IKKE i Microsoft 365 Copilot (bruker annen filtreringsstack) eller Copilot Studio (krever separat Content Safety-ressurs for custom filtering).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/foundry-models/concepts/content-filter\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/content-safety-filter-calibration.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md",
|
||
"claim_count": 27,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#1\",\n \"claim\": \"Purview DLP-handlingen Processing prompts, som forhindrer Copilot i å returnere svar når prompten inneholder sensitive data, er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#2\",\n \"claim\": \"Purview DLP-handlingen Performing Web Searches er GA og blokkerer ekstern websøk som grounding-kilde når prompten inneholder sensitive information types, mens Copilot fortsatt svarer fra interne M365-kilder brukeren har tilgang til.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#3\",\n \"claim\": \"Betingelsen Email is received from > External users er i preview og ekskluderer ekstern e-post fra grounding, summarisering og citation i Microsoft 365 Copilot; kun avsender-metadata evalueres, ikke e-postkroppen.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#4\",\n \"claim\": \"Støttede lokasjoner for Purview DLP mot Copilot-prompts er: Microsoft 365 Copilot og Copilot Chat inkludert pre-built agents | Copilot in Word | Copilot in Excel | Copilot in PowerPoint.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#5\",\n \"claim\": \"Policy-lokasjonen for Microsoft 365 Copilot er kun tilgjengelig i Custom-policymalen, og alle andre lokasjoner i policyen deaktiveres når denne lokasjonen velges.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#6\",\n \"claim\": \"DLP-policyoppdateringer for Microsoft 365 Copilot tar opptil 4 timer før de trer i kraft.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#7\",\n \"claim\": \"Støttede filtyper for sensitivity label-basert blokkering i Copilot er: Word (.docx/.docm) | Excel (.xlsx/.xlsm/.xlsb) | PowerPoint (.pptx/.ppsx) | PDF-filer ved aktivert PDF-støtte.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#8\",\n \"claim\": \"Kun e-poster sendt på eller etter 1. januar 2025 dekkes av sensitivity label-basert blokkering for Copilot.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#9\",\n \"claim\": \"Sensitivity label-basert blokkering for Copilot omfatter kun filer lagret i SharePoint Online og OneDrive for Business.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#10\",\n \"claim\": \"Etiketter med bruker-definerte tillatelser støttes nå for search, DLP og eDiscovery, men kun for nyopplastede eller redigerte filer.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/purview/dlp-microsoft365-copilot-location-learn-about\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#11\",\n \"claim\": \"Konfigurasjon av restrictOutboundNetworkAccess på Microsoft.CognitiveServices/accounts gjøres mot api-version=2024-10-01.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-data-loss-prevention\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#12\",\n \"claim\": \"allowedFqdnList for Azure AI Services DLP kan inneholde maksimum 1000 URL-er.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-data-loss-prevention\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#13\",\n \"claim\": \"Det tar opptil 15 minutter før en oppdatert allowedFqdnList trer i kraft.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-data-loss-prevention\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#14\",\n \"claim\": \"Tjenester som støtter outbound URL-restriksjon (Azure AI Services DLP) er: Azure OpenAI | Microsoft Foundry (Foundry-baserte prosjekter) | Azure Vision | Content Moderator | Custom Vision | Face API | Document Intelligence | Speech Services | QnA Maker.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-data-loss-prevention\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#15\",\n \"claim\": \"Alle Azure OpenAI-data krypteres i ro med FIPS 140-2-kompatibel 256-bit AES-kryptering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#16\",\n \"claim\": \"Microsoft Purview Endpoint DLP støtter handlingene: block paste | block upload | warn with override.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#17\",\n \"claim\": \"Endpoint DLP mot tredjeparts generative AI-nettsteder støttes kun på Windows-maskiner med Endpoint DLP-agent installert.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#18\",\n \"claim\": \"Purview Insider Risk Management har policy-malene: DSPM for AI - Detect risky AI usage | DSPM for AI - Unethical behavior in AI apps | DSPM for AI - Protect sensitive data from Copilot processing.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#19\",\n \"claim\": \"Azure OpenAI støtter ikke persistent prompt caching på tvers av brukere; hvert API-kall er stateless med mindre samtalehistorikk sendes eksplisitt i requesten.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#20\",\n \"claim\": \"Azure API Management kan nå også sikre Model Context Protocol (MCP) server-endepunkter.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#21\",\n \"claim\": \"Azure API Management tilbyr SKU-en Developer, angitt med --sku-name Developer i az apim create.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#22\",\n \"claim\": \"Retention for Microsoft Purview Audit er konfigurerbar fra 90 dager til 10 år.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#23\",\n \"claim\": \"DSPM for AI (classic) er generelt tilgjengelig (GA).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#24\",\n \"claim\": \"DSPM er i preview som ny versjon med utvidet AI activities-fane.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#25\",\n \"claim\": \"Sentrale DLP-cmdlets i ExchangePowerShell er: New-DlpCompliancePolicy | New-DlpComplianceRule | Get-DlpCompliancePolicy | Set-DlpPolicy | Get-Label.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#26\",\n \"claim\": \"Microsoft Defender for AI Services detekterer: prompt injection | model manipulation | jailbreak-forsøk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md#27\",\n \"claim\": \"Defender-planen for AI aktiveres som pricing-navn AI med tier Standard via az security pricing create.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/data-leakage-prevention-ai.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/defender-threat-protection-ai-services.md",
|
||
"claim_count": 16,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/defender-threat-protection-ai-services.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/defender-threat-protection-ai-services.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#1\",\n \"claim\": \"Gjeldende offisielle navn er funksjonen «AI threat protection» under planen «Defender for AI Services», og dette erstatter den tidligere betegnelsen «Threat protection for AI workloads».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#2\",\n \"claim\": \"AI threat protection for AI applications (Azure OpenAI + Azure AI Model Inference) er GA og produksjonsklart for kommersiell Azure.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#3\",\n \"claim\": \"AI threat protection for AI agents (Foundry Agent Service) er i Preview fra 2026-02-02, med varsler under utrulling.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#4\",\n \"claim\": \"AI models (skanning av opplastede modeller) er i Preview, f.eks. varselet «Malicious content in uploaded AI model».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#5\",\n \"claim\": \"Ressurser som dekkes av AI threat protection: Azure OpenAI Service (alle støttede modeller) | Azure AI Model Inference (Foundry-modeller).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#6\",\n \"claim\": \"Kun tekst-tokens skannes av AI threat protection; bilde-tokens og lyd-tokens skannes ikke.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#7\",\n \"claim\": \"AI threat protection er tilgjengelig i kommersiell Azure, men ikke i Azure Government og ikke i Azure operated by 21Vianet.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#8\",\n \"claim\": \"Varslene i AI threat protection er gruppert i tre kategorier: AI applications | AI agents | AI models, og hvert varsel er kartlagt til MITRE ATT&CK-taktikker.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/alerts-ai-workloads\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#9\",\n \"claim\": \"Varsel-IDer for AI applications omfatter: AI.Azure_Jailbreak.ContentFiltering.BlockedAttempt | AI.Azure_Jailbreak.ContentFiltering.DetectedAttempt | AI.Azure_ASCIISmuggling | AI.Azure_CredentialTheftAttempt | AI.Azure_MaliciousUrl.ModelResponse | AI.Azure_AccessFromAnonymizedIP | AI.Azure_AccessFromSuspiciousIP | AI.Azure_DOWDuplicateRequests | AI.Azure_DOWVolumeAnomaly | AI.Azure_AnomalousToolInvocation | AI.Azure_LLMReconnaissance.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/alerts-ai-workloads\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#10\",\n \"claim\": \"Varselet «LLM Reconnaissance-forsøk» (AI.Azure_LLMReconnaissance) er i Preview, med MITRE-taktikk Reconnaissance og alvorlighet Low.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/alerts-ai-workloads\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#11\",\n \"claim\": \"For AI agents (Foundry Agent Service) finnes tilsvarende varsler, alle i Preview, samt «instruction prompt leak» og «agent reconnaissance attempt».\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/alerts-ai-workloads\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#12\",\n \"claim\": \"Defender for Cloud kan vise «user prompt evidence» — utdrag av bruker-prompt og modellrespons i Defender-portalen; sensitive data redigeres automatisk, og evidens kan skrus av per abonnement.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-onboarding\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#13\",\n \"claim\": \"Aktivering skjer i Defender for Cloud → Environment settings → abonnement → planen «AI services» → toggle On.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-onboarding\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#14\",\n \"claim\": \"Aktivering av planen krever rollen Owner eller Contributor på abonnementsnivå.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-onboarding\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#15\",\n \"claim\": \"Prøveperioden er 30 dager gratis og begrenset til 75 milliarder skannede tokens; fakturering starter hvis grensen nås innenfor perioden.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/ai-onboarding\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/defender-threat-protection-ai-services.md#16\",\n \"claim\": \"Release notes for februar 2026 beskriver at agent-trusselbeskyttelsen adresserer «high-impact, actionable threats aligned with OWASP guidance for LLM and agentic AI systems».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/defender-for-cloud/release-notes\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/defender-threat-protection-ai-services.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/jailbreak-prevention-production.md",
|
||
"claim_count": 16,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/jailbreak-prevention-production.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/jailbreak-prevention-production.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#1\",\n \"claim\": \"Referansens emne (Prompt Shields / jailbreak-deteksjon i Azure AI Content Safety) har status GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#2\",\n \"claim\": \"Microsoft tilbyr Prompt Shields som del av Azure OpenAI content filtering-systemet og Azure AI Content Safety-tjenesten, som ett unified API som detekterer og blokkerer adversarial user input attacks før innhold genereres.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-filter-prompt-shields\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#3\",\n \"claim\": \"User Prompt Attacks (direkte jailbreak-angrep) deles i fire hovedkategorier: Attempt to change system rules | Embedding a conversation mockup | Role-Play | Encoding Attacks.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#4\",\n \"claim\": \"Document Attacks (indirekte angrep, også kalt Indirect Prompt Attacks / Cross-Domain Prompt Injection Attacks) har ni hovedkategorier, angitt i filen som: Manipulated Content | Infrastructure Access | Information Gathering | Availability | Fraud | Malware | Attempt to change system rules | Embedding a conversation mockup | Role-Play | Encoding Attacks.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#5\",\n \"claim\": \"Prompt Shields består av to shields for ulike angrepstyper: Prompt Shields for User Prompts | Prompt Shields for Documents.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-filter-prompt-shields\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#6\",\n \"claim\": \"Prompt Shields for User Prompts het tidligere «Jailbreak risk detection».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#7\",\n \"claim\": \"Azure AI Content Safety REST-endepunktet for Prompt Shields kalles som POST <endpoint>/contentsafety/text:shieldPrompt med api-version=2024-09-01, og tar feltene userPrompt og documents.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#8\",\n \"claim\": \"Spotlighting er i preview, er en sub-feature av Prompt Shields, transformerer dokumentinnhold med Base-64-encoding, og er slått av som standard (turned off by default).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-filter-prompt-shields\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#9\",\n \"claim\": \"Default safety policies for Azure OpenAI tekstmodeller: Hate and Fairness | Violence | Sexual | Self-Harm filtreres på prompts og completions med severity threshold Medium; User prompt injection attack (Jailbreak) på prompts har terskel N/A og handling Detect and block; Protected Material – Text og Protected Material – Code på completions har terskel N/A og handling Annotate/Filter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/default-safety-policies\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#10\",\n \"claim\": \"Asynchronous filtering (asynkron kjøring av innholdsfiltre for bedre latency i streaming-scenarioer) er tilgjengelig for alle Azure OpenAI-kunder.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/default-safety-policies\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#11\",\n \"claim\": \"Microsofts profanity blocklist er tilgjengelig out-of-the-box og dekker engelsk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#12\",\n \"claim\": \"Azure AI Content Safety custom categories kalles via POST <endpoint>/contentsafety/text:analyzeCustomCategory med api-version=2024-09-01 og feltene text, categoryName og version.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#13\",\n \"claim\": \"API Managements llm-content-safety-policy støtter attributtene shield-prompt | enforce-on-completions | window-size | window-overlap-size samt elementene <categories> og <blocklists>; enforce-on-completions=\\\"true\\\" i inbound validerer også LLM-responser (chat completions) og ignoreres i outbound.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/llm-content-safety-policy\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#14\",\n \"claim\": \"window-size i llm-content-safety-policy har default 10 000 tegn og er kun konfigurerbar for responser; for requests brukes alltid default-vinduet.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/llm-content-safety-policy\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#15\",\n \"claim\": \"llm-content-safety-policy støtter kategoriene Hate | SelfHarm | Sexual | Violence, og kan brukes i både inbound og outbound og defineres flere ganger i samme policy definition.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/llm-content-safety-policy\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/jailbreak-prevention-production.md#16\",\n \"claim\": \"llm-content-safety-policy kan håndheve content safety-sjekker på requests og responses også for MCP-verktøy og A2A Agent-APIer som administreres i API Management.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/llm-content-safety-policy\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/jailbreak-prevention-production.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/model-fingerprinting-watermarking.md",
|
||
"claim_count": 21,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/model-fingerprinting-watermarking.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/model-fingerprinting-watermarking.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#1\",\n \"claim\": \"Azure OpenAI (DALL-E 3 og GPT-image-1) legger automatisk Content Credentials (C2PA) på alle genererte bilder.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#2\",\n \"claim\": \"Azure Text to Speech Avatar merker video-output med content credentials, men kun for formatet .mp4.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech-avatar/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#3\",\n \"claim\": \"Microsoft 365 Copilot kan merke AI-generert innhold — bilder | video | lyd — med watermarks, styrt av policy.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/watermarks\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#4\",\n \"claim\": \"C2PA-manifestet fra Microsoft inneholder feltene description (\\\"AI Generated Image\\\") | softwareAgent (\\\"Azure OpenAI DALL-E\\\" eller \\\"Azure OpenAI ImageGen\\\") | when (timestamp) | generator (\\\"Microsoft Azure Txt to Speech Avatar Service\\\").\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#5\",\n \"claim\": \"C2PA-manifestet er kryptografisk signert med et sertifikat som spores tilbake til Microsoft, noe som gjør metadataene tamper-evident.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#6\",\n \"claim\": \"Content Credentials i Azure OpenAI er alltid aktivert og krever ingen konfigurasjon; metadata legges automatisk til i alle støttede formater.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#7\",\n \"claim\": \"Administratorer aktiverer Microsoft 365-watermarking via Cloud Policy-innstillingen \\\"Include a watermark when content from Microsoft 365 is generated or altered by AI\\\".\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/watermarks\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#8\",\n \"claim\": \"Microsoft 365-watermark-policyen gjelder video og lyd, for eksempel Clipchamp-video og Copilot-audioresume, mens bilder er brukerstyrt og aktiveres i myaccount.microsoft.com/privacy.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/watermarks\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#9\",\n \"claim\": \"Microsoft 365-watermarks er ikke-muterbare (kan ikke fjernes eller modifiseres av brukeren) | persistente (vises også ved printing og screenshots) | MIP-labeling aware (støtter sensitivity-labeled PDF-er).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/watermarks\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#10\",\n \"claim\": \"Selv om watermark er deaktivert, legges C2PA-metadata (modell, app, timestamp) til i alle AI-genererte filer fra Microsoft 365.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/watermarks\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#11\",\n \"claim\": \"Relevante trusselteknikker er AML.T0050: Backdoor Model | AML.T0020: Compromise Model Supply Chain | T1195: Supply Chain Compromise.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/engineering/threat-modeling-aiml\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#12\",\n \"claim\": \"Ved modellregistrering i Azure Machine Learning Model Registry får hver modell en unik ID, et versjonsnummer og metadata.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-management-and-deployment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#13\",\n \"claim\": \"Azure ML fanger automatisk opp metadata ved modellregistrering: training script snapshot | training data lineage (hvilke datasett ble brukt) | training metrics og hyperparametere | hvem som trente modellen, når og hvor | eksperiment-ID (MLflow eller Azure ML experiment tracking).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-management-and-deployment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#14\",\n \"claim\": \"Approval workflow etter AI-1 i Microsoft Security Benchmark består av sentralisert model registry | automatisert sikkerhetsvalidering (hash-verifisering, backdoor-skanning, adversarial testing) | RBAC | multi-stage godkjenning | audit trails via Azure Monitor.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#15\",\n \"claim\": \"Azure Policy-innstillingen \\\"[Preview]: Azure Machine Learning Deployments should only use approved Registry Models\\\" er i Preview og er en BuiltIn-policy som blokkerer deployment av modeller utenfor godkjent liste (Deny effect).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-management-and-deployment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#16\",\n \"claim\": \"Unity Catalog gir runtime lineage ned til kolonnenivå på tvers av notebooks | jobs | dashboards, sporer model-to-dataset (upstream datasett når en modell trenes på en tabell) og deler lineage på tvers av workspaces i samme metastore.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/data-lineage\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#17\",\n \"claim\": \"Lineage vist i Catalog Explorer beholdes på ubestemt tid for data fra 1. september 2024 og senere, mens lineage-systemtabellene system.access.table_lineage og column_lineage har et rullende 1-års-vindu.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/data-lineage\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#18\",\n \"claim\": \"Unity Catalog bruker et three-level namespace: Catalog → Schema → Table/View/Model.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/data-lineage\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#19\",\n \"claim\": \"Azure OpenAI-bildegenerering med Content Credentials kalles med api_version \\\"2024-05-01-preview\\\".\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/content-credentials\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#20\",\n \"claim\": \"Azure Policy-parameteren allowedPublisherNames har defaultValue [\\\"Microsoft\\\", \\\"OpenAI\\\", \\\"Meta\\\"].\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-management-and-deployment\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/model-fingerprinting-watermarking.md#21\",\n \"claim\": \"Azure Policy-parameteren effect har defaultValue \\\"Deny\\\" og allowedValues Audit | Deny | Disabled.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-model-management-and-deployment\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/model-fingerprinting-watermarking.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/norwegian-content-safety.md",
|
||
"claim_count": 23,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/norwegian-content-safety.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/norwegian-content-safety.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#1\",\n \"claim\": \"Azure AI Content Safety: text moderation og Prompt Shields er GA, mens Groundedness detection og Custom Categories er i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#2\",\n \"claim\": \"Azure AI Content Safety klassifiserer innhold i fire skadekategorier: hate | sexual | violence | self-harm, og fire alvorlighetsgrader: safe | low | medium | high.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#3\",\n \"claim\": \"Azure AI Content Safety erstatter Azure Content Moderator, som ble deprecated i mars 2024.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#4\",\n \"claim\": \"Tekstmodereringsmodellen i Azure AI Content Safety er spesialtrent og testet på kinesisk | engelsk | fransk | tysk | spansk | italiensk | japansk | portugisisk; norsk (no) er støttet, men ikke blant de spesialtrente språkene.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/language-support\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#5\",\n \"claim\": \"Severity levels i Azure AI Content Safety tekstmoderering går fra 0 til 6, og skalaen er konsistent på tvers av språk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#6\",\n \"claim\": \"Prompt Shields dekker to angrepstyper: User Prompt Attacks (jailbreak-forsøk) | Document Attacks (skadelig innhold innebygd i dokumenter).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#7\",\n \"claim\": \"Prompt Shields-modellene er trent og testet på zh | en | fr | de | es | it | ja | pt; andre språk kan fungere, men med varierende kvalitet.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#8\",\n \"claim\": \"Azure AI Content Safety-modellene for protected material | groundedness detection | custom categories (standard) fungerer kun på engelsk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/language-support\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#9\",\n \"claim\": \"Groundedness correction krever Azure OpenAI GPT-4o i versjon 0513 eller 0806.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#10\",\n \"claim\": \"Groundedness detection tilbys i to moduser: reasoning mode (gir forklaringer for ungrounded segmenter) | non-reasoning mode (raskere, uten forklaringer).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#11\",\n \"claim\": \"Groundedness detection støtter domenevalg med verdiene MEDICAL | GENERIC.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#12\",\n \"claim\": \"Protected Material for Text dekker kjent opphavsrettsbeskyttet innhold som sanger | artikler | oppskrifter | nettinnhold.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/protected-material\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#13\",\n \"claim\": \"Protected Material for Code er basert på GitHub-repositories indeksert til og med april 2023.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/protected-material\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#14\",\n \"claim\": \"Custom Categories (standard) krever minst 50 treningseksempler og støtter maksimalt 3 kategorier.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/custom-categories\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#15\",\n \"claim\": \"Custom categories (standard) API støtter kun engelsk, mens custom categories (rapid) API støtter alle språk som Content Safety text moderation støtter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/custom-categories\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#16\",\n \"claim\": \"Microsofts forhåndsdefinerte profanity blocklist er kun engelskspråklig, mens custom blocklists er språkuavhengige og støtter regex.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/quickstart-blocklist\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#17\",\n \"claim\": \"Blocklist-API-et for Azure AI Content Safety kalles med api-version=2024-09-01 (PUT /contentsafety/text/blocklists/{name} og POST :addOrUpdateBlocklistItems).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/quickstart-blocklist\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#18\",\n \"claim\": \"Free tier for Azure AI Content Safety gir 5000 transaksjoner per måned for Text Moderation | Image Moderation | Prompt Shields | Protected Material | Custom Categories (rapid); Groundedness Detection er ikke tilgjengelig i free tier.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#19\",\n \"claim\": \"Azure Translator har en free tier på 2 millioner tegn per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#20\",\n \"claim\": \"Azure AI Content Safety faktureres etter volum og tilbys i tier-strukturen F0 (Free) | S0 (Standard).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/faq\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#21\",\n \"claim\": \"F0 (Free)-tier for Azure AI Content Safety har 5 RPS for Text/Image Moderation, Prompt Shields og Custom Categories (rapid), mens Groundedness ikke er tilgjengelig (N/A).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#22\",\n \"claim\": \"S0 (Standard)-tier for Azure AI Content Safety har 1000 RP10S for Text/Image Moderation, Prompt Shields og Custom Categories (rapid), og 50 RPS for Groundedness.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/norwegian-content-safety.md#23\",\n \"claim\": \"Azure AI Content Safety kan deployes i regionene Norway East eller West Europe for data residency i EU/EØS.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/norwegian-content-safety.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/output-validation-grounding-verification.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/output-validation-grounding-verification.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/output-validation-grounding-verification.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#1\",\n \"claim\": \"Groundedness-deteksjon i Azure AI Content Safety oppgis med status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#2\",\n \"claim\": \"Groundedness Detection API i Azure AI Content Safety har to deteksjonsmoduser: Non-reasoning mode (rask deteksjon, optimalisert for online-applikasjoner) | Reasoning mode (detaljerte forklaringer på ugrunnede segmenter).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#3\",\n \"claim\": \"Reasoning mode i Groundedness Detection API krever Azure OpenAI GPT-4o.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#4\",\n \"claim\": \"Groundedness Detection API støtter to domener: MEDICAL (medisinsk domene med spesialisert deteksjon) | GENERIC (generisk domene).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#5\",\n \"claim\": \"Groundedness Detection API støtter to oppgavetyper: QnA (Question & Answer) | Summarization (sammendrag).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#6\",\n \"claim\": \"Citations i RAG-systemer med Azure AI Search eller Microsoft Foundry Agents bruker formatene [message_idx:search_idx†source] (standard citation-format) | url_citation-annotasjoner (URL-baserte referanser i streaming-responser).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/search/transparency-note\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#7\",\n \"claim\": \"Microsoft Foundry Agents og Bing Grounding-tools følger en firetrinns prosess: Query formulation | Search execution | Information synthesis | Source attribution.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/bing-tools\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#8\",\n \"claim\": \"GroundednessEvaluator i Azure AI Evaluation SDK bruker en skala fra 1 til 5, og terskelverdien (threshold) settes innenfor denne skalaen (f.eks. 3.0).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/how-to/develop/evaluate-sdk\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#9\",\n \"claim\": \"Groundedness-endepunktet kalles som POST https://<resource>.cognitiveservices.azure.com/contentsafety/text:detectGroundedness med api-version=2024-09-15-preview.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/quickstart-groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#10\",\n \"claim\": \"Groundedness Detection API garanterer kvalitet kun for engelsk språk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/quickstart-groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#11\",\n \"claim\": \"Groundedness Detection API har en tekstgrense på maks 7500 tegn.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/quickstart-groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#12\",\n \"claim\": \"Groundedness-deteksjon i Azure AI Content Safety har regional begrensning — tilgjengelighet per region er dokumentert i region-availability-oversikten for Content Safety.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/overview#region-availability\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#13\",\n \"claim\": \"Azure OpenAI «On Your Data» har tre groundedness-relaterte konfigurasjonsvalg: Strictness-parameter (skala 1-5) | Limit responses to data content | Number of retrieved documents.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#14\",\n \"claim\": \"Copilot Studio har innebygd grounding via Dataverse-integrasjon (automatisk grounding mot organisasjonsdata) | SharePoint/Web search med konfigurerbare kildefiltre | Citation tracking med synlige kilder i chatbot-svar.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#15\",\n \"claim\": \"Power Platform AI Builder har ingen native groundedness-API; groundedness må integreres via custom connector til Azure AI Content Safety eller Power Automate-flow.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#16\",\n \"claim\": \"Groundedness API er inkludert i Azure AI Services commitment (Foundry-lisenser) og tilbys consumption-based (pay-as-you-go).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#17\",\n \"claim\": \"Groundedness API er ikke inkludert i Microsoft 365 Copilot-lisenser; M365 Copilot har egne groundedness-mekanismer.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#18\",\n \"claim\": \"Grounding with Bing Search har eget prisnivå og er ikke dekket av Azure Data Protection Addendum (dataflyt utenfor Azures compliance boundary).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/output-validation-grounding-verification.md#19\",\n \"claim\": \"I Groundedness API er parameteren correction omdøpt til mitigating, og responsfeltet correctedText er omdøpt til correctionText.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/quickstart-groundedness\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/output-validation-grounding-verification.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/owasp-llm-top10-azure-mitigations.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/owasp-llm-top10-azure-mitigations.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/owasp-llm-top10-azure-mitigations.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#1\",\n \"claim\": \"Azure Policy «ML Deployments should only use approved Registry Models» kan settes til Deny for å blokkere ikke-godkjente modeller i Azure ML model registry.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#2\",\n \"claim\": \"Microsoft Entra Agent ID er GA og gir hver agent egen identitet med scoped OAuth 2.0-tillatelser, der Entra blokkerer høyprivilegerte roller.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/entra/agent-id/what-is-microsoft-entra-agent-id\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#3\",\n \"claim\": \"Entra Agent Identity Blueprints støtter håndheving av Conditional Access-policyer per blueprint, og deaktivering av en blueprint deaktiverer alle tilhørende agenter umiddelbart.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/entra/agent-id/authorization-agent-id\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#4\",\n \"claim\": \"Copilot Studio-agenter styres av DLP-policyer i Power Platform Admin Center kombinert med granulære connector-scopes, slik at agenten kun gis de connector-operasjonene den faktisk bruker.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/admin-data-loss-prevention\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#5\",\n \"claim\": \"Microsoft Agent Framework støtter human-in-the-loop via innstillingen approval_mode: always_require, og Copilot Studio skriver aldri til Dataverse uten brukergodkjenning.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/guidance/autonomous-agents\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#6\",\n \"claim\": \"Document-level access control (security trimming) i Azure AI Search er GA for filtre, mens ACL fra ADLS Gen2 er i preview; permission-metadata lagres i indeksen og håndheves ved query-tid via Entra-token med headeren x-ms-query-source-authorization.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/search-document-level-access-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#7\",\n \"claim\": \"ACL-basert document-level access control fra ADLS Gen2 i Azure AI Search bruker API-versjon 2026-05-01-preview.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/search-document-level-access-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#8\",\n \"claim\": \"Managed Identity for outbound-autentisering i Azure AI Search (uten hardkodede nøkler) er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/search-security-best-practices\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#9\",\n \"claim\": \"Private Endpoints for Azure AI Search er GA og deaktiverer public network access slik at trafikken går over Microsoft backbone.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/service-create-private-endpoint\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#10\",\n \"claim\": \"Multi-tenant-isolasjonsmønstre for Azure AI Search er GA og består av index-per-tenant | service-per-tenant | hybrid, der dedikert service per tenant gir sterkest separasjon.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/search-modeling-multitenant-saas-applications\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#11\",\n \"claim\": \"Purview sensitivity labels på indeks i Azure AI Search (label-basert tilgangskontroll oppdaget under indeksering) er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/search/search-security-best-practices\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#12\",\n \"claim\": \"Groundedness Detection i Azure AI Content Safety er i public preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#13\",\n \"claim\": \"Groundedness Detection tilbyr modusene deteksjon av om svar er forankret i kildedokumenter | Reasoning-mode som gir rotårsak | Correction som korrigerer ungrounded tekst automatisk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/concepts/groundedness\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#14\",\n \"claim\": \"Citation/reference-mønsteret i Azure OpenAI-webapp er GA og bruker strukturerte citations-objekter med feltene title | filepath | url | content, som gir superscript-lenker til kilde i UI.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/use-web-app\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#15\",\n \"claim\": \"Copilot Studio «Grounding in Trusted Data» — svar forankret i datakilder brukeren har tilgang til — er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/system-service-card-copilot-studio\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#16\",\n \"claim\": \"AI Foundry Evaluation med metrikkene groundedness og completeness for end-to-end LLM-evaluering er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#17\",\n \"claim\": \"Prompt Shields, som blokkerer indirekte prompt injection, er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/owasp-llm-top10-azure-mitigations.md#18\",\n \"claim\": \"Microsoft Defender for AI Services er ikke tilgjengelig i Azure Government.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/owasp-llm-top10-azure-mitigations.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/pii-detection-norwegian-context.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/pii-detection-norwegian-context.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/pii-detection-norwegian-context.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#1\",\n \"claim\": \"Azure AI Language PII-deteksjon støtter norsk språk via `language: \\\"no\\\"` og kan detektere både generelle PII-kategorier (navn, e-post, telefon) og nordiske ID-numre (NOIdentityNumber).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#2\",\n \"claim\": \"Azure AI Language grupperer PII i tre feature-typer: Text PII | Conversation PII | Document-based PII (tidligere «Native Document PII»).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#3\",\n \"claim\": \"Document-based PII (tidl. Native Document PII) støtter filformatene `.pdf` | `.docx` | `.txt`, kjører asynkront og lagringsbasert, og bevarer dokumentstruktur samt JSON-metadata.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#4\",\n \"claim\": \"Azure AI Language PII gjenkjenner entitetskategoriene NOIdentityNumber | Person | Email | PhoneNumber | Address | Organization | EUPassportNumber | EUDriversLicenseNumber | InternationalBankingAccountNumber.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/concepts/entity-categories\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#5\",\n \"claim\": \"Azure detekterer norske fødselsnummer under den dedikerte kategorien `NOIdentityNumber` («Norway Identity Number»), og `language: \\\"no\\\"` må spesifiseres for optimal deteksjon.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/concepts/entity-categories\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#6\",\n \"claim\": \"GA-API-versjonen for Text PII er `2026-05-01`, og den anbefales for produksjon.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#7\",\n \"claim\": \"Nyeste preview-API-versjon for Text PII er `2026-05-15-preview`.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#8\",\n \"claim\": \"Preview-funksjonene redaction-policies | `confidenceScoreThreshold` | `disableEntityValidation` | `entitySynonyms` | `valueExclusionPolicy` ble først introdusert i API-versjon `2025-11-15-preview` og er dokumentert under gjeldende preview-versjon.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#9\",\n \"claim\": \"Azure AI Language tilbyr fire redaction policies (`redactionPolicies`): CharacterMask | EntityMask | SyntheticReplacement | NoMask.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#10\",\n \"claim\": \"CharacterMask er default redaction policy og støtter valgfri `redactionCharacter` (f.eks. `-`); NoMask returnerer ingen `redactedText` i responsen.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#11\",\n \"claim\": \"Text PII støtter per-entitet policy-overrides i samme request, med én `defaultRedactionPolicy` og entitetsspesifikke overrides.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#12\",\n \"claim\": \"`disableEntityValidation` lar deg deaktivere streng entitetsvalidering og har default-verdi `false`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#13\",\n \"claim\": \"Preview-API-et støtter `confidenceScoreThreshold` med global `default` og overrides per entitet og per språk (felt: `value`, `entity`, `language`).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#14\",\n \"claim\": \"`modelVersion: \\\"latest\\\"` kan brukes for nyeste modell; GA-API `2026-05-01` velges for produksjon og `2026-05-15-preview` for nye preview-funksjoner.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/how-to/redact-text-pii\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#15\",\n \"claim\": \"Microsoft Foundry (new) med Foundry-prosjekter og Foundry (classic) er begge tilgjengelige via `https://ai.azure.com/`, og Language-ressurs opprettes via «Azure Language in Foundry Tools».\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#16\",\n \"claim\": \"Azure AI Language støtter batch processing med opptil 5000 dokumenter per request.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#17\",\n \"claim\": \"Free-tieren (F0) for Azure AI Language gir 5000 text records per måned gratis og inkluderer PII detection, NER og sentiment.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#18\",\n \"claim\": \"Standard-tieren (S) for Azure AI Language inkluderer alle features og har SLA 99,9 %.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/language-service/personally-identifiable-information/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/pii-detection-norwegian-context.md#19\",\n \"claim\": \"Ett text record i Azure AI Language tilsvarer opptil 5120 tegn.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/language-service/personally-identifiable-information/overview\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/pii-detection-norwegian-context.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/prompt-injection-defense-patterns.md",
|
||
"claim_count": 14,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/prompt-injection-defense-patterns.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/prompt-injection-defense-patterns.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#1\",\n \"claim\": \"Prompt injection-angrep deles i tre hovedtyper: Direkte (jailbreaking) | Indirekte (ondsinnet innhold skjult i eksterne dokumenter/data) | Encoding-basert (koding, transformasjoner eller språkvarianter for å omgå filtre).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#2\",\n \"claim\": \"Direct Prompt Injection (jailbreaking) har undertypene: Attempt to change system rules | Embedding conversation mockup | Role-play attacks | Encoding attacks.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#3\",\n \"claim\": \"Indirect Prompt Injection (cross-domain-angrep) har undertypene: Manipulated content | Infrastructure access | Information gathering | Availability attacks | Fraud | Malware.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#4\",\n \"claim\": \"Prompt Shields i Azure AI Content Safety har kapabilitetene: User Prompt Attack Detection | Document Attack Detection | Real-time analysis (blokkerer angrep før de når modellen).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#5\",\n \"claim\": \"Prompt Shields-endepunktet kalles som POST <endpoint>/contentsafety/text:shieldPrompt med api-version=2024-09-01, og tar feltene userPrompt og documents.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#6\",\n \"claim\": \"Azure AI Content Safety har kategoriene: Hate and Fairness | Violence | Sexual content | Self-Harm | Protected Material (Text and Code) | Groundedness detection (for RAG-scenarioer), med severity threshold Medium for de fire første.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/foundry-models/concepts/content-filter\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#7\",\n \"claim\": \"Azure-rollen «Azure AI Services OpenAI User» tildeles en Managed Identity (SystemAssigned) med scope mot en Microsoft.CognitiveServices/accounts-ressurs for minste-privilegium-tilgang.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#8\",\n \"claim\": \"Microsoft Defender for Cloud — threat protection for AI services er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#9\",\n \"claim\": \"Defender for Cloud threat protection for AI services dekker sanntidsdeteksjon av: data leakage | data poisoning | jailbreak | credential theft, med Defender XDR-integrasjon for sentralisert hendelseskorrelering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/defender-for-cloud/ai-threat-protection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#10\",\n \"claim\": \"Microsofts red teaming-verktøy for prompt injection-forsvar er: PyRIT (automated adversarial testing) | Azure AI Red Teaming Agent (simulering av angrepsscenarioer).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#11\",\n \"claim\": \"ContentSafetyClient.analyze_text støtter kategorien «Jailbreak» og output_type «FourSeverityLevels», og resultatet eksponerer jailbreak_analysis.attack_detected.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#12\",\n \"claim\": \"Eksempelet på Azure OpenAI med safety meta-prompt bruker modellen gpt-4o via chat.completions.create.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/ai-services/content-safety/overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#13\",\n \"claim\": \"Prompt Shields i Azure AI Content Safety (jailbreak detection) er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/prompt-injection-defense-patterns.md#14\",\n \"claim\": \"Azure AI Content Safety API-versjon 2024-09-01 er GA.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/prompt-injection-defense-patterns.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/security-copilot-integration.md",
|
||
"claim_count": 34,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/security-copilot-integration.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/security-copilot-integration.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#1\",\n \"claim\": \"Security Copilot er inkludert i lisensen for Microsoft 365 E5/E7 inclusion-kunder og auto-provisjoneres etter en 7-dagers forhåndsvarsling fra Microsoft; ingen SCU-kjøp er nødvendig for grunnfunksjonalitet, og aktivering skjer per tenant.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/get-started-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#2\",\n \"claim\": \"Inclusion-modellen for Security Copilot er ikke designet for US Government-skyer: GCC | GCC High | DoD | Azure Government.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/get-started-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#3\",\n \"claim\": \"Security Copilot embedded-opplevelse finnes i disse portalene: Microsoft Defender XDR | Microsoft Sentinel | Microsoft Intune | Microsoft Entra | Microsoft Purview.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/experiences-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#4\",\n \"claim\": \"Phishing Triage Agent i Defender XDR er i Public Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/defender-xdr/phishing-triage-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#5\",\n \"claim\": \"Alert Triage for DLP i Microsoft Purview er i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#6\",\n \"claim\": \"Alert Triage for Insider Risk Management i Microsoft Purview er i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#7\",\n \"claim\": \"Threat Intelligence Briefing Agent i standalone-portalen er i Public Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#8\",\n \"claim\": \"Conditional Access Optimization Agent i Microsoft Entra er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#9\",\n \"claim\": \"Vulnerability Remediation Agent i Microsoft Intune er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/intune/agents/\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#10\",\n \"claim\": \"Access Review Agent i Microsoft Entra + Teams er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#11\",\n \"claim\": \"Security Copilot-agentene for endepunktadministrasjon i Intune består av: Change Review Agent | Device Offboarding Agent | Policy Configuration Agent.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/intune/agents/\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#12\",\n \"claim\": \"Ved agentoppsett velger administrator mellom to identitetstyper: Lag agentidentitet (kun for Microsoft-bygde agenter, oppretter dedikert Entra Agent ID med scoped tillatelser) | Koble til eksisterende brukerkonto (agenten arver brukerens credentials og tillatelser).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/agents-overview\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#13\",\n \"claim\": \"Fra 18. november 2025 er Security Copilot inkludert i Microsoft 365 E5- og E7-lisenser uten ekstra kostnad.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#14\",\n \"claim\": \"Inkludert kapasitet er 400 SCU (Security Compute Units) per måned per 1 000 betalte brukerlisenser, og skalerer proporsjonalt (400 lisenser gir 160 SCU/mnd, 4 000 lisenser gir 1 600 SCU/mnd).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#15\",\n \"claim\": \"Maksimalt inkludert kapasitet i inclusion-modellen er 10 000 SCU per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#16\",\n \"claim\": \"SCU-er nullstilles månedlig; ubrukte SCU-er overføres ikke til neste måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#17\",\n \"claim\": \"Kunder mottar 30-dagers forhåndsvarsel, deretter auto-provisjoneres Security Copilot uten Azure-oppsett eller manuell SCU-tildeling (zero-click activation).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#18\",\n \"claim\": \"Bruk utover inkludert SCU-kapasitet throttles, med fremtidig mulighet for pay-as-you-go-overskridelse (30-dagers forhåndsvarsel gis).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#19\",\n \"claim\": \"«Default Security Copilot Capacity» er den automatisk opprettede inklusjonstildelingen i tenanten: den kan ikke modifiseres, deles på tvers av alle brukere og opplevelser, og faktureres ikke per time.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#20\",\n \"claim\": \"Sentinel-scenariet er inkludert for M365 E5-kunder som også bruker Sentinel.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#21\",\n \"claim\": \"Inkluderte developer experiences omfatter: Agent Builder | APIer for tilpassede agenter | promptbooks | integrasjoner via MCP og Graph APIer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#22\",\n \"claim\": \"Følgende er IKKE inkludert i M365 E5-inklusjonen: Sentinel data lake-kostnader | Azure Logic Apps-kostnader | non-agentic Data Security Investigations i Purview | partner-built agent-lisenser kjøpt via Security Store | enkelte agenter med forutsetninger utenfor M365 E5.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/security-copilot-inclusion\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#23\",\n \"claim\": \"Phishing Triage Agent i Defender krever Microsoft Defender for Office 365 Plan 2 i tillegg til Security Copilot.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/defender-xdr/phishing-triage-agent\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#24\",\n \"claim\": \"Security Copilot integrerer med Microsoft Sentinel via to plugins: Microsoft Sentinel Plugin | Natural Language to KQL for Microsoft Sentinel.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/sentinel/sentinel-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#25\",\n \"claim\": \"Natural Language to KQL for Microsoft Sentinel er i Preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/sentinel/sentinel-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#26\",\n \"claim\": \"Natural Language to KQL er tilgjengelig i standalone-portalen og i Advanced Hunting-seksjonen i Defender-portalen; ikke alle Sentinel-tabeller støttes ennå.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/sentinel/sentinel-security-copilot\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#27\",\n \"claim\": \"Tilpassede Security Copilot-plugins finnes i typene: API-plugin (OpenAPI-spec-wrapper rundt REST API) | KQL-plugin (egendefinerte KQL-spørringer mot Sentinel/Defender) | OpenAI-format (ChatGPT-kompatibelt plugin-format) | Egendefinert agent.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#28\",\n \"claim\": \"Agent Builder i standalone-portalen er tilgjengelig for M365 E5-kunder.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#29\",\n \"claim\": \"Plugin-manifestet (YAML eller JSON) har obligatoriske felter: Descriptor (Name, DisplayName, Description) | SkillGroups.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#30\",\n \"claim\": \"Grenseverdier for plugin-manifest: name_for_model maks 100 tegn, name_for_human maks 40 tegn, description_for_model maks 16 000 tegn.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#31\",\n \"claim\": \"OpenAPI v3.0 eller 3.0.1 støttes for Security Copilot API-plugins.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#32\",\n \"claim\": \"For autentisering i plugin-manifestet er authorization_type begrenset til bearer; støtte for OAuth, api_key og AAD er under utvikling.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#33\",\n \"claim\": \"Tredjepartspluginer tilgjengelige via Security Store inkluderer: AbuseIPDB | Censys | CrowdSec CTI | CyberArk | Cybersixgill | Red Canary | Jamf.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/custom-plugins\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-copilot-integration.md#34\",\n \"claim\": \"Security Copilot er per 2026-02 kun tilgjengelig for kommersielle skytjenester og ikke tilgjengelig i GCC (Government Community Cloud) | GCC High | DoD | Microsoft Azure Government.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/security/get-started-security-copilot\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/security-copilot-integration.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/security-scoring-rubrics-6x5.md",
|
||
"claim_count": 21,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/security-scoring-rubrics-6x5.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/security-scoring-rubrics-6x5.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#1\",\n \"claim\": \"Lokal autentisering (API-nøkler) kan deaktiveres på Foundry Tools (tidligere Cognitive Services) og Azure OpenAI-ressurser via Azure Policy-egenskapen `disableLocalAuth = true`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#2\",\n \"claim\": \"Private Endpoints kan konfigureres for Azure AI-tjenestene Azure OpenAI | Azure AI Search | Azure Storage.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#3\",\n \"claim\": \"Azure AI-ressurser har egenskapen `publicNetworkAccess` som kan settes til `Disabled` for å slå av offentlig nettverkstilgang.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#4\",\n \"claim\": \"Private DNS-sonen for Azure OpenAI private endpoints er `privatelink.openai.azure.com`, og den må VNet-linkes.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/azure-openai-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#5\",\n \"claim\": \"Minimum TLS-versjon for Storage, SQL og AI-ressurser er TLS 1.2 (TLS 1.0/1.1 skal ikke brukes), og kan håndheves med Azure Policy.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#6\",\n \"claim\": \"Kryptering at rest med Customer-Managed Keys (CMK) via Azure Key Vault støttes for Azure AI-tjenester og storage accounts.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#7\",\n \"claim\": \"Azure-regionene Norway East (`norwayeast`) og Norway West (`norwaywest`) er tilgjengelige for provisjonering av AI-ressurser med data residency i Norge.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#8\",\n \"claim\": \"Azure AI-tjenester har innebygd data loss prevention i form av utgående URL-filtrering (outbound URL-liste), og Microsoft Purview sensitivity labels kan brukes på RAG-data.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#9\",\n \"claim\": \"Azure AI Content Safety tilbyr PII-deteksjon som kan aktiveres i AI-pipelinen.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#10\",\n \"claim\": \"Azure AI Content Safety content filters dekker fire harm-kategorier: hate | violence | sexual | self-harm, med severity-nivåer der medium og høyere kan konfigureres; konfigureres i AI Foundry under Guardrails.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#11\",\n \"claim\": \"Prompt Shields detekterer jailbreak-forsøk og indirekte prompt injection, og kan slås på for både user prompts og documents.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#12\",\n \"claim\": \"AI security posture management (AI SPM) / AI threat protection leveres i Defender for Cloud via Defender CSPM-planen, aktivert under Environment Settings.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#13\",\n \"claim\": \"Diagnostic settings for Azure AI-ressurser tilbyr logkategoriene `RequestResponse` | `Audit`, som kan sendes til et Log Analytics workspace.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/azure-openai-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#14\",\n \"claim\": \"Copilot Studio er en SaaS-tjeneste der private endpoints ikke er mulig, og NSG-regler ikke er relevante.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#15\",\n \"claim\": \"Microsoft 365 E5-lisensen gir tilgang til Purview DLP og Entra ID Conditional Access som kan brukes til å beskytte Copilot Studio-løsninger.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#16\",\n \"claim\": \"Microsoft Foundry støtter fine-tuning av GPT-4o, og Azure AI Search kan brukes som RAG-kilde i samme løsning.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/azure-ai-foundry-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#17\",\n \"claim\": \"Microsoft har omdøpt 'Cognitive Services' til 'Foundry Tools' i sikkerhetsbaselines; URL-en for cognitive-services-security-baseline er fortsatt aktiv, men omdirigerer til 'Azure security baseline for Foundry Tools'.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/cognitive-services-security-baseline\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#18\",\n \"claim\": \"Microsoft Cloud Security Benchmark (MCSB) v2 er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#19\",\n \"claim\": \"MCSB v2 er inndelt i security domains som omfatter IM (Identity Management) | PA (Privileged Access) | NS (Network Security) | DP (Data Protection) | AI (Artificial Intelligence Security) | GS (Governance and Strategy) | LT (Logging and Threat Detection) | IR (Incident Response), med kontroller som IM-1 | IM-3 | IM-7 | IM-8 | PA-1 | PA-7 | NS-1 | NS-2 | DP-1 til DP-6 | LT-1 | LT-4.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#20\",\n \"claim\": \"MCSB v2 AI Security-domenet består av kontrollene AI-1 til AI-7.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/mcsb-v2-artificial-intelligence-security\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/security-scoring-rubrics-6x5.md#21\",\n \"claim\": \"Det finnes en egen Microsoft Foundry security baseline publisert på learn.microsoft.com/security/benchmark/azure/baselines/azure-ai-foundry-security-baseline.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/security/benchmark/azure/baselines/azure-ai-foundry-security-baseline\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/security-scoring-rubrics-6x5.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md",
|
||
"claim_count": 17,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md`)\n\n[\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#1\",\n \"claim\": \"Microsoft Azure Security Benchmark klassifiserer AI supply chain-sikkerhet under kontrollen AI-1: Ensure use of approved models, som er merket som en «must have»-kontroll.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/security/benchmark/azure/mcsb-v2-artificial-intelligence-security#ai-1-ensure-use-of-approved-models\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#2\",\n \"claim\": \"Azure Machine Learning registries kan brukes som sentralisert modellregister på tvers av subscriptions og workspaces.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#3\",\n \"claim\": \"Dependency Scanning i Azure DevOps aktiveres via GitHub Advanced Security for Azure DevOps.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#4\",\n \"claim\": \"Azure Pipelines-oppgaven for avhengighetsskanning heter AdvancedSecurity-Dependency-Scanning@1 (versjon 1) og tar inputene scanMode og ecosystem (f.eks. «pip»).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#5\",\n \"claim\": \"Dependency scanning genererer alerts for tre kategorier: Direct vulnerabilities (pakker i requirements.txt) | Transitive vulnerabilities (pakker som direkte dependencies bruker) | CVE severity mapping.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#6\",\n \"claim\": \"CVE-alvorlighetsgradene mappes slik: Critical (CVSS ≥ 9.0) | High (7.0–9.0) | Medium (4.0–7.0) | Low (1.0–4.0).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#7\",\n \"claim\": \"Azure ML-miljøet kan bygges på base-imaget mcr.microsoft.com/azureml/openmpi4.1.0-ubuntu20.04.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#8\",\n \"claim\": \"Microsoft Defender for Containers gjør tre ting: genererer vulnerability assessments automatisk når et image pushes til Azure Container Registry | blokkerer deployment av images med kritiske sårbarheter (konfigurerbart via Azure Policy) | integrerer med Azure Monitor for alerting.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#9\",\n \"claim\": \"Microsoft tilbyr verifiserte modeller via to kanaler: Azure Machine Learning Model Catalog (kuraterte modeller med security attestation) | HuggingFace Registry i Azure (integrert med Azure ML, med provenance tracking).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#10\",\n \"claim\": \"Modellen gpt-35-turbo finnes i azureml-registeret med modellversjon 0301, referert som azureml://registries/azureml/models/gpt-35-turbo/versions/0301.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#11\",\n \"claim\": \"Managed online deployment i Azure ML kan bruke instance_type Standard_DS3_v2.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#12\",\n \"claim\": \"Azure Policy-definisjonen «[Preview]: Azure Machine Learning Deployments should only use approved Registry Models» er i preview og støtter effect «Deny» med parameterne allowedPublishers og approvedAssetIds.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#13\",\n \"claim\": \"Modellen Llama-2-7b finnes i registeret azureml-meta med versjon 18, referert som azureml://registries/azureml-meta/models/Llama-2-7b/versions/18.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#14\",\n \"claim\": \"Azure AI Anomaly Detector er tilgjengelig som tjeneste og kan deployes for å identifisere data poisoning i treningsdata (klienten AnomalyDetectorClient med detect_entire_series).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#15\",\n \"claim\": \"Azure ML gir delvis SBOM-funksjonalitet via tre mekanismer: Model Registry Metadata (navn, versjon, tags, properties, koblet treningsjobb) | Environment Registry (conda-avhengigheter, pip-pakker, Docker base image, kryptografisk hash av miljødefinisjonen) | Dataset Versioning (Azure ML Data Assets med versjonering og lineage tracking).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#16\",\n \"claim\": \"Azure DevOps-pipelinen bruker oppgavene AzureCLI@2 (versjon 2) og PublishBuildArtifacts@1 (versjon 1), og Azure CLI-kommandoen «az ml model download --name … --version … --download-path …».\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n },\n {\n \"id\": \"ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md#17\",\n \"claim\": \"Microsoft Defender for AI kan konfigureres for threat detection på AI-arbeidsbelastninger, sammen med Azure Monitor-varsler for model registry-hendelser.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-vulnerability-management\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/ai-security-engineering/supply-chain-security-ai-models.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/ai-builder-credits-transition.md",
|
||
"claim_count": 23,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/ai-builder-credits-transition.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/ai-builder-credits-transition.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#1\",\n \"claim\": \"Microsoft annonserte i oktober 2025 en progressiv avvikling av AI Builder credits til fordel for en felles kredittmodell basert på Copilot Credits.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/endofaibcredits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#2\",\n \"claim\": \"AI Builder capacity add-on er kun tilgjengelig for eksisterende kunder: salget stoppet 1. november 2025 og produktet når end-of-life 1. november 2026.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/endofaibcredits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#3\",\n \"claim\": \"Én AI Builder capacity add-on gir en kapasitet på 1 000 000 AI Builder credits per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#4\",\n \"claim\": \"Seeded AI Builder credits som er inkludert i premium-lisenser (varierer fra 250 til 20 000 per måned) fjernes 1. november 2026.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/endofaibcredits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#5\",\n \"claim\": \"Seeded AI Builder credits per måned per lisenstype: Power Apps Premium 500 | Power Apps per app 250 | Power Automate Premium 5 000 | Power Automate Process 5 000 | Power Automate Hosted RPA add-on 5 000 | Power Automate Unattended RPA add-on 5 000 | Dynamics 365 F&O 20 000 | Power Apps for Cloud for Sustainability USL Plus 25 000.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/administer-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#6\",\n \"claim\": \"Maksgrense på tenant-nivå for seeded AI Builder credits er 1 000 000 for Power Apps- og Power Automate-lisensene, 20 000 for Dynamics 365 F&O, og ingen grense for Power Apps for Cloud for Sustainability USL Plus.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/administer-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#7\",\n \"claim\": \"Copilot Credits er en felles valuta for AI-kapasitet på tvers av Copilot Studio | AI Builder | Microsoft 365 Copilot | Microsoft Foundry.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#8\",\n \"claim\": \"Fra 1. november 2025 kan nye kunder ikke kjøpe AI Builder capacity add-ons, og må i stedet kjøpe Copilot Credits.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/endofaibcredits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#9\",\n \"claim\": \"Copilot Credits kan kjøpes gjennom to modeller: prepaid pack subscription (månedlig kapasitetspakke) | pay-as-you-go meter (Azure-fakturering per forbruk).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#10\",\n \"claim\": \"AI Builder-funksjoner i Power Apps og Power Automate konsumerer AI Builder credits først, og faller deretter tilbake til Copilot Credits; Copilot Credits kan allokeres til spesifikke environments eller ligge uallokert på tenant-nivå.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#11\",\n \"claim\": \"AI Builder-funksjoner i Copilot Studio konsumerer kun Copilot Credits, uten fallback til AI Builder credits.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#12\",\n \"claim\": \"Forbruk nullstilles den 1. hver måned, og ubrukt kapasitet overføres ikke til neste måned — verken for AI Builder credits eller Copilot Credits.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#13\",\n \"claim\": \"AI Builder credit-rater per enhet: prompt (basic LLM) 1,2 per 1k tokens | prompt (standard LLM) 24 per 1k tokens | prompt (premium LLM) 182 per 1k tokens | receipt/invoice processing 32 per side | custom document processing 100 per side | text recognition (OCR) 3 per side | object detection 8 per bilde.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#14\",\n \"claim\": \"Copilot Credit-rater per enhet: prompt (basic LLM) 0,1 per 1k tokens | prompt (standard LLM) 1,5 per 1k tokens | prompt (premium LLM) 10 per 1k tokens | receipt/invoice processing 8 per side | custom document processing 8 per side | text recognition (OCR) 0,1 per side | object detection 8 per bilde.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#15\",\n \"claim\": \"Overage av kredittkapasitet håndteres som grace period og faktureres ikke, men kjøring blokkeres etter 125 % av kapasiteten.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#16\",\n \"claim\": \"AI Builder-funksjoner tilgjengelig i Power Apps: AI prompts (text generation, summarization) | document processing (invoice, receipt, identity document) | object detection | text recognition (OCR).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#17\",\n \"claim\": \"En app som bruker AI Builder-funksjoner blir «premium», og brukeren som kjører appen må ha Power Apps Premium-lisens.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/power-platform/admin/powerapps-flow-licensing-faq\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#18\",\n \"claim\": \"Prebuilt prompts i Power Automate: AISummarize | AIExtract | AIReply | AIClassify | AISentiment.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#19\",\n \"claim\": \"En cloud flow blir ikke «premium flow» selv om den bruker AI Builder actions, men appen blir premium hvis flowen kalles fra en app.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/power-platform/admin/powerapps-flow-licensing-faq\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#20\",\n \"claim\": \"AI Builder i Dataverse styres med rollebasert tilgangskontroll med rollene maker | user | admin.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/administer-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#21\",\n \"claim\": \"Seeded AI Builder credits er inkludert i Enterprise Agreement (EA)-lisenser fram til 1. november 2026, og fjernes da også for EA-kunder uten unntak.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/administer-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#22\",\n \"claim\": \"Verktøy for å monitorere AI Builder-forbruk: Power Platform admin center → Licensing → Capacity add-ons → Summary tab | AI Builder consumption report (nedlastbar fra admin center) | AI Builder Activity page (sanntidsprediksjoner) | Dataverse AI Event-tabell.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/administer-consumption-report\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ai-builder-credits-transition.md#23\",\n \"claim\": \"Gratis handlinger som ikke konsumerer credits: trening av modeller | testing av modeller i AI Models page | testing av prompts i prompt builder | preview-scenarier i AI Models (unntatt prompts).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/ai-builder-credits-transition.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/azure-cost-management-ai.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/azure-cost-management-ai.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/azure-cost-management-ai.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#1\",\n \"claim\": \"Azure Cost Management (kostnadsovervåking, budsjettering og optimalisering for Azure-ressurser) har status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#2\",\n \"claim\": \"Azure Cost Management tilbyr tre primære mekanismer for kostnadsovervåking: budget alerts (faktiske kostnader mot budsjett) | forecast alerts (prediktive varsler basert på trender) | anomaly detection (automatisk identifisering av uventede kostnadsmønstre).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-mgt-alerts-monitor-usage-spending\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#3\",\n \"claim\": \"Azure Cost Management-plattformen er gratis for alle Azure-kunder.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#4\",\n \"claim\": \"Kjernekomponentene i Azure Cost Management for AI-workloads er: Budget Alerts | Forecast Alerts | Anomaly Detection | Cost Analysis Views | Action Groups | Exports | Budgets API (REST API for programmatisk budsjettering og alert-konfigurasjon).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-mgt-alerts-monitor-usage-spending\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#5\",\n \"claim\": \"Forecast alerts i Azure Cost Management bygger på en 36-timers forecast-algoritme.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-mgt-alerts-monitor-usage-spending\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#6\",\n \"claim\": \"Anomaly detection i Azure Cost Management er ML-basert og bruker en baseline på 60 dagers historikk for å identifisere avvik.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/understand/analyze-unexpected-charges\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#7\",\n \"claim\": \"Action Groups i Azure Cost Management støtter integrasjon med Azure Logic Apps | Webhooks | Azure Functions for automatiserte responser på kostnadsvarsler.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/manage-automation\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#8\",\n \"claim\": \"Budget alerts (actual) evalueres 1 gang per dag, etter at all usage-data er tilgjengelig, og notifikasjon sendes innen 1 time etter evaluering; det samme gjelder forecast alerts.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-mgt-alerts-monitor-usage-spending\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#9\",\n \"claim\": \"Anomaly alerts evalueres 1 gang per dag, 36 timer etter at dagen er slutt (UTC), med auto-tunet konfidensintervall basert på 60 dagers historikk.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/understand/analyze-unexpected-charges\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#10\",\n \"claim\": \"Kostnadsdata i Azure Cost Management er normalt tilgjengelig innen 8–24 timer, og anomaly detection bruker normalisert usage (ikke kostnader) for å unngå prissvingninger.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/understand/analyze-unexpected-charges\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#11\",\n \"claim\": \"Budsjetter i Azure Cost Management kan filtrere bort kjøp (reservations/savings plans) med filteret `ChargeType != Purchase`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#12\",\n \"claim\": \"Integrasjonene mot Power BI og Microsoft Fabric er: Cost Management Connector (Power BI Desktop/Service) | FinOps Hub (open-source accelerator fra Microsoft, Data Factory + Fabric) | Azure Data Explorer (ADX) med KQL-spørringer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#13\",\n \"claim\": \"Microsoft anbefaler FOCUS (FinOps Open Cost and Usage Specification), et leverandør-agnostisk skjema, som eksport-template i Cost Management, med anbefalt pipeline Cost Management exports → ADLS Gen2 → Fabric Lakehouse → Power BI.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#14\",\n \"claim\": \"Azure Cost Management beholder kostnadsdata i 13 måneder; lengre historikk krever eksport til storage (cool/archive).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#15\",\n \"claim\": \"Cost Management har en tag inheritance-funksjon som propagerer tags fra subscription/resource group til individuelle ressurser i kostnadsrapporter.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#16\",\n \"claim\": \"Cost Analysis tilbyr en amortized view som fordeler reservation-/savings plan-kostnader over perioden, i tillegg til Exports for daglige kostnader og Invoice Reconciliation for samsvar mot faktura.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#17\",\n \"claim\": \"Budgets og alerts i Azure Cost Management er gratis, med ubegrenset antall budsjetter og alerts.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/azure-cost-management-ai.md#18\",\n \"claim\": \"Deling av Power BI-rapporter med kostnadsdata krever lisensen Power BI Pro eller Premium.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/azure-cost-management-ai.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/budget-forecasting-ai-projects.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/budget-forecasting-ai-projects.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/budget-forecasting-ai-projects.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#1\",\n \"claim\": \"Forecasting i Azure Cost Management deles i fire metoder: Native Cost Analysis Forecast (1-12 måneder) | AutoML-basert forecasting (3-24 måneder) | Manual projection (variabel horisont) | Hybrid approach (6-36 måneder).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/framework/quantify/forecasting\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#2\",\n \"claim\": \"AI-kostnader segmenteres langs fem budsjettdimensjoner: Compute (training GPU-timer, inference TPM/RPM, PTU-hosting) | Storage (treningsdata, modellartefakter, feature stores, logging) | Networking (dataoverføring, API-kall, kryss-regional replikering) | Licensing (modell-API-er med token-kostnad, fine-tuning, commitment tiers) | Operational (Monitoring, Log Analytics, Application Insights).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#3\",\n \"claim\": \"Azure-budsjetter opprettes med Bicep-ressurstypen Microsoft.Consumption/budgets på API-versjon 2023-11-01.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#4\",\n \"claim\": \"Hosting av fine-tunede Azure OpenAI-modeller medfører løpende timebasert kostnad 24/7 så lenge deploymentet eksisterer, uavhengig av faktisk bruk.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/fine-tuning-cost-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#5\",\n \"claim\": \"Azure Cost Management + Budgets tilbyr native forecasting (1-12 måneder) | budsjettvarsler på både faktiske og forecastede terskler | cost exports til Storage Account | ML-basert anomalideteksjon | tag-basert kostnadsoversikt.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#6\",\n \"claim\": \"Azure-budsjetter har ingen harde grenser — de gir kun varslinger, og håndheving krever egen custom automatisering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#7\",\n \"claim\": \"Forecast-baseline i Azure Cost Management krever minimum 10 dager med historikk.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#8\",\n \"claim\": \"Budsjetter i Azure Cost Management kan kun settes på subscription- eller resource group-scope.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#9\",\n \"claim\": \"Azure Cost Management sitt Python SDK eksponerer CostManagementClient fra pakken azure.mgmt.costmanagement, autentisert med DefaultAzureCredential fra azure.identity.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#10\",\n \"claim\": \"Azure ML SDK v2 (azure.ai.ml) tilbyr automl.forecasting(...) med set_forecast_settings som støtter parameterne time_column_name, forecast_horizon og country_or_region_for_holidays.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-auto-train-forecast\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#11\",\n \"claim\": \"Fra juni 2026 kan en AI-agent kobles direkte til FinOps hub-databasen via Azure MCP server og besvare naturlig-språk-spørsmål om allokering, forecasting, anomalier og rate-optimering.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/hubs/configure-ai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#12\",\n \"claim\": \"FinOps hub-agenten kan kjøres i GitHub Copilot (Agent mode i VS Code) med ferdige FinOps-instruksjoner | som Copilot Studio-agent publisert til Teams/M365 Copilot | via andre MCP-klienter som Claude, og forstår FinOps- og FOCUS-skjemaet.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/hubs/configure-ai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#13\",\n \"claim\": \"Azure commitment tiers (Provisioned Throughput Units) krever langsiktig binding på 1-3 år.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#14\",\n \"claim\": \"Azure Cost Management er gratis og inkludert i Azure-subscriptionen.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/budget-forecasting-ai-projects.md#15\",\n \"claim\": \"FinOps Hubs er gratis som verktøysett; kun underliggende infrastrukturkostnad påløper.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/hubs/configure-ai\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/budget-forecasting-ai-projects.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/cost-allocation-chargeback.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/cost-allocation-chargeback.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/cost-allocation-chargeback.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#1\",\n \"claim\": \"Cost allocation og chargeback-funksjonaliteten som beskrives i filen har status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/platform/governance\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#2\",\n \"claim\": \"Azure Cost Managements innebygde cost allocation rules støttes for kunder med Enterprise Agreement (EA) og Microsoft Customer Agreement (MCA).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#3\",\n \"claim\": \"En cost allocation rule har source og target som hver kan være ett av: subscription | resource group | tag.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#4\",\n \"claim\": \"Allocation percentage i en cost allocation rule kan settes manuelt eller automatisk basert på: compute | storage | network-forbruk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#5\",\n \"claim\": \"Cost allocation rules kjøres sekvensielt i opprettelsesrekkefølge, og det kan ta opptil 24 timer før en ny regel aktiviseres.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#6\",\n \"claim\": \"Cost allocation rules påvirker ikke Azure-fakturaen; de endrer kun hvordan kostnadene vises i Cost Analysis | budgets | eksportert data.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-allocation-introduction\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#7\",\n \"claim\": \"Azure Policy kan håndheve tagging-strategier, og tag inheritance propagerer tags fra subscription/resource group ned til child resources.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/enable-tag-inheritance\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#8\",\n \"claim\": \"Allokerte kostnader inkluderes i CSV-eksport fra Cost Management med kolonnen costAllocationRuleName.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#9\",\n \"claim\": \"Power BI App og Power BI Desktop Connector støtter ikke cost allocation.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#10\",\n \"claim\": \"Usage Details API støtter ikke cost allocation; Cost Details API må brukes i stedet.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#11\",\n \"claim\": \"Reservasjoner og Savings Plans støttes ikke for cost allocation.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/allocate-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#12\",\n \"claim\": \"Microsoft anbefaler FOCUS (FinOps Open Cost and Usage Specification) som eksport-template i Cost Management for standardiserte, leverandør-agnostiske eksporter (verifisert MCP 2026-06).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/well-architected/cost-optimization/collect-review-cost-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#13\",\n \"claim\": \"Azure Policy tilbyr tagging-policyer for cost allocation: Require tag and its value on resources | Inherit a tag from the resource group if missing | Add a tag to resources.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/enable-tag-inheritance\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#14\",\n \"claim\": \"PowerShell-cmdletene Get-AzResource (med -Tag) og New-AzTag brukes til å hente og sette tags for cost center-allokering.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/tag-resources-powershell\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#15\",\n \"claim\": \"FinOps Toolkits Power BI-rapporter omfatter: Cost Summary → Commitments (amortized cost for reservations og savings plans) | Rate Optimization → Chargeback (tabell på subscription/resource group/resource-nivå) | Governance → Summary (tagging compliance).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/power-bi/rate-optimization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#16\",\n \"claim\": \"Azure Cost Management er inkludert uten ekstra kostnad for alle EA-, MCA- og Pay-As-You-Go-kunder, og det er ingen lisenskostnad for cost allocation og chargeback-funksjonalitet.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-allocation-introduction\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#17\",\n \"claim\": \"Azure Cost Management inkluderer: Cost Analysis | Budgets og alerts | Cost allocation rules | Exports til storage account | Recommendations (Azure Advisor).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/cost-allocation-introduction\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#18\",\n \"claim\": \"Power BI-lisensnivåene for chargeback-rapportering er: Power BI Free (kan lese Cost Management-connector, kun personlig bruk) | Power BI Pro (kan dele rapporter med andre Pro-brukere) | Power BI Premium Per User (datamarts, deployment pipelines) | Power BI Premium Capacity (hele organisasjonen).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/power-bi/rate-optimization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/cost-allocation-chargeback.md#19\",\n \"claim\": \"Microsoft FinOps Toolkit er open source og gratis, og består av FinOps Hubs (ARM-template for datapipeline Cost Management → Storage → Data Explorer) | Power BI Reports (maler for cost summary, rate optimization, governance).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/cloud-computing/finops/toolkit/power-bi/rate-optimization\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/cost-allocation-chargeback.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/deterministic-cost-calculation-model.md",
|
||
"claim_count": 24,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/deterministic-cost-calculation-model.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/deterministic-cost-calculation-model.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#1\",\n \"claim\": \"Azure OpenAI Pay-as-You-Go (Global Standard) omfatter modellene GPT-5 | GPT-5-mini | GPT-5-nano | GPT-4o | GPT-4o-mini | o3-mini | GPT-4.1 | GPT-4.1-mini | GPT-4.1-nano | text-embedding-3-small | text-embedding-3-large.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#2\",\n \"claim\": \"Azure OpenAI tilbys i deployment-typene Global Standard | Regional deployment | Data Zone deployment.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#3\",\n \"claim\": \"GPT-5-mini ble lansert i august 2025.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#4\",\n \"claim\": \"Nyere GPT-5-generasjoner er gpt-5.2 | gpt-5.4 | gpt-5.5 (verifisert juni 2026).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#5\",\n \"claim\": \"Azure AI Search har tierne Free | Basic | Standard S1 | Standard S2 | Standard S3 | Storage Optimized L1 | Storage Optimized L2.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#6\",\n \"claim\": \"Lagring per partisjon i Azure AI Search: Free 50 MB | Basic 15 GB | Standard S1 160 GB | Standard S2 512 GB | Standard S3 1 TB | Storage Optimized L1 2 TB | Storage Optimized L2 4 TB.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#7\",\n \"claim\": \"Maks antall Search Units i Azure AI Search: Basic 9 (3 partisjoner x 3 replikaer) | Standard S1 36 (12 partisjoner x 12 replikaer) | Standard S2, S3, L1 og L2 36.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#8\",\n \"claim\": \"Azure AI Search semantic ranker: de første 1 000 forespørslene per måned er gratis, deretter faktureres per 1 000 forespørsler.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#9\",\n \"claim\": \"Fra 2025-09-01 erstattet Copilot Credits «messages» som felles valuta på tvers av Copilot Studio-kapabiliteter; antall per prepaid pack og pay-as-you-go-raten er uendret.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/billing-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#10\",\n \"claim\": \"Copilot Studio har kjøps-/lisensmodellene Pay-as-you-go (Azure-fakturert via billing policy) | Copilot Credit prepurchase plan (CCCU-pool kjøpt i Azure portal) | Capacity Pack | M365 Copilot-brukerrettighet.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/billing-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#11\",\n \"claim\": \"En Copilot Studio Capacity Pack inkluderer 25 000 Copilot Credits per pack per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/billing-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#12\",\n \"claim\": \"Copilot Studio-bruk som er inkludert i M365 Copilot-brukerrettigheten er underlagt en Fair Usage Limit.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/billing-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#13\",\n \"claim\": \"Copilot Credit-forbruk i Copilot Studio: standard melding (ikke-generativ AI) 1 credit | generativt AI-svar (GenAnswers, orchestration) 2 credits | agent flow action 1 credit.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/microsoft-copilot-studio/billing-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#14\",\n \"claim\": \"Microsoft 365 Copilot tilbys i lisensvariantene M365 Copilot (Enterprise, årlig fakturering) | M365 Copilot Business (SMB) | M365 Copilot Chat (forbruksbasert pay-as-you-go).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#15\",\n \"claim\": \"M365 Copilot Business (SMB) faktureres årlig og er begrenset til maks 300 brukere.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/partner-center/announcements/2025-november\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#16\",\n \"claim\": \"Azure AI Content Safety har kapabilitetene Text moderation (S0, per 1 000 text records) | Image moderation (S0, per 1 000 images) | Prompt Shields (per 1 000 requests) | Groundedness detection (per 1 000 requests) | Free tier (F0).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#17\",\n \"claim\": \"Azure AI Content Safety gratisnivå F0 gir 5 000 transaksjoner per 20 dager.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#18\",\n \"claim\": \"Azure AI Document Intelligence har modellene Read (OCR) | Prebuilt models (faktura, kvittering, ID) | Custom extraction | Free tier (F0), alle priset per 1 000 sider.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#19\",\n \"claim\": \"Azure AI Document Intelligence gratisnivå F0 gir 500 sider per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#20\",\n \"claim\": \"Application Insights / Log Analytics inkluderer 5 GB dataingest gratis per måned per faktureringskonto.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#21\",\n \"claim\": \"Azure Monitor-datalagring (retention) er inkludert uten ekstra kostnad i 0–90 dager; lagring utover 90 dager faktureres per GB per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#22\",\n \"claim\": \"Azure Monitor tilbyr commitment tier med fast dagspris fra 100 GB/dag.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#23\",\n \"claim\": \"Azure Blob Storage har lagringsnivåene Hot (egen pris for første 50 TB) | Cool | Archive.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/deterministic-cost-calculation-model.md#24\",\n \"claim\": \"Azure AI Search SLA på 99,9 % tilgjengelighet krever minimum 2 replikaer for lesing og 3 replikaer for lesing/skriving.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/copilot/microsoft-365/microsoft-365-copilot-licensing\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/deterministic-cost-calculation-model.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/inference-endpoint-cost-optimization.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/inference-endpoint-cost-optimization.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/inference-endpoint-cost-optimization.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#1\",\n \"claim\": \"Managed inference endpoints i Azure Machine Learning og Microsoft Foundry har status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#2\",\n \"claim\": \"Deployment-typer for inferens med hver sin prismodell: Managed Online Endpoint (VM-timer per instans per time) | Serverless API Endpoint (pay-per-token + pay-per-request) | Provisioned Throughput (PTU) (fast månedskostnad for reservert kapasitet) | Priority Processing (tier på Standard serverless, pay-per-token til priority-tier-rate) | Low-Priority VMs (rabattert mot dedikerte VM-er, med preemption-risiko).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#3\",\n \"claim\": \"Priority Processing er en tier på Standard serverless som prises per token til en egen priority-tier-rate og gir et definert latensmål (SLA) per modell.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/priority-processing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#4\",\n \"claim\": \"Priority processing aktiveres på GlobalStandard- og DataZoneStandard-deployments og krever modellversjon 2025-12-01 eller nyere.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/priority-processing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#5\",\n \"claim\": \"Priority processing er under utrulling, og enkelte dokumentasjonsflater (deployment-types-/enable-siden) markerer den fortsatt som preview eller invitasjonsbasert per 2026-06.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#6\",\n \"claim\": \"Autoscaling-konfigurasjonen for online endpoints består av parameterne: Minimum instances | Maximum instances | Default instances | Scale-out threshold | Scale-in threshold | Cooldown period | Idle time before scale-down.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-autoscale-endpoints?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#7\",\n \"claim\": \"Standardverdien for «idle time before scale-down» (sekunder før en idle node frigjøres) er 120 sekunder.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-autoscale-endpoints?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#8\",\n \"claim\": \"CPU-instanstyper for managed endpoints med spesifikasjoner: Standard_DS2_v2 (2 vCPU, 7 GB RAM, ingen GPU) | Standard_DS3_v2 (4 vCPU, 14 GB RAM, ingen GPU) | Standard_F2s_v2 (2 vCPU, 4 GB RAM, ingen GPU).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-manage-optimize-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#9\",\n \"claim\": \"Standard_NC4as_T4_v3 har 4 vCPU, 28 GB RAM og T4-GPU, og brukes til GPU-inferens for dype nevrale nett.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-manage-optimize-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#10\",\n \"claim\": \"Standard_NC6s_v3 har 6 vCPU, 112 GB RAM og V100-GPU, og brukes til høy-ytelses GPU-inferens.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-manage-optimize-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#11\",\n \"claim\": \"Metrikker tilgjengelige for autoscaling av online endpoints: CpuUtilizationPercentage (scope deployment) | RequestLatency (scope endpoint) | RequestsPerMinute (scope endpoint) | GpuUtilizationPercentage (scope deployment med GPU) | MemoryUtilizationPercentage (scope deployment).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-autoscale-endpoints?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#12\",\n \"claim\": \"Serverless endpoints i Microsoft Foundry provisjoneres enten via AI Foundry Portal eller via SDK-entiteten ServerlessEndpoint.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/how-to/deploy-models-serverless?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#13\",\n \"claim\": \"Serverless endpoints støtter både Microsoft-egne modeller (blant annet Phi-4-familien) og modeller fra Azure Marketplace.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/how-to/deploy-models-serverless?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#14\",\n \"claim\": \"Kostnadssporing for managed endpoints skjer via taggene azuremlendpoint og azuremldeployment.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-view-online-endpoints-costs?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#15\",\n \"claim\": \"I Azure Cost Management vises «Models sold by Azure» (inkludert Azure OpenAI/Microsoft) som meters under selve Foundry-ressursen, mens partner-/Marketplace-modeller vises under Global resources med formatet model-name-GUID.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#16\",\n \"claim\": \"Managed compute gir full kontroll over compute-region, slik at Norway East eller Norway West kan velges for datalagring i Norge, mens serverless har begrenset region-valg og krever verifisering av at modellene støtter norske regioner.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-manage-optimize-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#17\",\n \"claim\": \"Azure Reservations for managed VM-er tilbys med 1 til 3 års commitment og gir opptil 72 % rabatt.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-plan-manage-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#18\",\n \"claim\": \"Managed endpoints støtter private endpoints (VNet-integrasjon), mens serverless gir mindre kontroll over nettverksisolasjon.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/how-to-manage-optimize-cost?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/inference-endpoint-cost-optimization.md#19\",\n \"claim\": \"Standard serverless-deployments har typisk kvote på 200 000 tokens per minutt og 1 000 requests per minutt per deployment, og grensene varierer per modell, deployment-type og region.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/how-to/deploy-models-serverless?view=foundry-classic\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/inference-endpoint-cost-optimization.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/licensing-compliance-cost-avoidance.md",
|
||
"claim_count": 21,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/licensing-compliance-cost-avoidance.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/licensing-compliance-cost-avoidance.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#1\",\n \"claim\": \"Microsoft 365 Copilot er en add-on-lisens per bruker som krever base-lisens M365 E3/E5 eller Business Standard/Premium, samt Entra ID-konto og Exchange Online-postboks.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-minimum-requirements\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#2\",\n \"claim\": \"Microsoft 365 Copilot Chat er inkludert i base-lisensen for M365/O365 A1 | A3 | A5 | E1 | E3 | E5 | Business Basic | Business Standard | Business Premium.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-chat-requirements\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#3\",\n \"claim\": \"For Microsoft 365 Copilot Chat er web chat uten ekstra kostnad, mens work chat er metered (pay-as-you-go); Copilot Pages krever OneDrive-lisens og Copilot Notebooks krever M365 Copilot-lisens.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-chat-requirements\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#4\",\n \"claim\": \"Azure OpenAI Service er consumption-basert med Azure-subscription som base-krav, og faktureres token-basert (input/output) eller via PTU (provisioned throughput units).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#5\",\n \"claim\": \"Copilot Studio lisensieres standalone eller som add-on til M365 base-lisens, og måles message-basert via Copilot Credits eller pay-as-you-go.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#6\",\n \"claim\": \"AI Builder credits fases ut 1. november 2026, med overgang til Copilot Credits.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#7\",\n \"claim\": \"Seeded AI Builder credits som følger med Power Automate Premium og Power Apps-lisenser fjernes.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#8\",\n \"claim\": \"Copilot Credits blir standard metering unit på tvers av Copilot Studio | AI Builder | M365 Copilot Chat work data.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#9\",\n \"claim\": \"Power Platform har en tenant-innstilling «Block use of unallocated AI Builder credits» som blokkerer forbruk av utildelte AI Builder-credits på tenant-nivå; default tillater ukontrollert forbruk.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#10\",\n \"claim\": \"Ved over 150 millioner tokens per måned bør Provisioned Throughput Units (PTU) vurderes for Azure OpenAI ved stabil workload.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#11\",\n \"claim\": \"Power Automate Premium gir 5000 seeded AI Builder credits per lisens (gjelder før november 2026).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#12\",\n \"claim\": \"Flere Azure AI Services har free tier: 5000 transaksjoner per måned for Text Analytics og 20 transaksjoner per minutt for Translator.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#13\",\n \"claim\": \"Azure Cost Management + Billing krever rollen Cost Management Contributor eller Billing Reader.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/manage-automation\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#14\",\n \"claim\": \"Power Platform Admin Center tilbyr AI Builder capacity management (Licensing → Capacity add-ons) | environment-nivå capacity-allokering | AI Builder consumption-rapport per environment og datointervall | tenant-innstilling for å blokkere utildelte AI Builder credits.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#15\",\n \"claim\": \"Microsoft 365 E3 inkluderer M365 Copilot Chat (web) og seeded AI Builder-credits (til november 2026), og M365 Copilot må kjøpes som add-on.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#16\",\n \"claim\": \"Microsoft 365 Business Premium har samme AI-kapabiliteter som E3, men er begrenset til under 300 brukere.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-minimum-requirements\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#17\",\n \"claim\": \"Microsoft 365 Copilot add-on gir full Copilot i Word | Excel | Teams m.fl., work-grounded chat og Copilot Pages/Notebooks, mens extensibility (connectors og custom agents med work data) er metered.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/microsoft-365-copilot/extensibility/cost-considerations\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#18\",\n \"claim\": \"Copilot Studio standalone inkluderer 25 000 meldinger per måned per tenant, og overage faktureres i blokker på 10 000 meldinger.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#19\",\n \"claim\": \"AI Builder capacity Tier 1 add-on gir 1 000 000 AI Builder credits per måned, og overage går over til Copilot Credits hvis tilgjengelig.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/ai-builder/credit-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#20\",\n \"claim\": \"Azure AI Search S1 faktureres som fast månedsavgift, med ekstra kostnad for lagring over 100 GB.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/licensing-compliance-cost-avoidance.md#21\",\n \"claim\": \"Azure AI Search S1-tier dekker opptil 1 million dokumenter, mot S3-tier for større volumer.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/cloud-adoption-framework/scenarios/ai/plan\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/licensing-compliance-cost-avoidance.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/model-selection-price-performance.md",
|
||
"claim_count": 23,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/model-selection-price-performance.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/model-selection-price-performance.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#1\",\n \"claim\": \"GPT-5-generasjonen er utvidet til GPT-5.1 til GPT-5.5, slik at referansens modelleksempler (GPT-5 / GPT-4.1 / GPT-4o-mini) er én generasjon bak.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#2\",\n \"claim\": \"Modellklassen resonneringsmodeller i Azure AI består av GPT-5 | GPT-5-mini | GPT-5-nano.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/model-choice-guide?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#3\",\n \"claim\": \"Modellklassen store generelle modeller i Azure AI består av GPT-4.1 | GPT-4o.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/model-choice-guide?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#4\",\n \"claim\": \"Modellklassen små effektive modeller i Azure AI består av GPT-4.1-mini | GPT-4.1-nano | GPT-4o-mini.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/model-choice-guide?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#5\",\n \"claim\": \"Modellklassen spesialiserte modeller i Azure AI består av Embeddings | Whisper | DALL-E.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#6\",\n \"claim\": \"Azure OpenAI har deployment-typene Standard (pay-per-token) | Global Standard (pay-per-token, ingen data residency) | Provisioned Throughput (PTU, fast PTU-time-pris) | Developer Tier (fine-tuning).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#7\",\n \"claim\": \"Developer Tier for fine-tuning bruker pay-per-token uten hosting-fee, og deployments auto-slettes etter 24 timer.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/fine-tuning-cost-management?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#8\",\n \"claim\": \"1 PTU gir omtrent 5 400 input tokens per minutt for o4-mini.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#9\",\n \"claim\": \"1 PTU gir omtrent 3 000 input tokens per minutt for GPT-4.1.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#10\",\n \"claim\": \"Model Router er GA-funksjonalitet i Microsoft Foundry og velger automatisk modell basert på prompt-kompleksitet.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/control-plane/how-to-optimize-cost-performance?view=foundry\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#11\",\n \"claim\": \"GPT-5 har fire reasoning-nivåer: Minimal | Low | Medium (default) | High.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/model-choice-guide?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#12\",\n \"claim\": \"GPT-5 kalles via Responses API (client.responses.create) med parameteren reasoning_effort for å styre resonneringsnivå.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/model-choice-guide?view=foundry-classic\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#13\",\n \"claim\": \"Azure Cost Management tilbyr funksjonene Cost Analysis (per-modell kostnad via deployment tags) | Budgets + Alerts | Export til Storage.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#14\",\n \"claim\": \"AI Builder / prompt builder bruker GPT-4.1 mini som standardmodell for generative oppgaver.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#15\",\n \"claim\": \"I AI Builder brukes GPT-4o mini og GPT-4o nå kun i US government-regioner.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#16\",\n \"claim\": \"AI Builder-kostnader dekkes av AI Builder credits, med 500 credits per bruker per måned i premium-planer.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#17\",\n \"claim\": \"Copilot Studio bruker Azure OpenAI-modeller (GPT-4o eller GPT-4.1-serien).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#18\",\n \"claim\": \"Standard deployment (regional) gir garantert data residency i Norge, mens Global Standard ikke gir data residency-garanti.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#19\",\n \"claim\": \"Azure OpenAI lisensieres som pay-per-token eller PTU, og AI-kostnad er ikke inkludert i lisensen.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#20\",\n \"claim\": \"Copilot Studio lisensieres per bruker per måned, og inferenskostnader er inkludert i lisensen.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#21\",\n \"claim\": \"M365 Copilot lisensieres per bruker per måned med inferens inkludert og ingen ekstra kostnad.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#22\",\n \"claim\": \"Modeller i GA per 2026-06 er GPT-4.1-serien | GPT-4o/GPT-4o-mini | o-serien (o1 | o3 | o3-mini | o4-mini) | GPT-5-generasjonen.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/model-selection-price-performance.md#23\",\n \"claim\": \"GPT-5-generasjonens utgivelser: GPT-5/-mini/-nano (2025-08), GPT-5.1 (2025-11-13), GPT-5.2 (2025-12-11), GPT-5.3-codex/-chat (2026-02/03), GPT-5.4 + GPT-5.4-pro (2026-03-05), GPT-5.5 (2026-04-24), gpt-chat-latest (2026-05, preview).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/model-selection-price-performance.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/multi-model-strategy-costs.md",
|
||
"claim_count": 27,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/multi-model-strategy-costs.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/multi-model-strategy-costs.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#1\",\n \"claim\": \"Model Router versjon `2025-11-18` velger optimal modell fra et sett på 28 underliggende modeller (inkludert GPT-serien, Claude, DeepSeek, Llama og Grok), og settet oppdateres løpende.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#2\",\n \"claim\": \"Model Router `2025-11-18` er GA (generelt tilgjengelig) og er en trent LLM som ruter prompts til beste underliggende modell.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#3\",\n \"claim\": \"Model Router har tre routing-modi, alle GA: Quality (maks nøyaktighet) | Balanced (default) | Cost (maks besparelse).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#4\",\n \"claim\": \"Model Subset — egendefinert utvalg av underliggende modeller for routing — er GA i Model Router.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#5\",\n \"claim\": \"Model Router støtter to deployment-typer: Global Standard | Data Zone Standard.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#6\",\n \"claim\": \"Model Router-deployments er regionalt tilgjengelige i East US 2 og Sweden Central.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#7\",\n \"claim\": \"OpenAI-modellene som inngår i Model Router `2025-11-18` er: gpt-4o | gpt-4o-mini | gpt-4.1 | gpt-4.1-mini | gpt-4.1-nano | o4-mini | gpt-5-nano | gpt-5-mini | gpt-5 | gpt-5-chat | gpt-5.2 | gpt-5.2-chat | gpt-5.3-chat | gpt-5.4-nano | gpt-5.4-mini | gpt-5.4 | gpt-5.5.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#8\",\n \"claim\": \"Tredjeparts-modellene i Model Router `2025-11-18` er: DeepSeek-V3.1 | DeepSeek-V3.2 | gpt-oss-120b | Llama-4-Maverick-17B-128E-Instruct-FP8 | grok-4 | grok-4-fast-reasoning.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#9\",\n \"claim\": \"Claude-modellene i Model Router krever egen deployment og omfatter: claude-haiku-4-5 | claude-sonnet-4-5 | claude-opus-4-1 | claude-opus-4-6 | claude-opus-4-7.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#10\",\n \"claim\": \"Rate limits for Model Router skalerer med abonnementets kvotenivå (Quota Tier 1–6), ikke lenger med Default/Enterprise-nivåer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#11\",\n \"claim\": \"GlobalStandard-grenser for Model Router per kvotenivå: Tier 1 = 1 000 RPM / 1 000 000 TPM | Tier 2 = 2 000 / 2 000 000 | Tier 3 = 4 000 / 4 000 000 | Tier 4 = 7 000 / 7 000 000 | Tier 5 = 10 000 / 10 000 000 | Tier 6 = 15 000 / 15 000 000.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#12\",\n \"claim\": \"DataZoneStandard-grenser for Model Router per kvotenivå: Tier 1 = 300 RPM / 300 000 TPM | Tier 2 = 670 / 670 000 | Tier 3 = 1 000 / 1 000 000 | Tier 4 = 2 000 / 2 000 000 | Tier 5 = 3 000 / 3 000 000 | Tier 6 = 4 000 / 4 000 000.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#13\",\n \"claim\": \"Gateway-topologiene for flere Azure OpenAI-deployments er: Single Instance + Multiple Deployments | Multiple Instances (Same Region) | Multiple Instances (Multi-Region).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#14\",\n \"claim\": \"Azure API Managements AI gateway medierer også ikke-OpenAI-skjemaer — Anthropic Messages API og Google Vertex AI — i tillegg til Foundry-modeller, og tilbyr unified model API i preview (ett OpenAI-kompatibelt endepunkt på tvers av leverandører).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#15\",\n \"claim\": \"Model Routers context window er begrenset til den minste underliggende modellen — 128k for GPT-4.1-serien.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#16\",\n \"claim\": \"Model Router baserer routing kun på tekst-input, ikke bilder.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#17\",\n \"claim\": \"Kvoter for Standard-deployments er på abonnementsnivå (subscription-level), ikke på instansnivå.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#18\",\n \"claim\": \"GPT-5-serien støtter 400k context window.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#19\",\n \"claim\": \"Model Router kalles via OpenAI-kompatibelt v1-endepunkt med base_url `https://YOUR-RESOURCE.openai.azure.com/openai/v1/`.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#20\",\n \"claim\": \"Circuit breaker-policyen i Azure API Management (backend circuit-breaker rules) er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#21\",\n \"claim\": \"Foundry Model Catalog tilbyr modeller utenfor Azure OpenAI: Meta Llama | Mistral | Cohere | Phi-modeller.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#22\",\n \"claim\": \"Deployment-alternativene i Foundry Model Catalog er: Managed compute | Serverless API | Pay-as-you-go.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#23\",\n \"claim\": \"Data Zone Standard holder data innenfor en Microsoft-spesifisert data zone, f.eks. EU Data Boundary.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#24\",\n \"claim\": \"Underliggende modeller i Model Router må deployes i samme data zone, med unntak av Claude-modellene som krever separate deployments.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#25\",\n \"claim\": \"Azure Reservations for provisioned throughput kan kjøpes som 1-års eller 3-års avtale.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#26\",\n \"claim\": \"Input-gjennomstrømning per PTU: gpt-4.1-nano = 59 400 TPM | gpt-4.1-mini = 14 900 TPM | gpt-4.1 = 3 000 TPM | gpt-5-mini = 23 750 TPM | gpt-5 = 4 750 TPM | o4-mini = 5 400 TPM.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/multi-model-strategy-costs.md#27\",\n \"claim\": \"Azure OpenAI er nå tagget som «Foundry Tools / Azure OpenAI in Foundry Models» i Microsoft-dokumentasjonen.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/multi-model-strategy-costs.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/observability-cost-reduction.md",
|
||
"claim_count": 20,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/observability-cost-reduction.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/observability-cost-reduction.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#1\",\n \"claim\": \"Azure Monitor-økosystemet for observability-kostnad består av komponentene Application Insights | Log Analytics Workspace | Azure Monitor Metrics | Azure Monitor Logs.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/fundamentals/best-practices-cost\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#2\",\n \"claim\": \"Kostnadsmodellene for Log Analytics er Pay-as-you-go | Commitment Tiers | Basic Logs | Auxiliary Logs | Long-term Retention.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/cost-logs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#3\",\n \"claim\": \"Commitment Tiers er forhåndsbetalte daglige volumer med nivåene 100 GB | 200 GB | 500 GB (og flere).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/cost-logs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#4\",\n \"claim\": \"Basic Logs har redusert ingestion-pris, egen query-kostnad og begrenset query-tid på 8 dager.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#5\",\n \"claim\": \"Auxiliary Logs har lavest ingestion-pris og kan kun spørres via search jobs.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#6\",\n \"claim\": \"Long-term retention gir arkivering utover interactive retention i opptil 12 år.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/cost-logs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#7\",\n \"claim\": \"Sampling-strategiene i Application Insights er Adaptive Sampling | Fixed-rate Sampling | Rate-limited Sampling | Ingestion Sampling | Sampling Overrides.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/sampling-classic-api\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#8\",\n \"claim\": \"Adaptive sampling justerer automatisk basert på telemetri-volum med default 5 items/sec.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/sampling-classic-api\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#9\",\n \"claim\": \"Adaptive sampling gjelder kun Classic API SDK (ASP.NET, ASP.NET Core); OpenTelemetry-baserte distros har ikke adaptive sampling aktivert som default.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/opentelemetry-sampling\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#10\",\n \"claim\": \"Rate-limited sampling begrenser til maks N requests/sek (f.eks. 1.5 req/sec) og brukes for Java-applikasjoner.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/java-standalone-config#sampling\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#11\",\n \"claim\": \"Azure Monitor OpenTelemetry-distroen har ikke sampling på som default og støtter to strategier: fixed-rate (ratio 0–1, f.eks. 0.1 = ~10 %) | rate-limited (traces/sek, f.eks. 5.0 = fem traces/sek).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/opentelemetry-sampling\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#12\",\n \"claim\": \"Trace-based sampling for logs dropper logger knyttet til ikke-samplede traces og er på som default når sampling er aktivert (i støttede språk); distroens custom sampler bevarer hele traces og kreves for Live Metrics.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/opentelemetry-configuration#enable-sampling\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#13\",\n \"claim\": \"Rate-limited sampling med requestsPerSecond konfigureres i Java-agent versjon 3.7.5 eller nyere.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/app/java-standalone-config#sampling\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#14\",\n \"claim\": \"Log Analytics table plans er Analytics | Basic | Auxiliary.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#15\",\n \"claim\": \"Analytics-plan støtter alerts, Basic-plan støtter Simple Log Alerts, mens Auxiliary-plan ikke støtter alerts.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#16\",\n \"claim\": \"Retention per table plan: Analytics 30–730 dager | Basic 8 dager interactive pluss long-term | Auxiliary kun long-term.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#17\",\n \"claim\": \"Workspace replication og Customer Lockbox støttes for Analytics- og Basic-tabeller, men ikke for Auxiliary-tabeller.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#18\",\n \"claim\": \"Classic Application Insights er deprecated og støtter kun pay-as-you-go, mens workspace-based Application Insights lagrer data i Log Analytics workspace og kan bruke commitment tiers og Basic Logs.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/service-guides/application-insights#cost-optimization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#19\",\n \"claim\": \"For volumer over 1 TB/dag er dedicated cluster med cluster commitment tier aktuelt i Log Analytics.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/cost-logs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/observability-cost-reduction.md#20\",\n \"claim\": \"Auxiliary-plan støtter Microsoft Sentinel | Search jobs | Summary rules.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/azure-monitor/logs/data-platform-logs#table-plans\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/observability-cost-reduction.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/prompt-engineering-cost-reduction.md",
|
||
"claim_count": 20,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/prompt-engineering-cost-reduction.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/prompt-engineering-cost-reduction.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#1\",\n \"claim\": \"Prompt caching i Azure OpenAI krever minimum 1024 tokens i promptlengde, og de første 1024 tokenene må være identiske for å gi cache-treff.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#2\",\n \"claim\": \"Cache-granulariteten er 128 tokens: etter de første 1024 tokenene skjer cache-treff i inkrementer på hver 128 tokens.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#3\",\n \"claim\": \"In-memory prompt-cache tømmes typisk etter 5-10 minutter med inaktivitet, og alltid innen 1 time.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#4\",\n \"claim\": \"Parameteren `prompt_cache_retention: \\\"24h\\\"` (extended retention) holder cachede prefikser aktive i opptil 24 timer.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#5\",\n \"claim\": \"Cachede input-tokens gir rabatt på Standard-deployment (typisk 50 %) og opptil 100 % på Provisioned; prisen er den samme for begge retention-policyene.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#6\",\n \"claim\": \"In-memory prompt caching støttes av alle Azure OpenAI-modeller som er GPT-4o eller nyere.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#7\",\n \"claim\": \"Extended 24h-retention støttes av gpt-5 | gpt-5.1 | gpt-5.2 | gpt-5.4 (inkludert codex- og chat-varianter) | gpt-4.1.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#8\",\n \"claim\": \"Default retention-policy: gpt-5.4 og eldre modeller bruker `in_memory` (med `24h` som valgbart alternativ), mens nyere modeller har `24h` som default og ikke støtter `in_memory`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#9\",\n \"claim\": \"Overstiger samme prefix kombinert med `prompt_cache_key` omtrent 15 requests per minutt, flyter noen requests over til ekstra maskiner og cache-effektiviteten faller.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#10\",\n \"claim\": \"Prompt caching støtter system messages | user messages | tool definitions.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#11\",\n \"claim\": \"Azure OpenAI eksponeres via v1-API-endepunktet `https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/` som base_url i OpenAI-klienten.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#12\",\n \"claim\": \"Azure OpenAI API returnerer feltet `cached_tokens` under `usage.prompt_tokens_details` i responsen.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#13\",\n \"claim\": \"Prompt Flow støtter gjeldende modeller, blant annet GPT-4.1-serien | GPT-5-serien.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#14\",\n \"claim\": \"AI Foundry Model Catalog støtter in-memory prompt caching for GPT-4o (2024-11-20, 2024-08-06) | GPT-4o-mini (2024-07-18) | o1-serien | o3-mini | GPT-4.1-serien | GPT-5-serien (gpt-5/5.1/5.2/5.4 + codex-varianter).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#15\",\n \"claim\": \"I Copilot Studio er prompt caching ikke eksponert til brukeren, selv om tjenesten bruker underliggende Azure OpenAI.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#16\",\n \"claim\": \"Extended retention lagrer cache-data midlertidig på GPU-maskiner og holdes i-region kun ved deployment-typene Regional Standard | Regional Provisioned; ved Global- og DataZone-typer kan extended-cache-data forlate regionen.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#17\",\n \"claim\": \"Norske organisasjoner med residenskrav anbefales å bruke regionene Norway East | West Europe.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#18\",\n \"claim\": \"o1-preview er retired fra 2025-07 og kan ikke lenger deployes.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#19\",\n \"claim\": \"GPT-4 og GPT-4-32K er retired fra 2025-06 (kan ikke lenger deployes) og støtter ikke cached input-rabatt.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/prompt-engineering-cost-reduction.md#20\",\n \"claim\": \"Azure OpenAI On Your Data er deprecated og pensjoneres 2026-10-14; migrasjonsanbefalingen er Foundry Agent Service + Foundry IQ.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data#token-usage-estimation-for-azure-openai-on-your-data\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/prompt-engineering-cost-reduction.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/ptu-vs-paygo-economics.md",
|
||
"claim_count": 20,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/ptu-vs-paygo-economics.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/ptu-vs-paygo-economics.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#1\",\n \"claim\": \"Azure OpenAI tilbyr tre deployment-typer for provisioned throughput: Global Provisioned | Data Zone Provisioned | Regional Provisioned.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#2\",\n \"claim\": \"Microsoft Foundry skiller mellom fire deployment-kategorier: Standard (pay-per-token, ingen latency-SLA) | Priority processing (pay-per-token til priority-tier-rate, med definert latency-target per modell) | Provisioned (per PTU/time eller reservation, garantert throughput) | Batch (rabattert pay-per-token, asynkront).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#3\",\n \"claim\": \"Provisioned Throughput Unit (PTU) er en generisk, ikke modellspesifikk enhet for modellprosesseringskapasitet — samme PTU-quota kan brukes på tvers av Azure OpenAI-modeller og Foundry-modeller (DeepSeek, Llama).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#4\",\n \"claim\": \"Azure Reservations for PTU tilbys som 1-måneds eller 1-års commitment, og reservasjoner kjøpes i Azure Portal, ikke i AI Foundry.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#5\",\n \"claim\": \"Hver provisioned deployment-type (Global Provisioned, Data Zone Provisioned, Regional Provisioned) krever separat reservation — reservasjonene er ikke utbyttbare på tvers av typene.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#6\",\n \"claim\": \"Minimum PTU varierer per modell: GPT-4o 50 PTU regional og 15 PTU global | GPT-4o-mini 25 PTU regional og 15 PTU global | DeepSeek-R1 100 PTU global uten regional-alternativ.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#7\",\n \"claim\": \"For nyere modeller (GPT-4.1 og senere) oppgis separate input/output-TPM per PTU, og GPT-5 har 4750 input-TPM per PTU; output-tokens forbruker mer kapasitet enn input.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#8\",\n \"claim\": \"Azure Monitor-metrikken «Provisioned-Managed Utilization V2» måler PTU-utnyttelse, og ved 100 % utnyttelse returneres HTTP 429.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/provisioned-get-started\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#9\",\n \"claim\": \"Dynamic quota er i preview og lar standard-deployments opportunistisk bruke mer quota når det er tilgjengelig, uten ekstra konfigurasjon.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/dynamic-quota\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#10\",\n \"claim\": \"For standard-deployments definerer TPM-quota (Tokens Per Minute) maks throughput, og quota kan økes via quota-request.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#11\",\n \"claim\": \"Spillover er GA og støttes av alle Azure OpenAI-modeller med PTU, men ikke av Foundry-modeller fra andre leverandører (Azure DeepSeek, Meta Llama).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/spillover-traffic-management\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#12\",\n \"claim\": \"Azure Monitor eksponerer disse metrikkene for Azure OpenAI: Provisioned-Managed Utilization V2 (PTU) | Processed Prompt Tokens | Generated Completion Tokens | Azure OpenAI Requests.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#13\",\n \"claim\": \"PTU-kapasitetskalkulatoren tar inn modell og versjon | peak calls per minute (RPM) | gjennomsnittlig antall tokens i prompt | gjennomsnittlig antall tokens i modellsvar, og returnerer estimert PTU avrundet til deployment-inkrement samt rått PTU-estimat.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#14\",\n \"claim\": \"Regional Provisioned gir data residency i valgt region (for eksempel Norway East eller West Europe).\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#15\",\n \"claim\": \"Data Zone Provisioned gir data residency innenfor EU data zone, som omfatter 12 regioner.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#16\",\n \"claim\": \"Global Provisioned bruker multi-region routing og gir ingen garanti for data residency.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#17\",\n \"claim\": \"Azure Reservations for Azure OpenAI kan kjøpes i hvilken som helst region eller subscription scope, og valget påvirker ikke data residency.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#18\",\n \"claim\": \"PTU-reservasjoner kan kjøpes på subscription- eller management group-nivå og deles på tvers av prosjekter og team.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#19\",\n \"claim\": \"Med prompt caching får PTU 100 % rabatt på cachede tokens i utilization-beregningen.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/ptu-vs-paygo-economics.md#20\",\n \"claim\": \"Spillover ruter automatisk trafikk fra PTU-deployment til standard-deployment ved kapasitetsgrense (HTTP 429/500/503), og kan konfigureres per deployment eller per request via headeren x-ms-spillover-deployment.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/spillover-traffic-management\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/ptu-vs-paygo-economics.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/rag-query-cost-reduction.md",
|
||
"claim_count": 27,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/rag-query-cost-reduction.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/rag-query-cost-reduction.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#1\",\n \"claim\": \"Microsoft Learns token-estimater for Azure OpenAI On Your Data er basert på standardkonfigurasjon med 5 hentede dokumenter, strictness=3 og chunk size 1024.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#2\",\n \"claim\": \"gpt-35-turbo-16k forbruker i snitt 5 799 tokens per On Your Data-query (generation prompt 4 297, intent prompt 1 366, response output 111, intent output 25).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#3\",\n \"claim\": \"gpt-4-0613 forbruker i snitt 5 518 tokens per On Your Data-query (generation prompt 3 997, intent prompt 1 385, response output 118, intent output 18).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#4\",\n \"claim\": \"gpt-4-1106-preview forbruker i snitt 5 495 tokens per On Your Data-query (generation prompt 4 538, intent prompt 811, response output 119, intent output 27).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#5\",\n \"claim\": \"gpt-35-turbo-1106 forbruker i snitt 6 362 tokens per On Your Data-query (generation prompt 4 854, intent prompt 1 372, response output 110, intent output 26).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#6\",\n \"claim\": \"Microsofts token-tall for On Your Data er målt med 191 samtaler, 250 spørsmål, 10 tokens per spørsmål i snitt og 4 samtale-turns per samtale.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#7\",\n \"claim\": \"Azure AI Search Basic-tier har 1 partisjon, 3 replikaer og 15 GB lagring per partisjon (eldre tjenester: 2 GB).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#8\",\n \"claim\": \"Azure AI Search S1-tier har opptil 12 partisjoner, 12 replikaer og 160 GB lagring per partisjon.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#9\",\n \"claim\": \"Azure AI Search S2-tier har opptil 12 partisjoner, 12 replikaer og 512 GB lagring per partisjon.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#10\",\n \"claim\": \"Azure AI Search S3-tier har opptil 12 partisjoner, 12 replikaer og 1 024 GB lagring per partisjon.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#11\",\n \"claim\": \"Azure AI Search tilbyr to prismodeller: Dedicated (provisjonert, fast pris per Search Unit) | Serverless (Preview, forbruksbasert med Compute Units/time + per-GB/mnd lagring, ingen compute-kost ved idle).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#12\",\n \"claim\": \"Per juni 2026 er Serverless-prismodellen for Azure AI Search i preview, støtter ikke migrering til/fra dedicated og er ikke anbefalt for produksjon.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#13\",\n \"claim\": \"Serverless (Preview) for Azure AI Search er kun tilgjengelig i West Central US | Switzerland North | Japan East.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#14\",\n \"claim\": \"De første 1 000 semantic ranking-queriene per måned er inkludert i Azure AI Search Basic-tier eller høyere; påfølgende queries får per-query-avgift.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#15\",\n \"claim\": \"Deler av agentic retrieval i Azure AI Search er generelt tilgjengelig (GA) via REST API, mens Azure-portalen og Microsoft Foundry-portalen kun gir preview-tilgang til agentic retrieval-funksjonene.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#16\",\n \"claim\": \"Agentic retrieval er GA i REST API-versjon 2026-04-01.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#17\",\n \"claim\": \"Agentic retrieval har to planer: Free plan (default, månedlig gratis token-kvote inkludert) | Standard plan (pay-as-you-go etter at gratiskvoten er brukt), mens Azure OpenAI faktureres separat for query planning og answer synthesis.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#18\",\n \"claim\": \"Agentic retrieval inkluderer 50 millioner gratis tokens per måned.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#19\",\n \"claim\": \"Parameteren reasoning_effort i agentic retrieval kan settes til minimal | low | medium.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#20\",\n \"claim\": \"Standardverdier i Azure OpenAI On Your Data: topNDocuments=5 | strictness=3 | chunk_size=1024 | inScope=true | max_tokens=800.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#21\",\n \"claim\": \"Vektorkomprimering i Azure AI Search med scalar/binary quantization kan redusere vector-lagring med opptil 92,5 %.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-performance-tips\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#22\",\n \"claim\": \"Anbefalt startpunkt for TPM-kvote ved lastbalansering av Azure OpenAI bak Azure Container Apps er 30K TPM per instans.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/developer/python/get-started-app-chat-scaling-with-azure-container-apps\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#23\",\n \"claim\": \"Deploy til Copilot Studio er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#24\",\n \"claim\": \"En Copilot Studio-deployment kan gjenbrukes på flere kanaler: Teams | web | Dynamics 365.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#25\",\n \"claim\": \"Azure AI Search og Azure OpenAI kan deployes i regionene Norway East | Norway West for å holde data innenfor EU/EØS.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#26\",\n \"claim\": \"gpt-4o er multimodal, men støtter kun tekst i Azure OpenAI On Your Data.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/rag-query-cost-reduction.md#27\",\n \"claim\": \"text-embedding-ada-002 er den eneste støttede embedding-modellen for vector search i Azure OpenAI On Your Data.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/use-your-data\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/rag-query-cost-reduction.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/request-batching-aggregation.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/request-batching-aggregation.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/request-batching-aggregation.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#1\",\n \"claim\": \"Request batching og response aggregation for Microsoft AI-stakken (Azure OpenAI Batch API, Azure ML batch endpoints, Microsoft Graph JSON batching) har status GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/concept-endpoints-batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#2\",\n \"claim\": \"Azure OpenAI Batch API gir 50 % kostnadsreduksjon sammenlignet med standard global deployments, med separert token quota og 24-timers SLA.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#3\",\n \"claim\": \"Global-Batch er en dedikert deployment-type i Azure OpenAI med 50 % lavere pris enn global standard.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#4\",\n \"claim\": \"Data Zone Batch (SKU `DataZoneBatch`) er en batch-deployment-type med samme 50 % rabatt som Global-Batch, men med inferens-prosessering begrenset til Microsofts definerte datasone (EU eller US).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#5\",\n \"claim\": \"Azure OpenAI Batch API bruker en separat «enqueued token quota» som ikke forstyrrer online workloads.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#6\",\n \"claim\": \"Azure OpenAI Batch API har et 24-timers completion window som er en target-SLA — jobber kan ta lenger tid, men utløper ikke.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#7\",\n \"claim\": \"I utvalgte regioner feiler nye Azure OpenAI batch-jobber raskt (fail fast) når enqueued-token-grensen overskrides, slik at klienten kan kø-stille og retry-e med exponential backoff.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#8\",\n \"claim\": \"Microsoft Graph JSON batching bruker OData-standardens URL path segment `$batch`, tilgjengelig som `/v1.0/$batch` eller `/beta/$batch`.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#9\",\n \"claim\": \"I Microsoft Graph JSON batching består requests-arrayen av feltene id | method | url | headers | body, og responses-arrayen av feltene id | status | headers | body.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#10\",\n \"claim\": \"Microsoft Graph JSON batching støtter `dependsOn`-property for sekvensielle dependencies mellom requests (valgfritt).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#11\",\n \"claim\": \"Microsoft Graph JSON batching har en batch-størrelsesgrense på maksimalt 20 requests per batch.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#12\",\n \"claim\": \"Azure Machine Learning batch endpoints omfatter komponentene: batch endpoint (asynkron inferencing med auto-scaling compute clusters) | pipeline component deployments | low-priority VMs | scale-to-zero clusters | parallelisering over flere filer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-endpoints-batch?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#13\",\n \"claim\": \"Regioner med fail-fast-kø-støtte for Azure OpenAI batch er: australiaeast | eastus | eastus2 | germanywestcentral | italynorth | northcentralus | polandcentral | swedencentral | switzerlandnorth | westus.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#14\",\n \"claim\": \"Microsoft Graph JSON batching har en maksimal URL-lengde per request på omtrent 2000 tegn.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#15\",\n \"claim\": \"Throttling i Microsoft Graph gjelder fortsatt per individuell request inne i en JSON-batch.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/graph/json-batching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#16\",\n \"claim\": \"Low-priority VMs i Azure Machine Learning gir 60–80 % kostnadsreduksjon sammenlignet med standard VMs, og AML batch endpoints har auto-recovery ved deallokering.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/machine-learning/concept-endpoints-batch?view=azureml-api-2\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#17\",\n \"claim\": \"For Azure OpenAI gjelder at data lagret at rest forblir i den angitte Azure-geografien, mens data kan prosesseres for inferens i enhver Azure OpenAI-lokasjon.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/request-batching-aggregation.md#18\",\n \"claim\": \"Azure OpenAI deployment-typene i prismodellen er: Global Standard (100 % baseline) | Global Batch (50 %) | Data Zone Batch (50 %, inferens innenfor EU/US-datasone) | Provisioned Throughput (reservasjonsbasert).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/request-batching-aggregation.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/reserved-capacity-planning.md",
|
||
"claim_count": 26,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/reserved-capacity-planning.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/reserved-capacity-planning.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#1\",\n \"claim\": \"Bindingstid for Azure Reservations (PTU) er 1 måned eller 1 år, mens Commitment Tier Pricing er 1 måned (web/connected) eller 1 år (disconnected containers).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#2\",\n \"claim\": \"Deployment types for PTU-reservasjoner er Regional | Data Zone | Global Provisioned, mens commitment tier-typene er Web API | Connected containers | Disconnected containers.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#3\",\n \"claim\": \"Regional Provisioned (quota-navn: Regional Provisioned Throughput Unit) har minimum 50 PTU og scale increment 50 (25 for mini/nano-modeller).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#4\",\n \"claim\": \"Data Zone Provisioned (quota-navn: Data Zone Provisioned Throughput Unit) har minimum 15 PTU og scale increment 5.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#5\",\n \"claim\": \"Global Provisioned (quota-navn: Global Provisioned Throughput Unit) har minimum 15 PTU og scale increment 5.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#6\",\n \"claim\": \"Reservasjoner for Regional, Data Zone og Global Provisioned er ikke utskiftbare — det må kjøpes separate reservasjoner for hver deployment type.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#7\",\n \"claim\": \"Minimum PTU og scale increment varierer per modell og versjon: gpt-5-mini har 25 PTU minimum regional, og enkelte Foundry Models (DeepSeek/Fireworks) har 100+ PTU minimum.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-throughput-sizing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#8\",\n \"claim\": \"Reservation scopes for PTU-reservasjoner er Single resource group | Single subscription | Management group | Shared (billing account).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#9\",\n \"claim\": \"Fra mai 2025 støtter PTU-reservasjoner automatisk cross-model sharing, slik at én reservasjon kan dekke både Azure OpenAI og Foundry Models (DeepSeek, Llama), med matching per time på aggregert PTU-forbruk.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#10\",\n \"claim\": \"Commitment tier pricing gjelder kun single-service resources, ikke multi-service eller Foundry multi-service-ressurser.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/commitment-tier\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#11\",\n \"claim\": \"Tjenester som støtter commitment tier pricing er Speech to Text (Standard) | Text to Speech (Neural) | Text Translation (Standard) | Language Understanding (LUIS) | Azure Language (Sentiment, Key Phrase, NER, Language Detection) | Vision OCR | Document Intelligence (Custom/Invoice).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/commitment-tier\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#12\",\n \"claim\": \"Speech to Text (Standard) og Text to Speech (Neural) støtter commitment-typene Web | Connected | Disconnected, mens Text Translation og Vision OCR støtter Web | Connected, og LUIS, Azure Language og Document Intelligence kun støtter Web.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/commitment-tier\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#13\",\n \"claim\": \"Commitment-typene er Web (1 måned, månedlig fakturering, første måned pro-rated) | Connected container (1 måned, månedlig fakturering) | Disconnected container (1 år, årlig fakturering med fullt beløp ved kjøp).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/ai-services/commitment-tier\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#14\",\n \"claim\": \"Legacy resource-bound commitments for Azure OpenAI var begrenset til modellene gpt-4o og gpt-4o-mini.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/provisioned-migration\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#15\",\n \"claim\": \"Nye påmeldinger (enrollments) til den gamle resource-bundne commitment-modellen for Azure OpenAI ble stoppet 1. august 2024.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/provisioned-migration\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#16\",\n \"claim\": \"Den nye hourly + Azure Reservation-modellen dekker alle modeller, inkludert gpt-5.1 og o-serien.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry-classic/openai/concepts/provisioned-migration\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#17\",\n \"claim\": \"Priority processing er pay-per-token med latensmål og kan ikke reserveres.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#18\",\n \"claim\": \"Reservasjoner støttes ikke for serverless SKU-er som Azure SQL Serverless og Cosmos DB Serverless — kun pay-as-you-go.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#19\",\n \"claim\": \"Autorenew for reservasjoner er opt-in og ikke på som standard; en erstatningsreservasjon har autorenew av som standard, mens fornyelse kan skje på samme reservasjons-ordre-ID.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#20\",\n \"claim\": \"PTU Capacity Calculator er innebygd i Microsoft Foundry og tilgjengelig i deployment workflow, med input Input TPM | Output TPM | Peak calls per minute | Tokens per prompt | Tokens per response og output Recommended PTUs (avrundet til scale increment) | Raw PTUs.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-throughput-sizing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#21\",\n \"claim\": \"For gpt-5.1 teller output-tokens 8x input-tokens i PTU-sizing-ratioen.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-throughput-sizing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#22\",\n \"claim\": \"PTU-sizing regner 4750 input-TPM per PTU (300K normalisert TPM / 4750 = ~63 rå PTU).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-throughput-sizing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#23\",\n \"claim\": \"Regional Provisioned PTU garanterer region for data residency, Data Zone Provisioned holder data innenfor EU Data Boundary, og med Global Provisioned kan data forlate EU.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#24\",\n \"claim\": \"RBAC-roller for reservasjoner: Reservation Purchaser (kjøpe reservasjoner) | Owner på subscription (administrere reservations scope) | Billing Account Admin (EA) (aktivere Reserved Instances-policy).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/azure-openai\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#25\",\n \"claim\": \"Azure OpenAI og Cognitive Services krever ingen egen lisensiering (consumption-based); reservasjoner anvendes automatisk basert på scope for Azure OpenAI, mens commitment tier kjøpes per ressurs for Cognitive Services.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/reserved-capacity-planning.md#26\",\n \"claim\": \"M365 Copilot krever M365 E3/E5 + Copilot-lisens og har ikke PTU-modell eller reservasjoner — kapasitet er inkludert i per-user-lisensen.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/reserved-capacity-planning.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/semantic-caching-patterns.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/semantic-caching-patterns.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/semantic-caching-patterns.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#1\",\n \"claim\": \"Semantic caching for LLM-APIer i Azure API Management har status GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#2\",\n \"claim\": \"Embeddings-modeller som brukes til semantic caching via Azure OpenAI Embeddings API er text-embedding-3-large | text-embedding-ada-002.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#3\",\n \"claim\": \"Vector-database-alternativer for semantic caching på Azure er Azure Managed Redis (RediSearch) | Azure Cache for Redis Enterprise | Azure AI Search.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/cache-overview-vector-similarity\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#4\",\n \"claim\": \"Azure Redis støtter disse similarity-metrikkene for vektorsøk: COSINE (cosine similarity) | L2 (euclidean distance) | IP (inner product).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/cache-overview-vector-similarity\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#5\",\n \"claim\": \"APIM-policyene llm-semantic-cache-lookup og azure-openai-semantic-cache-lookup bruker score-threshold som en semantisk avstand (prompts med score over terskelen bruker ikke cachen), og Microsofts eget eksempel bruker score-threshold=\\\"0.15\\\".\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching#configure-semantic-caching-policies\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#6\",\n \"claim\": \"Alle Azure API Management-tiers støtter semantic caching-mønsteret med Azure Managed Redis.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#7\",\n \"claim\": \"Azure Managed Redis med RediSearch-modulen er påkrevd for semantic caching i APIM-mønsteret.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/tutorial-semantic-cache\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#8\",\n \"claim\": \"APIM semantic cache støtter multi-tenant partisjonering via vary-by på subscription | header | claim.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#9\",\n \"claim\": \"Embeddings-modellenes dimensjoner: text-embedding-3-small = 1536 | text-embedding-3-large = 3072 | text-embedding-ada-002 = 1536.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#10\",\n \"claim\": \"text-embedding-ada-002 er legacy og bør unngås for nye prosjekter.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#11\",\n \"claim\": \"Azure Managed Redis for semantic caching opprettes med SKU Enterprise_E10 og modulen RediSearch (az redis create --sku Enterprise_E10 --redis-module RediSearch).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/tutorial-semantic-cache\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#12\",\n \"claim\": \"Parametrene i APIM semantic cache-policyene er score-threshold | embeddings-backend-id | embeddings-backend-auth | ignore-system-messages | max-message-count | vary-by | duration.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching#configure-semantic-caching-policies\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#13\",\n \"claim\": \"Microsoft Foundry-modeller (via Model Inference API) støttes i APIM med de generiske policyene llm-semantic-cache-lookup og llm-semantic-cache-store.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#14\",\n \"claim\": \"Azure Managed Redis og Azure OpenAI er tilgjengelig i regionene Norway East | Norway West, slik at data kan forbli i Norge/EU.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#15\",\n \"claim\": \"Azure Managed Redis-tiers for semantic caching er Memory Optimized 1GB | Memory Optimized 10GB | Memory Optimized 50GB | Compute Optimized.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/tutorial-semantic-cache\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#16\",\n \"claim\": \"Azure OpenAI lisensieres pay-per-token som PTU eller Consumption, uten ekstra lisenser for caching.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/genai-gateway-capabilities\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#17\",\n \"claim\": \"Azure API Management inkluderer semantic cache-policyene fra tier Basic v2 og oppover.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/api-management/azure-openai-enable-semantic-caching\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/semantic-caching-patterns.md#18\",\n \"claim\": \"Azure Managed Redis faktureres pay-per-hour per tier, og RediSearch er inkludert i Enterprise-tier.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/redis/tutorial-semantic-cache\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/semantic-caching-patterns.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/small-language-models-economics.md",
|
||
"claim_count": 29,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/small-language-models-economics.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/small-language-models-economics.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#1\",\n \"claim\": \"Eksempler på SLM-er er Phi-4-mini (3.8B) | Phi-3-small (7B) | Falcon-7B, mens eksempler på LLM-er er GPT-4o | Llama-3.3-70B.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/concepts-ai-ml-language-models\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#2\",\n \"claim\": \"Microsofts Phi-serie av små språkmodeller består av Phi-4-mini | Phi-4-multimodal | Phi-3-small | Phi-3-medium | Phi-2.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#3\",\n \"claim\": \"Phi-4-mini har 3,8 milliarder parametere og støtter en input-lengde på 131 072 tokens.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#4\",\n \"claim\": \"Phi-4-mini er GA i Azure med Global Standard-deployment.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#5\",\n \"claim\": \"Phi-4-multimodal støtter 131 072 tokens input som kombinerer tekst, bilde og lyd.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#6\",\n \"claim\": \"Phi-4-multimodal er GA i Microsoft Foundry.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#7\",\n \"claim\": \"Phi-3-small har 7 milliarder parametere og støtter 128 000 tokens input-lengde.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#8\",\n \"claim\": \"Phi-3-small, Phi-3-medium og Phi-2 er alle GA i Azure.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#9\",\n \"claim\": \"Phi-3-medium har 14 milliarder parametere og støtter 128 000 tokens input-lengde.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#10\",\n \"claim\": \"Phi-2 har 2,7 milliarder parametere og støtter 2 048 tokens input-lengde.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#11\",\n \"claim\": \"Deployment-alternativene for SLM-er i Azure er Microsoft Foundry (Serverless) | Azure App Service Sidecar | Azure Kubernetes Service (AKS) + KAITO | On-premises (Ollama, ONNX Runtime) | Edge/IoT.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#12\",\n \"claim\": \"Azure App Service Sidecar-deployment av SLM krever P3MV3-tier.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/app-service/scenario-ai-local-small-language-model\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#13\",\n \"claim\": \"Azure App Service støtter Phi-4 sidecar extensions direkte via portalen, med OpenAI-kompatibelt API på localhost:11434.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/app-service/scenario-ai-local-small-language-model\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#14\",\n \"claim\": \"Sidecar-utvidelsen for SLM i Azure App Service velges som «AI: phi-4-q4-gguf (Experimental)» i Deployment Center, og SLM-en eksponeres på http://localhost:11434/v1/chat/completions.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/app-service/scenario-ai-local-small-language-model\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#15\",\n \"claim\": \"Azure App Service Phi-4 sidecar er GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/app-service/scenario-ai-local-small-language-model\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#16\",\n \"claim\": \"Azure App Service Phi-4 sidecar støtter rammeverkene ASP.NET Core | FastAPI | Spring Boot | Express.js.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/app-service/tutorial-ai-slm-dotnet\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#17\",\n \"claim\": \"Microsoft Foundry tilbyr deployment-typene Serverless API | Managed Online Endpoints | Global Standard, der Global Standard gir fungibel kvote på tvers av regioner.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#18\",\n \"claim\": \"Managed Online Endpoints i Microsoft Foundry krever dedikert VM av typen Standard_DS3_v2 eller bedre.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#19\",\n \"claim\": \"Phi-4-mini støtter 131 072 tokens input og 4 096 tokens output.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#20\",\n \"claim\": \"Phi-4-mini på AKS krever GPU av typen T4 eller A100, der T4 anbefales av kostnadshensyn.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#21\",\n \"claim\": \"Phi-3-small på AKS krever A100-GPU.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#22\",\n \"claim\": \"A100-GPU for SLM-deployment er regionalt tilgjengelig i West US | West US 3 | Sweden Central | Australia East, mens T4 er tilgjengelig i West Europe.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#23\",\n \"claim\": \"KAITO (Kubernetes AI Toolchain Operator) støtter Phi-4-mini med automatisk GPU-node-provisjonering på AKS.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#24\",\n \"claim\": \"Ollama på Azure VM anbefales kjørt på Standard_D4s_v3 eller bedre.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#25\",\n \"claim\": \"Phi-3 er tilgjengelig som ONNX-modell (phi-3-mini-4k-instruct-onnx) på Hugging Face og kan kjøres på CPU via ONNX Runtime.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#26\",\n \"claim\": \"Azure har Norge-regioner i Oslo og Stavanger (norwayeast) der SLM kan deployes via Azure App Service eller AKS.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#27\",\n \"claim\": \"AKS-hosting av SLM med T4-GPU bruker VM-SKU-en Standard_NC4as_T4_v3.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#28\",\n \"claim\": \"AKS-hosting av SLM med A100-GPU bruker VM-SKU-en Standard_NC24ads_A100_v4.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/aks/ai-toolchain-operator\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/small-language-models-economics.md#29\",\n \"claim\": \"Modell-lisensene er Phi-4-mini: MIT | Phi-4-multimodal: MIT | Phi-3 (alle): MIT | Phi-2: MIT | Falcon-7B: Apache 2.0 | Llama-3.3-70B: Meta (custom, ingen redistribusjon uten avtale).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/small-language-models-economics.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/cost-optimization/vector-storage-cost-optimization.md",
|
||
"claim_count": 36,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/cost-optimization/vector-storage-cost-optimization.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/cost-optimization/vector-storage-cost-optimization.md`)\n\n[\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#1\",\n \"claim\": \"text-embedding-ada-002 genererer vektorer på 1536 dimensjoner, mens text-embedding-3-large gir opptil 3072 dimensjoner, der hver dimensjon lagres som 32-bit flyttall (float32).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#2\",\n \"claim\": \"Embedding-modellenes standarddimensjoner: text-embedding-ada-002 = 1536 | text-embedding-3-small = 1536 (default) | text-embedding-3-large = 3072 (default).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#3\",\n \"claim\": \"text-embedding-3-small og text-embedding-3-large støtter MRL, og text-embedding-3-large støtter truncation av dimensjoner.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-truncate-dimensions\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#4\",\n \"claim\": \"Azure AI Search støtter tre komprimeringsmetoder for vektorer: scalar quantization (float32 → int8, 4x reduksjon) | binary quantization (float32 → 1 bit, opptil 28x reduksjon) | float16 (float32 → float16, 2x reduksjon).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#5\",\n \"claim\": \"Scalar quantization krever original float32-vektorer for rescoring, mens binary quantization kan bruke dot-product til rescoring og float16 ikke krever rescoring.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#6\",\n \"claim\": \"Azure AI Search-benchmark: baseline float32 gir 21.36 MB storage og 4.83 MB vector index; scalar quantization gir 17.76 MB storage og 1.22 MB vector index (75 % reduksjon); binary quantization gir 4.92 MB storage og 1.22 MB vector index (77 % total reduksjon).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-configure-compression-storage\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#7\",\n \"claim\": \"Alle komprimeringsteknikker kombinert gir 4.92 MB storage og 1.22 MB vector index, tilsvarende 92,5 % reduksjon i vector index-størrelse.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-configure-compression-storage\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#8\",\n \"claim\": \"Matryoshka Representation Learning (MRL) er innebygd i text-embedding-3-modellene, slik at dimensjoner kan trunkeres fra 3072 → 1024 eller 1536 → 512.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-truncate-dimensions\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#9\",\n \"claim\": \"MRL-trunkering: text-embedding-3-large 3072 → 1024 gir 3x lagringsreduksjon (~95 % av original MTEB-score) og text-embedding-3-small 1536 → 512 gir 3x lagringsreduksjon (~92 % av original).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-truncate-dimensions\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#10\",\n \"claim\": \"Anbefalt minstegrense ved bruk av binary quantization er 1024 dimensjoner; under 1000 dimensjoner gir merkbar kvalitetsforringelse.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#11\",\n \"claim\": \"Azure AI Search lagrer vektorer i to kopier: index copy (i minne, brukt til query execution) | stored copy (på disk, brukt til retrieval i query response).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-storage-options\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#12\",\n \"claim\": \"Ved å sette `stored: false` i Azure AI Search kan man spare opptil 50 % disklagring, men mister muligheten til å returnere vektorer i query-responser.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-storage-options\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#13\",\n \"claim\": \"Azure AI Search tilbyr to vector index-algoritmer: HNSW (høyt minnekrav, graf i minne, 20-50 ms query-latens på standard tier) | Exhaustive KNN (lavt minnekrav, paged loading, høyere latens).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#14\",\n \"claim\": \"HNSW krever at hele grafen ligger i minne og driver opp vector quota-forbruk, mens Exhaustive KNN laster data on-demand og teller ikke mot vector quota.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#15\",\n \"claim\": \"Azure AI Search vectorSearch-compression med `kind: binaryQuantization` støtter feltene rescoringOptions.enableRescoring | rescoringOptions.defaultOversampling | rescoringOptions.rescoreStorageMethod (`discardOriginals`) | truncationDimension.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-configure-compression-storage\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#16\",\n \"claim\": \"Azure AI Search støtter felttypen `Collection(Edm.Half)` for float16-vektorer og `Collection(Edm.Single)` for float32-vektorer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#17\",\n \"claim\": \"Scalar quantization i Azure AI Search konfigureres med `kind: scalarQuantization` og scalarQuantizationParameters.quantizedDataType = `int8`, med rescoreStorageMethod `preserveOriginals`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#18\",\n \"claim\": \"Default oversampling i Azure AI Search er 4; anbefalt verdi for binary quantization er 10-20.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#19\",\n \"claim\": \"Azure AI Search-tjenester opprettet etter april 2024 har høyere vector quotas enn eldre tjenester.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#20\",\n \"claim\": \"Vector quantization i Azure AI Search har vært GA siden api-version 2024-07-01.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#21\",\n \"claim\": \"Azure AI Search REST-API for oppretting av indeks og vektorsøk bruker `api-version=2025-09-01`.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-configure-compression-storage\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#22\",\n \"claim\": \"HNSW-algoritmen i Azure AI Search konfigureres med hnswParameters: m | efConstruction | metric (f.eks. `cosine`).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#23\",\n \"claim\": \"Azure OpenAI embeddings med dimensions-parameter (MRL) brukes med api_version `2024-02-01`.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#24\",\n \"claim\": \"Cosmos DB for MongoDB vCore støtter HNSW og IVF vector indexing med half-precision (float16) via cosmosSearchOptions med kind `vector-hnsw` og compression `half`.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#25\",\n \"claim\": \"Vector extension i Azure SQL Database er i preview, støtter float32-vektorer, men ikke native quantization.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#26\",\n \"claim\": \"Azure AI Search støtter regionene Norway East og Norway West med full data residency.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#27\",\n \"claim\": \"Azure OpenAI i Norway East støtter text-embedding-3-modellene for embedding-generering.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#28\",\n \"claim\": \"Vector quota per partition for Azure AI Search-tjenester opprettet etter april 2024: Basic = 5 GB | S1 = 35 GB | S2 = 150 GB | S3 = 300 GB.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#29\",\n \"claim\": \"Azure AI Search tilbyr de dedikerte tierne Basic | S1 | S2 | S3 for vektorindekser.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#30\",\n \"claim\": \"Azure AI Search tilbyr en Serverless-prismodell (forbruksbasert: Compute Units/time + per-GB/mnd lagring) ved siden av de dedikerte tierne, og den er i preview per juni 2026 uten SLA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#31\",\n \"claim\": \"Azure AI Search Serverless er i preview i regionene West Central US | Switzerland North | Japan East.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#32\",\n \"claim\": \"Azure AI Search Serverless mangler funksjonene index aliases | debug sessions | shared private link, og støtter ikke migrering til eller fra Dedicated.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#33\",\n \"claim\": \"Bruk av vector search som grunnlag for Copilot for Microsoft 365 krever Microsoft 365 E3/E5 pluss Copilot-lisens, og Azure AI Search er ikke inkludert i Copilot-lisensen.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#34\",\n \"claim\": \"Microsoft Foundry unified billing er i preview (februar 2026).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/foundry/concepts/manage-costs\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#35\",\n \"claim\": \"Vector quota i Azure AI Search kan monitoreres via Azure Portal eller `Get Index Statistics`-API-et, og indeksering blokkeres ved quota-overskridelse.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n },\n {\n \"id\": \"ms-ai-security/cost-optimization/vector-storage-cost-optimization.md#36\",\n \"claim\": \"HNSW-overhead utgjør 1-20 % av rå vektorstørrelse, avhengig av dimensjoner og `m`-parameteren.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/en-us/azure/search/vector-search-index-size\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/cost-optimization/vector-storage-cost-optimization.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/async-processing-patterns.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/async-processing-patterns.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/async-processing-patterns.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#1\",\n \"claim\": \"Azure OpenAI tilbyr flere innebygde asynkrone mekanismer: Batch API for store volum | Background Tasks i Responses API for langvarige oppgaver | Webhooks for hendelsesbasert leveranse.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#2\",\n \"claim\": \"Egne asynkrone AI-arkitekturer kan bygges med disse Azure-mellomlagene: Azure Service Bus | Azure Queue Storage | Azure Event Hubs.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#3\",\n \"claim\": \"Kjernekomponentene for asynkron AI-prosessering er: Azure Service Bus (enterprise message broker med køer og topics) | Azure Queue Storage (enkel meldingskø) | Azure Event Hubs (høy-throughput event streaming) | Azure Functions (serverless compute for kø-triggered prosessering) | Batch API (Azure OpenAI) | Background Tasks i Responses API (Azure OpenAI) | Webhooks (Azure OpenAI).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#4\",\n \"claim\": \"Modellnavnet gpt-4o brukes som deployment/modell i Azure OpenAI chat completions-kall.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#5\",\n \"claim\": \"Azure OpenAI-klienten instansieres med api_version=\\\"2024-10-21\\\".\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#6\",\n \"claim\": \"o3 er en reasoning-modell som kan kalles via Azure OpenAI Responses API (client.responses.create) og kan ta flere minutter.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#7\",\n \"claim\": \"Azure OpenAI Responses API støtter parameteren background=True for å kjøre en forespørsel asynkront (background mode).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#8\",\n \"claim\": \"En background-task hentet med client.responses.retrieve har statusverdier som inkluderer: completed | failed.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#9\",\n \"claim\": \"Azure OpenAI webhook-kall sender HTTP-headerne Webhook-Signature | Webhook-ID, som brukes til signaturverifisering og idempotens.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/webhooks\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#10\",\n \"claim\": \"Azure OpenAI webhook-hendelser inkluderer event-typene batch.completed | fine_tuning.job.succeeded.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/webhooks\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#11\",\n \"claim\": \"Microsoft dokumenterer to primære topologier for event-drevet arkitektur: broker-topologi (events publiseres direkte til broker, f.eks. Azure Event Hubs + Service Bus) | mediator-topologi (sentral mediator koordinerer workflow, f.eks. Azure Durable Functions).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#12\",\n \"claim\": \"Rollefordelingen mellom Azure-meldingstjenestene er: Azure Event Hubs = durable event stream (log) | Azure Event Grid = publish-subscribe, reaktiv | Azure Service Bus = message queue med garantert levering, retry og dead-letter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#13\",\n \"claim\": \"Azure OpenAI Batch API gir 50 % kostnadsreduksjon sammenlignet med ordinær (ikke-batch) prosessering.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#14\",\n \"claim\": \"Background Tasks API er den anbefalte asynkrone mekanismen for reasoning-modellene o3 og o1.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/async-processing-patterns.md#15\",\n \"claim\": \"Azure Service Bus gir funksjoner som Queue Storage ikke har: sessions | dead letter queues | transaksjonsstøtte.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/guide/architecture-styles/event-driven\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/async-processing-patterns.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/batch-api-usage-optimization.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/batch-api-usage-optimization.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/batch-api-usage-optimization.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#1\",\n \"claim\": \"Azure OpenAI Batch API har status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#2\",\n \"claim\": \"Batch API har en målsatt leveringstid (completion window) på 24 timer.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#3\",\n \"claim\": \"Maksimal filstørrelse for en batch-fil er 200 MB ved direkte opplasting og 1 GB via Azure Blob Storage.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#4\",\n \"claim\": \"Maksimalt antall requests per batch-fil er 100 000.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#5\",\n \"claim\": \"Maks antall batch-filer per ressurs er 500 uten utløpsdato og 10 000 med utløpsdato.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#6\",\n \"claim\": \"Batch API bruker en separat enqueued token-kvote, atskilt fra online/standard TPM-kvote.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#7\",\n \"claim\": \"Modeller som støttes av Batch API: GPT-4o | GPT-4o mini | GPT-4.1 | o3-mini (m.fl.).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#8\",\n \"claim\": \"Deployment-typene for Batch API er GlobalBatch | DataZoneBatch.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#9\",\n \"claim\": \"Batch API støtter endepunktene /v1/chat/completions og det nyere /v1/responses (Responses API-format) i JSONL-requestene.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#10\",\n \"claim\": \"Batch-operasjoner (filopplasting og batch-jobber) mot Azure OpenAI bruker api-version 2025-03-01-preview.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#11\",\n \"claim\": \"Filer over 200 MB og opptil 1 GB må lastes opp via Azure Blob Storage (BYOS) i stedet for direkte opplasting.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#12\",\n \"claim\": \"Batch-inputfiler kan enten lagres uten utløp eller settes med utløpstid på 14–30 dager (expires_after).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#13\",\n \"claim\": \"Batch-statusflyten består av statusene validating | in_progress | completed | failed | cancelled | expired.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#14\",\n \"claim\": \"Global Batch kan prosessere data i enhver Azure OpenAI-region, DataZoneBatch begrenser prosesseringen til EU-regioner, og Regional Batch brukes ved strengeste datasuverenitetskrav.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/batch-api-usage-optimization.md#15\",\n \"claim\": \"Completion window på 24 timer er et mål og ikke en garanti; jobber som tar lengre tid utløper ikke, men kan kanselleres med resultater for allerede fullført arbeid.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch-blob-storage\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/batch-api-usage-optimization.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/concurrent-request-optimization.md",
|
||
"claim_count": 7,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/concurrent-request-optimization.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/concurrent-request-optimization.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#1\",\n \"claim\": \"Referansefilen angir status GA for concurrent request optimization i Azure OpenAI (Azure AI Foundry).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#2\",\n \"claim\": \"Azure OpenAI tilbyr to deployment-typer som er relevante for samtidighet: Standard | PTU (provisioned throughput units).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/quota\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#3\",\n \"claim\": \"Kvote for Azure OpenAI tildeles som TPM (tokens per minutt) og RPM (requests per minutt).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/quota\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#4\",\n \"claim\": \"For Standard-deployments bestemmer RPM-kvoten den harde grensen for antall samtidige forespørsler per minutt.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/quota\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#5\",\n \"claim\": \"For PTU-deployments er grensen definert av utilization: når prosessert kapasitet nærmer seg 100 % av tildelte PTUs, begynner tjenesten å returnere 429-feil.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#6\",\n \"claim\": \"Modellen gpt-4o er tilgjengelig som deployment-/modellnavn i Azure OpenAI chat completions.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/concurrent-request-optimization.md#7\",\n \"claim\": \"Microsoft Learn-dokumentasjonen «Performance and latency» for Azure OpenAI dekker samtidige forespørsler (concurrent requests) og throughput.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/concurrent-request-optimization.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/connection-pooling-patterns.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/connection-pooling-patterns.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/connection-pooling-patterns.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#1\",\n \"claim\": \"Referansefilens emne (connection pooling mot Azure AI Services / Azure OpenAI) er merket med status GA.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#2\",\n \"claim\": \"SocketsHttpHandler er den underliggende socket-håndtereren med connection pool i .NET 6+.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/dotnet/fundamentals/networking/http/httpclient-guidelines\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#3\",\n \"claim\": \"Azure OpenAI SDK-en har innebygd connection management og leveres som pakken azure-ai-openai.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#4\",\n \"claim\": \"SocketsHttpHandler eksponerer pool- og keep-alive-egenskapene MaxConnectionsPerServer | PooledConnectionLifetime | PooledConnectionIdleTimeout | KeepAlivePingPolicy | KeepAlivePingDelay | KeepAlivePingTimeout | EnableMultipleHttp2Connections | AutomaticDecompression.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/dotnet/fundamentals/networking/http/httpclient-guidelines\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#5\",\n \"claim\": \"HttpKeepAlivePingPolicy har verdiene WithActiveRequests | Always.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/dotnet/fundamentals/networking/http/httpclient-guidelines\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#6\",\n \"claim\": \"IHttpClientFactory-registrering støtter SetHandlerLifetime for å rotere handleren (handler lifetime).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/aspnet/core/performance/performance-best-practices\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#7\",\n \"claim\": \"Azure OpenAI .NET-klienten heter Azure.AI.OpenAI.AzureOpenAIClient og konstrueres med Uri + Azure.AzureKeyCredential.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#8\",\n \"claim\": \"Azure OpenAI støtter HTTP/2, som muliggjør multipleksing av flere forespørsler over én TCP-forbindelse.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#9\",\n \"claim\": \"Azure OpenAI Python SDK bruker api_version-strengen \\\"2024-10-21\\\".\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#10\",\n \"claim\": \"Azure OpenAI returnerer HTTP 429 med retry-after ved kvoteoverskridelse, og 5xx-statuskoder ved backend-feil.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#11\",\n \"claim\": \"APIM-policyen set-backend-service kan peke på en backend via backend-id-attributtet.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#12\",\n \"claim\": \"APIM <retry>-policyen støtter attributtene condition | count | interval | first-fast-retry, og kan inneholde <forward-request timeout=\\\"…\\\" />.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#13\",\n \"claim\": \"Microsoft anbefaler fire APIM-backendtopologier for connection pooling mot Azure OpenAI: Single backend | Multi-backend single region (weighted round-robin) | Multi-subscription | Multi-region.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#14\",\n \"claim\": \"Azure Norway East kan brukes som primær region for Azure OpenAI med failover til Sweden Central, i tråd med Schrems II-kravene.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/connection-pooling-patterns.md#15\",\n \"claim\": \"DNS TTL for privatelink-soner (Private Link) er typisk 10 sekunder.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/connection-pooling-patterns.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/gpu-compute-sizing.md",
|
||
"claim_count": 20,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/gpu-compute-sizing.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/gpu-compute-sizing.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#1\",\n \"claim\": \"Azure tilbyr GPU-akselererte VM-serier for AI: NC-serien (NVIDIA T4) for inferens | ND-serien (NVIDIA A100/H100) for trening | NV-serien for visualisering.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#2\",\n \"claim\": \"Azure VM-SKU NC4as_T4_v3 har 1x NVIDIA T4 med 16 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#3\",\n \"claim\": \"Azure VM-SKU NC24ads_A100_v4 har 1x NVIDIA A100 med 80 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#4\",\n \"claim\": \"Azure VM-SKU NC96ads_A100_v4 har 4x NVIDIA A100 med totalt 320 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#5\",\n \"claim\": \"Azure VM-SKU ND96asr_v4 har 8x NVIDIA A100 (40 GB hver) med totalt 320 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#6\",\n \"claim\": \"Azure VM-SKU ND96isr_H100_v5 har 8x NVIDIA H100 med totalt 640 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#7\",\n \"claim\": \"Azure VM-SKU NC40ads_H100_v5 har 1x NVIDIA H100 med 80 GB GPU-minne.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#8\",\n \"claim\": \"Azure ML-instanstypen Standard_NC48ads_A100_v4 gir 2x A100 80 GB og brukes til modeller på rundt 70B parametere.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/how-to-deploy-online-endpoints\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#9\",\n \"claim\": \"Modellen Llama-3.3-70B-Instruct er tilgjengelig i Azure ML-registeret azureml-meta (azureml://registries/azureml-meta/models/Llama-3.3-70B-Instruct).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/how-to-deploy-online-endpoints\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#10\",\n \"claim\": \"For gpt-4o gir én PTU 2 500 input-TPM.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#11\",\n \"claim\": \"For gpt-4.1 gir én PTU 3 000 input-TPM.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#12\",\n \"claim\": \"Minste PTU-deployment er 50 PTU-enheter.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#13\",\n \"claim\": \"Azure ML Online Endpoints har to deployment-typer: Managed Online Endpoint (Azure-administrert infrastruktur) | Kubernetes Online Endpoint (kundeeid K8s-kluster).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/machine-learning/how-to-deploy-online-endpoints\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#14\",\n \"claim\": \"Azure ML GPU-instanstypen Standard_NC6s_v3 har 1x V100 med 16 GB VRAM.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#15\",\n \"claim\": \"Azure ML GPU-instanstypen Standard_NC24s_v3 har 4x V100 med totalt 64 GB VRAM.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#16\",\n \"claim\": \"Azure ML GPU-instanstypen Standard_ND96amsr_A100_v4 har 8x A100 med totalt 640 GB VRAM.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#17\",\n \"claim\": \"Azure Reserved Instances tilbys med 1-3 års binding og gir 40-60 % besparelse på forutsigbare GPU VM-workloads.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#18\",\n \"claim\": \"GPU VM-er er tilgjengelige i Azure-regionen Norway East for self-hosted modeller.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/virtual-machines/sizes-gpu\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#19\",\n \"claim\": \"Azure OpenAI PTU-deployments finnes i variantene regional | data zone | global.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/gpu-compute-sizing.md#20\",\n \"claim\": \"gpt-4.1-nano gir 59 400 input-TPM per PTU.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/gpu-compute-sizing.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/load-testing-ai-services.md",
|
||
"claim_count": 13,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/load-testing-ai-services.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/load-testing-ai-services.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#1\",\n \"claim\": \"Microsoft tilbyr to offisielle verktøy for lasttesting av Azure AI Services: Azure Load Testing (JMeter-basert managed service) | azure-openai-benchmark (CLI-verktøy spesifikt for Azure OpenAI).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#2\",\n \"claim\": \"Faktisk throughput for Provisioned Throughput Units (PTU) avhenger av workload shape, som består av forholdet mellom input- og output-tokens | call rate | cache match rate.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#3\",\n \"claim\": \"Kjernekomponentene for lasttesting av AI-tjenester er Azure Load Testing (JMeter) | azure-openai-benchmark | Azure Monitor | Application Insights | Performance Optimizer.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#4\",\n \"claim\": \"Azure Load Testing inkluderer Performance Optimizer for ytelsesoptimalisering av Azure Functions.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#5\",\n \"claim\": \"Azure Load Testing bruker YAML-konfigurasjonsfil med skjemaversjon v0.1 (felt: testId | testPlan | engineInstances | configurationFiles | failureCriteria | env | secrets).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#6\",\n \"claim\": \"Azure OpenAI API-versjonen 2024-10-21 brukes for chat completions mot en gpt-4o-deployment.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#7\",\n \"claim\": \"TPM per PTU ifølge Microsoft-dokumentasjonen: gpt-4o = 2500 | gpt-4o-mini = 37000 | gpt-4.1 = 3000 | gpt-4.1-mini = 14900 | gpt-4.1-nano = 59400 | o3 = 3000.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#8\",\n \"claim\": \"Azure OpenAI provisioned-deployments finnes i typene Global Provisioned | Data Zone Provisioned | Regional Provisioned, der Global og Data Zone Provisioned er anbefalt default.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#9\",\n \"claim\": \"Global og Data Zone Provisioned krever minimum 15 PTU med økning i trinn på 5 PTU for alle GPT- og o-modeller.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#10\",\n \"claim\": \"Regional Provisioned krever minimum 50 PTU for gpt-4.1, o3 og gpt-4o, og 25 PTU for mini-/nano-modeller.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#11\",\n \"claim\": \"Det offisielle benchmarking-verktøyet azure-openai-benchmark installeres som pip-pakken azure-openai-benchmark.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#12\",\n \"claim\": \"azure-openai-benchmark støtter kommandolinjeflaggene --api-key | --api-base-endpoint | --deployment | --shape-profile (f.eks. balanced) | --clients | --duration | --output-format | --output | --context-tokens | --max-tokens | --rate.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/provisioned-get-started#run-a-benchmark\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/load-testing-ai-services.md#13\",\n \"claim\": \"Azure OpenAI tilbyr deployment-typen Global Standard, som er kostnadseffektiv for lasttesting i separate testdeployments.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/load-testing-ai-services.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/model-distillation-performance.md",
|
||
"claim_count": 18,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/model-distillation-performance.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/model-distillation-performance.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#1\",\n \"claim\": \"Model distillation i Azure OpenAI / Microsoft Foundry er GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#2\",\n \"claim\": \"Typiske teacher-modeller for distillasjon i Azure OpenAI er GPT-4o | o3, og typiske student-modeller er GPT-4o-mini | GPT-4.1-nano.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#3\",\n \"claim\": \"Microsoft Foundry tilbyr en integrert distillation-pipeline via Stored Completions-funksjonen, der produksjonsforespørsler og -svar lagres automatisk og konverteres til fine-tuning-datasett.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/stored-completions\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#4\",\n \"claim\": \"Kjernekomponentene i distillasjonsflyten er Stored Completions (Microsoft Foundry) | Fine-tuning API (Azure OpenAI, LoRA-basert) | Evaluation Framework (Microsoft Foundry Evaluations) | Teacher Model (GPT-4o, o3, GPT-5) | Student Model (GPT-4o-mini, GPT-4.1-nano).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/stored-completions\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#5\",\n \"claim\": \"Azure OpenAI-klienten bruker api_version «2024-12-01-preview» for å aktivere stored completions (store=True) mot teacher-modellen.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/stored-completions\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#6\",\n \"claim\": \"Distillasjon krever minimum 10 stored completions som treningsdata, mens Microsoft anbefaler 500-1000+ (hundrevis til tusenvis) for best resultat.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/stored-completions\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#7\",\n \"claim\": \"GPT-5 har 4 750 input-TPM per PTU og et latens-mål på 50 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#8\",\n \"claim\": \"GPT-4.1 har 3 000 input-TPM per PTU og et latens-mål på 80 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#9\",\n \"claim\": \"GPT-4o har 2 500 input-TPM per PTU og et latens-mål på 25 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#10\",\n \"claim\": \"GPT-4.1-mini har 14 900 input-TPM per PTU og et latens-mål på 90 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#11\",\n \"claim\": \"GPT-4o-mini har 37 000 input-TPM per PTU og et latens-mål på 33 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#12\",\n \"claim\": \"GPT-4.1-nano har 59 400 input-TPM per PTU og et latens-mål på 100 TPS.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#13\",\n \"claim\": \"Treningsdata brukt til fine-tuning kan ikke eksporteres fra Microsoft Foundry.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/fine-tuning\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#14\",\n \"claim\": \"Hosting av en fine-tuned modell i Azure OpenAI faktureres per time uavhengig av bruk, i motsetning til standard pay-per-token.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/fine-tuning\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#15\",\n \"claim\": \"Microsoft dokumenterer 10 seleksjonskriterier for valg av AI-modell: Task fit | Routing strategy | Cost | Context window | Security | Region | Deployment | Domain | Performance | Tunability.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/choose-ai-model\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#16\",\n \"claim\": \"Azure OpenAI-modeller kan deployes som PTU eller Standard, og studentmodellen deployes oftest som Standard til å begynne med.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/choose-ai-model\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#17\",\n \"claim\": \"Routing-eksempelet mot Microsoft Foundry-endepunktet bruker api_version «2024-10-21».\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/fine-tuning-considerations\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/model-distillation-performance.md#18\",\n \"claim\": \"Tunability (støtte for direkte fine-tuning) per modell: GPT-4.1-nano Ja | GPT-4o-mini Ja | GPT-4.1-mini Ja | GPT-4.1 Nei | GPT-4o Nei.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/fine-tuning\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/model-distillation-performance.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/performance-benchmarking-frameworks.md",
|
||
"claim_count": 13,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/performance-benchmarking-frameworks.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/performance-benchmarking-frameworks.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#1\",\n \"claim\": \"Microsoft tilbyr et offisielt benchmarking-verktøy kalt azure-openai-benchmark spesifikt for Azure OpenAI.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#2\",\n \"claim\": \"Microsoft tilbyr Azure Load Testing for bredere lasttesting.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#3\",\n \"claim\": \"Microsoft Foundry tilbyr innebygde evalueringsverktøy som kan brukes til å måle modellkvalitet.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/how-to/evaluate-generative-ai-app\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#4\",\n \"claim\": \"Kjernekomponentene i et benchmarking-rammeverk for Azure AI Services er: azure-openai-benchmark | Azure Load Testing | Microsoft Foundry Evaluations | Azure Monitor | Application Insights | Custom Benchmark Suite.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#5\",\n \"claim\": \"Azure Load Testing er en managed lasttestingtjeneste basert på JMeter.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/load-testing/overview-what-is-azure-load-testing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#6\",\n \"claim\": \"Azure Monitor brukes til metrikk-innsamling og visualisering for Azure OpenAI, og Application Insights til end-to-end request tracing.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/monitor-openai\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#7\",\n \"claim\": \"Azure OpenAI svarer med HTTP-statuskode 429 på forespørsler som strupes (throttles) ved overskredet kvote/rate limit.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#8\",\n \"claim\": \"Azure OpenAI støtter prompt cache, slik at en andel av input-tokens kan treffe cachen (prompt cache hit rate).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#9\",\n \"claim\": \"2024-10-21 er en gyldig api-version for Azure OpenAI.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#10\",\n \"claim\": \"SDK-versjon 1.x brukes for Azure OpenAI-klienten (OpenAI Python SDK 1.x).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#11\",\n \"claim\": \"«standard» er en gyldig deployment-type for Azure OpenAI-deployments.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#12\",\n \"claim\": \"norwayeast er en Azure-region som kan brukes for Azure OpenAI-deployment.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/performance-benchmarking-frameworks.md#13\",\n \"claim\": \"Det offisielle azure-openai-benchmark-verktøyet brukes for PTU-dimensjonering (Provisioned Throughput Units) i Azure OpenAI.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/performance-benchmarking-frameworks.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/prompt-caching-performance.md",
|
||
"claim_count": 12,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/prompt-caching-performance.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/prompt-caching-performance.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#1\",\n \"claim\": \"Prompt caching i Azure OpenAI har status GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#2\",\n \"claim\": \"Prompt caching utløses når de første 1024+ tokens i en prompt er identiske med en tidligere forespørsel.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#3\",\n \"claim\": \"Cached tokens faktureres med rabatt for Standard-deployments og med opptil 100 % rabatt for Provisioned (PTU)-deployments.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#4\",\n \"claim\": \"Prompt caching er automatisk aktivert for alle støttede modeller (GPT-4o og nyere) uten ekstra konfigurasjon.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#5\",\n \"claim\": \"Cachen er basert på en hash av de første ~256 tokens, krever minimum 1024 identiske tokens for å trigge, og etter den initiale 1024-token-terskelen caches ytterligere identiske tokens i blokker på 128.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#6\",\n \"claim\": \"Cacher tømmes typisk innen 5-10 minutter uten aktivitet og alltid innen 24 timer.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#7\",\n \"claim\": \"Azure OpenAI API eksponerer prompt_cache_key (valgfri parameter for cache-routing) | cached_tokens (felt i prompt_tokens_details i API-responsen som viser cache hits).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#8\",\n \"claim\": \"Prompt-cache deles ikke mellom Azure-abonnement; cachen er isolert per abonnement og deles ikke mellom kunder.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#9\",\n \"claim\": \"Modeller som støtter prompt caching: gpt-4o-* | gpt-4o-mini-* | gpt-4.1-* | gpt-4.1-mini-* | gpt-4.1-nano-* | o1-* | o3-* | o3-mini-*.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#10\",\n \"claim\": \"Operasjoner som støtter prompt caching: chat-completions | completions | responses | real-time.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#11\",\n \"claim\": \"Azure OpenAI-klienten bruker api_version «2024-12-01-preview».\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/prompt-caching-performance.md#12\",\n \"claim\": \"Mer enn ca. 15 RPM med samme prefix og samme prompt_cache_key kan overflow til andre maskiner og redusere cache-effektiviteten.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/prompt-caching-performance.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/rate-limit-management.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/rate-limit-management.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/rate-limit-management.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#1\",\n \"claim\": \"Azure OpenAI bruker to rate limit-mekanismer: Tokens-per-Minute (TPM) | Requests-per-Minute (RPM).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#2\",\n \"claim\": \"Når en TPM- eller RPM-grense overskrides, returnerer Azure OpenAI HTTP 429 (Too Many Requests) med en Retry-After-header som angir hvor mange sekunder klienten bør vente.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#3\",\n \"claim\": \"For Standard deployments er rate limits direkte koblet til den tildelte kvoten, mens Provisioned Throughput (PTU) deployments returnerer 429 når utilization overstiger 100 %.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#4\",\n \"claim\": \"Microsofts offisielle Azure OpenAI-SDK-er for Python | JavaScript har innebygd retry-logikk med eksponentiell backoff.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/supported-languages\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#5\",\n \"claim\": \"Python-SDK-en (AzureOpenAI-klienten) har max_retries med standardverdi 2, som kan overstyres per klient eller per forespørsel.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/supported-languages\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#6\",\n \"claim\": \"Azure OpenAI data plane-API-versjonen 2024-10-21 brukes i Python-SDK-klienten (api_version=\\\"2024-10-21\\\").\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#7\",\n \"claim\": \"Azure Management REST API for Microsoft.CognitiveServices-deployments (liste og oppdatere deployments) bruker api-version=2023-05-01.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/quota\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#8\",\n \"claim\": \"Ved kvotejustering av en Azure OpenAI-deployment settes SKU-navnet \\\"Standard\\\" med capacity angitt i tusen-enheter TPM (new_tpm // 1000).\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry-classic/openai/how-to/quota\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#9\",\n \"claim\": \"Azure OpenAI-endepunkter for multi-region failover konfigureres i regionene norwayeast | swedencentral | westeurope, med Norway East som førsteprioritet for norske kunder.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#10\",\n \"claim\": \"Azure Monitor-metrikker for Azure OpenAI under ResourceProvider MICROSOFT.COGNITIVESERVICES omfatter AzureOpenAIRequests | ProcessedPromptTokens | GeneratedCompletionTokens.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#11\",\n \"claim\": \"Standard-kvote for Azure OpenAI er på subscription-nivå, ikke på instansnivå; load balancing mellom standard-instanser i samme subscription gir ikke høyere gjennomstrømning — reell kvoteutvidelse krever separate subscriptions eller global/data zone deployments.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#12\",\n \"claim\": \"Kvote-kapasitet per gateway-topologi: Single instance = baseline TPM | Multi-backend, single region = 2-5x baseline | Multi-subscription = 5-20x baseline | Multi-region = nær ubegrenset.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#13\",\n \"claim\": \"Azure API Management har policyen azure-openai-token-limit med attributtene counter-key | tokens-per-minute | estimate-prompt-tokens | tokens-consumed-variable-name | remaining-tokens-variable-name for token-basert rate limiting foran Azure OpenAI.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#14\",\n \"claim\": \"Standard Azure OpenAI-deployments har ingen latens-SLA, og 429-feil er forventet atferd under høy belastning.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/rate-limit-management.md#15\",\n \"claim\": \"PTU-deployment (Provisioned Throughput) gir garantert kapasitet og eliminerer rate limiting innenfor tildelt kapasitet.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/quotas-limits\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/rate-limit-management.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/regional-deployment-latency.md",
|
||
"claim_count": 15,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/regional-deployment-latency.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/regional-deployment-latency.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#1\",\n \"claim\": \"Azure OpenAI er tilgjengelig i Azure Norway East, som anbefales som primærregion for norsk offentlig sektor, med Sweden Central som sekundærregion.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#2\",\n \"claim\": \"Azure OpenAI tilbyr seks deployment-typer: Global Standard | Data Zone Standard | Regional Standard | Global Provisioned | Data Zone Provisioned | Regional Provisioned.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#3\",\n \"claim\": \"Global Standard-deployment kan plassere data i hvilken som helst Azure-region og ruter automatisk til datasentre med ledig kapasitet.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#4\",\n \"claim\": \"Data Zone Standard-deployment holder data innenfor en geografisk sone (EU eller US) og ruter automatisk innen sonen.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#5\",\n \"claim\": \"Aktuelle Azure-regioner for norske virksomheter er norwayeast (data i Norge) | swedencentral (EU/EØS) | westeurope | northeurope (EU/EØS) | eastus | eastus2 | westus (US, utenfor EU/EØS).\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#6\",\n \"claim\": \"Azure Front Door har SKU-en Premium_AzureFrontDoor, som angis med --sku ved `az afd profile create`.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/frontdoor/front-door-overview\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#7\",\n \"claim\": \"Azure OpenAI REST API-kall mot /openai/deployments bruker api-version=2024-10-21.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#8\",\n \"claim\": \"For kravet «data prosesseres i EU» gjelder: Global Standard = nei (global), Data Zone (EU) = ja, Regional (Norway East) = ja.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#9\",\n \"claim\": \"Kun Regional deployment i Norway East lagrer data i Norge; Global Standard og Data Zone (EU) gjør det ikke.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#10\",\n \"claim\": \"Azure Front Door har 118+ edge-lokasjoner på tvers av 100+ metroområder globalt.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/frontdoor/front-door-overview\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#11\",\n \"claim\": \"Azure Front Door Premium støtter Private Link til origins, slik at trafikk kan rutes til Azure OpenAI uten offentlig eksponering av backend.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/frontdoor/front-door-overview\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#12\",\n \"claim\": \"Azure Front Door Premium inkluderer innebygd Web Application Firewall med managed rule sets | bot manager foran backend.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/frontdoor/front-door-overview\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#13\",\n \"claim\": \"Microsoft dokumenterer fire formelle gateway-topologier for Azure OpenAI: Multiple model deployments, single instance | Single region, multiple instances | Single region, multiple subscriptions | Multiple regions.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#14\",\n \"claim\": \"Topologien «Single region, multiple subscriptions» brukes til kvote-utvidelse via flere Azure-subscriptions ved høye TPM-kvotekrav.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/architecture/ai-ml/guide/azure-openai-gateway-multi-backend\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/regional-deployment-latency.md#15\",\n \"claim\": \"Data Zone-deployments finnes i variantene Standard og Provisioned og gir automatisk EU-routing med data residency-garanti.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/foundry-models/concepts/deployment-types\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/regional-deployment-latency.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/response-chunking-strategies.md",
|
||
"claim_count": 14,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/response-chunking-strategies.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/response-chunking-strategies.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#1\",\n \"claim\": \"Response chunking / streaming for Azure OpenAI er merket som GA (generelt tilgjengelig).\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/application-gateway/use-server-sent-events\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#2\",\n \"claim\": \"Når `stream: true` settes i Azure OpenAI API-kallet, returnerer tjenesten delta-oppdateringer som Server-Sent Events (SSE) ettersom tokens genereres.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#3\",\n \"claim\": \"Kjernekomponenter for response chunking er: Server-Sent Events (SSE, HTTP SSE) | stream_options (Azure OpenAI API) | Application Gateway (Azure App Gateway, SSE proxy og load balancing) | API Management (Azure APIM, SSE-støtte med policy-basert routing) | SignalR (Azure SignalR, real-time push til web-klienter).\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/application-gateway/use-server-sent-events\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#4\",\n \"claim\": \"Azure API Management støtter Server-Sent Events med policy-basert routing.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/api-management/how-to-server-sent-events\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#5\",\n \"claim\": \"Azure OpenAI-klienten (AzureOpenAI i Python) brukes med api_version «2024-10-21».\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#6\",\n \"claim\": \"Modellstrengen «gpt-4o» brukes som standard deployment/modell for chat completions-streaming.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#7\",\n \"claim\": \"Parameteret `stream_options={\\\"include_usage\\\": true}` gjør at token-bruk returneres i siste chunk av en streaming-respons.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#8\",\n \"claim\": \"Usage-objektet i en Azure OpenAI chat completions-respons inneholder feltene prompt_tokens | completion_tokens | total_tokens.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#9\",\n \"claim\": \".NET-streaming mot Azure OpenAI bruker namespacene Azure.AI.OpenAI og OpenAI.Chat med klassen AzureOpenAIClient og metoden GetChatClient(deploymentName).\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#10\",\n \"claim\": \"I .NET-biblioteket settes maksimalt antall output-tokens via egenskapen MaxOutputTokenCount på ChatCompletionOptions.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#11\",\n \"claim\": \"Metoden chatClient.CompleteChatStreamingAsync(messages, options) returnerer oppdateringer som itereres med await foreach, der hver update har ContentUpdate-deler med Text.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#12\",\n \"claim\": \"Streaming-chunks fra Azure OpenAI inneholder feltet finish_reason på choices, som angir hvorfor genereringen stoppet.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/responses\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#13\",\n \"claim\": \"For SSE-støtte gjennom Azure Application Gateway eller API Management må response buffering deaktiveres.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/application-gateway/use-server-sent-events\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/response-chunking-strategies.md#14\",\n \"claim\": \"Application Gateway for Containers støtter Server-Sent Events.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/application-gateway/for-containers/server-sent-events\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/response-chunking-strategies.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/throughput-optimization-strategies.md",
|
||
"claim_count": 10,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/throughput-optimization-strategies.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/throughput-optimization-strategies.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#1\",\n \"claim\": \"Referansen angir status GA for throughput-optimalisering i Azure OpenAI / Azure AI Services.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#2\",\n \"claim\": \"Azure OpenAI måler throughput i tokens per minutt (TPM) og forespørsler per minutt (RPM).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#3\",\n \"claim\": \"For Standard deployments setter den tildelte TPM-kvoten en øvre grense for gjennomstrømming, mens Provisioned Throughput Units (PTU) gir dedikert kapasitet der throughput avhenger av workload shape (forholdet mellom input- og output-tokens).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#4\",\n \"claim\": \"Kjernekomponentene for throughput-optimalisering er: Token quota (TPM/RPM) via Azure OpenAI Quota | Provisioned Throughput Units (PTU) | Batch API (Azure OpenAI Global Batch) | Azure Load Testing | Azure Monitor | azure-openai-benchmark.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#5\",\n \"claim\": \"Azure OpenAI Global Batch (Batch API) gir 50 % rabatt for asynkrone batch-jobber sammenlignet med standard prosessering.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#6\",\n \"claim\": \"Azure OpenAI chat completions kalles med api_version «2024-10-21».\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#7\",\n \"claim\": \"En batch-forespørsel i JSONL-filen til Azure OpenAI Batch API består av feltene custom_id | method | url | body, der body inneholder model | messages | max_tokens.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#8\",\n \"claim\": \"Azure OpenAI Batch API-jobber opprettes med completion_window «24h», altså 24 timers behandlingsvindu.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#9\",\n \"claim\": \"Provisioned throughput (PTU) gir latens-SLA (99 % over N tokens per sekund per PTU), mens Standard deployments ikke har latens-SLA.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/throughput-optimization-strategies.md#10\",\n \"claim\": \"Global Batch behandler data i Azure OpenAI-lokasjoner globalt, mens Data Zone Batch holder databehandlingen innenfor EU/EØS.\",\n \"claim_type\": \"region\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/batch\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/throughput-optimization-strategies.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
},
|
||
{
|
||
"file": "skills/ms-ai-security/references/performance-scalability/token-per-second-optimization.md",
|
||
"claim_count": 19,
|
||
"prompt": "# Per-claim groundedness judge — bake-off **v3.1** (recall-hardened over v3's 3 confirmed FNs)\n\nv3.1 of `judge-claim-prompt-v3.md`. Same blind, per-claim, one-subagent-per-file\ndesign, same three verdicts, same output schema, same evidence discipline. v3.1\nchanges only **three reasoning rules** (R1, R7, and a new R8), each fixing one of the\n**3 false negatives v3 still carried** (the judge-vs-gold disagreements G5 confirmed\nwere genuine judge misses, not stale gold).\n\n**Why v3.1 exists — and what it is NOT (transparent, not p-hacking).** The G5b\nfreshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent *false\npositives* against live Microsoft Learn. **All 4 turned out to be stale gold — v3\nflagged every one correctly** (`adr-template#1` \"zero permission management\" is\ncontradicted; `multi-region#2` lists retired `gpt-35-turbo`; `network-resilience#4`\noverstates \"recommended\" as \"obligatorisk\"; `vector-storage#7` cites the wrong GA\ndate). So **v3 has zero real false positives** (P = 100% on corrected gold), and the\nprecision-side \"FP-vakt\" originally planned for v3.1 is **dropped — there is nothing to\ndefend.** v3.1 is therefore a **pure recall hardening**: it tightens three rules so the\njudge catches 3 documented failure modes it currently misses, without touching the\nprecision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).\n\n**Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /\nR 92.9% on corrected gold), not v2.** v3 sits at the precision ceiling, so v3.1 can only\nbe adopted if it **holds P = 100% AND lifts R above 92.9%** (catches FNs without\nintroducing a single new false positive). Any new FP drops P below 100% and fails the\ngate — keep v3. Recall rules are double-edged over the full population (v3's own\nbake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay\nfrozen; v3.1 writes to `judge-bakeoff-results-v3.1.json` and is graded against the\ncorrected `gold-correctness-set.json`.\n\n---\n\nYou are a correctness judge for Microsoft AI reference documentation. You verify\nfactual claims against **live, official Microsoft Learn** (`learn.microsoft.com`).\nBe strict and adversarial — do not give the benefit of the doubt, do not pad, do not\ninfer a value the source does not state.\n\nYou are judging claims extracted from `skills/ms-ai-security/references/performance-scalability/token-per-second-optimization.md`. For EACH claim in the batch below,\ndecide whether the cited Microsoft Learn source **grounds** the claim.\n\n## The three verdicts (exhaustive, mutually exclusive)\n\n- **`grounded`** — you fetched a `learn.microsoft.com` page that states the claimed\n value(s). The page supports the claim. (Maps to gold `correct`.)\n- **`not_grounded`** — you fetched a `learn.microsoft.com` page that states a\n **different / contradicting / superseded** value for what the claim asserts. The\n claim disagrees with the source. (Maps to gold `outdated` + `wrong`.)\n- **`source_silent`** — you fetched the cited page (and searched as a fallback) but\n **no** `learn.microsoft.com` page states the claimed value at all. You cannot\n confirm or refute it. (Maps to gold `unsourced`.) Pricing on JS-rendered Azure\n pages typically lands here — that is expected, not a failure. **Exception: existence\n claims — see Rule R2.**\n\n## ⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)\n\nA claim is `grounded` ONLY if the fetched page states the **exact** asserted value(s).\nVerifying that the page \"is about\" the SKU/model/feature is **not** enough — the\nspecific number, name, date, tier, dimension, or status must match. If the claim\nasserts value **X** and the page states a **different** value **Y** (even if adjacent\nor plausible), the verdict is **`not_grounded`**. This rule does NOT lower the bar for\n`not_grounded`: you still need a fetched quote stating the **differing** value.\n\nApplies with special force to `sku`, `taxonomy`, `version`, `tpm`, `region`, `status`.\n\n---\n\n## CALIBRATION RULES — read all eight before judging\n\nThe exact-value rule is a blunt instrument. The 8 rules below sharpen it on both\nedges: **R1–R4 and R8 catch real errors** (more `not_grounded`); **R5–R7 stop\nover-flagging where the core is grounded** (more correctly `grounded`). When a rule\nbelow conflicts with a literal reading of the exact-value rule, the rule below governs\n— it is the more precise standard.\n\n### Recall side — flag these as `not_grounded`\n\n**R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper).**\nA claim may assert a **bound**. Direction matters — judge it by which side the bound\nconstrains:\n\n- **Lower bound** (\"100+\", \"200k+\", \"at least N\", \"minimum N\"): do not auto-`grounded`\n it just because the true value satisfies the inequality. Apply the **lower-bound\n policy:** if the true current value **grossly exceeds** the stated bound — roughly\n **>2× and decision-changing** — the bound materially misleads → `not_grounded`. A\n *tight* lower bound (true value within the same order of magnitude) stays `grounded`.\n *Example (FN2): \"200k+ context\" while the page states 1,047,576 (~1M) — ~5× → `not_grounded`.*\n- **Upper bound** (\"up to N\", \"opptil N\", \"maximum N\", \"no more than N\", \"as many as N\"):\n this is a **ceiling**, not a floor. The lower-bound leniency does **NOT** apply. The\n exact-value rule governs: if the live page states a current maximum **higher** than N,\n the stated ceiling is **superseded** → `not_grounded` — **regardless of ratio** (even\n a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it\n is really higher. *Example (v3-FN): claim \"up to 18 underlying models\" while the page\n states 28 → the ceiling has moved → `not_grounded`.* (Only `grounded` if the true\n maximum is N or the claim's ceiling still binds.)\n\n**R2 — `source_silent` does NOT excuse an existence claim (fixes FN3, FN5).** When the\nclaim asserts that a named entity **exists / is offered / is in a list** (\"X is a\nbuilt-in judge\", \"feature Y is available\", \"tier Z exists\"), and you fetch the\nauthoritative page that *would* enumerate it and the entity is **absent**, that absence\nis **evidence the claim is wrong** — return `not_grounded`, not `source_silent`. Reserve\n`source_silent` for values a page would not be expected to enumerate (e.g. JS-rendered\nprices). State in `reason` that you checked the canonical enumerating page and the\nentity was not present. *Example (FN5): claim \"99.99% SLA tier\" while the reliability\npage lists only 99.9% → absence of any 99.99% tier = `not_grounded`.*\n\n**R3 — Frame/unit replacement (fixes FN4).** A claim's **organizing frame or unit** can\nbe superseded even when derived ratios survive. If the page shows the claim's framing\nhas been **replaced** (e.g. \"1 Unit Capacity\" → \"Quota Tiers\"; a renamed/retired\nmetric), the claim is `not_grounded` even if some embedded numbers still appear\nsomewhere — the claim describes a world that no longer exists. Check that the *unit and\nstructure* the claim assumes still match the current page, not just the digits.\n\n**R4 — Current row, never a legacy row (fixes FN6).** Pages often carry historical or\neffective-dated rows (\"Before April 3, 2024\", \"Legacy\", \"Retiring\"). A claim is\n`grounded` only if it matches the **current/effective** row. Matching a clearly\ntime-stamped *past* row is `not_grounded` (the value has since changed). Always locate\nthe row that applies *today*. *Example (FN6): storage limits matching only the\n\"Before April 3, 2024\" row while current limits differ → `not_grounded`.*\n\n**R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the\n`ai-foundry-dr#9` FN).** A single claim often bundles **several load-bearing\nsub-assertions** (a status AND a region; a capability AND a named target; a date AND a\nGA level). Verify **each load-bearing part separately**. If **any one** load-bearing\npart is contradicted by the source, the whole claim is `not_grounded` — even when the\nother parts check out. Do not let a correct first half earn a `grounded` for a wrong\nsecond half. *Example (v3-FN): \"Global training (Public Preview), cheaper, no data\nresidency; use regional in Norway East\" — the GA-vs-Preview part and the \"no residency\"\npart hold, but **Norway East is a Global (non-residency) training region, not a regional\none** → one load-bearing part is wrong → `not_grounded`.* (R8 is the mirror of R6:\nR6 forgives an **omitted, non-load-bearing** detail; R8 condemns a **stated,\nload-bearing** part that is wrong. Decide first whether the part is load-bearing — if\nthe claim *asserts* it and a reader would act on it, it is.)\n\n### Precision side — keep these `grounded` (do not over-flag)\n\n**R5 — Documented theoretical↔benchmark equivalence (fixes FP3).** Do not flag a\nnumeric claim merely because the exact string is not verbatim, when the asserted value\nis the **documented theoretical or benchmark equivalent** of what the page states and\nboth trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the\nsame technique, same order of magnitude, same direction). The exact-value rule targets\n*drifted/contradicting* values — not two Microsoft-sourced expressions of the same\nfact. If the page substantiates the magnitude and the technique, keep `grounded` and\nnote the equivalence in `reason`.\n\n**R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4).** Distinguish \"the\nclaim's **core** assertion is grounded but it omits a sub-category\" from \"the core is\nungrounded.\" If the page confirms the claim's **central** behavior/categorization and\nthe only gap is an *unstated additional* case the claim did not deny, that is\n`grounded` (the claim is incomplete, not wrong). Reserve `not_grounded` for when the\npage **maps the core differently** or the claim **asserts** something the page\ncontradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;\na *stated* but wrong load-bearing part is not — that is R8's domain.)\n\n**R7 — Follow the capability to its canonical page; don't punish illustrative numbers\n(fixes FP5; v3.1 sharpens the load-bearing carve-out).** If a claim asserts a **real\ncapability** and the cited `evidence_url` does not foreground it, search for the\n**canonical** page that documents the capability before judging — do not return\n`not_grounded` merely because the *cited* page was a weak choice. And when a capability\nis solidly grounded, do **not** flag it over an *illustrative* attached number (e.g.\n\"~0 RTO/RPO\", \"≈15 min\") that the claim offers as an order-of-magnitude illustration\nrather than a cited spec. Judge the **capability**; treat an illustrative figure as\ngrounded if the capability is.\n\n> **⚠️ Load-bearing carve-out (v3.1, fixes the `token-usage#3` FN).** R7's leniency\n> covers only *illustrative* values. It does **NOT** cover a value or **exact string**\n> that **IS the assertion** — a metric name, an API field, an SDK identifier, an enum\n> value, a specific date/version. When the claim's load-bearing content is the literal\n> name/string itself (e.g. \"the metrics are `PromptTokens` and `CompletionTokens`\"),\n> the exact-value rule applies in full: if the live page names them differently\n> (`ProcessedPromptTokens` / `InputTokens` / `GeneratedTokens` / `OutputTokens`), the\n> claim is `not_grounded`. \"Follow to the canonical page\" means find the **right\n> names**, not rescue wrong ones. A reader would copy that string into code; an\n> illustrative magnitude they would not.\n\n---\n\n## Procedure (per claim)\n\n1. **Identify the volatile assertion(s)** in the claim text — and when the claim\n bundles several (R8), enumerate **each load-bearing part**. The `claim_type` tells\n you what to check:\n - `version` → model/API version, GA date, context window, max output, training cutoff\n - `tpm` → tokens-per-minute / throughput / quota numbers\n - `sku` → SKU name, tier, PTU minimums, deployment type\n - `region` → regional availability\n - `status` → GA / preview / retirement / deprecation status\n - `taxonomy` → categorization, capability mapping, which-feature-does-what\n2. **Fetch the cited source** with `microsoft_docs_fetch` on the claim's\n `evidence_url`. If the claim has no `evidence_url`, or the fetched page does not\n address the assertion, run `microsoft_docs_search` to find the authoritative page.\n **Under R2/R7, actively seek the canonical enumerating/capability page** — a weak\n cited URL is not the last word.\n3. **Exact-value entailment check** each checkable value (and each load-bearing part\n under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any,\n govern the claim. For R1, first decide whether the bound is a **lower** bound (floor)\n or an **upper** bound (ceiling) — they invert.\n4. **Strict evidence rule:** a `grounded` or `not_grounded` verdict REQUIRES a verbatim\n quote you actually fetched from a `learn.microsoft.com` URL. For R2 (existence\n absence), the quote is the canonical enumeration in which the entity does **not**\n appear — quote the enumeration and state the entity is absent. No quote → `source_silent`.\n\n## Hard rules\n\n- Verify against the fetched page only. Do not rely on prior knowledge of model\n specs / prices — those are exactly what may have drifted.\n- Stable identifiers are not volatile and are not your job to refute: regulation year\n (2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10\n 2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on\n whatever volatile value it carries, else `source_silent`.\n- One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.\n- `evidence_quote` = the verbatim sentence/value from the fetched page that drove the\n verdict (empty string for `source_silent`). `evidence_url` = the page you actually\n used (may differ from the cited one if you fell back to search).\n- `rule` = which calibration rule governed, if any (`R1`–`R8`), else empty.\n\n## Batch to judge (from `skills/ms-ai-security/references/performance-scalability/token-per-second-optimization.md`)\n\n[\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#1\",\n \"claim\": \"Azure OpenAI tilbyr latens-mål per PTU som varierer fra 25 TPS (o1) til 100 TPS (gpt-4.1-nano).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/latency\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#2\",\n \"claim\": \"Predicted Outputs i Azure OpenAI er i preview.\",\n \"claim_type\": \"status\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/predicted-outputs\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#3\",\n \"claim\": \"Prompt caching krever at det statiske innholdet plasseres først og utgjør minimum 1024 tokens; caching gjelder identiske prefikser.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#4\",\n \"claim\": \"Azure Monitor-metrikker for Azure OpenAI-throughput omfatter: GeneratedTokens | ProvisionedManagedUtilizationV2.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#5\",\n \"claim\": \"gpt-5.2 har 3 400 input TPM per PTU, latens-mål 99 % > 50 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#6\",\n \"claim\": \"gpt-5.1 har 4 750 input TPM per PTU, latens-mål 99 % > 50 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#7\",\n \"claim\": \"gpt-5 har 4 750 input TPM per PTU, latens-mål 99 % > 50 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#8\",\n \"claim\": \"gpt-5-mini har 23 750 input TPM per PTU, latens-mål 99 % > 80 TPS, minimum 15 PTU (Global) og 25 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#9\",\n \"claim\": \"gpt-4.1 har 3 000 input TPM per PTU, latens-mål 99 % > 80 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#10\",\n \"claim\": \"gpt-4.1-mini har 14 900 input TPM per PTU, latens-mål 99 % > 90 TPS, minimum 15 PTU (Global) og 25 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#11\",\n \"claim\": \"gpt-4.1-nano har 59 400 input TPM per PTU, latens-mål 99 % > 100 TPS, minimum 15 PTU (Global) og 25 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#12\",\n \"claim\": \"o3 har 3 000 input TPM per PTU, latens-mål 99 % > 80 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#13\",\n \"claim\": \"o4-mini har 5 400 input TPM per PTU, latens-mål 99 % > 90 TPS, minimum 15 PTU (Global) og 25 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#14\",\n \"claim\": \"gpt-4o har 2 500 input TPM per PTU, latens-mål 99 % > 25 TPS, minimum 15 PTU (Global) og 50 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#15\",\n \"claim\": \"gpt-4o-mini har 37 000 input TPM per PTU, latens-mål 99 % > 33 TPS, minimum 15 PTU (Global) og 25 PTU (Regional).\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#16\",\n \"claim\": \"Azure OpenAI-API-versjonen 2024-12-01-preview brukes for Predicted Outputs via prediction-parameteren.\",\n \"claim_type\": \"version\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/predicted-outputs\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#17\",\n \"claim\": \"Responsfeltet completion_tokens_details for Predicted Outputs inneholder: accepted_prediction_tokens | rejected_prediction_tokens.\",\n \"claim_type\": \"taxonomy\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/predicted-outputs\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#18\",\n \"claim\": \"Prompt caching gir opptil 100 % rabatt på cached input-tokens for PTU-deployments.\",\n \"claim_type\": \"sku\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/how-to/prompt-caching\"\n },\n {\n \"id\": \"ms-ai-security/performance-scalability/token-per-second-optimization.md#19\",\n \"claim\": \"PTU-latensmålet er definert som «99 % > N TPS» beregnet som p50 over 5-minutters vinduer.\",\n \"claim_type\": \"tpm\",\n \"evidence_url\": \"https://learn.microsoft.com/azure/foundry/openai/concepts/provisioned-throughput-billing\"\n }\n]\n\n## Output (strict JSON, no fence)\n\n```\n{\"file\":\"skills/ms-ai-security/references/performance-scalability/token-per-second-optimization.md\",\"results\":[\n {\"id\":\"<claim id>\",\"judge_verdict\":\"grounded|not_grounded|source_silent\",\"rule\":\"<R1-R8 or empty>\",\"evidence_url\":\"<url actually used>\",\"evidence_quote\":\"<verbatim quote or empty>\",\"reason\":\"<one sentence: what the source said vs the claim>\"}\n]}\n```\n"
|
||
}
|
||
]
|