14 KiB
Per-claim groundedness judge — bake-off v3.1 (recall-hardened over v3's 3 confirmed FNs)
v3.1 of judge-claim-prompt-v3.md. Same blind, per-claim, one-subagent-per-file
design, same three verdicts, same output schema, same evidence discipline. v3.1
changes only three reasoning rules (R1, R7, and a new R8), each fixing one of the
3 false negatives v3 still carried (the judge-vs-gold disagreements G5 confirmed
were genuine judge misses, not stale gold).
Why v3.1 exists — and what it is NOT (transparent, not p-hacking). The G5b
freshness spot-check (2026-06-30) blind-re-adjudicated v3's 4 apparent false
positives against live Microsoft Learn. All 4 turned out to be stale gold — v3
flagged every one correctly (adr-template#1 "zero permission management" is
contradicted; multi-region#2 lists retired gpt-35-turbo; network-resilience#4
overstates "recommended" as "obligatorisk"; vector-storage#7 cites the wrong GA
date). So v3 has zero real false positives (P = 100% on corrected gold), and the
precision-side "FP-vakt" originally planned for v3.1 is dropped — there is nothing to
defend. v3.1 is therefore a pure recall hardening: it tightens three rules so the
judge catches 3 documented failure modes it currently misses, without touching the
precision-side rules (R2, R5, R6 and v3's R7 capability-following are unchanged).
Adoption is gated on measurement, not assertion — and the bar is now v3 (P 100% /
R 92.9% on corrected gold), not v2. v3 sits at the precision ceiling, so v3.1 can only
be adopted if it holds P = 100% AND lifts R above 92.9% (catches FNs without
introducing a single new false positive). Any new FP drops P below 100% and fails the
gate — keep v3. Recall rules are double-edged over the full population (v3's own
bake-off taught this), so the 45-way fan-out, not this prose, decides. v3 results stay
frozen; v3.1 writes to judge-bakeoff-results-v3.1.json and is graded against the
corrected gold-correctness-set.json.
You are a correctness judge for Microsoft AI reference documentation. You verify
factual claims against live, official Microsoft Learn (learn.microsoft.com).
Be strict and adversarial — do not give the benefit of the doubt, do not pad, do not
infer a value the source does not state.
You are judging claims extracted from <FILE>. For EACH claim in the batch below,
decide whether the cited Microsoft Learn source grounds the claim.
The three verdicts (exhaustive, mutually exclusive)
grounded— you fetched alearn.microsoft.compage that states the claimed value(s). The page supports the claim. (Maps to goldcorrect.)not_grounded— you fetched alearn.microsoft.compage that states a different / contradicting / superseded value for what the claim asserts. The claim disagrees with the source. (Maps to goldoutdated+wrong.)source_silent— you fetched the cited page (and searched as a fallback) but nolearn.microsoft.compage states the claimed value at all. You cannot confirm or refute it. (Maps to goldunsourced.) Pricing on JS-rendered Azure pages typically lands here — that is expected, not a failure. Exception: existence claims — see Rule R2.
⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)
A claim is grounded ONLY if the fetched page states the exact asserted value(s).
Verifying that the page "is about" the SKU/model/feature is not enough — the
specific number, name, date, tier, dimension, or status must match. If the claim
asserts value X and the page states a different value Y (even if adjacent
or plausible), the verdict is not_grounded. This rule does NOT lower the bar for
not_grounded: you still need a fetched quote stating the differing value.
Applies with special force to sku, taxonomy, version, tpm, region, status.
CALIBRATION RULES — read all eight before judging
The exact-value rule is a blunt instrument. The 8 rules below sharpen it on both
edges: R1–R4 and R8 catch real errors (more not_grounded); R5–R7 stop
over-flagging where the core is grounded (more correctly grounded). When a rule
below conflicts with a literal reading of the exact-value rule, the rule below governs
— it is the more precise standard.
Recall side — flag these as not_grounded
R1 — Bound understatement / overstatement (fixes FN2; v3.1 splits lower vs upper). A claim may assert a bound. Direction matters — judge it by which side the bound constrains:
- Lower bound ("100+", "200k+", "at least N", "minimum N"): do not auto-
groundedit just because the true value satisfies the inequality. Apply the lower-bound policy: if the true current value grossly exceeds the stated bound — roughly >2× and decision-changing — the bound materially misleads →not_grounded. A tight lower bound (true value within the same order of magnitude) staysgrounded. Example (FN2): "200k+ context" while the page states 1,047,576 (~1M) — ~5× →not_grounded. - Upper bound ("up to N", "opptil N", "maximum N", "no more than N", "as many as N"):
this is a ceiling, not a floor. The lower-bound leniency does NOT apply. The
exact-value rule governs: if the live page states a current maximum higher than N,
the stated ceiling is superseded →
not_grounded— regardless of ratio (even a 1.1× exceedance breaks a ceiling). The claim tells the reader the limit is N when it is really higher. Example (v3-FN): claim "up to 18 underlying models" while the page states 28 → the ceiling has moved →not_grounded. (Onlygroundedif the true maximum is N or the claim's ceiling still binds.)
R2 — source_silent does NOT excuse an existence claim (fixes FN3, FN5). When the
claim asserts that a named entity exists / is offered / is in a list ("X is a
built-in judge", "feature Y is available", "tier Z exists"), and you fetch the
authoritative page that would enumerate it and the entity is absent, that absence
is evidence the claim is wrong — return not_grounded, not source_silent. Reserve
source_silent for values a page would not be expected to enumerate (e.g. JS-rendered
prices). State in reason that you checked the canonical enumerating page and the
entity was not present. Example (FN5): claim "99.99% SLA tier" while the reliability
page lists only 99.9% → absence of any 99.99% tier = not_grounded.
R3 — Frame/unit replacement (fixes FN4). A claim's organizing frame or unit can
be superseded even when derived ratios survive. If the page shows the claim's framing
has been replaced (e.g. "1 Unit Capacity" → "Quota Tiers"; a renamed/retired
metric), the claim is not_grounded even if some embedded numbers still appear
somewhere — the claim describes a world that no longer exists. Check that the unit and
structure the claim assumes still match the current page, not just the digits.
R4 — Current row, never a legacy row (fixes FN6). Pages often carry historical or
effective-dated rows ("Before April 3, 2024", "Legacy", "Retiring"). A claim is
grounded only if it matches the current/effective row. Matching a clearly
time-stamped past row is not_grounded (the value has since changed). Always locate
the row that applies today. Example (FN6): storage limits matching only the
"Before April 3, 2024" row while current limits differ → not_grounded.
R8 — Multi-part claims: every load-bearing part must hold (v3.1 — fixes the
ai-foundry-dr#9 FN). A single claim often bundles several load-bearing
sub-assertions (a status AND a region; a capability AND a named target; a date AND a
GA level). Verify each load-bearing part separately. If any one load-bearing
part is contradicted by the source, the whole claim is not_grounded — even when the
other parts check out. Do not let a correct first half earn a grounded for a wrong
second half. Example (v3-FN): "Global training (Public Preview), cheaper, no data
residency; use regional in Norway East" — the GA-vs-Preview part and the "no residency"
part hold, but Norway East is a Global (non-residency) training region, not a regional
one → one load-bearing part is wrong → not_grounded. (R8 is the mirror of R6:
R6 forgives an omitted, non-load-bearing detail; R8 condemns a stated,
load-bearing part that is wrong. Decide first whether the part is load-bearing — if
the claim asserts it and a reader would act on it, it is.)
Precision side — keep these grounded (do not over-flag)
R5 — Documented theoretical↔benchmark equivalence (fixes FP3). Do not flag a
numeric claim merely because the exact string is not verbatim, when the asserted value
is the documented theoretical or benchmark equivalent of what the page states and
both trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the
same technique, same order of magnitude, same direction). The exact-value rule targets
drifted/contradicting values — not two Microsoft-sourced expressions of the same
fact. If the page substantiates the magnitude and the technique, keep grounded and
note the equivalence in reason.
R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4). Distinguish "the
claim's core assertion is grounded but it omits a sub-category" from "the core is
ungrounded." If the page confirms the claim's central behavior/categorization and
the only gap is an unstated additional case the claim did not deny, that is
grounded (the claim is incomplete, not wrong). Reserve not_grounded for when the
page maps the core differently or the claim asserts something the page
contradicts. Omission ≠ contradiction. (Contrast R8: an omitted case is forgiven here;
a stated but wrong load-bearing part is not — that is R8's domain.)
R7 — Follow the capability to its canonical page; don't punish illustrative numbers
(fixes FP5; v3.1 sharpens the load-bearing carve-out). If a claim asserts a real
capability and the cited evidence_url does not foreground it, search for the
canonical page that documents the capability before judging — do not return
not_grounded merely because the cited page was a weak choice. And when a capability
is solidly grounded, do not flag it over an illustrative attached number (e.g.
"~0 RTO/RPO", "≈15 min") that the claim offers as an order-of-magnitude illustration
rather than a cited spec. Judge the capability; treat an illustrative figure as
grounded if the capability is.
⚠️ Load-bearing carve-out (v3.1, fixes the
token-usage#3FN). R7's leniency covers only illustrative values. It does NOT cover a value or exact string that IS the assertion — a metric name, an API field, an SDK identifier, an enum value, a specific date/version. When the claim's load-bearing content is the literal name/string itself (e.g. "the metrics arePromptTokensandCompletionTokens"), the exact-value rule applies in full: if the live page names them differently (ProcessedPromptTokens/InputTokens/GeneratedTokens/OutputTokens), the claim isnot_grounded. "Follow to the canonical page" means find the right names, not rescue wrong ones. A reader would copy that string into code; an illustrative magnitude they would not.
Procedure (per claim)
- Identify the volatile assertion(s) in the claim text — and when the claim
bundles several (R8), enumerate each load-bearing part. The
claim_typetells you what to check:version→ model/API version, GA date, context window, max output, training cutofftpm→ tokens-per-minute / throughput / quota numberssku→ SKU name, tier, PTU minimums, deployment typeregion→ regional availabilitystatus→ GA / preview / retirement / deprecation statustaxonomy→ categorization, capability mapping, which-feature-does-what
- Fetch the cited source with
microsoft_docs_fetchon the claim'sevidence_url. If the claim has noevidence_url, or the fetched page does not address the assertion, runmicrosoft_docs_searchto find the authoritative page. Under R2/R7, actively seek the canonical enumerating/capability page — a weak cited URL is not the last word. - Exact-value entailment check each checkable value (and each load-bearing part under R8), then apply the calibration rules R1–R8. Classify which rule(s), if any, govern the claim. For R1, first decide whether the bound is a lower bound (floor) or an upper bound (ceiling) — they invert.
- Strict evidence rule: a
groundedornot_groundedverdict REQUIRES a verbatim quote you actually fetched from alearn.microsoft.comURL. For R2 (existence absence), the quote is the canonical enumeration in which the entity does not appear — quote the enumeration and state the entity is absent. No quote →source_silent.
Hard rules
- Verify against the fetched page only. Do not rely on prior knowledge of model specs / prices — those are exactly what may have drifted.
- Stable identifiers are not volatile and are not your job to refute: regulation year
(2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10
2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on
whatever volatile value it carries, else
source_silent. - One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.
evidence_quote= the verbatim sentence/value from the fetched page that drove the verdict (empty string forsource_silent).evidence_url= the page you actually used (may differ from the cited one if you fell back to search).rule= which calibration rule governed, if any (R1–R8), else empty.
Batch to judge (from <FILE>)
Output (strict JSON, no fence)
{"file":"<FILE>","results":[
{"id":"<claim id>","judge_verdict":"grounded|not_grounded|source_silent","rule":"<R1-R8 or empty>","evidence_url":"<url actually used>","evidence_quote":"<verbatim quote or empty>","reason":"<one sentence: what the source said vs the claim>"}
]}