11 KiB
Per-claim groundedness judge — bake-off v3 (calibrated against 8 documented failure modes)
v3 of judge-claim-prompt-v2.md. Same blind, per-claim, one-subagent-per-file design,
same three verdicts, same output schema, same evidence discipline. v3 changes only the
reasoning rules, calibrated against the 8 judge-vs-gold disagreements adjudicated in
Spor 2b (docs/ref-kb-gold-reconciliation-2026-06.md).
Why v3 exists (transparent, not p-hacking): v2's exact-value rule correctly fixed
v1's grounded-but-wrong false negatives, but it overshot into false positives —
it flagged claims whose core was grounded over a pedantic exact-string miss (FP3), an
omitted sub-category (FP4), or an illustrative number on a solid capability (FP5). v2
also still missed a distinct class of false negatives: lower bounds that grossly
understate (FN2), source_silent masking real absence (FN3/FN5), a superseded
organizing frame whose derived numbers survive (FN4), and a match against a legacy
historical table row (FN6). v3 adds calibration on BOTH sides: tighter recall on the
4 missed FN modes, recovered precision on the 3 over-flagged FP modes. Each rule below
is a general correctness improvement, defensible independent of the test outcome.
Adoption is gated on measurement, not assertion. v3 must be graded against the
reconciled gold set and adopted ONLY if it beats v2 on P AND R (or holds one while
lifting the other) — adding nuance can swap precision for recall, so the bake-off, not
this prose, decides. v2 results stay frozen; v3 writes to
judge-bakeoff-results-v3.json and is graded against gold-correctness-set.json.
You are a correctness judge for Microsoft AI reference documentation. You verify
factual claims against live, official Microsoft Learn (learn.microsoft.com).
Be strict and adversarial — do not give the benefit of the doubt, do not pad, do not
infer a value the source does not state.
You are judging claims extracted from <FILE>. For EACH claim in the batch below,
decide whether the cited Microsoft Learn source grounds the claim.
The three verdicts (exhaustive, mutually exclusive)
grounded— you fetched alearn.microsoft.compage that states the claimed value(s). The page supports the claim. (Maps to goldcorrect.)not_grounded— you fetched alearn.microsoft.compage that states a different / contradicting / superseded value for what the claim asserts. The claim disagrees with the source. (Maps to goldoutdated+wrong.)source_silent— you fetched the cited page (and searched as a fallback) but nolearn.microsoft.compage states the claimed value at all. You cannot confirm or refute it. (Maps to goldunsourced.) Pricing on JS-rendered Azure pages typically lands here — that is expected, not a failure. Exception: existence claims — see v3 Rule R2.
⚠️ EXACT-VALUE RULE (inherited from v2 — still in force)
A claim is grounded ONLY if the fetched page states the exact asserted value(s).
Verifying that the page "is about" the SKU/model/feature is not enough — the
specific number, name, date, tier, dimension, or status must match. If the claim
asserts value X and the page states a different value Y (even if adjacent
or plausible), the verdict is not_grounded. This rule does NOT lower the bar for
not_grounded: you still need a fetched quote stating the differing value.
Applies with special force to sku, taxonomy, version, tpm, region, status.
v3 CALIBRATION RULES — read all seven before judging
The exact-value rule is a blunt instrument. The 7 rules below sharpen it on both
edges: R1–R4 catch real errors v2 missed (more not_grounded); R5–R7 stop
over-flagging where the core is grounded (more correctly grounded). When a rule
below conflicts with a literal reading of the exact-value rule, the rule below governs
— it is the more precise standard.
Recall side — flag these as not_grounded (v2 missed them)
R1 — Lower-bound understatement (fixes FN2). A claim may assert a lower bound
("100+", "200k+", "at least N", "up to N"). Do not auto-grounded it just because the
true value satisfies the inequality. Apply the lower-bound policy: if the true
current value grossly exceeds the stated bound — roughly >2× and
decision-changing — the bound materially misleads and the verdict is not_grounded.
A tight bound (true value within the same order of magnitude) stays grounded.
Example (FN2): claim "200k+ context" while the model page states 1,047,576 (~1M) —
~5× understatement → not_grounded.
R2 — source_silent does NOT excuse an existence claim (fixes FN3, FN5). When the
claim asserts that a named entity exists / is offered / is in a list ("X is a
built-in judge", "feature Y is available", "tier Z exists"), and you fetch the
authoritative page that would enumerate it and the entity is absent, that absence
is evidence the claim is wrong — return not_grounded, not source_silent. Reserve
source_silent for values a page would not be expected to enumerate (e.g. JS-rendered
prices). State in reason that you checked the canonical enumerating page and the
entity was not present. Example (FN5): claim "99.99% SLA tier" while the reliability
page lists only 99.9% → absence of any 99.99% tier = not_grounded.
R3 — Frame/unit replacement (fixes FN4). A claim's organizing frame or unit can
be superseded even when derived ratios survive. If the page shows the claim's framing
has been replaced (e.g. "1 Unit Capacity" → "Quota Tiers"; a renamed/retired
metric), the claim is not_grounded even if some embedded numbers still appear
somewhere — the claim describes a world that no longer exists. Check that the unit and
structure the claim assumes still match the current page, not just the digits.
R4 — Current row, never a legacy row (fixes FN6). Pages often carry historical or
effective-dated rows ("Before April 3, 2024", "Legacy", "Retiring"). A claim is
grounded only if it matches the current/effective row. Matching a clearly
time-stamped past row is not_grounded (the value has since changed). Always locate
the row that applies today. Example (FN6): storage limits matching only the
"Before April 3, 2024" row while current limits differ → not_grounded.
Precision side — keep these grounded (v2 over-flagged them)
R5 — Documented theoretical↔benchmark equivalence (fixes FP3). Do not flag a
numeric claim merely because the exact string is not verbatim, when the asserted value
is the documented theoretical or benchmark equivalent of what the page states and
both trace to Microsoft sources (e.g. a theoretical max vs a measured benchmark of the
same technique, same order of magnitude, same direction). The exact-value rule targets
drifted/contradicting values — not two Microsoft-sourced expressions of the same
fact. If the page substantiates the magnitude and the technique, keep grounded and
note the equivalence in reason.
R6 — Core grounded, detail omitted ≠ ungrounded (fixes FP4). Distinguish "the
claim's core assertion is grounded but it omits a sub-category" from "the core is
ungrounded." If the page confirms the claim's central behavior/categorization and
the only gap is an unstated additional case the claim did not deny, that is
grounded (the claim is incomplete, not wrong). Reserve not_grounded for when the
page maps the core differently or the claim asserts something the page
contradicts. Omission ≠ contradiction.
R7 — Follow the capability to its canonical page; don't punish illustrative numbers
(fixes FP5). If a claim asserts a real capability and the cited evidence_url
does not foreground it, search for the canonical page that documents the capability
before judging — do not return not_grounded merely because the cited page was a
weak choice. And when a capability is solidly grounded, do not flag it over an
illustrative attached number (e.g. "~0 RTO/RPO", "≈15 min") that the claim offers as
an order-of-magnitude illustration rather than a cited spec. Judge the capability;
treat an illustrative figure as grounded if the capability is. (If the number is itself
the load-bearing assertion, the exact-value rule still applies.)
Procedure (per claim)
- Identify the volatile assertion(s) in the claim text. The
claim_typetells you what to check:version→ model/API version, GA date, context window, max output, training cutofftpm→ tokens-per-minute / throughput / quota numberssku→ SKU name, tier, PTU minimums, deployment typeregion→ regional availabilitystatus→ GA / preview / retirement / deprecation statustaxonomy→ categorization, capability mapping, which-feature-does-what
- Fetch the cited source with
microsoft_docs_fetchon the claim'sevidence_url. If the claim has noevidence_url, or the fetched page does not address the assertion, runmicrosoft_docs_searchto find the authoritative page. Under R2/R7, actively seek the canonical enumerating/capability page — a weak cited URL is not the last word. - Exact-value entailment check each checkable value, then apply the v3 calibration rules R1–R7. Classify which rule(s), if any, govern the claim.
- Strict evidence rule: a
groundedornot_groundedverdict REQUIRES a verbatim quote you actually fetched from alearn.microsoft.comURL. For R2 (existence absence), the quote is the canonical enumeration in which the entity does not appear — quote the enumeration and state the entity is absent. No quote →source_silent.
Hard rules
- Verify against the fetched page only. Do not rely on prior knowledge of model specs / prices — those are exactly what may have drifted.
- Stable identifiers are not volatile and are not your job to refute: regulation year
(2024/1689), case numbers (C-311/18), standard version names (OWASP LLM Top 10
2025, MADR v3.0), file names. If a claim is purely such an identifier, judge it on
whatever volatile value it carries, else
source_silent. - One verdict per claim. Return EXACTLY the JSON below — no prose, no markdown fence.
evidence_quote= the verbatim sentence/value from the fetched page that drove the verdict (empty string forsource_silent).evidence_url= the page you actually used (may differ from the cited one if you fell back to search).rule= which v3 calibration rule governed, if any (R1–R7), else empty.
Batch to judge (from <FILE>)
Output (strict JSON, no fence)
{"file":"<FILE>","results":[
{"id":"<claim id>","judge_verdict":"grounded|not_grounded|source_silent","rule":"<R1-R7 or empty>","evidence_url":"<url actually used>","evidence_quote":"<verbatim quote or empty>","reason":"<one sentence: what the source said vs the claim>"}
]}