ms-ai-architect/docs/r11-pilot-results.md
Kjell Tore Guttormsen d8ce788709 docs(ms-ai-architect): R11 §10 måling #2 ført inn — O2 er triage, ikke automatisering [skip-docs]
§9 + appendiks C i pilot-resultatene, og kontrakten oppdatert der den fortsatt
sa at dette var umålt (§3, §5, §4.1, §4.2).

Målt over de 46 R8 ∧ MULTI_PART_CLAIM: 17 O2-kandidater, 29 O3.
- ALLE 29 felles av betingelse 3 — kilden leverer en korrigert verdi, så fiksen
  er swap/rewrite og subtraksjon ville ødelagt sann informasjon. Betingelse 1
  («strengt mindre») feller bare 5, aldri alene. R8s sviktende multipart-påstander
  er overveiende en FEILVERDI-klasse, ikke en overflødig-spesifisitet-klasse.
  Det er F2 reprodusert i skala, og det bekrefter at en mekanisk O2-driver ville
  vært feil å bygge.
- Kun 2 av 46 klarerer begge menneske-dømte betingelser bekreftende; 15 er
  merket human_must_confirm. Gjentakende grunn: det fjernede er ofte SANT om noe
  ANNET, bare ikke om radens eget subjekt — flytting kan slå sletting.
- Ekstrapolering til korpus (~60 kandidater) er merket som ekstrapolering, ikke
  måling.

Presisering under skriving: triage-fordelingen er ikke måling #4 (review-
throughput krever menneskelige review-økter som ikke er kjørt) — den er inputen
#4 trenger.

Suite 1021/1021.
2026-08-03 17:43:24 +02:00

30 KiB
Raw Blame History

R11 pilot results — measured, 2026-08-03

The §10 acceptance measurement of docs/r11-tiered-fix-design.md, run against the live ledger. No KB file was edited and no ledger record was written.

Instrument: scripts/kb-eval/lib/fix-op.mjs (+ tests/kb-eval/test-fix-op-classify.test.mjs, 49 tests after §4b) driven by scripts/kb-eval/classify-fix-ops.mjs. The classifier is the O1 driver with writes disabled — it constructs the swap and checks the §4 invariant, so measurements 1 and 3 come out of the mechanism that would later touch the corpus, not out of a proxy heuristic.

Artefact: scripts/kb-eval/data/r11-pilot-classification.json (untracked, regenerable; per-flag records so the run can be re-analysed without re-running). It holds the pilot run — node scripts/kb-eval/classify-fix-ops.mjs --write. Every corpus-wide figure below is from --threshold 1, and the per-table reproduce command is stated where it is used.

Two different 202s. This population is 202 flags. §3's "202 flags whose claim and quote contain a numeric token" is a different 202, measured over the full 712-flag population. They are unrelated.


1. The headline

Nine provable, correct value swaps exist in the entire 776-flag not_grounded population — 1.2 %. The machine half of the R11 tiering buys nine edits. Every other flag needs a human.

Population Files Flags O1 admitted O1 hand-verified correct
Pilot (not_grounded ≥ 7) 24 202 2 2
Whole not_grounded corpus 218 776 15 9

These are numeric-path figures, and they stay that way. §4b (the status synonym table) was implemented afterwards and adds a separate class with its own hand-verification — see §8. A run today prints O1 = 7 (pilot) and 23 (corpus) because the status proposals are included in the total; the numeric line above is unchanged and is still what s4_as_written compares against.

Reproduce: node scripts/kb-eval/classify-fix-ops.mjs (pilot) and node scripts/kb-eval/classify-fix-ops.mjs --threshold 1 (corpus). The nine are enumerated with verdicts in appendix A — that hand-verification is the only thing separating 9 from 15, so it is recorded rather than left in a session transcript.

This is the answer §10 asked for, and it is materially worse than the design assumed: "If the split is materially worse than assumed, that is known after one session rather than after ten."

The O2 half, measured afterwards (§9), does not rescue the number. Of the 46 R8 multi-part claims that have the O2 shape, 17 are candidates and 2 are clean subtractions on the available evidence; all 29 non-candidates fail because the source supplies a corrected value, which makes them swaps or rewrites. O2's value in R11 is triage — it tells a human which 17 to look at first — not automation.

2. §4 as written is not sufficient — measured, not argued

§4 claims its invariant is "deliberately stronger than human review at scale." It is not. Run exactly as specified over the pilot, it admitted 6 swaps, of which 4 were wrong — precision 2/6:

Proposed swap Why it is wrong
30-dagers24-dagers Unit crossing. The quote says 24 hours.
3000 requests/sekund50 Metric crossing. The quote is a query throttle per index; the claim is an indexing rate per replica.
Microsoft Agent 365Agent 7 Identifier mutilated. The 7 was harvested out of E7.
text-embedding-ada-002ada-2 Identifier mutilated. The 2 came from a dimensions column.

The defect is structural, not incidental. §4 constrains where the new value came from (verbatim in the cited quote) and what the edit looks like (one line, rest byte-identical). It constrains nothing about whether the two tokens denote the same quantity. Same-type-and-provenance is not same-referent.

The added condition

contextCorresponds() requires the token to sit under the same label or the same trailing unit on both sides. It is deliberately lexical, with no translation table: dokumenter is not taught to equal documents, because a synonym table introduces a new fact source and is an operator decision, not an engineering one. Consequence, measured: a swap is provable essentially only where the context is language-neutral — a URL, a code sample, a parameter key.

3. The condition is necessary but still not sufficient

Corpus-wide the context condition cut 38 admissions to 15. Hand-verifying all 15 splits them cleanly by token type:

Reproduce: node scripts/kb-eval/classify-fix-ops.mjs --threshold 1 (the persisted artefact is the pilot run — the corpus-wide tables in §1 and §3 come from this threshold-1 run). Per-proposal verdicts: appendix A.

Token type Proposals Correct Wrong Unverified Failure mode
iso_date 9 9 0 0 — every one is an api-version= bump in a URL or code sample
number 5 0 3 2 AI-900AI-901, gpt-4ogpt-5.1o (×2)
version 1 0 1 0 Java agent 3.7.53.4.0 — a downgrade

A matching identifier prefix (AI-, gpt-) satisfies the context condition while the digit is part of a name, not a quantity. Only iso_date survives hand-verification, and the report marks it as the sole recommended class (o1_recommended). number and version proposals must not be applied.

4. The four §10 measurements

  1. O1 / O2 / O3 split. O1 = 2/202 on the pilot (9/776 corpus-wide, safe class only). O2 is undetermined — it does not exist as a class until §5 is ratified, so every non-O1 item is O3 by design. §5 was ratified 2026-08-03, and O2 has since been measured over the class where it can exist at all — see measurement 2 and §9. It remains unmeasured outside R8 ∧ MULTI_PART_CLAIM; the classifier still routes every non-O1 item to O3, so O3 ≥ 200/202 stands as the machine's own partition.

  2. How much of R8 resolves as O2. MEASURED 2026-08-03 — see §9. R8 is 87/202 on the pilot (366/776 corpus-wide) and yields zero O1. Of the pilot's 87, 46 are structural enumerations (R8 ∧ MULTI_PART_CLAIM) — the O2 candidate shape. Which of them subtract cleanly turns on the judge's prose reason, and no regex reads prose, so this was done by prose classification: 17 of 46 (37 %) are O2 candidates, 29 are O3. All 29 are foreclosed by condition 3.

  3. O1 abort rate: 99 % (200/202). Typed, because "99 %" alone is not actionable:

    Code Pilot Class
    MULTI_PART_CLAIM 96 (47.5 %) intrinsic — not a value swap at all
    MULTI_VALUE_TOKEN 29 intrinsic
    NO_VALUE_TOKEN 28 intrinsic — the claim asserts prose
    STATUS_SYNONYM 15 operator question (§6.2)
    NOT_VERBATIM 15 intrinsic
    LOCATOR_AMBIGUOUS 7 fixable engineering gap
    MULTI_REPLACEMENT 6 intrinsic
    CONTEXT_MISMATCH 4 intrinsic — these are the 4 wrong edits above

    Only 7 of 200 aborts (3.5 %) are a fixable engineering gap. More locator engineering cannot move the O1 number materially.

  4. Review throughput per class. NOT MEASURED. It requires human review sessions, which have not happened. Recording it as measured would be false.

5. Two further findings

F1 — subtraction can leave a misleading remainder. §5 argues O2 "cannot introduce a new error, because it asserts strictly less." True of the sentence, false of the reader's inference. Real case: *"Deep Research tool (o3-deep-research

  • Bing) er GA (juni 2025)"* where the source says the tool is deprecated. Subtracting er GA (juni 2025) leaves the tool standing in a list of available tools. Strictly less asserted; still misleading. O2 therefore still requires a human to look at the remainder — cheaper than O3 (no fact-finding) but not mechanical.

F2 — subtraction can destroy true information. Real case: a list of seven prebuilt model IDs where the judge found six correct and prebuilt-check wrong — the real ID is prebuilt-check.us. Subtraction drops a model that exists; the correct fix is a swap. Subtraction is not the safe default everywhere.

F3 — disposition carries zero information. It is outdated on 202 of 202 flags. docs/r11-flag-format-2026-07.md specifies not_grounded → {outdated, wrong} with "the human assigns which at R11", but the pass hard-assigned outdated. Do not use it as a classifier signal. Spec/data divergence, recorded.

F4 — claims are not file text. claim is an LLM-extracted, translated restatement: 0 of 202 match their file line verbatim, and 188 share no 40-char run with it. For table claims, line points at the header, not the value. This is why the locator exists at all, and why it searches the enclosing block rather than the line.

6. Operator decisions — ALL THREE RATIFIED 2026-08-03

All three were put to the operator with the recommendations below and all three were accepted as recommended. The contract text now lives in docs/r11-tiered-fix-design.md §4a/§4b/§5; this section records what was asked and what the answer was.

None of the three is implemented yet. The classifier still aborts STATUS_SYNONYM and still routes every non-O1 item to O3. A later session builds against the ratified contract — it must not assume the code already honours it.

  1. Ratify O2 (§5)?RATIFIED, with the remainder check (not as a blanket rule), exactly as F1/F2 above argued. Contract: design doc §5, three conditions, human-confirmed.
  2. Amend §4 with a ratified synonym table?RATIFIED, narrow and closed. Contract: design doc §4b — four label rows, closed table, complete-label-only, corpus-side value with the file's own markup preserved. Unlocks up to 54 corpus-wide flags.
  3. Is O1 worth building at all?KEPT, locked to iso_date. Contract: design doc §4a condition 5 — a driver may apply iso_date proposals and must never apply number or version ones. Nine edits corpus-wide.

7. What this does not change

The design's core reading survives: the expensive half (locating the source, reading it, extracting the deciding passage) was already paid for by the judge pass, and §7's freshness re-fetch per distinct URL is untouched. What the pilot falsifies is the assumption that a meaningful share of that evidence converts into machine-provable edits. It does not. R11 is a human review programme with a nine-item machine assist, and its leverage lies entirely in the O2 decision.


8. §4b implemented — the status class measured, 2026-08-03

The ratified synonym table (docs/r11-tiered-fix-design.md §4b) is implemented in fix-op.mjs and the class is measured. Reproduce: node scripts/kb-eval/classify-fix-ops.mjs --threshold 1 (status_synonym block).

Population STATUS_SYNONYM flags Proven by §4b Still aborting
Pilot (not_grounded ≥ 7) 15 5 10
Whole not_grounded corpus 54 8 46

Why the other 46 abort, corpus-wide — this sub-distribution is the actionable part, because the top-level STATUS_SYNONYM count alone says nothing:

Reason N What it means
NO_COMPLETE_FILE_LABEL 27 the file writes the status inside a sentence — (preview) in a list item, **DSPM (preview):**, "[Preview]: …" in a JSON string. Constraint 2 refuses these, correctly.
NO_SOURCE_STATUS 15 the cited quote carries no listed lifecycle phrasing at all — the flag was never a status swap.
SOURCE_STATUS_AMBIGUOUS 2 the quote asserts two different rows (e.g. "…is now generally available. Partner solutions remain in preview.").
FILE_ALREADY_MATCHES 2 file and source agree; the mismatch was in the LLM-extracted claim, not in the corpus.

The class is REVIEW-grade, not apply-grade — 5 of 8 correct. All eight were hand-judged against the cited source (appendix B). Three defects, all one family: §4b binds the table, the completeness of the file label and the written value, and nothing about whether the source phrasing refers to the row's own subject. That is the same provenance-without-referent defect that falsified §4 (§2), now reproduced in the status class. status is therefore deliberately absent from o1_recommended: the machine writes nothing, and every proposal reaches a human.

Two candidate conditions, costed over the eight

Neither is implemented — extending a table the operator ratified as closed is an operator decision, exactly as condition 5 was in §4a. Both are pure gain on this population (they kill wrong proposals and no correct one), which is the number the decision needs:

Candidate Kills Correct proposals lost
A — count a bare GA in the quote as a GA-row source phrasing, so a quote saying both GA and "public preview" becomes ambiguous 1 (onelake:198) 0
B — abort when the quote is a multi-entity enumeration (≥ 2 pipes, or a numbered list) 2 (onelake:198, owasp:79) 0

B subsumes A on these eight. Neither catches security-copilot-integration.md:94, where the quote is prose and the (Preview) marker simply belongs to a different agent. A stricter referent-name condition (the row's subject must appear in the quote) would catch it — and would also kill two correct proposals (:83, :93), where the source names the capability rather than the agent. That trade is real and is why this is put to the operator rather than shipped.

9. §10 measurement #2 — how much of R8 resolves as O2, measured 2026-08-03

The last machine-answerable pilot measurement. §4.2 recorded it as "not answered, and not answerable by machine" — true of a regex, not of prose classification, which is what this ran. Reproduce the verification and the tally: node scripts/kb-eval/check-o2-returns.mjs.

Population. The 46 pilot flags that are R8 ∧ MULTI_PART_CLAIM — the O2 candidate shape, re-derived from the ledger, not read from a plan (classify-fix-ops.mjs --threshold 7). 17 distinct files.

Method. Eight subagents, six items each, classifying against §5's three conditions. Constraints, all deliberate: read-only (no writes, no commits); no web or MCP lookups — O2 is defined by requiring no new fact-finding, so the evidence_quote is the only source evidence a classifier may use; and every proposal must be obtainable from the file text by deleting characters only. Because claim matches its file line verbatim in 0 of 202 cases (F4), each classifier had to open the actual file and locate the real text rather than edit the restatement. All 46 located it; locator_failed is 0.

Verdict N Share
O2 candidate 17 37 %
O3 29 63 %

The result that matters is not the split — it is what blocks the other 29.

Blocking condition(s) N
condition 2 + condition 3 21
condition 3 alone 3
all three 5

All 29 are foreclosed by condition 3: the source supplies a corrected value, so the fix is a swap or a rewrite and subtraction would destroy true information. Condition 1 — "asserts strictly less", the one that sounds like the hard one — blocks only 5, and never alone. This is F2 (§5) reproduced at scale: prebuilt-check.us, prebuilt-mortgage.us.closingDisclosure, Set-DlpCompliancePolicy, jensen_shannon_distance, F300 = 384 GB, DurationMs / ResultSignature, Claude 4.5 → 4.6. R8's failing multi-part claims are predominantly a wrong-value class, not a surplus-specificity class. The design's reading of R8 in §3 — "where the grounded part stands on its own, the fix is O2" — holds for a minority of the class.

Machine verification of the returns (V1/V2/V2b, §9.1). 46 of 46 pass V1: the quoted file text occurs verbatim in the named file, æ/ø/å and markup intact. 16 of the 17 O2 proposals are deletion-only; one is flagged (feedback-loops-continuous-improvement.md:555, where Automatically addAdd recapitalises rather than merely deletes). That is a text change, not a subtraction, and it goes to a human as such.

The 17 are candidates, not admitted edits — and the split inside them is the honest number:

N
conditions 2 and 3 both affirmatively yes 2 (idx 8, 17)
at least one condition marked human_must_confirm 15
classifier confidence high 1

So the machine's own reading is that 2 of 46 (4 %) are clean subtractions on the evidence available, and 15 more are worth a human's time. This is a triage, not a machine assist. It is not measurement #4 — review throughput still requires human review sessions that have not happened (§4.4) — but it is the input #4 needs: it says how many items enter review and in what state, which is the half of throughput that does not require a stopwatch. The recurring reason for human_must_confirm on condition 3 is structural and worth recording: the removed material is often true of something else (Purview really does classify data; the Communication Compliance template really exists; Redis really is in Norway West) — it is merely false of the subject the row names. Deleting it is defensible; relocating it may be better. That is a judgement about the corpus, not about the source, which is exactly why §5 put conditions 2 and 3 in human hands.

Extrapolation, flagged as such. MULTI_PART_CLAIM is 161 corpus-wide under R8. At the pilot's 37 % that is ~60 O2 candidates and ~6 clean ones. This is an extrapolation from one measured sample, not a measurement, and the pilot was deliberately drawn from the densest files.

9.1 What the machine checks, and what it deliberately does not

scripts/kb-eval/lib/o2-return-check.mjs (25 tests). The checks do not decide O2 — conditions 2 and 3 stay human by ratified contract. They bound the two failure modes a human reviewing 46 proposals cannot catch cheaply:

  • V1 — the quoted file text must occur verbatim in the file. Catches invented text and silent æ/ø/å transliteration. Applied to every row, not just the O2 ones: an O3 verdict resting on invented text is equally wrong, merely wrong in the safe direction.
  • V2 — the remainder must be obtainable by deleting characters only.
  • V2b — word-level and case-sensitive, because V2 alone is too weak: deleting a leading word and recapitalising the next passes the character test, since the capital already existed inside the deleted word. This check was added after wave 1 produced exactly that case.
  • V3 — schema completeness and verdict/condition coherence.

The raw returns are committed under scripts/kb-eval/data/r11-o2-returns/ — they are evidence, not regenerable output, same discipline as appendices A and B.

Appendix A — the 15 admitted proposals, hand-verified

Every proposal the classifier (§4 + context condition) admitted over the whole not_grounded population, with the verdict that produced §3's table. A later run that admits a 16th can diff against this list; without it, "9 of 15" is an unreproducible claim.

# File:line Swap Type Verdict
1 agent-orchestration/agent-evaluation-testing-frameworks.md:56 4.15 (gpt-4.1-minigpt-5-mini) number unverified — model identifier; the result is a real model name, but not checked against the source. Not applied.
2 api-management/logging-analytics-ai-traffic.md:49 2023-09-012025-09-01 iso_date correct — ARM loggers@ api-version bump
3 azure-ai-services/translator-document-translation.md:162 4010 (MB) number unverified — matched on the unit MB, but sync/async limits differ; metric-crossing risk. Not applied.
4 monitoring-observability/log-analytics-kql-ai-queries.md:617 2025-09-012026-04-01 iso_date correctapi-version= inside a KQL string literal
5 responsible-ai/responsible-ai-training-awareness.md:77 900901 (AI-900AI-901) number wrong — certification identifier mutilated
6 bcdr/cost-analysis-dr-configurations.md:120 45.1 (GPT-4oGPT-5.1o) number wrong — model identifier mutilated
7 bcdr/multi-region-azure-openai-deployment.md:316 2024-06-012024-10-01 iso_date correctapi-version= in a management URL
8 ai-security-engineering/ai-prompt-shield-network.md:309 2024-09-012024-09-15 iso_date correct — Content Safety api-version
9 ai-security-engineering/content-safety-filter-calibration.md:277 2024-10-012024-10-21 iso_date correct — Azure OpenAI api-version in a curl sample
10 ai-security-engineering/jailbreak-prevention-production.md:305 2024-09-012024-09-15 iso_date correct — Content Safety api-version in a curl sample
11 cost-optimization/observability-cost-reduction.md:114 3.7.53.4.0 (Java Agent) version wrong — a downgrade; the quote's version is not the claim's referent
12 cost-optimization/vector-storage-cost-optimization.md:266 2025-09-012026-04-01 iso_date correct — AI Search api-version
13 cost-optimization/vector-storage-cost-optimization.md:318 2024-02-012024-10-21 iso_date correct — embeddings api-version
14 performance-scalability/response-chunking-strategies.md:56 45.1 (gpt-4ogpt-5.1o) number wrong — model identifier mutilated
15 performance-scalability/token-per-second-optimization.md:295 2024-12-012025-01-01 iso_date correct — Azure OpenAI api-version

9 correct · 4 wrong · 2 unverified. All nine correct are iso_date; every wrong one is a digit inside a product, model or certification identifier, where a matching prefix (AI-, gpt-, Agent ) satisfies the context condition while the digit is part of a name rather than a quantity. The two unverified are also number and are excluded by the same class rule — verifying them costs a source fetch each and would move the total to at most 11.

Appendix B — the 8 §4b status proposals, hand-verified

Every status proposal the classifier admits over the whole not_grounded population, judged against the cited source. Same discipline as appendix A: a later run that admits a ninth can diff against this list, and "5 of 8" is otherwise an unreproducible claim.

Reproduce: node scripts/kb-eval/classify-fix-ops.mjs --threshold 1 --write, then read the proposal.type === 'status' items in scripts/kb-eval/data/r11-pilot-classification.json.

# File:line Swap Verdict
1 agent-orchestration/foundry-agent-service-ga.md:68 **Preview****GA** correct — quote: "hosted agents are generally available"; the row's subject is Hosted agents
2 agent-orchestration/foundry-agent-service-ga.md:72 **GA****Preview** correct — quote: "Trigger an agent by using Logic Apps (preview)"; the row's subject is the Logic Apps trigger
3 ai-security-engineering/security-copilot-integration.md:83 Public PreviewGA correct — quote: "Email and collaboration alert triage capabilities are already generally available (GA)"; the row is the phishing/email triage agent
4 ai-security-engineering/security-copilot-integration.md:93 GAPreview correct — quote is from the agent's own doc page: "This feature is in public preview"
5 ai-security-engineering/entra-agent-id-zero-trust.md:439 Public PreviewGA correct — quote: "The Microsoft Entra Agent ID platform is now generally available"; the row's subject is Entra Agent ID (kjerne)
6 ai-security-engineering/security-copilot-integration.md:94 GAPreview unproven — the quote's (Preview) marker belongs to Identity Risk Management Agent, not to Access Review Agent. The judge's prose reason does support preview from a what's-new post, so the outcome is plausibly right; the cited evidence does not establish it. Not applied.
7 data-engineering/onelake-data-strategy.md:198 GAPreview wrong — the quote says Lakehouse | Yes | GA. "Public preview" in the same quote belongs to Eventhouse. Killed by candidate A and B.
8 ai-security-engineering/owasp-llm-top10-azure-mitigations.md:79 GAPreview wrong — the source marks only Response Completeness as preview; the row covers the groundedness/completeness pair, so the edit makes the groundedness half false. The correct fix is to split the row (O2/O3). Killed by candidate B.

5 correct · 1 unproven · 2 wrong. All five correct ones carry the source phrasing on the row's own subject; all three defects are the referent gap described in §8. No proposal was applied — §4b output is a human review list.

Appendix C — the 17 O2 candidates

Every proposal the prose classification admitted over the 46, with the two human-judged conditions as the classifier left them. Same discipline as appendices A and B: without this list, "17 of 46" is an unreproducible claim. confirm = the classifier marked the condition human_must_confirm, i.e. it could not settle it on the evidence available and is handing it over — not a defect, it is the contract.

The full records, including each proposal's verbatim file text and the exact remainder, are in scripts/kb-eval/data/r11-o2-returns/.

# File:line Failing sub-assertion Cond 2 Cond 3 Confidence
7 ms-ai-engineering/azure-ai-services/document-intelligence-prebuilt-models.md:79 The third table row presenting prebuilt-document (General Document) as a current basic model — the source states the general document model is no l… yes confirm medium
8 ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md:191 Two sub-assertions: (a) eller managed compute cluster as an alternative compute option — the how-to page and the monitor schema require a Spark poo… yes yes high
9 ms-ai-engineering/mlops-genaiops/data-drift-monitoring-detection.md:218 The second sentence, Støtter også drift detection for grounding data i RAG scenarios. — the canonical observability page lists only Evaluation, Mon… yes confirm medium
14 ms-ai-engineering/mlops-genaiops/feedback-loops-continuous-improvement.md:555 Two sub-assertions: 'SharePoint' as a feedback storage service, and the word 'Automatically' in 'Automatically add reviewed samples to training set' … confirm yes medium
17 ms-ai-engineering/rag-architecture/rag-caching-optimization.md:254 The three-band rubric (0.1-0.2 strict / 0.3-0.5 balanced / 0.6-0.8 liberal) — undocumented, and the two upper bands contradict the source's warning t… yes yes medium
18 ms-ai-engineering/rag-architecture/rag-caching-optimization.md:29 The list item 'Azure AI Search (built-in caching av search results)' — the source states each query operates on the current index view with no cachin… confirm yes medium
19 ms-ai-engineering/rag-architecture/rag-caching-optimization.md:297 The bullet "Automatic indexing av vectors" — the judge states vector indexes must be declared explicitly in the indexing policy (only at container cr… confirm yes medium
26 ms-ai-governance/responsible-ai/transparency-documentation-standards.md:117 Items 4 (Error analysis) and 5 (Counterfactual analysis) are listed as Responsible AI Scorecard components, but the canonical scorecard segment enume… confirm yes medium
27 ms-ai-governance/responsible-ai/transparency-documentation-standards.md:426 The 'Chat interface' row (a "Powered by AI" badge in the chat window) and the 'Plugin actions' row (confirmation prompts before sensitive actions) ar… yes confirm medium
28 ms-ai-governance/responsible-ai/transparency-documentation-standards.md:83 The second and third bullets — that Hugging Face model cards are synchronised automatically, and that a template exists for generating model cards fo… yes confirm medium
31 ms-ai-security/ai-security-engineering/ai-incident-response-procedures.md:139 The legalHold object is given a field named "enabled"; the Storage API's LegalHold model exposes tags and hasLegalHold, so the literal field name "en… yes confirm medium
33 ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md:211 Two parts attributed to Defender for Cloud AI Security Posture Management that the AISPM page does not support: the discovery mechanism "(via Azure R… yes confirm medium
36 ms-ai-security/ai-security-engineering/ai-threat-modeling-stride.md:38 The "/indiscriminate" qualifier, which extends the Tampering placement and the Critical severity to indiscriminate data poisoning; the source gives t… yes confirm medium
38 ms-ai-security/ai-security-engineering/data-leakage-prevention-ai.md:396 The listing of 'DSPM for AI - Unethical behavior in AI apps' and 'DSPM for AI - Protect sensitive data from Copilot processing' as Insider Risk Manag… yes confirm medium
40 ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md:133 The third bullet 'CVE severity mapping' presented as a category of alert that dependency scanning generates; severity is a property of an alert, … yes confirm medium
42 ms-ai-security/ai-security-engineering/supply-chain-security-ai-models.md:200 The second bullet presenting the HuggingFace Registry as a Microsoft channel for verified models with provenance tracking; the source calls it a comm… yes confirm medium
45 ms-ai-security/cost-optimization/semantic-caching-patterns.md:436 The '/West' half of the region pair, i.e. the standing implication that Azure OpenAI can be deployed in Norway West. yes confirm medium

2 affirmative on both conditions (8, 17) · 15 needing a human call. Item 14 additionally carries a machine flag: its remainder recapitalises rather than deletes (§9.1, V2b), so it is a text change and must be reviewed as one.

Reproduce the tally and the machine checks: node scripts/kb-eval/check-o2-returns.mjs.