docs(ms-ai-architect): cond2 ratifisert som hele-fila-lesning; verifisert skår 1 av 46 [skip-docs]
This commit is contained in:
parent
03caa72d58
commit
e1d344307c
1 changed files with 31 additions and 32 deletions
|
|
@ -382,11 +382,14 @@ remainder must not be misleading. **Its scope was never fixed:**
|
|||
file is a **separate ungrounded claim**, to be flagged on its own, not a
|
||||
defeater of this subtraction.
|
||||
|
||||
The analysis below takes the **broad** reading. It is not ratified, and it is
|
||||
the operator's to settle — the same operator who ratified conditions 2 and 3 as
|
||||
human-judged. Under the narrow reading the classifier did exactly what cond 2
|
||||
asks and idx 8's cond 2 holds; the finding then becomes "lines 263-265 are a
|
||||
second ungrounded claim" rather than "the classifier was wrong".
|
||||
**The operator ratified the broad reading, 2026-08-03.** Cond 2 is measured
|
||||
against the whole file: a remainder that contradicts the file it sits in is
|
||||
misleading, whatever the source says. The stated ground is that these files are
|
||||
publicly distributed and read as wholes — a self-contradicting file is a trust
|
||||
defect regardless of which half is wrong. The narrow reading (cond 2 scoped to
|
||||
the edited passage, with contradictions elsewhere handled as separate ungrounded
|
||||
claims) was considered and rejected. Recorded here because the fork was real and
|
||||
a later run must not silently re-open it.
|
||||
|
||||
The classifier justified cond 2 on one of the two sub-deletions and never
|
||||
checked the other:
|
||||
|
|
@ -396,20 +399,16 @@ checked the other:
|
|||
(`**Azure Monitor + Application Insights** (Verified)` / "Drift metrics
|
||||
emitteres til Application Insights"), so removing the prerequisite bullet
|
||||
leaves no false implication.
|
||||
- Sub-deletion (a), `eller managed compute cluster`: **fails under the broad
|
||||
reading, holds under the narrow one.** The same file documents
|
||||
- Sub-deletion (a), `eller managed compute cluster`: **fails.** The same file
|
||||
documents
|
||||
`**Managed Compute Cluster** (for store volumer)` as a real compute option at
|
||||
lines 263-265, with pricing and a usage recommendation. Deleting the
|
||||
alternative from the prerequisites leaves the file asserting serverless Spark
|
||||
as the only compute requirement seventy lines above a cost section that prices
|
||||
the alternative. Broad: that is a misleading remainder. Narrow: the remainder
|
||||
states what the source grounds, and 263-265 is a separate claim needing its
|
||||
own verdict.
|
||||
the alternative. That is a misleading remainder under the ratified reading.
|
||||
|
||||
Either way the candidate does not go in as written — under the broad reading
|
||||
because (a) must be dropped from the subtraction, under the narrow reading
|
||||
because lines 263-265 must be adjudicated first and may themselves need an edit.
|
||||
The remedies are different and the operator picks which applies.
|
||||
The candidate therefore does not go in as written: **(a) must be dropped from
|
||||
the subtraction**, leaving (b) alone.
|
||||
|
||||
There is a further limit the O2 envelope cannot resolve: cond 3 passes for (a)
|
||||
only because the source does not *mention* managed compute — not-mentioned
|
||||
|
|
@ -424,20 +423,19 @@ string** than the one V2b attested, so it is not machine-clean until
|
|||
amendment, including dropping idx 17's trailing colon: **amended remainder →
|
||||
re-run the check before it counts as verified.**
|
||||
|
||||
**The generalisable finding — conditional on the broad reading.** The classifier
|
||||
judged cond 2 against the *source* and against citations it chose itself. It did
|
||||
not systematically judge it against **the rest of the same file**. Under the
|
||||
broad reading that is a defect: internal consistency is a cond-2 dimension the
|
||||
wave prompts never assigned, and idx 8 — the single `high` confidence record in
|
||||
the set — is the proof that it bites. Under the narrow reading it is not a
|
||||
defect at all; the same sweep is still worth running, but its output is a list
|
||||
of *additional* ungrounded claims rather than a list of corrections to the
|
||||
existing verdicts.
|
||||
**The generalisable finding.** The classifier judged cond 2 against the *source*
|
||||
and against citations it chose itself. It did not systematically judge it against
|
||||
**the rest of the same file**. Under the ratified reading that is a defect:
|
||||
internal consistency is a cond-2 dimension the wave prompts never assigned, and
|
||||
idx 8 — the single `high` confidence record in the set — is the proof that it
|
||||
bites. The remaining 15 have not had the check. It must be run per candidate
|
||||
before any of them reaches a ratifier, and its output corrects the existing
|
||||
cond-2 verdicts rather than merely adding to them; expect it to move some.
|
||||
|
||||
Either way the remaining 15 have not had the check, and it should be run per
|
||||
candidate before any of them reaches a ratifier. What changes with the reading
|
||||
is what the results *mean*, not whether to gather them — which is why the sweep
|
||||
can proceed before the scope question is settled, and the verdicts cannot.
|
||||
**Score after hand-verification: 1 of 46 clears both conditions, not 2.** Only
|
||||
idx 17 survives. That is the number §10 measurement #2 should be read with — the
|
||||
classifier's own "2 of 46" counted idx 8 on a cond-2 justification that was
|
||||
half-unchecked.
|
||||
|
||||
## Appendix A — the 15 admitted proposals, hand-verified
|
||||
|
||||
|
|
@ -537,10 +535,11 @@ deletes (§9.1, V2b), so it is a text change and must be reviewed as one.
|
|||
Hand-checking the two affirmative rows (§9.2) showed that half of idx 8's cond-2
|
||||
justification was never checked, and that the same file documents the deleted
|
||||
alternative in its cost section. Read every `yes` in this column as "the
|
||||
classifier believed this". Only **idx 17** has been hand-verified so far, and
|
||||
idx 8 is not ratifiable as written under either reading of cond 2 — but **which
|
||||
reading applies is an open operator decision** (§9.2), and it governs both the
|
||||
count and what a sweep of the remaining 15 would mean.
|
||||
classifier believed this". Only **idx 17** has been hand-verified. Idx 8 is not
|
||||
ratifiable as written: cond 2 is measured against the whole file (operator
|
||||
ratification, 2026-08-03, §9.2), and the file contradicts the remainder at lines
|
||||
263-265. **The verified score is 1 of 46, not 2.** The other 15 rows have not had
|
||||
the whole-file check.
|
||||
|
||||
Reproduce the tally and the machine checks:
|
||||
`node scripts/kb-eval/check-o2-returns.mjs`.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue