docs(ms-ai-architect): R11 §9.6 — idx 27 felt av kildeevidens; idx 26 skjerpet; G7 målt til fem, én er en live regresjon [skip-docs]
Hentet førstepartskildene for de to tilbakeholdte kandidatene. Ingen edit
anvendt; §9.6 er evidens.
idx 27 FELT. Copilot Studio-FAQ-en som den siterte hub-siden lenker til sier
"Makers can require user confirmation before executing tools that modify
data". Kilden etablerer altså mekanismen, så en sletting ville fjerne
kilde-bekreftet informasjon -> cond 3 feiler -> idx 27 forlater O2. Dette er
en felling, ikke en klarering, så asymmetri-regelen tillater den uten
ratifikator. Feilen er modalitet (built-in vs maker-konfigurerbar) = en
erstatning, ikke en subtraksjon.
Rammings-korreksjon: §9.4s reduksjon snevret idx 27 til nøyaktig den raden
cond 3 var uavklart for. Reduksjonen fortynnet ikke tvilen, den konsentrerte
den til 100 %.
idx 26 BEKREFTET og skjerpet. De kanoniske scorecard-segmentene er hentet;
Error analysis og Counterfactual analysis er dashboard-komponenter, ikke
scorecard-segmenter. Begge er gale. Men reduksjonen beholder item 4 (linje
300 hevder den også), så den tilbakeholdte halvdelen er nå MÅLT falsk, ikke
bare uavklart. Operatørspørsmålet er derfor to-delt, ikke ett.
REGRESJON funnet i 957ebef: idx 17-subtraksjonen etterlot en ledetekst som
ender i kolon med ingenting etter, rett før et **Verified**-stempel
(rag-caching-optimization.md:253). Subtraksjonen var korrekt; avsnittet den
etterlot er ikke. Kolon-valget ble tatt på maskin-grunnlag, men V1/V2/V2b/V3
er streng-invarianter og kan ikke se dokument-koherens.
G7 MÅLT, ikke ekstrapolert: 2 av 4 anvendte subtraksjoner etterlot en rest.
Medlemstall fem (17, 33, 26, 36, 18). To av dem (17, 33) er erstatninger,
ikke fler-lokator — som en delete-orientert O4-klasse ikke ville fikset.
Målingen taler for kandidatform (b), navngitt kø inn i menneskelig review.
Skår uendret 4 av 46. Suite 1032/1032.
This commit is contained in:
parent
a1295a97a2
commit
57a491ab66
1 changed files with 100 additions and 0 deletions
|
|
@ -736,6 +736,106 @@ against that number. This is the first cond-2 evidence gathered at corpus scope
|
||||||
rather than file scope, and whether corpus scope becomes a standing third reading
|
rather than file scope, and whether corpus scope becomes a standing third reading
|
||||||
of cond 2 is unratified and deliberately left open.
|
of cond 2 is unratified and deliberately left open.
|
||||||
|
|
||||||
|
### 9.6 The two held-back candidates resolved against first-party evidence
|
||||||
|
|
||||||
|
Run 2026-08-03, the session after the first corpus edit. §9.4 left idx 26 and 27
|
||||||
|
each carrying **one** human condition, and STATE framed both as operator calls.
|
||||||
|
Fetching the sources changed the answer for one of them and sharpened the other.
|
||||||
|
Neither was applied; §9.6 is evidence, not an edit.
|
||||||
|
|
||||||
|
**A framing correction that had to come first.** §9.4's reduction narrows idx 27's
|
||||||
|
subtraction to **the Plugin-actions row alone** — and that is precisely the row the
|
||||||
|
classifier's cond 3 said a human must confirm. So the reduction does not dilute the
|
||||||
|
open condition, it **concentrates it**: after reduction, 100 % of the edit is the
|
||||||
|
part nobody had cleared. The original two-row form had a clean half; the reduced
|
||||||
|
form has none. "One condition remains" understated it.
|
||||||
|
|
||||||
|
**idx 27 — FAILED, and the checker may say so unilaterally.** The file cites
|
||||||
|
`microsoft-copilot-studio/responsible-ai-overview` for the whole section. That page
|
||||||
|
is a hub; it establishes nothing itself, but it links the FAQ set, and
|
||||||
|
`faqs-generative-orchestration` states plainly:
|
||||||
|
|
||||||
|
> "Makers can require user confirmation before executing tools that modify data."
|
||||||
|
|
||||||
|
The source therefore **does** establish a confirmation mechanism for Copilot Studio.
|
||||||
|
Cond 3 asks whether the subtraction destroys source-confirmed information, and the
|
||||||
|
answer is yes — so **cond 3 fails and idx 27 leaves O2 entirely.** This is a kill,
|
||||||
|
not a clearance, so §9.3's asymmetry rule permits it without a ratifier.
|
||||||
|
|
||||||
|
The row is not fabricated; its **modality** is wrong. The file asserts confirmation
|
||||||
|
prompts as a *built-in disclosure*, whereas the source makes them a maker-configured
|
||||||
|
option. The correct repair is a replacement, not a deletion — which is an O3-class
|
||||||
|
fix, outside the delete-only envelope. Same for the Chat-interface row the reduction
|
||||||
|
kept: the FAQ documents a default transparency **message** ("Just so you are aware,
|
||||||
|
I sometimes use AI to answer your questions."), not the "Powered by AI" **badge** the
|
||||||
|
file claims. That row is not in this subtraction, but it is now on the record as
|
||||||
|
imprecise.
|
||||||
|
|
||||||
|
**idx 26 — confirmed, and its held-back half is now positively false.** The canonical
|
||||||
|
scorecard segments, from `how-to-responsible-ai-scorecard`, are: summary/model
|
||||||
|
overview, data analysis, model performance, cohorts, top important factors, fairness
|
||||||
|
insights, causal insights. `concept-responsible-ai-dashboard` lists Error analysis
|
||||||
|
and Counterfactual analysis as **dashboard** components. The classifier was right
|
||||||
|
about **both** items 4 and 5.
|
||||||
|
|
||||||
|
That matters because §9.4's reduction deletes item 5 **and deliberately keeps item
|
||||||
|
4**, since line 300 asserts Error analysis as scorecard content too. Before this
|
||||||
|
session item 4 was merely *uncleared*; it is now **measured false in two places**
|
||||||
|
(117 and 300). So the operator question is not one part but two:
|
||||||
|
|
||||||
|
1. accept the renumbering artifact (`1,2,3,4,6,7` — delete-only cannot renumber), and
|
||||||
|
2. accept that a **known-false** claim stays in a publicly distributed file, with the
|
||||||
|
117+300 pair booked to G7.
|
||||||
|
|
||||||
|
Presenting only (1) would let the whole-file reading of cond 2 quietly convert a
|
||||||
|
defect into a permanent resident.
|
||||||
|
|
||||||
|
**A regression this session found in the edit already shipped.** §9.4 recorded idx
|
||||||
|
17's residue as a `**Verified**` stamp sitting outside the verbatim block. The live
|
||||||
|
file shows something worse. `957ebef` deleted the three bands that followed the lead-in,
|
||||||
|
leaving (`rag-caching-optimization.md:253`):
|
||||||
|
|
||||||
|
```
|
||||||
|
**Score Threshold Tuning** (APIM `score-threshold` er en DISTANSE: …likhet):
|
||||||
|
|
||||||
|
**Verified** (Microsoft Learn - Enable semantic caching for LLM APIs)
|
||||||
|
```
|
||||||
|
|
||||||
|
A lead-in ending in a colon, promising an enumeration that no longer exists, followed
|
||||||
|
by a verification stamp. **The subtraction was correct; the paragraph it left is not.**
|
||||||
|
The colon was kept as an operator choice on the grounds that both variants were
|
||||||
|
machine-clean — and they were. V1/V2/V2b/V3 are string invariants over the deleted
|
||||||
|
text; **none of them can see document coherence.** Dropping the colon would not have
|
||||||
|
saved it either, since the lead-in is empty in both variants. This is a genuine
|
||||||
|
defect introduced by our own edit into public material, and it belongs to the
|
||||||
|
**replacement** class, not the subtraction class. The file already carries the
|
||||||
|
correct guidance twice (164 and 429: "Start med 0.15, tune opp basert på metrics").
|
||||||
|
|
||||||
|
**G7 measured rather than extrapolated.** Of the four subtractions applied in
|
||||||
|
`957ebef`, **two left a residue** (17 — the dangling lead-in; 33 — the CAF-attributed
|
||||||
|
survival at 357 plus five cross-file assertions) and two were clean (19, 14). Adding
|
||||||
|
the held-back set, G7's membership is now **five**, and one is a live regression:
|
||||||
|
|
||||||
|
| idx | residue | class |
|
||||||
|
|---|---|---|
|
||||||
|
| 17 | dangling lead-in at 253 + `**Verified**` stamp | replacement — **shipped, live** |
|
||||||
|
| 33 | line 357 CAF attribution, stamped `Verified MCP 2026-04`, + 5 cross-file | replacement / corpus-scope |
|
||||||
|
| 26 | item 4 at 117 kept, asserted again at 300 — **measured false** | multi-locator |
|
||||||
|
| 36 | companion edit at 310 | multi-locator |
|
||||||
|
| 18 | whole section 303-318 + `**Verified**` row 510 | multi-locator |
|
||||||
|
|
||||||
|
**This is input the §9.4 write-up did not have, and it tilts the (a)/(b) choice.**
|
||||||
|
A 50 % residue rate on applied edits means residues are not an exception to be
|
||||||
|
queued; they are the **normal by-product** of a delete-only envelope. A named queue
|
||||||
|
into human review (b) absorbs a steady stream. An explicit multi-locator class with
|
||||||
|
its own return contract (a) would have to be built for the common case, not the
|
||||||
|
edge — and note that two of the five (17, 33) are not multi-locator at all but
|
||||||
|
**replacements**, which an O4 deletion-oriented class would not fix. On this
|
||||||
|
measurement (b) is the better fit, and (a) would be mis-sized against the evidence.
|
||||||
|
|
||||||
|
**Verified score is unchanged at 4 of 46.** idx 27 leaves O2 without becoming a
|
||||||
|
subtraction; idx 26 remains available to a ratifier as a two-part accept.
|
||||||
|
|
||||||
## Appendix A — the 15 admitted proposals, hand-verified
|
## Appendix A — the 15 admitted proposals, hand-verified
|
||||||
|
|
||||||
Every proposal the classifier (§4 + context condition) admitted over the whole
|
Every proposal the classifier (§4 + context condition) admitted over the whole
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue