docs(f16): the 1.18.0 assessment measured -- 0 of 19 forms moved, three features assessed and none built

The F15 report's shape repeated: result table, denominator discipline, one
command per claim. What the measurement changed against the order's own
framing:

- orchestrations CANNOT be lifted (1.1.1 is still the head on PyPI), so F15's
  "the two floors are lifted together, there is no partial bump" is measurably
  weaker after F16 -- and eight of the nineteen probe checks live in that
  unmoved package, which is a property of the denominator, not a strength of
  the probe.
- foundry/openai were NOT forced up; the coupling the order read belongs to the
  LATEST releases, not the pinned 1.8.2. They were lifted for a measured reason
  instead (two BREAKING bullets name core AND foundry in the same line).
- ONE of the order's transcriptions deviated from the source: #8219 lazy
  loading is foundry/foundry-hosting/openai -- NOT core. Consequence, not
  pedantry: had the floors stayed at 1.8.2, the change the order named as the
  prime suspect for golden stderr would never have reached po at all.
- (a) touches NO U row. The wall-clock timeout is B4's own sketch word and G1;
  and the primitive is `max_duration_seconds` degrading gracefully via
  `budget_state["truncated"]` + a log line -- NOT a typed stop reason, so it
  does not satisfy B4's "every breach -> a structured event, never a silent
  stop". Recommendation: do not adopt it as B4's half. STATE's prohibition
  stands, untouched.
- (b) does not move U13: `approval_mode` is 0 in src/ and 0 in tests/ (102 in
  the venv). #7988 hardens the cooperative-marker path G4 already refused.
- (c) does not break: po's own parameters are typed `Sequence[Any] | None` and
  all seven callsites pass lists, so sequence-only inputs are the only form po
  can produce. 74 passed across the five middleware-bearing suites; no fix was
  needed and none was made.

My own probe was WRONG FIRST and that is written down: five checks read
CHANGED/MISSING in BOTH versions because my queries sliced a Protocol stub,
counted `self`, and demanded two names be adjacent. Running it against BOTH
versions is what made the instrument failure visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Kjell Tore Guttormsen 2026-09-12 17:18:46 +02:00
commit 9f14c642c8
2 changed files with 524 additions and 1 deletions

View file

@ -367,7 +367,7 @@ en EGEN økt etter denne planen (bonus: `InvocationsHostServer` krever core ≥1
| 4 | **U15** — EGEN sammenligning på golden-bundelen: utforskning vs. debatt på tre tall (validerte kandidater, tokens, runder), skrevet til `docs/` som måling | U15 | JA for «finner den bedre tiltak» — men MAF-evals AVVISES (eksperimentell i kode, «wrong shape», kapabilitetskart :58) | 59 | Opus 5/high |
| — | U1 `ConcurrentBuilder` | U1 | NEI — bølgemodellen er målt, bærer budsjett-reservasjon (S3.4); Concurrent har ingen runder å reservere mot | etter planen, om noensinne | — |
| — | U2 `WorkflowBuilder` | U2 | NEI — Magentic ER graf-laget; direkte adopsjon gir null ny kapabilitet | nei | — |
| — | U5 `SkillsProvider` | U5 | NEI — `ExperimentalFeature.SKILLS`; egen loader virker; commons eier innholdet | etter planen | — |
| — | U5 `SkillsProvider` | U5 | NEI — egen loader virker og commons eier innholdet; eksperimentalitets-grunnen er BORTE (målt mot 1.18.0, F16: `SkillsProvider`/`FileSkillsSource` bærer ingen `@experimental`-markør, de tre i `_skills.py` er alle `ExperimentalFeature.MCP_SKILLS`, og MAFs egen `FileSkillsSource` laster våre to skills 2 av 2) | etter planen | — |
| — | U11 MAF RAG-providere | U11 | NEI — «navigate, never stuff» forbyr query-time retrieval; sitatbæring finnes | avvist med begrunnelse | — |
Seks økter (5459); faller S4, sparer pkt. 3 én økt.