docs(changelog): the entry belongs in the untagged 0.10.1 section, not under [Unreleased]
Found by the clean-extraction run, not by inspection: `test_the_changelog_heads_with_the_packaged_version` requires the first `## [x]` heading to be the packaged version with a date, and 0.10.1 is the version being prepared and not yet tagged. The content is unchanged; the two subsections say when they were written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
62c82a3435
commit
f5bf5fccbf
1 changed files with 59 additions and 61 deletions
120
CHANGELOG.md
120
CHANGELOG.md
|
|
@ -5,67 +5,6 @@ All notable changes to this project will be documented in this file.
|
||||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||||
|
|
||||||
## [Unreleased]
|
|
||||||
|
|
||||||
### Changed
|
|
||||||
|
|
||||||
- **A withheld concept now carries the rule that actually decided it.** The
|
|
||||||
source quota filters the WHOLE ranked candidate list rather than the top
|
|
||||||
`k`, so every over-quota candidate came back `source_quota_exceeded` —
|
|
||||||
including the ones the RANK had already put outside `k`, which the quota
|
|
||||||
only reached because it ran first. Measured on 25 real misses 2026-09-17:
|
|
||||||
**13 of them** were labelled by the quota and decided by the rank.
|
|
||||||
`consume._fates_without_quota` asks the same cut what would have become of
|
|
||||||
each candidate with no quota in force, and the drop keeps THAT rule; only a
|
|
||||||
candidate the quota-off cut would have delivered is named as the quota's.
|
|
||||||
The budget step is lifted into `consume._pack` and used by both, so the
|
|
||||||
quota-off fate is decided by the code the run itself uses. This moves the
|
|
||||||
`rule` string a consumer reads for some withheld entries; no delivery, no
|
|
||||||
rank and no excerpt byte moves, and no committed payload in this repository
|
|
||||||
changed. The retrieval gate's row 3 goes **2 of 5 RED to 5 of 5 GREEN**.
|
|
||||||
|
|
||||||
### Added
|
|
||||||
|
|
||||||
- **The payload says what of the question it reached** — a new top-level
|
|
||||||
`coverage` member carrying three lists: the terms the pre-pass read the
|
|
||||||
question as, the terms no concept in the bundle answers, and the terms no
|
|
||||||
delivered excerpt answers. Without it a reader holding eight excerpts cannot
|
|
||||||
tell a bundle that ANSWERED its question from one that merely ranked
|
|
||||||
something; the two payloads have the same shape. Documented as SS 8 point 7
|
|
||||||
of `docs/consumption-contract.md`, and the generated consumption skill is
|
|
||||||
told to read it.
|
|
||||||
- **Facts, and no verdict, which is a measurement rather than caution.** Two
|
|
||||||
readings were built and both falsified over **81 questions** (16 synthetic,
|
|
||||||
65 across three real gold sets, 2026-09-20): the share of a question's
|
|
||||||
terms a delivered excerpt answers separates the synthetic controls at 0.33
|
|
||||||
against 0.50 and then REVERSES on real data, where covered questions run
|
|
||||||
down to 0.27 while one genuinely uncovered question sits at 0.71; and the
|
|
||||||
share of a bundle tying the best lexical match is ~0.00 for every question
|
|
||||||
in a large bundle, covered or not. Question style dominates the first and
|
|
||||||
corpus size the second, so a pre-pass emitting a verdict would assert
|
|
||||||
across corpora what was measured on one.
|
|
||||||
- **Contract change, and the cost to a consumer is a re-run.** Every payload
|
|
||||||
grows the member; the checker does not read it, so a third-party pre-pass
|
|
||||||
that omits it stays conformant. The SS 7.4 known-positive moves with the
|
|
||||||
document it is measured on (14 721 / 375 → 16 389 / 417).
|
|
||||||
- **The retrieval gate is measurable where it was assertable**
|
|
||||||
(`tools/okf_retrieval_gate.py`, not shipped in the wheel):
|
|
||||||
- Row 8 prints the identity of every bundle it measured — path,
|
|
||||||
`bundle_id` and content ref — beside the set's sha256. Measured the same
|
|
||||||
day: two wiki bundles carrying the one `bundle_id` `claude-code-wiki` at
|
|
||||||
different refs score **6 of 20** and **5 of 20** on the same pinned set.
|
|
||||||
- `REAL_SET_PINS` states what each of the three real sets IS — questions,
|
|
||||||
fasit entries, controls and sha256 — so a self-written file in the right
|
|
||||||
shape is refused instead of reading `1 of 1 | 3 of 3 | GREEN`.
|
|
||||||
- Row 5 reads the hold-out threshold as a number in [0, 1] and RUNS the
|
|
||||||
registered set against the registered bundle, printing
|
|
||||||
`answered of asked = share against threshold`. `bool(threshold)` was the
|
|
||||||
whole check, so `report-only; any number is acceptable for v1` passed it.
|
|
||||||
- Row 4's marking reads the payload's `coverage`: `UNANSWERED_BAR = 2/3`,
|
|
||||||
swept and collapsing at both ends (at 0.50 eleven real covered questions
|
|
||||||
are marked; at 0.70 the row falls to 5 of 6). The margin is thin — 0.6087
|
|
||||||
against 0.6667 — and what it does not catch is published with it.
|
|
||||||
|
|
||||||
## [0.10.1] — 2026-09-19
|
## [0.10.1] — 2026-09-19
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|
@ -633,6 +572,65 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||||
"degenerate" means zero characters after stripping whitespace and a pointer
|
"degenerate" means zero characters after stripping whitespace and a pointer
|
||||||
block is text.
|
block is text.
|
||||||
|
|
||||||
|
### Changed after the 0.10.1 notes were written (2026-09-20, untagged)
|
||||||
|
|
||||||
|
- **A withheld concept now carries the rule that actually decided it.** The
|
||||||
|
source quota filters the WHOLE ranked candidate list rather than the top
|
||||||
|
`k`, so every over-quota candidate came back `source_quota_exceeded` —
|
||||||
|
including the ones the RANK had already put outside `k`, which the quota
|
||||||
|
only reached because it ran first. Measured on 25 real misses 2026-09-17:
|
||||||
|
**13 of them** were labelled by the quota and decided by the rank.
|
||||||
|
`consume._fates_without_quota` asks the same cut what would have become of
|
||||||
|
each candidate with no quota in force, and the drop keeps THAT rule; only a
|
||||||
|
candidate the quota-off cut would have delivered is named as the quota's.
|
||||||
|
The budget step is lifted into `consume._pack` and used by both, so the
|
||||||
|
quota-off fate is decided by the code the run itself uses. This moves the
|
||||||
|
`rule` string a consumer reads for some withheld entries; no delivery, no
|
||||||
|
rank and no excerpt byte moves, and no committed payload in this repository
|
||||||
|
changed. The retrieval gate's row 3 goes **2 of 5 RED to 5 of 5 GREEN**.
|
||||||
|
|
||||||
|
### Added after the 0.10.1 notes were written (2026-09-20, untagged)
|
||||||
|
|
||||||
|
- **The payload says what of the question it reached** — a new top-level
|
||||||
|
`coverage` member carrying three lists: the terms the pre-pass read the
|
||||||
|
question as, the terms no concept in the bundle answers, and the terms no
|
||||||
|
delivered excerpt answers. Without it a reader holding eight excerpts cannot
|
||||||
|
tell a bundle that ANSWERED its question from one that merely ranked
|
||||||
|
something; the two payloads have the same shape. Documented as SS 8 point 7
|
||||||
|
of `docs/consumption-contract.md`, and the generated consumption skill is
|
||||||
|
told to read it.
|
||||||
|
- **Facts, and no verdict, which is a measurement rather than caution.** Two
|
||||||
|
readings were built and both falsified over **81 questions** (16 synthetic,
|
||||||
|
65 across three real gold sets, 2026-09-20): the share of a question's
|
||||||
|
terms a delivered excerpt answers separates the synthetic controls at 0.33
|
||||||
|
against 0.50 and then REVERSES on real data, where covered questions run
|
||||||
|
down to 0.27 while one genuinely uncovered question sits at 0.71; and the
|
||||||
|
share of a bundle tying the best lexical match is ~0.00 for every question
|
||||||
|
in a large bundle, covered or not. Question style dominates the first and
|
||||||
|
corpus size the second, so a pre-pass emitting a verdict would assert
|
||||||
|
across corpora what was measured on one.
|
||||||
|
- **Contract change, and the cost to a consumer is a re-run.** Every payload
|
||||||
|
grows the member; the checker does not read it, so a third-party pre-pass
|
||||||
|
that omits it stays conformant. The SS 7.4 known-positive moves with the
|
||||||
|
document it is measured on (14 721 / 375 → 16 389 / 417).
|
||||||
|
- **The retrieval gate is measurable where it was assertable**
|
||||||
|
(`tools/okf_retrieval_gate.py`, not shipped in the wheel):
|
||||||
|
- Row 8 prints the identity of every bundle it measured — path,
|
||||||
|
`bundle_id` and content ref — beside the set's sha256. Measured the same
|
||||||
|
day: two wiki bundles carrying the one `bundle_id` `claude-code-wiki` at
|
||||||
|
different refs score **6 of 20** and **5 of 20** on the same pinned set.
|
||||||
|
- `REAL_SET_PINS` states what each of the three real sets IS — questions,
|
||||||
|
fasit entries, controls and sha256 — so a self-written file in the right
|
||||||
|
shape is refused instead of reading `1 of 1 | 3 of 3 | GREEN`.
|
||||||
|
- Row 5 reads the hold-out threshold as a number in [0, 1] and RUNS the
|
||||||
|
registered set against the registered bundle, printing
|
||||||
|
`answered of asked = share against threshold`. `bool(threshold)` was the
|
||||||
|
whole check, so `report-only; any number is acceptable for v1` passed it.
|
||||||
|
- Row 4's marking reads the payload's `coverage`: `UNANSWERED_BAR = 2/3`,
|
||||||
|
swept and collapsing at both ends (at 0.50 eleven real covered questions
|
||||||
|
are marked; at 0.70 the row falls to 5 of 6). The margin is thin — 0.6087
|
||||||
|
against 0.6667 — and what it does not catch is published with it.
|
||||||
|
|
||||||
## [0.10.0] — 2026-09-17
|
## [0.10.0] — 2026-09-17
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue