docs(limitations): the entropy headroom is 0.099, not 0.36 -- measured in the field
ms-ai-architect re-measured after retracting their entropy refutation (their original finding scored whole filenames as one unit, so it was never a finding about our rule). The retraction confirms the branch; the re-measurement produced a number we did not have, and it goes against us. The sharpest legitimate token in their 2400-URL corpus is a 78-character percent-escaped lovdata title at H=4.301. That is 0.099 from the 4.4 floor where this bullet previously implied 0.36 -- the mechanism was already ours, but the worked example understated how thin the margin gets in real Norwegian government paths. Verified against our own code before recording, not reconstructed: our real _URL_TOKEN_RE emits it as a single 78-character token, our shannon_entropy returns 4.301, and is_ordinary_url returns False. Their three calibration examples also reproduce to three decimals here. Threshold stays. Their own conclusion too -- moot per URL, since the % rule already disqualifies it before entropy is consulted. Records the 2400-vs-2401 provenance split explicitly. The counts come from two different harvesters over the same knowledge base and differ by one URL; the neighbouring hex bullet cites 2401. Left unstated, a later reader would reasonably merge them into one corpus figure, which is exactly the prose-vs-measurement conflation this file exists to prevent.
This commit is contained in:
parent
434b74e766
commit
05ee26aec8
1 changed files with 11 additions and 0 deletions
|
|
@ -169,6 +169,17 @@ items; this is the full list, each with the mechanism.
|
|||
short ones. Measured at H=4.04 — under the 4.4 floor, and moot in practice because
|
||||
the `%` rule already disqualifies the URL. It is recorded because it means the
|
||||
length floor does less work in non-ASCII paths than the calibration assumed.
|
||||
**A field re-measurement (2026-07-31) sharpens this by roughly 3x.** The
|
||||
highest-entropy *legitimate* token in a consumer's 2400-URL corpus is not the one
|
||||
above but a 78-character percent-escaped lovdata title, `Fra%20%C3%A5ndsverk%20…`,
|
||||
at H=4.301 — leaving **0.099 of headroom to the 4.4 floor, not 0.36**. Re-run here
|
||||
through the real `_URL_TOKEN_RE` and `shannon_entropy` rather than a reconstruction:
|
||||
our tokenizer does emit it as one 78-character token. Still moot per URL
|
||||
(`is_ordinary_url` is False on the `%` rule) and still not a reason to move the
|
||||
threshold — but the margin this bullet reports is much thinner than 4.04 implies.
|
||||
Note the corpus count: 2400 is a rebuilt harvester's, where the 2401 in the next
|
||||
bullet is the earlier run's. Same knowledge base, different scripts, one URL apart —
|
||||
they are not interchangeable and must not be summed or quoted as one figure.
|
||||
- **Legitimate CDN content-asset ids trip the hex branch, permanently.** A ≥32-char
|
||||
hex path segment is opaque by design, and several public CMSes mint asset URLs that
|
||||
way (measured: 9 distinct on `regjeringen.no`, `ks.no`, `datatilsynet.no`). This is
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue