Commit graph

5 commits

Author SHA1 Message Date
8e9f385d17 test(spec): make the pulled method spec load-bearing at rule level, not just section level
The spec arrives by subtree pull from commons, so drift is something we RECEIVE
rather than author. The existing guard only asserted the section HEADINGS were
present: a pull could empty a section of its normative content and stay green.
Measured, not assumed — with `escape, not depth` deleted from Step 1, the old
skeleton test still passes (M5 below).

Two new guards, one file:

- Rule level (test 6): 25 normative rules, each bound to the section that OWNS
  it, so a phrase surviving elsewhere in the document does not count as the spec
  still stating the rule. Selection is principled rather than taste — every entry
  anchors a §11 conformance seam or a CLAUDE.md invariant. Matching is
  whitespace- and emphasis-insensitive: re-wrapping a paragraph upstream is an
  editorial change, not a rule change, and must not cost a false red.
- §11 table (test 7): the conformance table still names the seams this suite
  implements, and every test it cites exists here. One-directional by design —
  the repo carries 26 load-bearing tests, the spec anchors 10, because the spec
  governs the METHOD, not this repo.

Five measurements, all as expected: M1 delete a rule → RED · M2 RELOCATE a rule
out of its section while leaving it in the document → RED (this is what the
section binding buys) · M3 drop a seam row from §11 → RED · M4 harmless re-wrap
+ de-bolding → GREEN · M5 the old skeleton test does NOT catch M1 → GREEN.
Spec restored from a byte copy after each mutation, shasum verified.

Known gap, deliberately not patched here: §11 has no row for the portfolio-wide
budget seam (S3.4/F10) — the pulled spec predates it. That is a commons
amendment; a consumer editing shared/ is exactly the drift this guards against.

553 → 555 tests, ruff + mypy green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aw9TECznT5b5H6374CBTvV
2026-08-01 19:57:33 +02:00
0a11af74a4 refactor(ingest): adopt shared llm-ingestion-okf v0.3.1 behind a thin adapter
Door A (manifest -> connector -> deterministic materialization -> index) is no
longer implemented here. src/portfolio_optimiser/ingest.py becomes a thin
consumer seam over the shared library, git-pinned to v0.3.1 on the same Forgejo
channel portfolio-optimiser-claude uses. Net -626/+385; ingest.py 599 -> 145 lines.

shared/ingest-spec.md remains the normative spec: the library implements it, it
does not replace it. Spec changes continue to go via commons.

Acceptance criterion met and proven: all three golden bundles (file/sql/http)
are byte-exact before and after, including the idempotence re-run. examples/ and
shared/ carry ZERO modifications -- the fasit was not adjusted to fit.

The rejection set was verified equivalent, not assumed: all 22 malformations the
repo's pydantic models refused are refused by the library, with typed codes
(okf_type_reserved, credential_embedded, extraction_id_duplicate, ...).

Test rebinding (invariants preserved, vehicle changed): the library has zero
runtime dependencies by design, so pydantic is unavailable to it.
ManifestV1.model_validate(dict) -> load_manifest_bytes(bytes); ValidationError ->
ManifestError; model_fields -> dataclasses.fields; PathSecurityError ->
SourceError(path_escape); ValueError -> MaterializationError(ingested_at_invalid).
Tests now also pin the refusal `code`, the library's documented stability
contract -- a sharper assertion than "some validation error was raised".

Two accepted behavioural deltas, recorded rather than silently dropped:
- Title whitespace is stored verbatim instead of collapsed at validation, so the
  frontmatter title and the index label are no longer guaranteed identical for
  irregular whitespace. Both behaviours are spec-conformant (the spec is SILENT;
  the old one was a repo-local pinned decision). Queued as a commons-amendment
  candidate so both stacks pin the same answer. Goldens unaffected.
- The section 8 audit log moves to logger llm_ingestion_okf.materialize. Nothing
  in the repo consumed the old channel.
Also: the `type` discriminator is no longer a dataclass field, so the spec
cross-check asserts it explicitly -- without that line the swap would have
silently narrowed the test.

New tests/test_ingest_library_seam.py pins the seam itself: the restated section 5
stamp formula against the stamp the library actually writes (the one place the
adapter does not purely delegate, since v0.3.1 exposes no stamp helper), the
local-only allow_network default, the list[Path] unwrapping, and a guard that the
adapter never regrows local Door A machinery. All four verified RED when detached,
as were both golden regressions under a byte-level render mutation.

Door A is UNGATED: it calls no guard before writing to disk. Gating untrusted
content remains the caller's responsibility (guard wiring still planned).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B4jNN186eVqfe1x5DnTU6r
2026-07-20 07:47:55 +02:00
dbf5dc5c13 test(spec-guard): ingest contract field cross-check vs spec §12 (I2) 2026-07-03 18:38:35 +02:00
7ba0fae933 test(spec-guard): extend framework guard + structure test to ingest-spec.md (I1)
The neutrality guard now covers BOTH shared specs explicitly (a new spec
file is never guarded implicitly), and a new structure test pins the
ingest spec's normative skeleton incl. the verdict-layer reservation and
the explicit-timestamp field. Proven RED against a throwaway copy with a
framework name injected and with a required section dropped; field-level
cross-check vs contract code arrives with the I2 contracts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaQCFnfsh3tfq1VfzdJpoi
2026-07-03 14:05:39 +02:00
5de1c93b69 feat(spec): S2 — normative framework-neutral method spec as the fourth shared artifact
Author shared/method-spec.md: the 8-step loop (normative, RFC-2119), the
verdict JSON contract incl. the id-minting algorithm and the chosen conflict
semantics, the inbox/outbox folder contract, the fail-closed promotion-gate
semantics, the IR projection + golden suite as the only ground truth (incl.
the reproducible Monte Carlo procedure), and the budget/stop, provenance and
startup-contract requirements — every normative claim cross-checked against
the load-bearing tests/code. The sibling implementation builds from this spec
alone.

Load-bearing trio (tests/test_method_spec_loadbearing.py, persona-trio
style): required structure, a name-shaped framework-neutrality guard over the
spec + the persona skill tree, and a cross-check-completeness test driven
from the REAL artifacts and the REAL verdict serializer (red on code drift).
All three detach points proven RED (missing file / framework name / dropped
field). shared/README.md: the "(planned)" line replaced with the real entry.

Suite 152 -> 155 passed / 4 skipped; ruff check+format clean; mypy src clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaQCFnfsh3tfq1VfzdJpoi
2026-07-03 00:50:42 +02:00