feat(check): the checker and the contract read a folder's reply
`okf check --payload` takes the reply to one call over a folder as well as a single payload: every bundle's payload is held to all 19 rules on its own, a finding is named with its bundle, one every payload carries alike is reported once, an answer labelled with a bundle its payload does not describe is `answer_misattributed`, and a reply with no answer is `payload_invalid`. No rule is added, and a single payload's report is unchanged. Contract SS 2.5.4 names the folder run and SS 8.11 fixes the reply; the known-positive moves to 24 620 / delta 592. The skill text follows: the working method's steps 1 and 4 name the folder, and the generic skill says to use the server's tools first where they are registered, with the skill as the supplement. The folder is an instruction in both generators, never a path: the bundle's parent written absolute named this checkout, and the test holding generated commands to no repository path fell on it. v1.1 order F, part F4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
570496470b
commit
21f9241712
10 changed files with 240 additions and 29 deletions
16
CLAUDE.md
16
CLAUDE.md
|
|
@ -1540,6 +1540,19 @@ R761 **8** (S1-S6 + KP + KN), vegnormal **32** questions / **43**
|
||||||
of them, so over a folder it is REFUSED by name with exit 2
|
of them, so over a folder it is REFUSED by name with exit 2
|
||||||
(`consume.FOLDER_FLAGS` is the allowlist), never dropped; `--bundle-id` on a
|
(`consume.FOLDER_FLAGS` is the allowlist), never dropped; `--bundle-id` on a
|
||||||
bundle path is refused the same way. A bundle path reads exactly as before.
|
bundle path is refused the same way. A bundle path reads exactly as before.
|
||||||
|
- **`okf check` READS A FOLDER'S REPLY AND THE SKILL TEXT SAYS SO (v1.1 F4).**
|
||||||
|
`contract_check.check_reply`: a reply carrying `answers` and no `bundle` is
|
||||||
|
one payload per bundle, each held to all 19 rules on its own; a finding is
|
||||||
|
named `[bundle_id]`, one every payload carries alike is reported once
|
||||||
|
unnamed (it is the SKILL's), an answer whose label is not its payload's
|
||||||
|
bundle is `answer_misattributed`, no answer at all is `payload_invalid`. No
|
||||||
|
rule was added to `RULES` -- the count stays 19 and a single payload's report
|
||||||
|
is byte-for-byte as before; a folder's report says `over N payloads`.
|
||||||
|
Contract SS 2.5.4 names the folder run, SS 8.11 fixes the reply. The
|
||||||
|
template's step 1 and 4 name the folder (`<FOLDER>`: both generators fill
|
||||||
|
a lower-case instruction, never a path -- the bundle's parent written
|
||||||
|
absolute named a checkout, and `test_the_generated_commands_name_this_repository_nowhere` fell on it); the generic header says the server comes first and the skill
|
||||||
|
is the supplement, and that `--ref` belongs to one bundle.
|
||||||
- **THE SERVER IS THE STANDARD WAY IN AND THE SKILL THE SUPPLEMENT (v1.1 F3,
|
- **THE SERVER IS THE STANDARD WAY IN AND THE SKILL THE SUPPLEMENT (v1.1 F3,
|
||||||
operator 2026-09-21).** `okf project`'s closing text and README's first
|
operator 2026-09-21).** `okf project`'s closing text and README's first
|
||||||
screen say it in that order: register `okf mcp --root` once (every project,
|
screen say it in that order: register `okf mcp --root` once (every project,
|
||||||
|
|
@ -1601,7 +1614,8 @@ R761 **8** (S1-S6 + KP + KN), vegnormal **32** questions / **43**
|
||||||
form (`okf-consumption/2`'s `withheld` mapping, `absent_terms`/`weak`,
|
form (`okf-consumption/2`'s `withheld` mapping, `absent_terms`/`weak`,
|
||||||
`passage`, `questions`/`subquestions`, `own_title`); `okf check` gains
|
`passage`, `questions`/`subquestions`, `own_title`); `okf check` gains
|
||||||
`passage_malformed` and `subquestions_unindexed` (19 rules). Editing the
|
`passage_malformed` and `subquestions_unindexed` (19 rules). Editing the
|
||||||
contract moved the known-positive to 23 672 / delta 580.
|
contract moved the known-positive to 23 672 / delta 580, and v1.1 F4's
|
||||||
|
SS 2.5.4 / SS 8.11 edit to **24 620 / delta 592** (`wc -c` 24 028).
|
||||||
What follows describes the fusion.
|
What follows describes the fusion.
|
||||||
- Consume a bundle: `okf consume <bundle> --question "<q>"
|
- Consume a bundle: `okf consume <bundle> --question "<q>"
|
||||||
[--k N] [--limit N] [--out PATH] [--ref IDENTITY]` — the **pre-pass**
|
[--k N] [--limit N] [--out PATH] [--ref IDENTITY]` — the **pre-pass**
|
||||||
|
|
|
||||||
|
|
@ -1429,6 +1429,12 @@ The flags that change how ONE bundle is cut (`--ref`, `--ranking`,
|
||||||
than dropped, because the server takes none of them; point at one bundle to use
|
than dropped, because the server takes none of them; point at one bundle to use
|
||||||
them.
|
them.
|
||||||
|
|
||||||
|
`okf check --payload` takes that reply as well as a single payload: every
|
||||||
|
bundle's payload is held to every rule on its own, a finding is named with its
|
||||||
|
bundle, and an answer labelled with a bundle its payload does not describe is a
|
||||||
|
finding (`answer_misattributed`). The generic skill tells its reader both
|
||||||
|
forms, and says to use the server's tools first where they are registered.
|
||||||
|
|
||||||
`okf skill <bundle> --for-bundle` still writes the per-bundle form, with the
|
`okf skill <bundle> --for-bundle` still writes the per-bundle form, with the
|
||||||
identity and the numbers measured into the text — which is exactly what makes
|
identity and the numbers measured into the text — which is exactly what makes
|
||||||
that file stale the moment the bundle is rebuilt. It refuses out loud when it
|
that file stale the moment the bundle is rebuilt. It refuses out loud when it
|
||||||
|
|
|
||||||
|
|
@ -78,8 +78,9 @@ searches — and MUST NOT state one that stops at a single run.
|
||||||
`withheld` near misses (§ 5.3) and § 2.2 exist so that the second run can
|
`withheld` near misses (§ 5.3) and § 2.2 exist so that the second run can
|
||||||
be aimed.
|
be aimed.
|
||||||
4. Where more than one bundle is in scope, it MUST tell its reader to run the
|
4. Where more than one bundle is in scope, it MUST tell its reader to run the
|
||||||
same sub-questions against each and to keep each piece of material
|
same sub-questions against each — in ONE run over the folder that holds
|
||||||
attributed to its bundle.
|
them where the pre-pass takes a folder (§ 8.11) — and to keep each piece of
|
||||||
|
material attributed to its bundle.
|
||||||
5. It MUST tell its reader to assemble ONE answer — ordered by sub-question,
|
5. It MUST tell its reader to assemble ONE answer — ordered by sub-question,
|
||||||
stating which source holds where sources disagree and with which version,
|
stating which source holds where sources disagree and with which version,
|
||||||
and saying what the bundle does not cover.
|
and saying what the bundle does not cover.
|
||||||
|
|
@ -360,6 +361,17 @@ are permitted; the checker reads only the members this section names.
|
||||||
as `title` the title of the concept it stands under in the same document,
|
as `title` the title of the concept it stands under in the same document,
|
||||||
and then MUST carry the file's own title as `own_title`, so the name shown
|
and then MUST carry the file's own title as `own_title`, so the name shown
|
||||||
is never mistaken for the one in the file.
|
is never mistaken for the one in the file.
|
||||||
|
11. A pre-pass MAY take a FOLDER of bundles and ask every bundle under it in
|
||||||
|
one run. Its reply is then not a payload but a list of them: `asked` (the
|
||||||
|
bundle ids, in order), `budget_per_bundle`, and `answers`, one
|
||||||
|
`{bundle_id, payload}` per bundle, each payload conformant on its own and
|
||||||
|
cut to its share of the budget; `question` or `questions` as point 9. The
|
||||||
|
reply carries no `bundle` of its own, which is how a reader tells the two
|
||||||
|
apart. The checker holds every payload to every rule, names a finding with
|
||||||
|
the bundle whose payload carries it, reports once a finding every payload
|
||||||
|
carries alike, and refuses an answer labelled with a bundle its payload
|
||||||
|
does not describe (`answer_misattributed`) — a claim is attributed to the
|
||||||
|
label — and a reply with no answer at all (`payload_invalid`).
|
||||||
|
|
||||||
## 9. Prohibitions
|
## 9. Prohibitions
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -59,7 +59,9 @@ asked in the wrong words reaches the wrong concepts however good the ranking is.
|
||||||
|
|
||||||
**1. Understand the question first.** Read the bundle's `map` before you search
|
**1. Understand the question first.** Read the bundle's `map` before you search
|
||||||
it — `okf card <BUNDLE_ROOT>` prints it: one line per document with its section
|
it — `okf card <BUNDLE_ROOT>` prints it: one line per document with its section
|
||||||
titles, a series of like-named documents as one line. Then put the question
|
titles, a series of like-named documents as one line. Pointed at a FOLDER of
|
||||||
|
bundles, the same command lists every bundle under it, each with its map, so
|
||||||
|
you see what each one covers before you choose. Then put the question
|
||||||
into the bundle's own words: a bundle written in one language and a question
|
into the bundle's own words: a bundle written in one language and a question
|
||||||
asked in another share few tokens, and the pre-pass matches tokens. Take the
|
asked in another share few tokens, and the pre-pass matches tokens. Take the
|
||||||
terms from the map's titles, not from your vocabulary.
|
terms from the map's titles, not from your vocabulary.
|
||||||
|
|
@ -88,10 +90,19 @@ is worth carrying. When `coverage.weak` is true — a word of yours the bundle
|
||||||
holds in no form (`coverage.absent_terms`), or nothing came back — rephrase in
|
holds in no form (`coverage.absent_terms`), or nothing came back — rephrase in
|
||||||
the bundle's own words, and if it stays weak, say the bundle does not cover it.
|
the bundle's own words, and if it stays weak, say the bundle does not cover it.
|
||||||
|
|
||||||
**4. Several bundles, same method.** When more than one bundle could answer,
|
**4. Several bundles, one run.** When more than one bundle could answer, give
|
||||||
run the same sub-questions against each, and keep track of which bundle each
|
the pre-pass the FOLDER that holds them instead of one bundle: it asks every
|
||||||
piece of material came from. A claim is attributed to its bundle as well as its
|
bundle under the folder with the same sub-questions in ONE run, splits the
|
||||||
concept — two bundles can hold the same sentence with different authority.
|
budget between them, and names the bundle on every answer and every excerpt.
|
||||||
|
`--bundle-id` narrows it to one of them.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
okf consume <FOLDER> --question "first sub-question" --question "second sub-question" --out /tmp/p1.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Keep track of which bundle each piece of material came from. A claim is
|
||||||
|
attributed to its bundle as well as its concept — two bundles can hold the same
|
||||||
|
sentence with different authority.
|
||||||
|
|
||||||
**5. Put it together.** Order the material by sub-question, not by rank. Where
|
**5. Put it together.** Order the material by sub-question, not by rank. Where
|
||||||
sources disagree, decide what holds NOW: the newest documentation or the
|
sources disagree, decide what holds NOW: the newest documentation or the
|
||||||
|
|
|
||||||
|
|
@ -70,7 +70,9 @@ asked in the wrong words reaches the wrong concepts however good the ranking is.
|
||||||
|
|
||||||
**1. Understand the question first.** Read the bundle's `map` before you search
|
**1. Understand the question first.** Read the bundle's `map` before you search
|
||||||
it — `okf card examples/ingest-golden-segmented-okf-v0-2/expected-bundle` prints it: one line per document with its section
|
it — `okf card examples/ingest-golden-segmented-okf-v0-2/expected-bundle` prints it: one line per document with its section
|
||||||
titles, a series of like-named documents as one line. Then put the question
|
titles, a series of like-named documents as one line. Pointed at a FOLDER of
|
||||||
|
bundles, the same command lists every bundle under it, each with its map, so
|
||||||
|
you see what each one covers before you choose. Then put the question
|
||||||
into the bundle's own words: a bundle written in one language and a question
|
into the bundle's own words: a bundle written in one language and a question
|
||||||
asked in another share few tokens, and the pre-pass matches tokens. Take the
|
asked in another share few tokens, and the pre-pass matches tokens. Take the
|
||||||
terms from the map's titles, not from your vocabulary.
|
terms from the map's titles, not from your vocabulary.
|
||||||
|
|
@ -99,10 +101,19 @@ is worth carrying. When `coverage.weak` is true — a word of yours the bundle
|
||||||
holds in no form (`coverage.absent_terms`), or nothing came back — rephrase in
|
holds in no form (`coverage.absent_terms`), or nothing came back — rephrase in
|
||||||
the bundle's own words, and if it stays weak, say the bundle does not cover it.
|
the bundle's own words, and if it stays weak, say the bundle does not cover it.
|
||||||
|
|
||||||
**4. Several bundles, same method.** When more than one bundle could answer,
|
**4. Several bundles, one run.** When more than one bundle could answer, give
|
||||||
run the same sub-questions against each, and keep track of which bundle each
|
the pre-pass the FOLDER that holds them instead of one bundle: it asks every
|
||||||
piece of material came from. A claim is attributed to its bundle as well as its
|
bundle under the folder with the same sub-questions in ONE run, splits the
|
||||||
concept — two bundles can hold the same sentence with different authority.
|
budget between them, and names the bundle on every answer and every excerpt.
|
||||||
|
`--bundle-id` narrows it to one of them.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
okf consume <the folder that holds the bundles> --question "first sub-question" --question "second sub-question" --out /tmp/p1.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Keep track of which bundle each piece of material came from. A claim is
|
||||||
|
attributed to its bundle as well as its concept — two bundles can hold the same
|
||||||
|
sentence with different authority.
|
||||||
|
|
||||||
**5. Put it together.** Order the material by sub-question, not by rank. Where
|
**5. Put it together.** Order the material by sub-question, not by rank. Where
|
||||||
sources disagree, decide what holds NOW: the newest documentation or the
|
sources disagree, decide what holds NOW: the newest documentation or the
|
||||||
|
|
@ -290,7 +301,7 @@ carries its denominator.
|
||||||
| Limit | `120000` |
|
| Limit | `120000` |
|
||||||
| Unit | `utf-8 bytes of emitted JSON` |
|
| Unit | `utf-8 bytes of emitted JSON` |
|
||||||
| Instrument | `okf_consume.measure (len of the ensure_ascii=False JSON encoding, utf-8)` |
|
| Instrument | `okf_consume.measure (len of the ensure_ascii=False JSON encoding, utf-8)` |
|
||||||
| Known-positive | `docs/consumption-contract.md, encoded as a JSON string` at `23672` |
|
| Known-positive | `docs/consumption-contract.md, encoded as a JSON string` at `24620` |
|
||||||
|
|
||||||
The instrument reproduces the known-positive figure before any of its own
|
The instrument reproduces the known-positive figure before any of its own
|
||||||
numbers are believed. Report what the run actually spent.
|
numbers are believed. Report what the run actually spent.
|
||||||
|
|
|
||||||
|
|
@ -11,10 +11,10 @@
|
||||||
"spent": 2289,
|
"spent": 2289,
|
||||||
"known_positive": {
|
"known_positive": {
|
||||||
"case": "docs/consumption-contract.md, encoded as a JSON string",
|
"case": "docs/consumption-contract.md, encoded as a JSON string",
|
||||||
"expected": 23672,
|
"expected": 24620,
|
||||||
"measured": 23672,
|
"measured": 24620,
|
||||||
"raw_bytes": 23092,
|
"raw_bytes": 24028,
|
||||||
"encoding_delta": 580
|
"encoding_delta": 592
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
"denominators": {
|
"denominators": {
|
||||||
|
|
|
||||||
|
|
@ -785,14 +785,14 @@ KNOWN_POSITIVE_CASE = "docs/consumption-contract.md, encoded as a JSON string"
|
||||||
|
|
||||||
#: `measure()`'s own answer for that file. Vacuous ALONE -- which is why the
|
#: `measure()`'s own answer for that file. Vacuous ALONE -- which is why the
|
||||||
#: delta below exists.
|
#: delta below exists.
|
||||||
KNOWN_POSITIVE_EXPECTED = 23_672
|
KNOWN_POSITIVE_EXPECTED = 24_620
|
||||||
|
|
||||||
#: The second, independent route. `wc -c` reports 23 092 raw bytes for the same
|
#: The second, independent route. `wc -c` reports 24 028 raw bytes for the same
|
||||||
#: file; the difference is this file's JSON quoting and escaping overhead. A
|
#: file; the difference is this file's JSON quoting and escaping overhead. A
|
||||||
#: reader can derive it without running `measure()` at all, and it moves the
|
#: reader can derive it without running `measure()` at all, and it moves the
|
||||||
#: moment `measure()` changes what it counts -- which is what stops
|
#: moment `measure()` changes what it counts -- which is what stops
|
||||||
#: `expected == measured` from proving nothing.
|
#: `expected == measured` from proving nothing.
|
||||||
KNOWN_POSITIVE_ENCODING_DELTA = 580
|
KNOWN_POSITIVE_ENCODING_DELTA = 592
|
||||||
|
|
||||||
#: The two places that file can be, resolved in this order.
|
#: The two places that file can be, resolved in this order.
|
||||||
#:
|
#:
|
||||||
|
|
|
||||||
|
|
@ -147,6 +147,9 @@ class Report:
|
||||||
#: What the payload says its withheld set holds. `None` when it states no
|
#: What the payload says its withheld set holds. `None` when it states no
|
||||||
#: total -- unmeasured, never zero.
|
#: total -- unmeasured, never zero.
|
||||||
withheld_total: int | None = None
|
withheld_total: int | None = None
|
||||||
|
#: How many payloads a FOLDER's reply carried (SS 8.11). `None` for a
|
||||||
|
#: single payload, whose report reads exactly as it always has.
|
||||||
|
payloads_examined: int | None = None
|
||||||
|
|
||||||
def render(self) -> str:
|
def render(self) -> str:
|
||||||
named = (
|
named = (
|
||||||
|
|
@ -154,8 +157,9 @@ class Report:
|
||||||
if self.withheld_total is None or self.withheld_total == self.withheld_examined
|
if self.withheld_total is None or self.withheld_total == self.withheld_examined
|
||||||
else f"{self.withheld_examined} of {self.withheld_total} withheld entries"
|
else f"{self.withheld_examined} of {self.withheld_total} withheld entries"
|
||||||
)
|
)
|
||||||
|
over = "" if self.payloads_examined is None else f"{self.payloads_examined} payloads, "
|
||||||
denominator = (
|
denominator = (
|
||||||
f"{self.rules_evaluated} rules over {self.excerpts_examined} excerpts and {named}"
|
f"{self.rules_evaluated} rules over {over}{self.excerpts_examined} excerpts and {named}"
|
||||||
)
|
)
|
||||||
if not self.findings:
|
if not self.findings:
|
||||||
return f"conformant: {denominator}, 0 findings"
|
return f"conformant: {denominator}, 0 findings"
|
||||||
|
|
@ -853,12 +857,87 @@ def check(skill_text: str, payload: object) -> Report:
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def is_folder_reply(payload: object) -> bool:
|
||||||
|
"""Whether `payload` is the reply to ONE call over a folder of bundles
|
||||||
|
(SS 8.11): `answers`, one per bundle, and no `bundle` of its own."""
|
||||||
|
return isinstance(payload, Mapping) and "answers" in payload and "bundle" not in payload
|
||||||
|
|
||||||
|
|
||||||
|
def check_reply(skill_text: str, reply: object) -> Report:
|
||||||
|
"""`check`, for a single payload or for a folder's reply.
|
||||||
|
|
||||||
|
A folder's reply is not a payload: it is one payload per bundle, and each
|
||||||
|
is held to every rule on its own -- the budget split between them makes
|
||||||
|
none of them a different kind of payload. A finding is named with the
|
||||||
|
bundle whose payload carries it; one that every answer carries
|
||||||
|
identically (a skill's missing section, say) is a fact about the SKILL and
|
||||||
|
is reported once, unnamed. An answer labelled with a bundle its payload
|
||||||
|
does not describe is `answer_misattributed`: the label is what a reader
|
||||||
|
attributes a claim to.
|
||||||
|
"""
|
||||||
|
if not is_folder_reply(reply):
|
||||||
|
return check(skill_text, reply)
|
||||||
|
assert isinstance(reply, Mapping)
|
||||||
|
answers = [_mapping(answer) for answer in _sequence(reply.get("answers"))]
|
||||||
|
if not answers:
|
||||||
|
return Report(
|
||||||
|
findings=(
|
||||||
|
Finding(
|
||||||
|
"payload_invalid",
|
||||||
|
"the folder's reply carries no answer, so there is no payload "
|
||||||
|
"to hold to the contract (SS 8.11)",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
rules_evaluated=len(RULES),
|
||||||
|
excerpts_examined=0,
|
||||||
|
withheld_examined=0,
|
||||||
|
payloads_examined=0,
|
||||||
|
)
|
||||||
|
reports = [check(skill_text, answer.get("payload")) for answer in answers]
|
||||||
|
common = set.intersection(
|
||||||
|
*({(finding.code, finding.message) for finding in report.findings} for report in reports)
|
||||||
|
)
|
||||||
|
findings: list[Finding] = [
|
||||||
|
finding for finding in reports[0].findings if (finding.code, finding.message) in common
|
||||||
|
]
|
||||||
|
for answer, report in zip(answers, reports):
|
||||||
|
label = _text(answer.get("bundle_id"))
|
||||||
|
declared = _text(_mapping(_mapping(answer.get("payload")).get("bundle")).get("bundle_id"))
|
||||||
|
if label != declared:
|
||||||
|
findings.append(
|
||||||
|
Finding(
|
||||||
|
"answer_misattributed",
|
||||||
|
f"an answer is labelled {label!r} and its payload describes "
|
||||||
|
f"{declared!r}; a claim is attributed to the label (SS 8.11)",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
findings.extend(
|
||||||
|
Finding(finding.code, f"[{label}] {finding.message}")
|
||||||
|
for finding in report.findings
|
||||||
|
if (finding.code, finding.message) not in common
|
||||||
|
)
|
||||||
|
totals = [report.withheld_total for report in reports]
|
||||||
|
return Report(
|
||||||
|
findings=tuple(findings),
|
||||||
|
rules_evaluated=len(RULES),
|
||||||
|
excerpts_examined=sum(report.excerpts_examined for report in reports),
|
||||||
|
withheld_examined=sum(report.withheld_examined for report in reports),
|
||||||
|
withheld_total=None if None in totals else sum(t for t in totals if t is not None),
|
||||||
|
payloads_examined=len(reports),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||||
parser = argparse.ArgumentParser(
|
parser = argparse.ArgumentParser(
|
||||||
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
|
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
|
||||||
)
|
)
|
||||||
parser.add_argument("--skill", type=Path, required=True, help="the SKILL.md to check")
|
parser.add_argument("--skill", type=Path, required=True, help="the SKILL.md to check")
|
||||||
parser.add_argument("--payload", type=Path, required=True, help="one pre-pass payload (JSON)")
|
parser.add_argument(
|
||||||
|
"--payload",
|
||||||
|
type=Path,
|
||||||
|
required=True,
|
||||||
|
help="one pre-pass payload (JSON), or the reply to one call over a folder of bundles",
|
||||||
|
)
|
||||||
return parser.parse_args(argv)
|
return parser.parse_args(argv)
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -877,7 +956,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
except json.JSONDecodeError as exc:
|
except json.JSONDecodeError as exc:
|
||||||
print(f"the payload is not readable JSON: {exc}")
|
print(f"the payload is not readable JSON: {exc}")
|
||||||
return 2
|
return 2
|
||||||
report = check(skill_text, payload)
|
report = check_reply(skill_text, payload)
|
||||||
print(report.render())
|
print(report.render())
|
||||||
return 1 if report.findings else 0
|
return 1 if report.findings else 0
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -553,6 +553,11 @@ def _rewrite(
|
||||||
# `<BUNDLE_ROOT>` would be the unfilled template's hole inside the one
|
# `<BUNDLE_ROOT>` would be the unfilled template's hole inside the one
|
||||||
# section that asks for a second run.
|
# section that asks for a second run.
|
||||||
("<BUNDLE_ROOT>", str(bundle_root)),
|
("<BUNDLE_ROOT>", str(bundle_root)),
|
||||||
|
# The folder form of step 4 (v1.1 F). An instruction, never a path:
|
||||||
|
# the bundle's parent directory is a path the caller never gave, and
|
||||||
|
# written absolute it names a checkout (the test holding generated
|
||||||
|
# commands to "no path into this repository" caught exactly that).
|
||||||
|
("<FOLDER>", GENERIC_FOLDER),
|
||||||
]
|
]
|
||||||
for old, new in replacements:
|
for old, new in replacements:
|
||||||
if old not in text:
|
if old not in text:
|
||||||
|
|
@ -820,6 +825,11 @@ CARD_COMMAND = "okf card"
|
||||||
|
|
||||||
GENERIC_BUNDLE = "<the bundle you were pointed at>"
|
GENERIC_BUNDLE = "<the bundle you were pointed at>"
|
||||||
|
|
||||||
|
#: Step 4's folder, in the generic skill. Lower-case on purpose, like
|
||||||
|
#: `GENERIC_BUNDLE`: it is an instruction to the reader, not a hole a
|
||||||
|
#: generator left.
|
||||||
|
GENERIC_FOLDER = "<the folder that holds the bundles>"
|
||||||
|
|
||||||
|
|
||||||
def render_generic() -> str:
|
def render_generic() -> str:
|
||||||
"""One installable skill for ANY bundle, carrying no bundle's numbers.
|
"""One installable skill for ANY bundle, carrying no bundle's numbers.
|
||||||
|
|
@ -844,11 +854,20 @@ def render_generic() -> str:
|
||||||
replacements: list[tuple[str, str]] = [
|
replacements: list[tuple[str, str]] = [
|
||||||
(
|
(
|
||||||
TEMPLATE_HEADER,
|
TEMPLATE_HEADER,
|
||||||
|
"**Use the server first.** When an `okf` MCP server is registered — its\n"
|
||||||
|
"tools `okf_describe` and `okf_ask` are then among yours — ask through it: it\n"
|
||||||
|
"is registered once, works from every project and reaches subagents, which\n"
|
||||||
|
"inherit tools and not skills. This skill is the supplement for a session\n"
|
||||||
|
"with no server. It runs the same code over the same bundles, so the two\n"
|
||||||
|
"cannot disagree about an answer, and neither has to be made again when a\n"
|
||||||
|
"bundle is added or rebuilt.\n\n"
|
||||||
"**This file is generic: it carries no bundle's identity and no bundle's\n"
|
"**This file is generic: it carries no bundle's identity and no bundle's\n"
|
||||||
"numbers,** and it is therefore never stale. It serves whichever bundle you\n"
|
"numbers,** and it is therefore never stale. It serves whichever bundle you\n"
|
||||||
"are pointed at. Before answering, read that bundle's own card:\n\n"
|
"are pointed at — or every bundle under a folder you are pointed at. Before\n"
|
||||||
|
"answering, read the card:\n\n"
|
||||||
"```sh\n"
|
"```sh\n"
|
||||||
f"{CARD_COMMAND} {GENERIC_BUNDLE}\n"
|
f"{CARD_COMMAND} {GENERIC_BUNDLE}\n"
|
||||||
|
f"{CARD_COMMAND} {GENERIC_FOLDER} # every bundle under it, each with its card\n"
|
||||||
"```\n\n"
|
"```\n\n"
|
||||||
"The card is DERIVED from the bundle on every run, never stored in it, so\n"
|
"The card is DERIVED from the bundle on every run, never stored in it, so\n"
|
||||||
"there is no second artefact that can disagree with the bytes. Its\n"
|
"there is no second artefact that can disagree with the bytes. Its\n"
|
||||||
|
|
@ -867,7 +886,9 @@ def render_generic() -> str:
|
||||||
" --out /tmp/payload.json\n"
|
" --out /tmp/payload.json\n"
|
||||||
"```\n\n"
|
"```\n\n"
|
||||||
"`--ref` is an **assertion**, never an override: the identity is computed\n"
|
"`--ref` is an **assertion**, never an override: the identity is computed\n"
|
||||||
"from the bytes either way, and a mismatch refuses. Read the pre-pass's\n"
|
"from the bytes either way, and a mismatch refuses. It belongs to one\n"
|
||||||
|
"bundle, so leave it out over a folder: each answer there carries its own\n"
|
||||||
|
"bundle's `ref`. Read the pre-pass's\n"
|
||||||
"own exit status, which carries three values: **0** a payload was written,\n"
|
"own exit status, which carries three values: **0** a payload was written,\n"
|
||||||
"**1** the run happened and refused, **2** the run did not happen at all.",
|
"**1** the run happened and refused, **2** the run did not happen at all.",
|
||||||
),
|
),
|
||||||
|
|
@ -948,6 +969,7 @@ def render_generic() -> str:
|
||||||
("<KNOWN_POSITIVE_CASE>", okf_consume.KNOWN_POSITIVE_CASE),
|
("<KNOWN_POSITIVE_CASE>", okf_consume.KNOWN_POSITIVE_CASE),
|
||||||
("<KNOWN_POSITIVE_EXPECTED>", str(okf_consume.KNOWN_POSITIVE_EXPECTED)),
|
("<KNOWN_POSITIVE_EXPECTED>", str(okf_consume.KNOWN_POSITIVE_EXPECTED)),
|
||||||
("<BUNDLE_ROOT>", GENERIC_BUNDLE),
|
("<BUNDLE_ROOT>", GENERIC_BUNDLE),
|
||||||
|
("<FOLDER>", GENERIC_FOLDER),
|
||||||
("<PAYLOAD_PATH>", "/tmp/payload.json"),
|
("<PAYLOAD_PATH>", "/tmp/payload.json"),
|
||||||
("<SKILL_PATH>", "this file"),
|
("<SKILL_PATH>", "this file"),
|
||||||
("<REF>", "the card's `ref`"),
|
("<REF>", "the card's `ref`"),
|
||||||
|
|
@ -964,9 +986,11 @@ def render_generic() -> str:
|
||||||
description = block_scalar(
|
description = block_scalar(
|
||||||
"Answer one question about ANY OKF bundle from a bounded payload assembled "
|
"Answer one question about ANY OKF bundle from a bounded payload assembled "
|
||||||
"by a deterministic pre-pass, marking every claim with its source, its title "
|
"by a deterministic pre-pass, marking every claim with its source, its title "
|
||||||
"and its provenance locator. Carries no bundle's identity: read the bundle's "
|
"and its provenance locator, over one bundle or every bundle under a folder. "
|
||||||
f"own card with `{CARD_COMMAND}` first. Use when the user asks a question of, "
|
"Carries no bundle's identity: read the card with "
|
||||||
"or states a hypothesis about, a corpus held as an OKF bundle."
|
f"`{CARD_COMMAND}` first. The supplement to the `okf` MCP server: use its tools "
|
||||||
|
"when they are registered, and this skill when they are not. Use when the user "
|
||||||
|
"asks a question of, or states a hypothesis about, a corpus held as OKF bundles."
|
||||||
)
|
)
|
||||||
header = f"---\nname: {block_scalar(GENERIC_NAME)}\ndescription: {description}\n---\n"
|
header = f"---\nname: {block_scalar(GENERIC_NAME)}\ndescription: {description}\n---\n"
|
||||||
return header + text
|
return header + text
|
||||||
|
|
|
||||||
|
|
@ -214,3 +214,57 @@ def test_one_bundle_is_read_as_before(folder: Path) -> None:
|
||||||
payload = json.loads(run.stdout)
|
payload = json.loads(run.stdout)
|
||||||
assert "answers" not in payload
|
assert "answers" not in payload
|
||||||
assert payload["bundle"]["bundle_id"] == "hage"
|
assert payload["bundle"]["bundle_id"] == "hage"
|
||||||
|
|
||||||
|
|
||||||
|
# --- F4: the checker reads the folder's reply --------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _check(tmp_path: Path, reply: object) -> subprocess.CompletedProcess[str]:
|
||||||
|
from llm_ingestion_okf import skill
|
||||||
|
|
||||||
|
skill_path = tmp_path / "SKILL.md"
|
||||||
|
skill_path.write_text(skill.render_generic(), encoding="utf-8")
|
||||||
|
payload_path = tmp_path / "reply.json"
|
||||||
|
payload_path.write_text(json.dumps(reply, ensure_ascii=False), encoding="utf-8")
|
||||||
|
return _okf("check", "--skill", str(skill_path), "--payload", str(payload_path))
|
||||||
|
|
||||||
|
|
||||||
|
def _reply(folder: Path) -> dict[str, object]:
|
||||||
|
return mcp_server.call_ask(_surface(folder), {"questions": list(QUESTIONS)})
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_generic_skill_is_conformant_on_a_folders_reply(folder: Path, tmp_path: Path) -> None:
|
||||||
|
run = _check(tmp_path, _reply(folder))
|
||||||
|
assert run.returncode == 0, run.stdout
|
||||||
|
assert run.stdout.startswith("conformant: ")
|
||||||
|
assert "over 2 payloads" in run.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_answer_labelled_with_another_bundle_is_a_finding(folder: Path, tmp_path: Path) -> None:
|
||||||
|
reply = _reply(folder)
|
||||||
|
answers = reply["answers"]
|
||||||
|
assert isinstance(answers, list)
|
||||||
|
answers[0]["bundle_id"] = "hage"
|
||||||
|
run = _check(tmp_path, reply)
|
||||||
|
assert run.returncode == 1
|
||||||
|
assert "answer_misattributed" in run.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_defect_in_one_answer_is_named_with_its_bundle(folder: Path, tmp_path: Path) -> None:
|
||||||
|
reply = _reply(folder)
|
||||||
|
answers = reply["answers"]
|
||||||
|
assert isinstance(answers, list)
|
||||||
|
del answers[1]["payload"]["contract"]
|
||||||
|
run = _check(tmp_path, reply)
|
||||||
|
assert run.returncode == 1
|
||||||
|
findings = [line for line in run.stdout.splitlines() if line.startswith(" ")]
|
||||||
|
assert findings == [
|
||||||
|
line for line in findings if line.startswith(" contract_unversioned: [hage]")
|
||||||
|
]
|
||||||
|
assert len(findings) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_reply_with_no_answer_is_a_finding_not_a_pass(tmp_path: Path) -> None:
|
||||||
|
run = _check(tmp_path, {"asked": [], "answers": []})
|
||||||
|
assert run.returncode == 1
|
||||||
|
assert "payload_invalid" in run.stdout
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue