The pilot invitations asked for feedback without specifying what to run or
what we expect to see. That returns a description, not a measurement: only a
stated expectation lets someone else's run falsify our model instead of
merely confirming it.
Each of the three tests now carries a numbered procedure, numbered expected
outcomes, and an explicit list of what would surprise us. Naming the
surprises is the load-bearing part -- a consumer who sees something odd but
passing otherwise has no reason to mention it.
Test A (producer, portfolio-optimiser-claude) runs one already-working
manifest twice with the same explicit ingested_at, once under DEFAULT as
baseline and once under OKF_V0_2, and diffs. A-E1 is the stop condition: if
the DEFAULT run is not byte-identical to their current pinned output, we have
broken a v0.1 consumer and the pilot ends there. The named surprises include
`at` differing from the ingested_at they passed, which would mean a
wall-clock crept in, and a collision refusal on a file that IS theirs, which
is the inverse of the fail-safe we predicted -- we expect foreign files to be
refused, so a false refusal of their own is the defect.
Test B (gate, catalog) runs their unmodified gate on the v0.2 fixture and on
a 0.1 control. B-E3 -- that no gate other than the version gate behaves
differently between the two -- is what converts our reading of their form
regex into a measurement. A membership list or equality comparison anywhere
in their chain is exactly what this is paid for.
Test C (expressiveness, wiki) is run by us, read-only at a recorded commit,
and the report goes to them. What we ask them to check is the part we cannot
measure from outside: whether a field we called an optional addition is in
fact load-bearing in their pipeline, and whether "no change required" holds
operationally rather than only formally.
Feedback gains a per-expectation verdict line so three independent runs are
comparable and a disagreement is located rather than merely reported. The
request also says outright that a wrong expectation is a better result than a
clean run, since a clean run only confirms what we already believed.
Inputs come from this repo at the pre-release tag; nothing is transported
through the mailbox except the specification.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A2aKJxLejT9S8jYwoZ9fut
Operator directive 2026-07-26: consumers are better served by getting the
latest version early and reporting back than by us holding it until we judge
it finished.
This closes a real gap. Every test in the plan -- including V-A8 against
upstream's reference implementation -- asks whether output is conformant.
None asks whether it is usable: whether a bundle is awkward to construct,
whether a rejection message is actionable, whether the profile can express
what a consumer's actual data needs. Only real data surfaces that.
Pilot set is three repos, one axis each, chosen for signal:
portfolio-optimiser-claude for the producer path (one real manifest run),
catalog for gate acceptance (their gate on our fixture, measured rather than
inferred from the regex), and claude-code-llm-wiki for expressiveness across
522 real documents.
The load-bearing design detail is that most of the pilot is read-only on our
side: we run the v0.2 reader over real consumer bundles and send the report.
That needs no adoption, no writes into their trees (O2 holds), and no change
to a contract their operator ratified. Only the producer axis asks a
consumer to do anything, and it asks for one run.
Excluded with reasons rather than silently: okr (Node side not yet lifted),
linkedin-studio (v0.2's provenance families would put implicit pressure on
the ingest/published carve-out we agreed not to normalize), commons (they
are deciding V1 -- a participant, not a test site).
Shipping a provisional surface without owing stability rests on three rules,
not on saying "provisional": OKF_LATEST does not point at v0.2 until GA, so
flipping it is the GA event rather than a merge side effect; the docstring
and CHANGELOG name the pilot repos; and breaking changes during the pilot
get no deprecation cycle. Stating that last one up front is what buys the
freedom to act on feedback -- discovering it later is what turns a pilot
into a de-facto release.
Feedback is requested in five named parts, because unstructured reports are
not comparable across three repos. The one that matters most is "what was
awkward but worked": workarounds are the highest-value signal and never
appear as a failure.
GA exit criteria are testable, and carry one honest limit: a three-repo
pilot exercises only what those three use, so `sources` with usage_window,
multi-verifier `verified`, and Attested Computation will likely go
unexercised. Those stay marked provisional at GA instead of being silently
promoted -- claiming otherwise would be the same unearned-claim pattern that
"conform first, claim after" exists to prevent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A2aKJxLejT9S8jYwoZ9fut
Operator directive 2026-07-26: the library always supports the current
latest version of Google OKF. v0.2 shipped 2026-07-25, so v0.2 support is
committed work rather than something a consumer has to request. This
overrides the previous default answer to open question V3 ("no until a
named consumer asks"), which is kept in the plan marked superseded so the
override reads as deliberate.
Support is additive: a new profile, never a migration of the existing two.
That single design choice is what makes an always-latest policy sustainable,
and it resolves the tension the directive would otherwise create with three
constraints that do not yield to it:
- DEFAULT states commons' ingest-spec section 5 layer, so its `generated`
shape is commons' call. Under the additive design this stops blocking us,
which takes commons off the critical path.
- STRICT_V1 mirrors the proving consumer's ratified contract; changing
another repo's contract from here would violate O2.
- v0.2 defers the attestation receipt and verdict wire formats upstream, so
the format is supportable and the unspecified runtime is not. It re-enters
scope when upstream specifies it.
This is also the first time the phase-3 profile abstraction is forced by
something outside this repo rather than by a second consumer, which is the
better test of whether the seam was cut in the right place.
Deliverables D1-D6 replace the earlier decision-round framing: a frontmatter
model that can carry block lists (`sources`, multi-verifier `verified`), an
OKF_V0_2 profile plus an OKF_LATEST alias whose moving-target tradeoff is
documented rather than hidden, Door C conformance against the consumer
tolerance rules, `Attested Computation` round-trip, v0.2 golden fixtures, and
a release-checklist re-check so the standing policy cannot decay silently.
Two new assumptions carry the weight. V-A7 forbids any profile from emitting
`timestamp` together with a malformed `generated`, since that combination
would have neither a valid `generated.at` nor an eligible section 13.1
fallback. V-A8 validates our own v0.2 fixture against upstream's reference
implementation, because every other test in the suite only asks whether we
agree with ourselves.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A2aKJxLejT9S8jYwoZ9fut
Everything built so far targets OKF v0.1. Upstream published v0.2 on
2026-07-25, so the plan records how the library relates to it and, more
importantly, who owns each decision.
Read from the spec itself rather than secondhand, which corrected two
readings that a summary had gotten wrong:
- STRICT_V1's `timestamp` is NOT a defect. Section 13.1 grants consumers a
documented fallback to legacy `timestamp` precisely when `generated` is
absent, and STRICT_V1 emits no `generated`. Nothing is asked of the wiki.
- The one measured shape problem is DEFAULT's `generated: "true"`, because
v0.2 requires `generated.by` within `generated`. That key was not reserved
in v0.1, so it was legal when written; v0.2 claimed the name. DEFAULT
states commons' ingest-spec section 5 layer, so the fix is commons' call
and is raised there as open question V1 rather than patched locally.
Two findings shrink the work. The canonical form for `generated` and a
single `verified` is an inline flow mapping, which the existing scalar
parser already round-trips as an opaque string, so block-list support is
only needed for `sources` and multi-verifier `verified` -- and only if a
named consumer asks. And the collision degrades safely: `_is_ingest_owned`
returns False for a v0.2 mapping, so a foreign concept is refused rather
than overwritten.
The track sits between Phases 3 and 4 because Phase 4 freezes the
cross-runtime contract. Freezing a v0.1 shape into two runtimes would let
the shared fixture suite certify the drift instead of catching it.
Self-imposed rule, since the spec does not require it: conform first, claim
after. Declaring `okf_version: "0.2"` is a MAY with no conformance
checkpoint, so claiming it early would be permitted -- and would be the same
class of true-sounding misleading claim as reporting a 0.2.0 measurement
under a 0.3.1 heading.
Also moves the guard 0.3.1 measurement procedure out of session state and
into execution-order.md, where it belongs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A2aKJxLejT9S8jYwoZ9fut