The plugin’s two layers get two different testing instruments, deliberately not unified into one suite. The Python script layer (vault.py, frontmatter.py, links.py, build_index.py, commit.py) has clear inputs and exact outputs, so it’s built test-first with pytest, including property tests (hypothesis) on the two functions that carry real correctness risk: frontmatter round-tripping and link-splicing on move. The agent layer (wiki-ingest, wiki-researcher) makes judgment calls that don’t have a single correct output, so it’s checked with evals against a fixed fixture vault instead — structural assertions wherever possible (does the page carry the expected fields? is the superseded page absent from the result set?), an LLM judge only for genuinely fuzzy questions, and a pass-rate over N runs rather than a single green, since agent behavior varies run to run.
The eval substrate is a small synthetic fixture, not a hand-authored vault of real knowledge (amended 2026-07-28, #7). The original decision called for a human-written “golden vault” of 15–30 real notes; that was shelved as disproportionate for a solo project — it billed the vault owner for exactly the hand-authoring the plugin exists to avoid. What replaced it is ~8 pages of deliberately fake content at wiki-plugin/evals/fixtures/ (a real vault root: wiki/ + raw/), with every property planted deliberately and the expected outcome written down in evals/PROPERTIES.md.
The property list is human-reviewed, not agent-certified — an agent that writes both the implementation and its own success criteria will converge on criteria its implementation already meets. Reducing the substrate from 15–30 real pages to ~8 fake ones changes what that review costs, not who owns the verdict: the agent may draft the fixture, but the human signs off on what counts as correct.
Retrieval evals run against the fixed fixture, never against ingestion’s own output, so an ingestion bug and a retrieval bug can’t cancel each other out and both show green. This is the clause that survives the amendment intact, and it’s why the live dogfooding vault — which is 100% ingestion output, and moving — cannot serve as the eval substrate no matter how realistic it is. It remains a complementary smoke surface: it catches reality a fixture didn’t imagine (as it did in #16), but nothing asserts against it.
Because a fake 8-page fixture is cheap to write and carries no maintenance burden, the expensive and irreducible part is the property list — deciding what the right answer is. That is why the substrate and the harness are separable, and are separately ticketed: the harness (headless runs, N-run pass rates) is mechanical and can be built later against a substrate that already exists.