wiki-knowledge

Lexical index in stdlib SQLite FTS5, not embeddings and not a new dependency

The plugin’s retrieval layer reads through a single artifact: a gitignored SQLite FTS5 db at .wiki-knowledge/index.db. It is the filter — narrow a big vault to the slice the summary-first frontier then works over. It is not an embedding store. The index mechanises three of retrieval’s hand-run rules and buys ADR-0002 more runway against the scale trigger the ADR itself names, without crossing the threshold that would make embeddings worth their dependency cost. A committed _index.md was removed (#117) once this filter became the sole read path — nothing consumed it, and the metadata table the db indexes is complete without it.

The db holds two tables: a page metadata table (kind, tags, source_date, git_date, volatility, supersedes, superseded_by, mtime_ns, size) and an FTS5 virtual table over title/summary/body, joined in one SQL statement. The motivating query shape — “pages updated in the last week, tagged foo, containing bar” — is therefore a single statement, not three sequential calls. A capability probe at open time falls back to a Python re backend when FTS5 isn’t compiled into the platform’s SQLite (the API is identical, ranking is degraded, not the contract).

The index is gitignored, decisively — measured 1.9 s cold rebuild at 2000 pages against the live dogfooding vault, weighed against 15 MB of binary churn per ingestion, poor git delta compression on a reshuffling b-tree, and unmergeable conflicts on a synced-folder vault. Two operational riders, both specific to this vault: no WAL (Resilio Sync + SQLite sidecar files is the textbook corruption case), and .wiki-knowledge/ must also be added to Resilio’s ignore list since .gitignore does not propagate to the syncer.

Correctness lives in an unconditional (mtime_ns, size) staleness scan on every search call (53 ms at 2000 pages, measured) — not in Vault.write’s inline index update, which is a latency optimisation. The scan catches git pull, Obsidian edits, and any other external write, so the index can never be wrong because someone forgot to call the right method. Schema versioning in a meta table; mismatch ⇒ delete and rebuild. Delete-and-rebuild is the migration strategy, deliberately — incremental migrations are not worth the correctness complexity when the cold rebuild is 2 seconds. This paragraph’s freshness mechanism is superseded by ADR-0015 — the index is a view of committed history, kept fresh by a HEAD-SHA watermark and built from git blobs, so the working-tree writes named above are deliberately not indexed until committed. Everything else here stands, including delete-and-rebuild as the migration strategy. (Two other details in this ADR did not survive into the current implementation independently of that: there is no re fallback backend, and Vault.write has no inline index update — see enchiridion-ts/src/searchindex.ts.)

The FTS5 query-syntax footgun (a hyphenated tag like wiki-knowledge is a MATCH syntax error) is closed at the API boundary: Vault.search tokenizes-and-phrase-quotes caller text by default, with raw=True as the escape hatch. The default protects Haiku-model agents from a class of crash that would only surface in production.

Consequences

The three retrieval hand-run rules that #16 caught in the live dogfooding vault are now structural, not prose. “Never grep a schema word” — deleted; structure is SQL columns now. “Sanity-check the git commit date before quoting a page’s age” — --date-field git_date against one batch git log pass, not per-page shell-outs. “Never cite a superseded page” — include_superseded=False is the default filter, with --include-superseded as the opt-in for history questions. Ranking on supersession/volatility rather than merely not getting it wrong stays #17.

This ADR strengthens ADR-0002 rather than amending it: that ADR’s own stated revisit trigger — “the map no longer fits a single read” — is answered by a metadata+FTS filter, which is precisely the non-embedding response to it. The filter has since become the sole read path: _index.md was removed (#117) rather than kept alongside it. The threshold for revisiting this decision is measured ranking failure — a labelled question set where FTS5 returns the right page outside the top 10 — not a vault-size number, and not an a-priori “if we ever hit 10k pages” guess. The eval substrate is #47; the revisit decision is owed once the property list has a question that exposes the ranking gap.

pyproject.toml is unchanged — sqlite3 is stdlib. That is the point. The plugin’s deterministic script layer has no new dependencies for this feature, and it should not. (commit.py did change once _index.md went away: its manifest’s extra_paths field existed only to stage the regenerated index, so it was dropped with it.)