wiki-knowledge

Consolidation is a lossless delete, not a supersession; tags get an index table

Context. wiki-lint gains a concept-fragmentation check: it finds clusters of small, closely-related concept pages that would be better as one page with sections, and proposes a consolidation (see CONTEXT.md; the operation was called a fold until #504 freed the word for line folding, ADR-0024). Consolidating raises two decisions the rest of the vault’s rules don’t settle on their own.

Decision 1 — a consolidation deletes the consolidated pages, and this does not violate the supersedes-not-overwrite invariant. Elsewhere the vault never destroys committed knowledge: on a contradiction, ingestion appends a new page and records supersedes, keeping both (ADR-0003 context; the invariant is stated in CONTEXT.md). A consolidation instead git rms the absorbed pages. The two are not in tension because supersession exists to preserve a conflicting claim — you keep the old page because it says something the new one denies. A consolidation has no conflict to preserve: it is lossless by construction, every absorbed page’s content becomes a section of the survivor, so deleting the now-empty original drops nothing. Recording the consolidation as supersedes instead (keeping the losers on disk) was rejected: it would defeat the point, since reducing total page count is part of the motivation (many vault operations are O(n) or worse). The safeguard is that losslessness is a property of the consolidation executor, not a promise the linter makes — the survivor must contain the losers’ content before the delete, in one atomic plan.

Decision 2 — tags are indexed in a dedicated page_tag(page_ref, tag) table in .wiki-knowledge/index.db, not stuffed into the FTS5 content column. The table already ships with the search index — added by the TypeScript port (#253), populated from the same frontmatter parse the reindex already runs — so concept-fragmentation consumes it rather than introducing it, and the name stays singular as shipped. The check’s candidate generation compares titles and tags across all pages. Doing that by re-parsing every page’s frontmatter on each run was the alternative; indexing tags is cheaper, scales, and is reusable by other checks (implicit-concepts, missing-cross-references). Tags are structured exact-match values, not prose, so they get their own table and candidate generation becomes a SQL self-join on shared tags (ranked by shared-tag count) plus an FTS5 MATCH on the already-indexed titles. Consequence: like all search, this is a view of HEAD (ADR-0015) — an uncommitted fragmented draft is not detected until committed, which is consistent with the rule that a page is only a page once committed.

Consequences