Operations
On-disk layout
Everything the service manages lives under .lcg/ in the workspace:
.lcg/├── wal/ # WAL root — one subdirectory per group_id (issue #378)│ └── liminis/ # the default group's stream: *.jsonl, .checkpoints/, .wal-bounds.json,│ # .wal-generation.json, .wal-ontology.json (issue #446)├── db/liminis.db # LadybugDB files — a derived index, rebuildable from the WAL├── ontology.yaml # optional workspace-wide extraction vocabulary (yours to edit)├── ontology/ # optional per-group extraction vocabulary (issue #446)│ └── <group_id>.yaml # one file per group_id, overrides ontology.yaml for that group├── ontology-hash.json # workspace-level ontology drift sidecar (issue #83)├── ontology-hash/ # per-group ontology drift sidecars (issue #451)│ └── <group_id>.json # one file per group_id, drift state for its resolved ontology├── identity-set/ # per-group identity-bearing type stamps (issue #616)│ └── <group_id>.json # which entity types were `identity: true` when the group's│ # entities were extracted; only written once a group uses the flag└── service.sock # JSON-RPC 2.0 endpoint while the service runsThe write-ahead log is the source of truth — and it’s just JSON. Every mutation is appended
to plain JSONL files in .lcg/wal/<group_id>/ before it touches the database. The WAL is
human-readable, append-only, and git-friendly: check it into the same repository as your notes or
documents, diff it, and carry it across machines. The database is a derived index — delete it and
knowledge_rebuild_from_wal reconstructs the entire graph from the log.
.lcg/wal/ is a WAL root, not a single stream (issue #378). Each group_id gets its own
subdirectory — its own *.jsonl files, its own .checkpoints/ store, its own
.wal-bounds.json manifest, its own .wal-generation.json identity, and its own independent
seq numbering starting at 0. A group’s subdirectory is created lazily on that group’s first
write; a group that has never been written to simply has no subdirectory yet. A single-group
deployment (the common case — everything under the default "liminis" group, no caller ever
passing a different group_id) behaves exactly as a pre-378 deployment did: one subdirectory, one
writer, one recorded position. An existing pre-378 .lcg/wal/ (loose
*.jsonl/.checkpoints//.wal-bounds.json directly under wal/, no liminis/ subdirectory) is
migrated automatically and idempotently on first boot under the upgraded binary — see
ADR-0378 for the migration mechanics; no
operator action is required. As of issue #431, this migration also mints a
.wal-generation.json for the group it relocates content into: a legacy flat WAL predates
generation identity (issue #387) entirely, and migration assumes it is locally owned — this is
an assumption, not something provable from the directory contents alone (see issue #431’s
## Assumptions for why it holds today and what would invalidate it) — rather than leaving it
with an unknown generation (see the unknown-generation refusal below, and
ADR-0414’s amendment note). No operator
action is required for this either — it happens as part of the same migration.
.wal-generation.json (issue #387) gives each group’s stream a stable identity, distinct from
its seq numbering. seq identifies a position within a stream; it says nothing about
which stream a position belongs to — so nothing distinguishes “the same stream, further along”
from “a different stream that happens to also number its lines from 0.” A publisher can
legitimately reset a group’s stream (re-extract a corpus and republish from seq: 0 with
entirely different content and entity identities); the generation is what lets a consumer tell
that apart from ordinary forward progress. It is minted once, the first time a group’s directory
is created with no prior content, and never changes for the life of that stream — appending never
changes it, and it is opaque (compared for equality only, never interpreted or ordered). The file
holds a single JSON object:
{"generation": "3f9a1c2e-4b7d-4e21-9c8a-1a2b3c4d5e6f"}Any string value works — lcg mints a UUID, but nothing requires that shape. This file is
publisher-writable: an external, non-lcg publisher (e.g. a distributed, git-published WAL model)
that creates a group’s stream directory directly, without going through lcg, MUST write this file
itself (a plain json.dump({"generation": <any unique string>}, f) from Python is sufficient) for
knowledge_rebuild_from_wal’s reset detection (below) to work against that stream — lcg never
retroactively mints one into a directory it didn’t create, and a directory with no
.wal-generation.json is treated as having an unknown generation (see generation_status in
the generation-scoped applied_seq fields below). Whether that
is silently tolerated or an outright failure depends on whether a position for the group has
already been recorded — see issue #414 below. Like .checkpoints/ and .wal-bounds.json, it is
invisible to every existing non-recursive *.jsonl scan.
Publishing a WAL stream (issue #414)
Publishing a group’s stream directory means copying the entire directory, dot-namespace
included — never a *.jsonl or wal/* glob. A shell glob does not match a leading dot by
default, so git add wal/*, cp wal/*.jsonl, rsync --include='*.jsonl', and tar wal/* all
silently drop every dotfile in the directory while appearing to publish the complete stream. This
was confirmed as the root cause of a real-world reset-detection outage (issue #414): a publisher’s
*.jsonl-only copy step dropped .wal-generation.json on every publish, so every consumer that
hydrated from it reported generation: null forever and knowledge_rebuild_from_wal’s reset
detection never once had a generation to compare.
Use a whole-directory copy instead — cp -R/rsync -a with no include-filter, or git add -A —
and know what each entry in the dot-namespace costs you if you omit it anyway:
| entry | requirement | consequence if dropped |
|---|---|---|
.wal-generation.json (issue #387) | MUST travel — load-bearing | reset detection can never run for this stream again; every consumer that already recorded a position for this group starts hard-failing knowledge_rebuild_from_wal (issue #414, below) until the stream is republished with its generation intact |
.wal-bounds.json (issue #375) | MAY be omitted | not wrong, just slow — a cache; the consumer regenerates it by rescanning every *.jsonl file on next read |
.wal-ontology.json (issue #446) | MAY be omitted — informational | not wrong, not slow either — replay and correctness are entirely unaffected; the consumer just loses the ability to see what vocabulary produced this group’s graph. Never applied to the consumer’s own extraction, validation, canonicalization, or reprocessing even when present (see Ontology) — it is provenance, not policy |
.checkpoints/ (issue #365) | MAY be excluded | local-only recovery state — omitting it is a legitimate choice, but make it an explicit, stated decision rather than an accident of the same glob that drops generation |
Only .wal-generation.json is load-bearing — every other entry is safe to omit deliberately, but
never safe to omit by accident as a side effect of a glob pattern that was only ever meant to
select *.jsonl files. .wal-bounds.json and .checkpoints/ degrade performance or local
recovery convenience if dropped; .wal-ontology.json degrades only documentation — a stream
published with it present must never have it change the consumer’s own behavior (issue #446).
A published stream carries no embedding vectors (issue #526). Every *.jsonl record’s source
text (name, fact, content, summary, …) still travels; the vector each of those texts
produced at write time does not. Hydrating a published stream into a database — via
knowledge_rebuild_from_wal or any other replay path — always recomputes every vector locally
from that source text, which requires a reachable embedder. This is not an extra prerequisite:
every lcg instance has one already (an unreachable embedder is fatal at startup), so a consumer
that can run the service at all can hydrate any published stream, including one written under a
different embedding model than the consumer’s own (the result is a database entirely in the
consumer’s model — see ADR-0526 for the full model). This
also means publishing carries no write-time embedding-model diagnostic any more: issue #440’s
.wal-embedding-model.json sidecar and its replay-time [WAL WARN] embedding-model mismatch: ...
check are both gone (issue #526) — there is nothing left for a WAL-side stamp to govern once no
stored vector is ever bound. The one surviving model-identity signal is knowledge_status’s
embedding_model_status (below), which compares the graph’s currently-applied vectors against
the running embedder — a database-side check, not a property of the WAL or the publish step.
WAL administration
- Rebuild one group’s data from its own WAL directory with
knowledge_rebuild_from_wal {group_id, ...}(group_iddefaults to"liminis", so a single-group deployment needs no change). Afrom_seq: 0(default) rebuild against a group that already has data in it fails fast with an explicit error instead of silently producing a duplicate-key failure per node — passforce_clear: trueto clear that group’s data automatically first (issue #378: this clears only the target group via the same primitiveknowledge_delete_by_groupuses, not the whole database file), or clear it yourself withknowledge_delete_by_groupbefore calling rebuild. Rebuilding one group never touches another group’sWalPosition, WAL directory, or data. A successful non-dry-run rebuild automatically rebuilds the entity/relationship search indices, soknowledge_find_entities/knowledge_find_relationshipsare immediately queryable afterward —knowledge_build_indicesis not normally required. - Unknown-generation refusal (issue #414). Before comparing anything,
knowledge_rebuild_from_walchecks whether the group already has a previously recorded position (applied_seqnot null — note aknowledge_statuscall can itself cause this to become true via its own backfill, so this can trip on what looks like the first explicit rebuild call ever made against a group) and whether the group’s current on-disk generation is unknown (missing or corrupt.wal-generation.json— the two are indistinguishable by design, see below). If both hold, the call fails outright with an explicit error naming the group and pointing at the publish contract above — replay does not proceed,from_seq/to_seq/force_clearare not applied, and this applies uniformly todry_run: trueas well (there is nothing safe to preview). No configuration flag, environment variable, or request parameter bypasses this check. The refusal is scoped to the affected group only — a sibling group sharing the same WAL root whose own generation is known remains independently replayable in the same or a later call. A group with no previously recorded position is unaffected: it performs ordinary first-time adoption, including adopting an unknown generation, exactly as before this issue. See ADR-0414 for the full rationale. A workspace migrated from a legacy flat WAL by a binary containing issue #431’s fix does not hit this refusal — migration itself stamps a generation, so the group’s current on-disk generation is never unknown afterward (see the migration paragraph above). If it still fires, the error message gives two possible remedies, since the two situations that can produce this state are indistinguishable on disk: republish the stream’s full directory if it was received from a publisher (above), or — for a local workspace with no publisher, e.g. one migrated by a binary older than issue #431’s fix — create.wal-generation.jsonin the group’s WAL directory by hand with any unique string value,{"generation": "<any unique string>"}, as a one-time, deliberate assertion of ownership. - Reset detection (issue #387). Once the check above has passed,
knowledge_rebuild_from_walcompares the group’s recorded generation against what’s currently on disk (.wal-generation.json). If they differ (both known and unequal — seewal.generation_statusbelow for the unknown-generation case, handled by the refusal above instead), the caller’sfrom_seq/to_seq/force_clearare overridden entirely: this is always a full, automatic self-heal — purge the group, replay it from scratch against the new generation, then re-bind any cross-group pointers into it — rather than silently replaying new-generation mutations on top of old-generation data (the corruption this issue exists to prevent; the two do not reconcile, since the native write path emitsCREATErather thanMERGE). The result reportsreset_detected: true,previous_generation,generation(the generation just replayed), andcross_group_rebind(the same countsknowledge_rebind_pointersreports), on both the streaming response and the background-job’s polledresult, so a caller can tell this apart from an ordinary incremental replay. Adry_run: truecall against a mismatched group reports the samereset_detected/previous_generation/generationfields but purges and replays nothing — report-only, like every other dry-run path in this codebase. - Bounded rebuild with
to_seq: pass an inclusive upper bound (from_seq <= seq <= to_seq) to exclude a known-bad mutation and everything after it — e.g. recovering from an operator mistake that is itself recorded in the WAL.knowledge_rebuild_from_wal {from_seq: 0, to_seq: <seq before the bad mutation>, force_clear: true}rebuilds the graph as it stood just before the mistake. This is not durable: WAL entries beyondto_seqare left on disk, unapplied — they are not truncated or archived. A later unbounded rebuild, or afrom_seqresume that covers the excluded range, reapplies everything that was excluded, including a previously-excluded bad mutation. Durable rollback (truncating/archiving the WAL tail) is not provided by this primitive. - Dump the database back to a compacted log with
knowledge_dump_wal— this is also the way to take a restore-point snapshot before a large or destructive operation, since WAL replay is forward-only. The output directory starts with no checkpoints: any WAL marks (below) recorded against the source directory are not carried forward, since dump_wal renumbers sequence numbers and a copied mark’sseqwould be meaningless against the new numbering. For the same reason, the output always gets a freshly minted generation (issue #387) — never the source’s: it is a new stream, not a copy of the source’s identity, so a consumer must not treat it as “the same stream” it was tracking before. - Strip embedding vectors from a pre-0.14 WAL with
knowledge_strip_wal_embeddings {group_id, dry_run}(issue #577). 0.14.0 stopped writing embedding vectors to the WAL and made replay ignore any vector it finds in an older WAL (see above) — this operation reclaims the space a WAL written before that change is still carrying, by rewriting every qualifying.jsonlfile in place to removeparamsentries keyed by an embedding-column name (name_embedding,fact_embedding,content_embedding,summary_embedding), leaving every other field, record ordering, and sequence number untouched. Because replay never reads a stored vector regardless of whether it’s present, this is pure reclamation: it cannot change what a rebuild produces.group_idscopes the operation to one group’s own WAL directory; omit it to process every group directory under the WAL root plus any legacy flat-layout files at the root itself. Idempotent: re-running against an already-stripped WAL (or one that never had vectors, i.e. anything written entirely by 0.14.x) is a zero-I/O no-op — no file is opened for writing, so bytes and mtime are left exactly as they were. Crash-safe per file: each file needing a rewrite is written to a temporary file in the same directory and atomically renamed over the original only once fully flushed, so an interrupted run leaves every already-rewritten file replaced, the in-flight file’s original intact, and a subsequent run resumes cleanly. A record whose embedding-vector value is malformed (not a well-formed JSON array of numbers) is reported as a per-file error — that file is left completely untouched, and every other file is still processed. Passdry_run: trueto preview the same statistics (files that would be rewritten, bytes that would be reclaimed, records that would be touched) without modifying anything. Reachable even when the database is degraded or unavailable, since it only touches the WAL directory on disk. Deliberately out of scope: a vector inlined as a raw Cypher literal rather than aparamsentry (see ADR-0526’s externally-produced-content edge case) has noparamskey to remove and is left as-is. See ADR-0577 for the full design. - Name a known-good position with
knowledge_wal_mark_create {name, group_id}(group_iddefaults to"liminis") — a lightweight alternative to a fullknowledge_dump_walsnapshot when all you need is a durable pointer back to “this group’s stream was good here,” not a materialized copy. Anamemust be 1-200 characters of[A-Za-z0-9_-], because it becomes a single directory name under that group’s own.checkpoints/. It records the target group’s currentapplied_seq, and the group’s current generation (issue #387), under<wal_root>/<group_id>/.checkpoints/, is O(1) (no WAL scan or replay), and fails if the position is unknown (applied_seqisnull) or the name is already in use by an active mark within that group — two different groups may each have an active mark of the same name, since each group’s checkpoint store is independent.knowledge_wal_mark_list {group_id}(also defaulting to"liminis", and always scoped to exactly one group — there is no cross-group aggregate listing) lists every active mark in that group with itsseq, itsgeneration, itswal_min_seq/wal_max_seq(the bounds of that group’s WAL content currently on disk), and whether it is currentlyreachable: this requires both the existing bounds check (wal_min_seq == 0— the WAL’s own prefix has not been externally truncated, e.g. by routine retention deleting old WAL files — andseq <= wal_max_seq) and, independently, that the mark’s recordedgenerationmatches the group’s current on-disk generation whenever both are known (issue #387, FR-007) — a mark taken against a generation that has since been reset is never reachable, even when itsseqstill falls comfortably inside[wal_min_seq, wal_max_seq](exactly the “looks like forward progress, isn’t” case issue #387 exists to close). Separately, on the bounds side, a mark whoseseqmerely falls inside[wal_min_seq, wal_max_seq]is still reported unreachable ifwal_min_seq > 0, since a restore would silently omit everything before it. Neither check detects a gap in the middle of that range.knowledge_wal_mark_delete {name, group_id}removes a mark from that group (recording a tombstone, never rewriting the original record) and frees the name for reuse within that group. To restore:knowledge_rebuild_from_wal {group_id, from_seq: 0, to_seq: <seq>, force_clear: true}for a mark with an integerseq, orknowledge_delete_by_group {group_ids: [group_id]}for a mark withseq: null(a genuinely empty group) — orknowledge_clear_allif you mean to reset every group, not just one. These tools are unrelated toknowledge_prepare_checkpointbelow — they name a WAL position, not flush a writer — and each group’s.checkpoints/store lives in its own subdirectory precisely so it is invisible to the WAL file scans that discover.jsonlmutation files (knowledge_dump_waland the replayer among them), and so it travels with that group’s WAL directory itself when checked into git. Exactly-one-wins under concurrentcreatefor the same name (within one group) relies on exclusive file creation (O_EXCL), a local-filesystem guarantee — not reliable on an NFS-mounted WAL directory (see ADR-0365). - Checkpoint before backups with
knowledge_prepare_checkpoint— this rotates and flushes every group’s live WAL writer (issue #378: an instance-wide operation now spans however many groups this process has written to, not one writer) so pending mutations are on disk before an external filesystem backup. It shares the word “checkpoint” withknowledge_wal_mark_*above by coincidence, not by relation, and takes nogroup_id— it is always whole-instance. - Rotation.
LCG_WAL_MAX_BYTES_PER_FILE(default 5 MB) andLCG_WAL_MAX_EVENTS_PER_FILE(default 10000) bound each WAL file’s size; rotation fires when either threshold is reached and emits awal_rotatedtelemetry event. - Failure reporting. Failure reports from replay dedupe by
(template, error), so a schema gap on one mutation type can no longer hide an unrelated failure category behind a wall of identical samples. UseLCG_REPLAY_FAILURE_SAMPLESto control how many distinct failing lines are retained per replay.
See Configuration for the full set of LCG_WAL_*/LCG_REPLAY_* environment
variables, and IPC & MCP Reference for the
knowledge_rebuild_from_wal non-empty-database refusal behavior in detail.
Self-healing and degraded mode
The service binds its socket before opening the database, so a corrupted store leaves it
reachable in degraded mode rather than dead (ADR-0009).
Legacy .graphiti/→.lcg/ workspace migration runs before the bind; the issue #378 WAL-root
relocation (migrate_wal_root_if_needed()) runs after the socket is already bound, in the same
pre-Db::open() window as the DB open itself. Autonomous startup recovery (ADR-0027)
then reopens at the last good checkpoint, replays the WAL tail, and rebuilds indices without
intervention. Recovery progress is observable via the wal_auto_recovery telemetry event,
whose phase field steps through corruption_detected → checkpoint_drop_complete →
cursor_derived → replay_complete → index_build_complete → recovery_complete (or
fallback_triggered, if automatic recovery gives up and manual intervention via
knowledge_recover/knowledge_recover_full is needed).
Readiness: a successful connect is not readiness. Because the socket is bound before the
database opens, a bare connect() to .lcg/service.sock can succeed while the service is still
relocating the WAL root or replaying WAL — before it has started actually serving graph requests.
(The process’s own accept loop, in run_socket_service, only starts after bootstrap_app_state()
resolves, so a request sent on such a connection queues in the kernel and is not read until
startup work has already finished — it does not race migration and get served with stale state.
The actual risk is a client that treats the connect() succeeding as sufficient evidence of
readiness by itself — e.g. proceeding to inspect on-disk WAL state, or reporting “ready” in its
own UI — without waiting for a health_check round-trip.) The correct readiness signal is a
health_check request/response round-trip reporting "healthy": handle_health_check can only
return healthy once Db::open() has succeeded, which is after both legacy-workspace migration
(which completes before the bind) and WAL-root migration (which runs after it) have finished.
Poll health_check until it reports healthy (or knowledge_status until connected and
queryable are both true and initializing is false — knowledge_status has no healthy
field of its own) before treating the service as ready.
busy is alive, not dead. health_check never waits on the write lock: while a write is
pending or in progress — above all a long knowledge_rebuild_from_wal — it answers immediately
with {"ok": true, "healthy": false, "state": "busy", "activity": "rebuilding" | "writing"}, and
for a rebuild job also job_id and progress (mutations replayed, WAL files processed/total,
elapsed seconds). Keep polling on busy; do not restart the service — a restart mid-rebuild
discards the replay and starts it over, so on a large corpus the rebuild never completes (#612).
Supervisor liveness probes should treat any response with ok: true as alive and gate readiness on
healthy: true. The four outcomes are: busy = alive, not ready; healthy = alive, ready;
degraded = alive but unusable; no answer = dead or wedged. Full contract:
IPC reference: Health check contract.
knowledge_status health fields
Beyond the ontology summary, knowledge_status reports:
dedup_mode (string, always present, including when degraded) — what protects the graph from
wrong embedding-path merges during extraction, beyond the cosine threshold (issue #652,
ADR-0652). "veto-only": the
deterministic identifier-mismatch veto (ADR-0650)
is the only check, so any candidate that passes it merges. "llm-verified": the configured
extractor additionally judges each surviving candidate (LCG_DEDUP_LLM, off by default). If
LCG_DEDUP_LLM is on but no extraction provider is configured the field reads "veto-only" and
startup logged a warning. The same mode is logged once at startup as dedup: mode=…; per-chunk
path counts are on knowledge_process_chunk’s dedup_paths.
The veto also keeps distinct short all-caps codes apart (issue #666,
ADR-0666): ACDS / ACDM, UK / USA
and US Army / UK Army never merge by embedding. A token counts as a code when it is 2–6
uppercase letters (not a Roman numeral) in the name as written, and the veto fires only when
each name has a code the other lacks, so acronym/expansion aliases (IBM /
International Business Machines) and case variants (NASA / nasa) still merge. Known gaps,
left to the embedding and LCG_DEDUP_LLM: a shout-case name against a lone mixed-case one
(Acme vs ACDM) and codes longer than six letters.
indices_built (boolean) — whether the entity/relationship FTS + HNSW search indices are
currently built and reflect the graph’s current contents. The service builds these indices
eagerly at startup — immediately after schema init on a fresh DB, or as part of
self-recovery after a WAL-corruption auto-heal — before the socket accepts any request, so
indices_built is normally true from the very first knowledge_status call onward (see
ADR-0036). A genuine build failure during that eager
startup build fails startup outright rather than silently leaving indices unbuilt.
A runtime recovery — any knowledge_recover strategy (drop_lbug_wal,
rebuild_from_workspace_wal, restore_from_backup) or knowledge_recover_full — also leaves
indices_built correctly true on success: drop_lbug_wal/restore_from_backup reopen an
already-indexed checkpoint or backup, while rebuild_from_workspace_wal/knowledge_recover_full
explicitly rebuild the indices before reporting success. Failure handling differs by strategy:
rebuild_from_workspace_wal and knowledge_recover_full invalidate indices as part of the
attempt, so a failure that aborts before the rebuild completes leaves the flag false rather than
reporting stale readiness; drop_lbug_wal and restore_from_backup never touch indices, so a
failed call leaves the flag at whatever it was before the attempt.
indices_built still goes back to false in narrower, later situations: after
knowledge_clear_all, or if a post-rebuild index build genuinely fails (as opposed to the
common, harmless “already built” case). In those cases false does not mean search or ingest
is broken — knowledge_find_entities/knowledge_find_relationships, and the ingest
hybrid-dedup path used once a group_id passes the dedup threshold, all auto-heal by
transparently rebuilding indices and retrying on their first call after a false state. The
field exists so a caller can observe readiness proactively instead of discovering it only via
a search or ingest attempt. The same field appears on knowledge_rebuild_from_wal’s result (and
on knowledge_rebuild_status’s result for the background-job path) for the specific rebuild
that produced it; it is omitted from dry-run rebuild results, since a dry run never touches
indices.
lookup_key_backfill_ok (boolean, issue #491) — appears alongside indices_built on the
same two rebuild result shapes (knowledge_rebuild_from_wal’s result and
knowledge_rebuild_status’s result), also omitted on dry-run for the same reason. It reports
whether that specific rebuild’s Entity.lookup_key backfill succeeded — independently of
indices_built, which tracks only the FTS/HNSW index build and never reflects a backfill
failure (see name_index_trusted below for why that scoping is deliberate). A rebuild can report
indices_built: true and lookup_key_backfill_ok: false in the same result: the index build and
the backfill are separate steps, and one can succeed while the other fails. Use this field to
learn a given rebuild’s backfill outcome directly from its own result, without a follow-up
knowledge_status call; use name_index_trusted (below) for the current global/cumulative
backfill health, independent of any specific rebuild.
Key format and the kind migration (issue #615). Entity.lookup_key is group_id ␟ kind ␟ lower(trim(name)) (U+001F separators; kind defaults to Entity), so same-named entities of
different kinds coexist (ADR-0615). It is a derived
value: it is stripped from every WAL record at write time and recomputed on replay, and any
lookup_key in an older WAL is ignored. On the first start after upgrading, schema::migrate adds
Entity.kind, and a one-time backfill (persisted under the SchemaState key
entity_kind_lookup_key_v2) sets every existing row’s kind to Entity and re-keys it — an O(N)
pass, one statement per row. The post-replay backfill after every rebuild/recovery covers
kind IS NULL OR lookup_key IS NULL. Downgrade is unsupported: a pre-#615 binary cannot replay
a WAL written after this change.
name_index_trusted (boolean) and name_index_fallback_scans (integer) — field names
kept for wire compatibility, but re-backed by the Entity.lookup_key secondary ART index
(ADR-0221, which supersedes the
in-process NameIndex accelerator these fields originally described
(ADR-0038)). Case-insensitive entity name lookups are now
served directly by the database, so there is no in-process copy to rebuild — a raw-Cypher
mutation via knowledge_query_cypher no longer affects either field, unlike the pre-#221
design. name_index_trusted instead reports whether the one-shot lookup_key backfill migration
completed without error (schema::migrate, and the equivalent post-WAL-rebuild/recovery backfill
passes); it is true by default (nothing to migrate on a fresh database) and only goes back to
true once a subsequent migration/backfill attempt succeeds. name_index_fallback_scans keeps
its prior meaning exactly: it counts how many times an endpoint-existence lookup (the “does this
entity exist anywhere in the group” check used during edge-endpoint resolution) missed the index
— now specifically signaling a lookup_key value that is NULL or stale, most likely from an
out-of-band write via the cypher MCP scope or a second process — and fell back to a bounded
database scan, which also self-heals the row’s lookup_key on a hit. Both fields are null
while the service is degraded (no connected database). A rising name_index_fallback_scans
count, or a name_index_trusted: false that doesn’t clear on its own, signals lookup_key
staleness worth investigating — see
ADR-0221 for the mechanism.
fts_repair_count (integer) and fts_last_repair_unix_ms (integer or null) — how many
times, since this process opened the database, a write hit FTS index '<idx>' is inconsistent and
the service rebuilt all three FTS indexes and retried that statement once, and when the last such
repair finished (Unix ms; null if none). Both are null while the service is degraded. A
non-zero count means a residual FTS inconsistency that the startup marker did not catch was
detected and repaired (the classifier matches the error text, not its cause, so it does not by
itself identify a version mismatch); it normally does not recur. The one-time startup rebuild (marker fts_built_by_lbug in
SchemaState) is not counted here; it is logged on stderr
(rebuilding full-text (FTS) indexes: they were not built by this lbug version) and its time is
proportional to corpus size. See
ADR-0649.
Troubleshooting FTS index '<idx>' is inconsistent. Raised by a delete/update on a row whose
non-ASCII term is missing from an FTS index built by lbug 0.20 (0.15.x and earlier). Releases after
0.16.2 rebuild the indexes automatically on first start. On 0.16.0–0.16.2, drop all three (not just
the named one) with CALL DROP_FTS_INDEX('Entity','node_name_and_summary'),
CALL DROP_FTS_INDEX('RelatesToNode_','edge_name_and_fact') and
CALL DROP_FTS_INDEX('Episodic','episode_content'), then call knowledge_build_indices.
wal_groups (issue #378) — an additive map, keyed by group_id, of every group that
currently has a WAL directory, each entry shaped like the flat wal object below
({applied_seq, max_seq, generation, generation_status, hydration_status, embedding_model, embedding_dim, embedding_model_status} — the last three added by issue #440, mirroring each
group’s own embedding identity the same way generation is already mirrored per group). This is
the multi-group view; the flat
wal.applied_seq/wal.max_seq/wal.generation/wal.generation_status/wal.hydration_status/wal.embedding_model/wal.embedding_dim/wal.embedding_model_status
fields described next remain present and
pinned specifically to the default "liminis" group, unchanged in meaning from a pre-378
single-group deployment — a caller that only reads the flat fields (e.g. an existing integration
written before this issue) needs no change. If the default group has no WAL directory at all
(e.g. a pure replica that has only ever hydrated non-default groups), the flat fields report
null/absent rather than an error — a documented signal that this instance has no default group,
not a broken or un-hydrated instance. Do not confuse “not in wal_groups” with “at position 0”: a
group present in the map with applied_seq: 0 has a directory and a known position; a group
absent from the map entirely has no WAL directory yet.
wal.applied_seq and wal.max_seq (issue #353; scoped to the default group by issue
#378) — let a caller decide, from a single knowledge_status call and an integer comparison,
whether its local DB is already consistent with the default group’s WAL, needs an incremental
resume, or needs a full rebuild. wal.applied_seq is read from a persisted DB row on every call —
never cached in memory, so the value survives a service restart. wal.max_seq always reports the
true highest seq actually present on disk for the default group (or
None/null if the WAL is empty or unconfigured); an externally-updated WAL (e.g. a distributed,
git-published WAL pulled by another process) is observed on the very next call, at worst after one
reconciling full scan (issue #375). In the common case it’s computed from a small manifest sidecar
(<wal_dir>/.wal-bounds.json) rather than by rereading every .jsonl file in the WAL directory on
every call — see ADR-0375 for the caching mechanism and
why an earlier “never cached” design was revised. The same manifest and fast path also back
wal_min_seq, so knowledge_wal_mark_list’s reachability check (below) does not scale with WAL
file count either.
wal.generation (issue #387; also scoped to the default group, and mirrored per-group inside
wal_groups) — the group’s current on-disk (source-side) generation, read from
.wal-generation.json alongside the same wal_max_seq machinery above, so reporting it costs
nothing beyond what applied_seq/max_seq already pay (no new full-directory scan). This is
deliberately the on-disk value, not lcg’s own DB-recorded consumer-side position — an external
consumer (e.g. orac) compares this against its own bookkeeping to answer “is this the same
stream I was tracking?”, the same on-disk-authoritative role max_seq already plays. null means
the stream currently has no generation recorded — its own generation_status (next) says whether
that is “no stream yet” or “unknown” (both used to collapse indistinguishably to this same null,
issue #414). Opaque: compare for equality only, never interpret or order it. lcg’s own
internally-recorded generation (paired with its own applied_seq, and what
knowledge_rebuild_from_wal’s reset detection actually compares against) is not surfaced by
knowledge_status at all — it is a purely internal bookkeeping value with no separate
consumer-facing use.
wal.generation_status (issue #414; also scoped to the default group, and mirrored per-group
inside wal_groups) — a sibling string field alongside generation, classifying why generation
reads the way it does, since generation: null alone cannot distinguish “no stream” from “stream,
but generation unknown.” Pure classification of max_seq/generation, no new I/O:
generation_status | meaning |
|---|---|
"not_applicable" | no WAL stream exists yet for this group (no *.jsonl content, no generation record) |
"unknown" | a stream exists (*.jsonl content is present) but its generation is currently unrecoverable — missing or corrupt .wal-generation.json, most commonly because a publish step dropped the dot-namespace (see Publishing a WAL stream above) |
"known" | a stream exists with a recorded generation — including a freshly-minted, still-empty stream (max_seq: null, generation non-null) |
generation_status: "unknown" is exactly the condition that makes knowledge_rebuild_from_wal
refuse once a position has been recorded for that group (see Unknown-generation refusal above) —
checking this field before calling rebuild lets an operator see the condition coming rather than
discovering it as an abrupt failure.
wal.hydration_status (issue #456; also scoped to the default group, and mirrored per-group
inside wal_groups) — a sibling string field alongside applied_seq/max_seq, classifying
whether the group’s database contents are caught up with its WAL, so a caller no longer needs to
compare the two fields itself to tell “genuinely empty” apart from “not yet hydrated.” Pure
comparison of applied_seq/max_seq, no new I/O — an absent or never-backfilled applied_seq is
treated as 0 for the comparison:
hydration_status | meaning |
|---|---|
"not_applicable" | the group has no WAL content at all (max_seq is zero or absent) — there is nothing to be behind on, regardless of applied_seq |
"wal_ahead" | the WAL holds content the database has not applied (max_seq is nonzero and exceeds the effective applied_seq) — this is the state that motivated the issue: a wiped or fresh database beside a populated WAL directory must not be mistaken for an authoritative empty corpus |
"hydrated" | the database is caught up with its WAL (applied_seq >= max_seq, and max_seq is nonzero) — this includes applied_seq > max_seq (e.g. following a generation reset elsewhere), which is deliberately classified as caught-up rather than as a distinct anomaly state |
Known narrow limitation: max_seq is 0-indexed (a group’s very first WAL write has seq: 0),
so a group whose entire WAL history is exactly one entry has max_seq == 0 — indistinguishable,
via max_seq alone, from “no content at all,” and reported as "not_applicable". This is the one
case where a wiped-DB-beside-a-populated-WAL condition (the state this field exists to surface) can
go unreported; it resolves itself once the group’s WAL receives a second write. See the
wal_hydration_status doc comment in handlers.rs for why this can’t be resolved by classifying
max_seq == 0 as content-bearing instead — applied_seq == 0 is itself an overloaded sentinel for
both “nothing ever applied” and “genuinely caught up through seq 0,” so doing so would trade this
narrow false "not_applicable" for an equally narrow but more actively misleading false
"hydrated" in the same colliding case.
hydration_status does not change health_check’s healthy/degraded determination in any
way: a wal_ahead group is a normal, fully-queryable state from the process’s own point of view
(it can still serve reads over whatever content it does hold) — the hydration question is per-group
data state, not process health, and is answered here rather than by health_check.
wal.embedding_model, wal.embedding_dim, and wal.embedding_model_status (issue #440;
scoped to the default group, and mirrored per-group inside wal_groups) — report the embedding
model identity under which the group’s currently-applied vectors were computed, alongside
applied_seq/generation in the same WalPosition row (no extra query). This is distinct from
replay reconstructing a graph from a WAL captured under a different embedder: replay
(knowledge_rebuild_from_wal, knowledge_recover with any strategy that replays WAL content, and
startup WAL-corruption self-recovery) always recomputes each embedding vector from its
co-located source text (name/fact/content/summary) using the currently running embedder
— it never binds a value found in the WAL record (issue #526), even for an older WAL captured
before this change: the presence of a stored vector in an older record changes nothing, since
replay never reads it. This is what makes upgrading the embedder self-healing (rebuild, and search
stays consistent) and a published stream’s vectors irrelevant to hydrating it (see “Publishing a
WAL stream” above).
A record with no co-located source text to recompute from — the one open gap this model
identifies explicitly rather than papering over (issue #526’s FR-005) — is handled one of two
ways, depending on the record’s shape: a mutation that exists for no purpose other than writing
that one vector (a “vector-only SET”) is skipped entirely, preserving whatever the entity’s own
create record already computed for that column, since overwriting it with a placeholder would only
degrade it (counted in ReplayStats::embeddings_skip_rows); any other record — most commonly the
node/edge’s own creation, which must still happen — gets a same-dimension zero vector instead, so
the entity is never silently dropped from the rebuilt graph. Recomputation failing for a record
that does have source text (embedder unreachable, or a recomputed vector whose length or
finiteness makes it unbindable) is treated the same way as missing text: skip-or-zero-fill per the
same rule. Both the no-text and the recompute-failure cases are counted in
ReplayStats::embeddings_recompute_skipped_no_text/embeddings_recompute_failed respectively, so
each is distinguishable from the other and from an ordinary clean recompute.
embedding_model_status classifies the comparison between the recorded identity and the
currently running embedder’s own (embedding_model, embedding_dim) (the top-level fields also
present on knowledge_status) — independent of whether a replay has happened in this session, so
a restart under a different embedder is caught regardless of whether the group has ever been
rebuilt. WalPosition.embedding_model/embedding_dim are re-derived from the running embedder
and re-stamped on every successful WAL-position advance — a full replay/rebuild, and every
ordinary write (add_episode, assertion/merge/rebind/correction handlers, backfill, canonicalize,
reprocess) alike — the same “re-derived and persisted on every write” treatment generation
already gets, not something limited to replay call sites. This is a best-effort marker, not a
full-graph audit: a write stamps the identity of the embedder that ran it, not a claim that
every vector currently in the group was computed under that identity — a group that changed
embedders mid-life without an intervening full rebuild can still carry some stale, un-recomputed
vectors from before the change even while embedding_model_status reads "match" (the status
reflects the most recent write’s embedder, and a delete/correction/relabel write stamps the
running identity the same way a content-embedding write does, even though it touched no vector
itself). The one case still uncovered is genuinely fresh: a group with an applied_seq recorded
before this issue shipped (or via a caller with recompute unavailable) shows "unknown" until its
next write or an explicit rebuild — never a false "match".
embedding_model_status | meaning |
|---|---|
"not_applicable" | nothing has ever been applied for this group (applied_seq itself is null) |
"unknown" | a position is recorded, but no embedding identity was recorded alongside it — a pre-#440 write, or a rebuild whose recompute attempts failed |
"match" | the recorded identity equals the running embedder’s (model, dim) |
"mismatch" | the recorded identity differs from the running embedder’s — by model name, by dimension (e.g. a LCG_EMBEDDING_DIM override under the same model name still counts), or both |
A "mismatch" is never a hard failure — it is deliberately self-healing, the same way a
generation mismatch triggers a full replay rather than refusing outright: the fix is to rebuild
(knowledge_rebuild_from_wal), which recomputes every vector it can under the now-running embedder
and updates the recorded identity to match — but only if no recompute attempt actually failed
during that rebuild (e.g. the embedder was unreachable partway through); if one did, the identity
is left unstamped ("unknown") rather than persisted as a "match" it can’t back up, so a rebuild
that didn’t fully succeed never reports a false confirmation. This remedy applies to a
same-dimension model-name change only. knowledge_rebuild_from_wal operates against the
already-created schema, so when the running embedder’s dimension genuinely differs from the one
the database was built with, a recomputed vector’s length won’t match the stored one, and the
rebuild keeps the stale, wrong-dimension vector (counted under embeddings_recompute_failed)
rather than fixing it — it will never report "match" for a real dimension change. For that case,
use knowledge_recover with {"strategy": "rebuild_from_workspace_wal"} instead, which recreates
the database schema at the new dimension before replaying; see
MCP client config recipes in Configuration for the
full explanation and worked examples. A row with no co-located source text to recompute from
(embeddings_recompute_skipped_no_text, issue #526’s FR-002/FR-005) does not by itself block the
"match" write — that outcome is normal, ongoing WAL shape (e.g. a targeted SET that updates
only a vector field, skipped rather than executed so it can’t overwrite a real vector with a
placeholder), not evidence the rebuild failed. Until a mismatch is resolved, it is a live signal
that vector search results may be degraded — the previously-active vectors were computed under a
different model than the one now serving queries.
This database-side comparison is the only embedding-model-identity mechanism as of issue #526:
an earlier, WAL-side sidecar (.wal-embedding-model.json, mirroring .wal-generation.json’s
pattern) and its own replay-time [WAL WARN] embedding-model mismatch: ... check existed
alongside it (issue #440) but were removed once replay stopped ever binding a stored vector — a
WAL’s own claimed write-time identity has nothing left to govern once its stored vectors are never
read. embedding_model_status above, comparing the graph’s currently-applied vectors against
the runner, is what remains; there is no longer a separate way to flag a mismatch at the
write-stream level, independent of whether that stream has actually been replayed into a
database (see “Publishing a WAL stream” above for what this means for a published stream).
The consumer decision, comparing the two fields — check both for null before any numeric
comparison. hydration_status above is a documented shortcut for the common case, but it treats
an absent/never-backfilled applied_seq the same as 0 (per FR-001(b)); it does not distinguish
that from the applied_seq: null “position unknown, full rebuild required” row below, which is a
more serious, overriding condition. A caller that needs to detect the unknown-position case
specifically must still check applied_seq for null itself — hydration_status alone is not a
complete substitute for this table:
applied_seq | max_seq | Meaning | Action |
|---|---|---|---|
null | any | position unknown | full rebuild |
| any | null | WAL empty or unconfigured | nothing to resume from; treat like an empty WAL |
N | N (equal) | DB is caught up | none |
N | M > N | DB is behind, as a forward extension | incremental resume from applied_seq + 1 (not applied_seq — replay’s from_seq filter keeps lines with seq >= from_seq, so resuming at applied_seq would re-replay the last-applied line) |
N | M < N | DB has advanced beyond what the currently-visible WAL contains (e.g. a corpus reset, or a stale copied-back WAL) — not a forward extension | full rebuild |
A bounded rebuild (to_seq set — see WAL administration above) is one
deliberate way to land in the “DB is behind, as a forward extension” row: applied_seq reports
the bounded landing point (<= to_seq), while max_seq still reflects the WAL’s true, unbounded
on-disk maximum. This is expected, not a fault to recover from automatically — an incremental
resume covering the gap (or a later unbounded rebuild) reapplies everything the bounded rebuild
excluded, including a previously-excluded bad mutation.
applied_seq has three distinct values, not two — treat them as different types, not points
on a number line:
null— unknown position. Reported when a pre-existing DB has no recorded position and the one-time backfill (below) fails to derive one: either a populated DB (hasEntityorEpisodiccontent) whose last episode’s uuid isn’t found in the WAL, or a DB with noEpisodicnodes but survivingEntity/relationship content (episode deletion removes only theEpisodicnode, never the entities it created, so a graph can be non-empty with zero episodes — there is nothing left to derive a position from, but real content to lose track of). The documented action is always a full rebuild.0(integer) — a known position: nothing has been applied yet. Reported for a fresh/cleared DB (including a pre-existing DB with noEpisodicnodes and noEntity/relationship content either — genuinely nothing to derive a position from and nothing to lose track of, so the backfill writes0directly without a WAL scan), or immediately afterknowledge_clear_all.- A positive integer — a known, applied WAL position.
Do not treat null as if it sorted below 0. This distinction is not just a Rust/Python
concern — it changes behavior across languages. null < 5 throws or is a type error in Rust
and Python (arithmetic on null/None isn’t defined), which tends to surface the bug
immediately. But in JavaScript, null < 5 coerces to true — a naive port of the “if behind,
resume” comparison silently takes the incremental resume branch on an unknown position,
skipping the full rebuild the null state actually calls for. The same footgun applies to a
null max_seq: 5 < null coerces to false in JavaScript, so a check written only as
applied_seq < max_seq silently falls through neither branch when the WAL is empty or
unconfigured. Check both fields for null explicitly, before doing any numeric comparison, in
every client language.
Upgrading an existing deployment: a DB populated before this feature existed has content but
no recorded position on its first boot under the new version. Rather than reporting null for
that (indistinguishable from a genuinely unknown position, and prone to a client either skipping
a needed rebuild or being unable to tell “empty” from “unknown”), the service backfills a
conservative position on first open, derived from the last Episodic node’s location in the WAL
(the retroactive episode-cursor mechanism from
ADR-0026; see ADR-0353
for why this issue persists a cursor for the fast path in addition to ADR-0026’s own recovery-time
use of the same mechanism). This backfill runs once at startup and is a no-op on every subsequent
boot once a position is recorded.
Streaming progress
Long operations accept a _progress_token and stream progress frames before the terminal
result — see Progress notifications for the MCP
bridge and the list of operations that support it.
Recovery and export tools
knowledge_dump_wal, knowledge_strip_wal_embeddings, knowledge_prepare_checkpoint,
knowledge_wal_mark_create, knowledge_wal_mark_list, knowledge_wal_mark_delete,
knowledge_rebuild_from_wal, knowledge_recover, and knowledge_recover_full are all
admin-scope IPC/MCP tools — see
Scopes for the full admin-scope list and the MCP --scope flag.
Documents liminis-context-graph v0.16.4.