Skip to content

IPC & MCP Reference

liminis-context-graph serves the same graph over two transport surfaces, both routed through the same core dispatch in crates/core/src/handlers.rs — no graph logic is duplicated between them:

  • JSON-RPC 2.0 over a local socket (default). Newline-delimited requests/responses over .lcg/service.sock — a Unix domain socket on macOS/Linux, a named pipe on Windows (see Windows: named pipe).
  • Model Context Protocol over stdin/stdout (--mcp-stdio). Any MCP client — Claude Code, Claude Desktop, other agents — can query and mutate the graph directly.

Windows: named pipe

On Windows the service serves the identical protocol over a named pipe, because AF_UNIX is not usable from Rust’s tokio, Node or CPython there (issue #581). LCG_SOCKET_PATH keeps its meaning — it still names the workspace’s endpoint — and maps to a pipe:

  • Discovery. At bind the service writes the pipe name to .lcg/service.endpoint (the socket path with its extension replaced). Clients should read that file.
  • Name. \\.\pipe\lcg-<16 hex digits>: FNV-1a 64-bit over the UTF-8 bytes of the absolute socket path, lower-cased, with / replaced by \. Stable per workspace, so a client that cannot read the file can compute it. A LCG_SOCKET_PATH (or --connect) that is already a \\.\pipe\… name is used as-is.
  • Access. Only the user the service runs as (and SYSTEM) can connect; remote clients are refused, and a second process cannot bind the same pipe name.
  • Signals. Closing the console, logoff/shutdown, Ctrl+Break and Ctrl+C all trigger the same graceful shutdown SIGTERM/SIGINT do on Unix.

A pipe opens like a file, so a client needs no socket library:

import json
endpoint = open(".lcg/service.endpoint").read().strip() # \\.\pipe\lcg-…
with open(endpoint, "r+b", buffering=0) as pipe:
pipe.write((json.dumps({"jsonrpc": "2.0", "id": 1, "method": "health_check", "params": {}}) + "\n").encode())
print(json.loads(pipe.readline()))

In Node, net.createConnection(endpoint) accepts the pipe name directly.

IPC methods (47)

The socket dispatch handles 47 methods: 46 knowledge_* methods plus health_check. health_check is the one method not prefixed knowledge_*, and it is the reason the IPC surface (47) and the MCP tool registry (46, below) differ by exactly one — health_check is not exposed as an MCP tool.

CategoryMethods
Healthhealth_check
Statusknowledge_status, knowledge_rebuild_status
Ingestionknowledge_process_chunk, knowledge_add_episode
Direct assertionknowledge_assert_entity, knowledge_assert_relationship
Searchknowledge_find_entities, knowledge_resolve_entity, knowledge_find_relationships, knowledge_search_passages, knowledge_query_cypher
Graph readsknowledge_get_episodes, knowledge_get_nodes_by_group, knowledge_get_edges_by_group, knowledge_get_edges_by_uuids, knowledge_list_entities, knowledge_list_relationships, knowledge_get_entity_neighbors, knowledge_get_entities_by_source
Deletionknowledge_delete_episode, knowledge_delete_by_source, knowledge_delete_chunk_episode, knowledge_delete_by_group, knowledge_clear_all
Curationknowledge_merge_entities, knowledge_validate_corrections, knowledge_apply_corrections, knowledge_reprocess_entity_types
Ontologyknowledge_reload_ontology
Relation typingknowledge_canonicalize_relations, knowledge_backfill_relation_types (deprecated), knowledge_reprocess_relation_types
Semantic search maintenanceknowledge_backfill_summary_embeddings
Cross-group pointersknowledge_add_cross_group_edge, knowledge_rebind_pointers
WAL administrationknowledge_dump_wal, knowledge_strip_wal_embeddings, knowledge_prepare_checkpoint, knowledge_wal_mark_create, knowledge_wal_mark_list, knowledge_wal_mark_delete, knowledge_rebuild_from_wal, knowledge_build_indices
Recovery / lifecycleknowledge_recover, knowledge_recover_full, knowledge_close

For request/response shapes and parameter details, the dispatch match arms in handlers.rs and their handler functions are the source of truth — this page is the method index, not a copy of each handler’s parameter parsing.

knowledge_status’s WAL fields — including wal.hydration_status (issue #456), which distinguishes a genuinely empty group from one whose WAL holds unapplied content — are documented field-by-field in Operations: knowledge_status health fields rather than here.

Every long-running method above (the six WAL/recovery/reclassification operations most likely to run for a while) accepts a _progress_token and streams {"type":"progress",...} frames before the terminal result — see Progress notifications below.

Readiness

A successful connection to .lcg/service.sock is not evidence the service is ready. The socket is bound before the database opens — deliberately, so health_check and recovery IPC stay reachable during degraded-mode recovery (ADR-0009) — and issue #378’s WAL-root migration also runs in that same pre-open window, after the bind. (Legacy .graphiti/→.lcg/ workspace migration runs earlier still, before the socket is even bound.) The process’s own accept loop only starts once startup work has fully resolved, so a request sent immediately after connect() queues in the kernel rather than racing the migration with stale state — the real risk is a client that treats connect() succeeding as readiness by itself and acts on that assumption (e.g. inspecting on-disk state) without waiting for a health_check round-trip.

The correct readiness signal is a health_check round-trip reporting "healthy": handle_health_check only returns healthy once Db::open() has succeeded, which is after migration has completed. Poll health_check until it reports healthy (or knowledge_status until connected and queryable are both true and initializing is false — knowledge_status has no healthy field of its own) before sending real work; see Operations: Self-healing and degraded mode for the full rationale. A degraded response after startup has otherwise settled is a legitimate outcome (e.g. unrecovered corruption) — not something to retry indefinitely.

A busy response (see Health check contract) is not a readiness signal: it means the service is alive but a write is holding or waiting for the write lock — most importantly a long knowledge_rebuild_from_wal. Pollers keep waiting on busy (gate on result.healthy == true) and must not treat it as failure.

Health check contract

health_check never waits on the write lock, so it answers promptly even during a long WAL rebuild (ADR-0628, #628). The response’s state is one of:

stateokhealthyMeaningSupervisor action
healthytruetrueAlive and ready: DB open, no write in progress or pending.Ready — send work.
busytruefalseAlive, not ready: a write is pending or in progress.Alive — keep waiting; do not restart.
degradedfalsefalseAlive but unusable: DB not loaded. reason says why.Recover (knowledge_recover) or investigate.
(no answer)——Dead or wedged.Restart.

busy carries an activity field. While a knowledge_rebuild_from_wal job is running:

{"ok": true, "healthy": false, "state": "busy", "activity": "rebuilding",
"job_id": "…", "progress": {"mutations_replayed": 48210, "wal_files_processed": 3,
"wal_files_total": 12, "elapsed_seconds": 91.4}}

progress uses the same values as knowledge_rebuild_status. For any other writer the response is {"ok": true, "healthy": false, "state": "busy", "activity": "writing"} with no job_id or progress. The busy path never connects to or queries the database.

Notes:

  • busy means “a write is pending or in progress”: a writer queued behind readers also makes the lock unavailable, so an otherwise idle service can briefly report busy.
  • Batched passes (ADR-0030) release the lock between batches, so during one the state alternates between busy and healthy.
  • A streaming rebuild (_progress_token) is not registered as a job: it reports activity: "writing" without progress.
  • Startup work (legacy and WAL-root migration, startup recovery) runs before the accept loop, so a probe sent then queues in the kernel and gets no answer until startup finishes.
  • healthy can be reported in the short window after a rebuild job is registered but before it takes the write lock.
  • busy has healthy: false on purpose: existing pollers that gate on healthy keep waiting. A liveness probe that only needs “is the process alive” should accept ok: true.

MCP-over-stdio transport

liminis-context-graph --mcp-stdio starts a native Model Context Protocol server over stdin/stdout, using the official Rust SDK (rmcp). Any MCP client (Claude Code, Claude Desktop, other agents) can query and mutate the knowledge graph directly — no Electron app, no Node, no custom JSON-RPC client required. This is an additional external-facing surface; the Unix-socket JSON-RPC protocol above is unchanged.

Every MCP tool is derived from the knowledge_* dispatch methods above — tool names match the IPC method names verbatim, and each tools/call is translated into an IpcRequest and routed straight through the same core dispatch the socket service uses. Tool descriptions and JSON schemas are maintained in the ToolSpec registry in crates/service/src/mcp/tools.rs — that file is the canonical source for per-tool descriptions; they are not duplicated here.

Flags

FlagDescription
--mcp-stdioStarts the MCP server over stdin/stdout instead of binding the Unix socket.
--scope=<list>Comma-separated list of scopes to advertise in tools/list (default all). See Scopes below.
--connect <path>Attached mode: forward every tools/call as JSON-RPC over the given Unix socket to an already-running service, instead of opening the database directly. By default the socket is dialled lazily, on the first tools/call — see DB-access modes below.
--connect-eagerAttached mode only: dial --connect’s socket at startup and exit immediately if it’s unreachable, restoring the behavior attached mode had before issue #575. No effect without --connect.
--allow-remote-closeAttached mode only: advertise and allow knowledge_close, forwarding the shutdown to the remote service. No effect in standalone mode (no --connect).

DB-access modes

  • Standalone (default, no --connect): the MCP process opens the .lcg database directly, reusing the same startup and self-recovery path as the socket service (ADR 0009). Zero-dependency — works with no other process running.
  • Attached (--connect <socket-path>): the MCP process never opens the database; it forwards each call over the given socket to a service that already has it open. Use this to add MCP access to a workspace where another socket-service instance is already running, without contending for lbug’s single-writer lock.
    • Lazy connect by default (issue #575). The socket is not dialled during startup — initialize and tools/list succeed immediately, serving the tool set from the compiled-in registry, regardless of whether the daemon at --connect’s path is up. The first tools/call triggers a dial; if it fails, that call returns a tool result with isError: true naming the socket path and suggesting you start the daemon, but the MCP process itself stays up and the client’s session is unaffected — no protocol error, no exit. This avoids two failure modes a client can otherwise get stuck in for the client’s own cached-failure window: registering a native MCP server before its daemon has started (first install, post-reboot), and the client caching a failed attach for several minutes even after the daemon comes up. Pass --connect-eager to restore the pre-#575 behavior of dialling at startup and exiting immediately (before initialize) if the socket is unreachable — useful if you rely on that fail-fast behavior as an external health check (e.g. a supervisor that restarts the front until the daemon is ready).
    • Idle timeout. LCG_ATTACHED_CALL_TIMEOUT_MS (default 30s) is a per-read-line idle timeout, not a whole-call timeout: it resets on every line read off the socket, including {"type":"progress"} lines. A call that keeps emitting progress is never bounded by it, no matter how long the call runs in total — only genuine silence (no output at all for the full timeout window) trips it. If the remote stops responding mid-call (e.g. it crashes), the attached client fails that call with a clean timeout error rather than blocking forever.
    • Reconnect and retry. If the connection to the remote breaks, the client transparently re-dials the same socket path rather than staying wedged. If the break is detected while writing the outgoing request — treated as safe to retry, since the write failing is the client’s best available signal that the request didn’t get through — the client automatically retries that request exactly once over the freshly-dialed connection. If the break is detected only after the request was fully written — while waiting for or reading the response — the call is not retried automatically, since the remote’s execution status is unknown and blind retry could double-apply a non-idempotent write (e.g. knowledge_add_episode); that call fails with a clear “connection lost mid-call” error, but the connection is marked dead so the next call reconnects fresh. If a reconnect attempt itself fails (no listener at that path), the call fails with a clear, descriptive error — never a hang — and a later call will try reconnecting again. See ADR-0040 for the full rationale.

Scopes

Scopes are additive and composable (e.g. --scope=read,admin). tools/list advertises the union of all active scopes.

ScopeMethods
readknowledge_status, knowledge_find_entities, knowledge_resolve_entity, knowledge_find_relationships, knowledge_get_episodes, knowledge_get_nodes_by_group, knowledge_get_edges_by_group, knowledge_get_edges_by_uuids, knowledge_search_passages, knowledge_list_entities, knowledge_list_relationships, knowledge_get_entity_neighbors, knowledge_get_entities_by_source, knowledge_rebuild_status, knowledge_validate_corrections
writeknowledge_process_chunk, knowledge_add_episode, knowledge_delete_episode, knowledge_delete_by_source, knowledge_delete_chunk_episode, knowledge_clear_all, knowledge_apply_corrections, knowledge_merge_entities, knowledge_reprocess_entity_types, knowledge_canonicalize_relations, knowledge_backfill_relation_types, knowledge_reprocess_relation_types, knowledge_add_cross_group_edge, knowledge_assert_entity, knowledge_assert_relationship
cypherknowledge_query_cypher
adminknowledge_dump_wal, knowledge_strip_wal_embeddings, knowledge_prepare_checkpoint, knowledge_wal_mark_create, knowledge_wal_mark_list, knowledge_wal_mark_delete, knowledge_rebuild_from_wal, knowledge_recover, knowledge_recover_full, knowledge_close, knowledge_build_indices, knowledge_rebind_pointers, knowledge_delete_by_group, knowledge_backfill_summary_embeddings, knowledge_reload_ontology
allevery scope above (default)

cypher is a power scope, not bundled into anything else. knowledge_query_cypher executes raw Cypher with no param interpolation or value coercion — despite being a “query” method, it can perform arbitrary mutations, and it bypasses the WAL-ordering and embedding invariants that the structured write tools maintain. It is never implicitly included in read, write, or admin; operators must opt in explicitly (or via all).

knowledge_close in attached mode is a footgun without --allow-remote-close. In standalone mode, knowledge_close is always advertised under admin scope and shuts down only this MCP process’s own DB connection. In attached mode, calling it would shut down the running remote service. Without --allow-remote-close, knowledge_close is omitted from tools/list entirely in attached mode (not merely rejected when called). Pass --allow-remote-close only when you specifically intend this MCP connection to be able to stop the remote service.

Recovery and export live under admin. knowledge_rebuild_from_wal (rebuild one group’s data from its own WAL directory — group_id, default "liminis"), knowledge_dump_wal (snapshot/export the graph into a fresh compacted WAL directory), knowledge_strip_wal_embeddings (rewrite existing WAL files in place to remove the embedding-vector params fields a pre-0.14 WAL still carries — see Operations), knowledge_wal_mark_create / _list / _delete (name a retained WAL position within one group’s stream, without a full snapshot — each also takes group_id, default "liminis", and _list reports only that one group’s marks, never an aggregate across groups), and knowledge_recover / knowledge_recover_full are all admin-scope tools — an attached client only sees them when launched with --scope=admin (or all). If a mutation goes wrong, this is the recovery path. See Operations for the recovery model in full. Note the WAL replays forward-only, so take periodic knowledge_dump_wal snapshots, or a lighter-weight knowledge_wal_mark_create named position, if you want restore points before large or destructive operations — a mark does not survive knowledge_dump_wal, since dump_wal renumbers sequence numbers and a copied mark’s seq would be meaningless against the new numbering.

knowledge_rebuild_from_wal refuses to run against a non-empty group, unless you ask it not to. Since issue #378, one instance holds an independent WAL directory and applied position per group_id; knowledge_rebuild_from_wal {group_id, ...} targets exactly one of them and never disturbs another group’s data or position. A from_seq: 0 (default) full rebuild against a group that already contains data fails fast with an explicit error rather than silently emitting a duplicate-primary-key failure for every existing Entity/Episodic/RelatesToNode_ row in that group — the native write path uses CREATE, not MERGE, for those labels. Pass force_clear: true to have the call clear that group’s data automatically before replaying (the same group-scoped purge knowledge_delete_by_group uses — this does not delete or reopen the database file, unlike the pre-378 whole-database force_clear behavior), or clear it yourself first with knowledge_delete_by_group {group_ids: [group_id]}. dry_run: true always fails fast on a non-empty group regardless of force_clear, since a dry run must never mutate the database — this lets a preview surface the problem before you commit to a real rebuild. None of this applies to an incremental from_seq > 0 resume, which intentionally targets a group that already has state.

to_seq bounds replay from the other end: from_seq <= seq <= to_seq. Pass an inclusive upper bound to exclude a mutation (and everything after it) from a rebuild — e.g. a WAL-recorded mutation that corrupted the graph. Omit to_seq for today’s unbounded behavior (replay to the end of the WAL); it must not be less than from_seq, or the call is rejected before any WAL line is read or the database is touched. A bounded rebuild is not durable: WAL entries past to_seq stay on disk, unapplied, not truncated or archived — a later unbounded rebuild, or a from_seq resume that covers the excluded range, reapplies them, including a previously-excluded bad mutation. to_seq bounds an endpoint; it does not add reverse/undo semantics to the forward-only replay noted above.

Search evidence and the similarity floor (find_entities, find_relationships)

Hybrid search fuses its candidate lists with Reciprocal Rank Fusion, which is rank-only: the top hit of a gibberish query scores almost the same as the top hit of a perfect one, and every query fills its top-k. To let a caller tell a real match from filler, each result of knowledge_find_entities and knowledge_find_relationships carries an additive search object (issue #629):

FieldOnMeaning
rrf_scorebothfused, rank-based score — ordering only, not a relevance measure
text_matchbothtrue iff the full-text (BM25) path retrieved the item
bm25_scorebothraw BM25 score, or null
name_similarity, summary_similarityentitiescosine similarity (1 - distance) on that vector path, or null
fact_similarityrelationshipscosine similarity (1 - distance) on the fact-vector path, or null

A field is null when that retrieval path did not return the item (text_match is then false). Values are raw and unnormalised; cosine similarity can be negative.

Both tools also accept an optional min_similarity (number). Vector candidates whose cosine similarity is below it are dropped before fusion, independently per vector path, so they never affect the ranking and are not reported for an item they did not qualify for. Full-text matches always stay eligible, so an exact-name query still returns its match while a gibberish query returns nothing. With the kind filter the floor applies on top of the kind path’s bounded nearest-neighbour probe.

How it differs from knowledge_search_passages’ min_score: the value is clamped to 0–1 the same way (a non-number is an invalid-params error), but min_similarity has no default — omitted or null means no floor and results are identical to before, apart from the added search object — and it acts before fusion and per path rather than as a post-filter on a single list.

group_ids semantics: omitted vs. empty

Every read tool that accepts group_ids treats an omitted or null group_ids uniformly as “all groups” (issue #413) — this holds across all 10 read tools: knowledge_find_entities, knowledge_find_relationships, knowledge_search_passages, knowledge_list_entities, knowledge_list_relationships, knowledge_get_entity_neighbors, knowledge_get_entities_by_source, knowledge_get_nodes_by_group, knowledge_get_edges_by_group, knowledge_get_episodes.

An explicit group_ids: [] does not mean the same thing on every tool, and this split is deliberate, not an inconsistency to be fixed here:

  • On knowledge_find_entities, knowledge_find_relationships, knowledge_get_nodes_by_group, and knowledge_get_edges_by_group, an explicit group_ids: [] is preserved as “exactly these groups” — i.e. zero rows, a filter matching nothing.
  • On the other six read tools (knowledge_search_passages, knowledge_list_entities, knowledge_list_relationships, knowledge_get_entity_neighbors, knowledge_get_entities_by_source, knowledge_get_episodes), an explicit group_ids: [] collapses to the same behavior as omitting it — all groups.

If you need “zero rows” as a filter result, use one of the four tools in the first list, or check the tool’s own ToolSpec description in crates/service/src/mcp/tools.rs before relying on [] to mean “nothing” — on the other six it doesn’t.

This section is the single, central statement of that contract; individual tool entries on this page don’t restate it.

Bulk reads: group scope, paging, projection, prefix (issue #667)

knowledge_get_episodes and knowledge_list_entities are the two bulk reads, and a large graph can overflow an MCP client’s tool-result limit if read in one call. Both accept the same optional, additive parameters; with none of them the response is exactly {<collection>, count} as before.

ParameterToolsMeaning
group_ids (array)bothGroups to read. knowledge_get_episodes also accepts the single-group alias group_id; both given ⇒ their deduplicated union. Neither given ⇒ every group (episodes used to default to liminis). An explicit [] also means all groups.
last_n / num_resultsepisodes / entitiesPage size (defaults 50 / 500, unchanged).
cursorbothOpaque keyset cursor. Send "" for the first page, then send back each response’s next_cursor until it is null.
fields (array)bothReturn only these keys per item. Episodes: uuid, name, group_id, created_at, source, source_description, content, valid_at, entity_edges, attributes. Entities: uuid, name, group_id, labels, kind, created_at, summary, attributes, episode_uuids, source_descriptions. Unknown or empty fields is an error; embeddings are never returned.
name_prefixbothNames starting with this prefix, matched literally. Case-insensitive for entities (via Cypher lower(), so non-ASCII folds only as far as lbug does); case-sensitive for episodes.

next_cursor appears only when the request contained cursor, fields or name_prefix; it is null on the last page. Order is deterministic (episodes newest first with a uuid tiebreaker; entities by uuid descending) and the cursor is a keyset position, so inserts and deletes between pages never repeat or skip an item that existed throughout. A cursor is bound to the tool and to its group_ids, kind and name_prefix; a malformed cursor or one replayed against a different query fails with invalid cursor: … (JSON-RPC -32000, like every other validation error here). count is always the number of items in the page. Filtering episodes by attributes is not supported yet. See ADR-0667.

knowledge_status: the flat wal block always describes the default group’s WAL stream. When a WAL root is configured it now also carries "scope": "default_group", "default_group": "liminis" and "see": "wal_groups", so exists: false there means the default group has no stream — read wal_groups for every group’s position. With no WAL root the block is unchanged.

Deletion (delete_chunk_episode, delete_by_source)

Breaking change in 0.13.2 (issue #406): group_ids is required and non-empty on both knowledge_delete_chunk_episode and knowledge_delete_by_source. A call that previously omitted group_ids deleted matching episodes across every group in the workspace; that was cross-group-unsafe, so as of 0.13.2 an omitted, null, or empty group_ids is rejected outright with an error, before any delete runs. There is no default to fall back to — a caller upgrading from a pre-0.13.2 client that relied on omission must start passing its own group(s) explicitly:

{"chunk_id": "notes-0001", "group_ids": ["liminis"]}
{"source_file": "notes.md", "group_ids": ["liminis"]}

This is a different contract from the read-tool one above — omitting group_ids here never means “all groups”; it means “reject the call.” A group_ids naming a group with no matching rows still succeeds, returning deleted_count: 0.

Direct assertion (assert_entity, assert_relationship)

knowledge_assert_entity and knowledge_assert_relationship (issue #379) are a direct write path: the caller already knows a fact and records it as a single entity or edge, without a prose round-trip through knowledge_process_chunk’s LLM-driven extraction. Use process_chunk when you have unstructured text and want the graph populated by extraction; use the assert tools when you already know exactly which entity or edge you want to write.

knowledge_assert_entity accepts name (required), entity_uuid (optional), kind (optional, see Entity kinds), labels, summary, attributes, and group_id (default liminis).

  • Upsert identity is (group_id, kind, name) (kind defaults to Entity), unless entity_uuid is supplied. entity_uuid, when given, is a strict, group-scoped lookup — the call fails if no entity with that UUID exists in group_id; there is no create-under-this-UUID fallback (ADR-0379 Decision 3). A caller cannot mint an entity at a UUID of its own choosing.
  • On an update, omitting summary or attributes clears the previously stored value rather than leaving it untouched — both fields are always overwritten with whatever the call supplies: summary defaults to "" if omitted, attributes defaults to "{}" (an empty JSON object, not an empty string) if omitted or non-object, matching how labels/name are handled.
  • Only name is embedded for semantic search (name_embedding). summary is stored and is full-text searchable (it’s part of Entity’s [name, summary] FTS index), but is not semantically searchable — a knowledge_find_entities vector query will not match on summary content until issue #470 lands. attributes is not indexed at all — full-text or semantic — and is retrievable only via direct UUID lookup or knowledge_query_cypher.
  • An update never re-embeds name_embedding — only a create does (ADR-0379 Decision 6, forced by a lbug constraint: the embedded column sits under an HNSW index once indexes are built, and lbug rejects a plain SET on an indexed column). Re-asserting an existing entity with a changed name updates the stored name, but name_embedding keeps reflecting the old name until the entity is deleted and recreated. There is currently no supported way to refresh it in place.

knowledge_assert_relationship accepts source_name, target_name, predicate (all required), source_kind / target_kind (optional, see Entity kinds), fact (optional — auto-derived as "{source_name} {predicate} {target_name}" if omitted), attributes, relation_type, valid_at, and group_id (default liminis).

  • Upsert identity is (source_node_uuid, predicate, target_node_uuid, group_id), resolved via find_active_relates_to_uuid (an invalidated edge is skipped, so re-asserting after an invalidation creates a fresh edge rather than resurrecting the old one).
  • Endpoint resolution is strictly scoped to the call’s own group_id — it never falls back to a cross-group search. If source_name or target_name doesn’t resolve to an entity already in that group, the call fails with an error naming knowledge_add_cross_group_edge as the tool to use for connecting entities across groups.
  • fact is the field embedded for semantic search (fact_embedding) — not name/predicate, the opposite of knowledge_assert_entity. fact is also part of RelatesToNode_’s [name, fact] FTS index, so it’s both full-text and semantically searchable; attributes is indexed neither way, same as on the entity side.
  • Same never-re-embed-on-update caveat as knowledge_assert_entity (ADR-0379 Decision 6): a re-assert that changes fact updates the stored fact text but leaves fact_embedding reflecting the prior text until the edge is deleted and recreated.

See ADR-0379 for the full rationale behind these upsert and embedding decisions.

Entity kinds (kind identity)

Entity identity is (group_id, kind, name) (issue #615, ADR-0615). kind is a single string classifying an entity (Topic, KnowledgeChannel, Team, …). It is stored in Entity.kind, is always one of the entity’s labels (appended after Entity), and is immutable. The default kind is exactly Entity; every entity written before kinds existed, and every extraction-created entity, is that kind. A kind is trimmed, case-sensitive, non-empty, must not contain U+001F, and must not be the reserved structural label Merged, else the call fails with a validation error. Every entity node in a response carries a kind field. So Topic "adr" and KnowledgeChannel "adr" are two distinct entities that coexist in one group.

Omitted kind: reads are broad, writes are scoped (the rule group_ids already follows):

Operationkind givenkind omitted
knowledge_assert_entity (write)resolves/creates exactly that kindacts on the default kind Entity only — never touches another kind’s same-named node
knowledge_assert_relationship endpoints (existing entities)that endpoint resolves exactly within the kindspans all kinds; more than one matching kind is an ambiguity error
knowledge_resolve_entity (name, group_id, kind?)exact lookup, never ambiguousspans all kinds; one match resolves, several is an ambiguity error, none returns {found: false}
knowledge_find_entities, knowledge_list_entitiesonly that kindall kinds, each node carrying its kind (multi-row reads never pick one, so never ambiguous)
knowledge_add_cross_group_edge source_kind/target_kind (foreign endpoints only)pointer pinned to that kind (endpoint_kind)pointer resolves across all kinds; ambiguous once a second kind shares the name
knowledge_merge_entities kindrestricts canonical_name/alias_namesthose names span all kinds; a multi-kind canonical_name is an ambiguity error

Ambiguity error. A name-only resolution that matches more than one kind fails with JSON-RPC error code -32002; error.data is {"reason": "ambiguous_entity", "name", "group_id", "candidates": [{"uuid", "kind"}, …]}. It never picks a candidate. Retry with an explicit kind.

Merges never cross kinds. knowledge_merge_entities refuses the whole call (dry_run included, nothing modified) if the canonical and any alias differ in kind, with an error naming the kinds; merge_all_by_name is limited to the canonical’s own kind.

Extraction and kinds. knowledge_add_episode/knowledge_process_chunk create kind-Entity entities, except that an entity whose primary type the ontology marks identity: true is created with that type as its kind and resolves only within it (see Identity-bearing types). Extraction never merges into an asserted kind it did not create. For a group whose identity-bearing set changed in a way that would reinterpret existing entities, knowledge_add_episode fails with JSON-RPC error code -32003; error.data is {"reason": "identity_set_change_refused", "group_id", "label", "entity_count"}, and knowledge_status lists the group under group_identity_refusals. Nothing is modified; restore the recorded set (and restart, or call knowledge_reload_ontology for the group) or re-ingest the group from source. knowledge_reprocess_entity_types never changes an entity’s kind and adds an additive kind_disagreements array (entity_id, entity_name, kind, classified_type) — report-only — to both its dry-run plan and its result. Ontology types flagged extract: false are never offered as reclassification targets by knowledge_reprocess_entity_types / knowledge_reprocess_relation_types (nor by knowledge_canonicalize_relations); see Assert-only types.

Ontology reload (knowledge_reload_ontology)

knowledge_reload_ontology {group_id} (admin scope) discards one group’s cached ontology resolution and re-resolves it, so an edited per-group ontology file takes effect without a service restart. It returns {group_id, previous_hash, new_hash, changed, drift, identity_refusal}; an identity-bearing-set change on a group that already holds entities of that type is reported in identity_refusal (and enforced as -32003 on knowledge_add_episode) rather than installed silently. Other groups are untouched and stored data is never modified. See Ontology → Reloading a group’s ontology.

Relation typing (canonicalize_relations, backfill_relation_types, reprocess_relation_types)

Three tools populate an edge’s relation_type, with different tradeoffs:

knowledge_canonicalize_relations maps each edge’s existing raw predicate onto your ontology’s declared relation_types. Its behavior has four caveats worth knowing before you rely on it:

  • group_id is required (issue #447) — candidate selection and the resulting WAL mutations are both restricted to that one group; no other group’s edges or WAL stream are ever touched. An omitted, null, or empty group_id is rejected before any candidate selection or write, rather than falling back to a database-wide rewrite or the default group — there is no supported way to canonicalize every group in one call.
  • The primary pass is lexical, over the predicate — not the fact. It matches the edge’s predicate / current relation_type against ontology type names, aliases, and keywords. It does not read the edge’s fact sentence, so an edge whose relation_type was cleared cannot be re-mapped from its fact by this pass.
  • embedding_threshold tunes only the fallback promoter (default 0.7). The fallback embeds each residual edge’s fact against the ontology types’ descriptions and force-assigns the single nearest type at or above the threshold. Lowering it types more edges, but by nearest-neighbor force-fit with no abstention — an idiosyncratic fact (e.g. “X is affiliated with Y”) can land on a wrong type (e.g. HOLDS).
  • Re-runs are only partly idempotent, and clearing relation_type can’t be undone by canonicalize. A re-run skips an edge only when it’s already at its target — a Mapped edge already equal to the canonical type, or a residual edge already UNCLASSIFIED; an edge whose classification changes is overwritten (including a previously-assigned type), while arrow-named “noise” edges keep any existing predicate. Critically, canonicalize’s only input is the edge’s existing predicate / relation_type — if you null that field to “start clean,” canonicalize has nothing to map from and cannot rebuild it. Snapshot with knowledge_dump_wal before such an operation.

knowledge_backfill_relation_types (DEPRECATED) does not classify at all — it mints uppercased fact-prefix pseudo-types (e.g. THE_SPECIFICATION_DOCUMENT_DEFINES) for edges with no relation_type, rather than matching against the ontology. Avoid it for building a typed taxonomy; prefer knowledge_reprocess_relation_types below, which supersedes it for that purpose. Like knowledge_canonicalize_relations, group_id is required (issue #447): candidate selection and WAL attribution are both restricted to that one group, and an omitted, null, or empty group_id is rejected rather than running database-wide or against the default group.

knowledge_reprocess_relation_types is the relation-side twin of knowledge_reprocess_entity_types: for each in-scope edge, it sends the edge’s fact and the ontology’s declared relation types (name + description) to the configured extraction LLM and asks it to pick exactly one type, or honestly abstain. This is the tool to reach for when you want genuine fact-based classification instead of lexical matching or pseudo-typing:

  • Always reads the fact, never the predicate. Unlike canonicalize_relations, classification is grounded in the edge’s natural-language fact sentence against the ontology’s declared menu — not lexical/alias/keyword matching on the existing predicate string.
  • A declared ontology relation-type menu is always required. Unlike knowledge_reprocess_entity_types (whose untyped scope works with no ontology via open-ended classification), every scope value (untyped, off_ontology, all) fails with a structured {success: false, error: ...} if the ontology declares no relation types — there is no open-ended fallback (see ADR-0037).
  • Abstention is an honest, real write of UNCLASSIFIED. If the LLM cannot map a fact to any declared type, the edge’s relation_type is set to the literal string UNCLASSIFIED — never a force-assigned nearest match. This differs from knowledge_reprocess_entity_types, where an unclassifiable entity is simply left unchanged (see ADR-0037).
  • scope controls candidates: "untyped" (default) — relation_type NULL/empty, the same predicate backfill_relation_types uses; "off_ontology" — untyped edges plus edges whose relation_type isn’t a declared type (this naturally covers prior UNCLASSIFIED sentinels and backfill_relation_types’s fact-prefix pseudo-types with no special-casing); "all" — every edge in the group.
  • Idempotent. An edge whose computed verdict already matches its current relation_type (including an edge already correctly UNCLASSIFIED) is left unchanged — no write, no WAL entry.
  • dry_run: true returns would_reclassify_count, a plan array of per-edge {edge_id, fact, old_type, new_type} entries, and a breakdown object counting edges per assigned new_type (including an UNCLASSIFIED count) — without mutating the graph.
  • dry_run: false (apply) returns reclassified_count, unchanged_count, and — since issue #332 — the same breakdown object as the dry-run path, so callers can see the per-type classification distribution (including how many candidates abstained to UNCLASSIFIED) without a separate dry-run call. plan and would_reclassify_count remain dry-run-only, since they describe a proposed mutation rather than one that already happened; breakdown is always present on a successful apply response, as {} when there were zero in-scope candidates.

Semantic search maintenance (backfill_summary_embeddings)

knowledge_backfill_summary_embeddings (issue #470) computes summary_embedding for every entity in group_id that has a non-empty summary, so entities created before this capability existed become semantically retrievable by summary paraphrase — not just by name-vector or full-text match. knowledge_find_entities already fuses a summary-vector match into its hybrid retrieval for entities embedded going forward (both knowledge_assert_entity and the extraction path embed summary on creation); this tool is what makes that retrieval available for entities that predate the capability.

  • Every candidate is unconditionally re-embedded on each call. There is no cheap way to tell “already has a real embedding” apart from “still carries the schema migration’s zero-vector placeholder” from a stored FLOAT[] value alone, so re-running this backfills the same rows again. This is safe (idempotent in effect) but not free (an embedder round-trip per candidate, same as knowledge_backfill_relation_types and subject to the same no-batching caveat, #445) — use dry_run: true first to see the candidate count.
  • Holds an exclusive lock for the whole run, blocking all other reads and writes. Unlike most admin operations, the summary-vector HNSW index is dropped for the run’s duration and rebuilt at the end — this is the only way an indexed embedding column can be refreshed for existing rows at all (a plain SET on an HNSW-indexed column is rejected once the index exists). Prefer running this at a low-traffic time, especially for a large group.
  • group_id is required (matching knowledge_backfill_relation_types’s convention): candidate selection and WAL attribution are both restricted to that one group; an omitted, null, or empty group_id is rejected rather than running database-wide or against the default group.
  • A partially-completed backfill is not an error. Entities not yet processed simply retrieve via existing name/lexical behavior, exactly as before this issue — running the tool again covers more of the group each time.
  • A knowledge_assert_entity re-assert does not refresh summary_embedding. A re-assert that changes an entity’s summary leaves its summary_embedding as it was — the vector reflects whichever summary was embedded last (at creation, at an extraction merge, or at the most recent backfill run), not necessarily the current summary text. Re-run this tool to bring a changed summary’s vector back in sync. (Extraction merges do refresh it: when extraction folds a new description into an existing entity the merged summary is consolidated into one bounded description of at most 600 characters and re-embedded in the same write — ADR-0651.)

Cross-group pointers (add_cross_group_edge, rebind_pointers)

knowledge_add_cross_group_edge creates an edge whose endpoint(s) may live in a group_id other than the edge’s own — the hub/layer-graph topology introduced by issue #369 (see ADR-0369). Every intra-group edge write (knowledge_add_episode, knowledge_process_chunk, and every other existing write path) is completely unaffected: pointers only ever exist on edges created through this tool.

  • Each endpoint is either a bare UUID or a name to resolve. {"uuid": "..."} names an entity already known to live in the edge’s own group_id — no resolution, no pointer. {"source_group_id": "...", "endpoint_name": "..."} names a foreign endpoint: it is resolved by case-insensitive name lookup against that group (the same authority get_entity_by_name_ci_with_scan_fallback uses for extraction-time endpoint resolution, per ADR-0283), and the edge carries a cross_group_pointers.{src,dst} object recording the assertion (source_group_id, endpoint_name, plus an optional endpoint_kind) and the resolution cache (resolved_uuid, bound_at_seq, binding_state). Top-level source_kind / target_kind (issue #615) pin a foreign endpoint to an entity kind: a kind-pinned pointer resolves within that kind only and stays bound when a same-named entity of another kind appears; a kind-less pointer resolves across all kinds and becomes ambiguous (no hop) once a second kind shares the name. A kind on a {uuid} endpoint is rejected. If a pinned kind no longer exists the pointer is unbound.
  • A foreign endpoint that doesn’t currently resolve is unbound, not dropped. Unlike ordinary extraction (which hard-drops an unresolvable endpoint at commit — see ADR-0051), the edge is still created; only the hop to that side is missing until a later knowledge_rebind_pointers call resolves it. A binding_state of ambiguous means more than one entity currently matches the name — also retained, also missing that hop, never a silently-guessed winner.
  • A bare-UUID endpoint that turns out to belong to a different group than the edge is rejected before any write happens — this is what keeps a cross-group edge from silently losing its pointer fields the first time a caller passes a UUID instead of a name.
  • knowledge_rebind_pointers ({"source_group_id": "..."}, required) re-resolves every pointer whose source_group_id matches, after that source group’s own hydration, incremental replay, or refresh cycle — including an ordinary knowledge_rebuild_from_wal targeting that one group, not only a full purge-and-rehydrate (issue #378). A pointer currently bound is skipped once its bound_at_seq is already at or past source_group_id’s own applied WAL position — never any other group’s, even when the edge carrying the pointer lives in a third, different group — this staleness gate is what makes a second call with no intervening change to source_group_id’s stream a true no-op for pointers that are already correct. A pointer currently unbound or ambiguous is always re-resolved regardless of bound_at_seq (issue #392): a known-broken pointer is repaired unconditionally, since the position comparison alone cannot tell “nothing changed” apart from “the source group was purged and then restored to the same position the pointer was originally bound at.” A resolution that would create a self-loop or duplicate an existing directed edge invalidates the edge instead of writing a broken or redundant one, reusing knowledge_merge_entities’s own self-loop/dedup handling rather than a new policy. Returns {checked, bound, unbound, ambiguous, invalidated_self_loop, invalidated_duplicate, staleness_skipped} — staleness_skipped counts pointers skipped by the gate above, distinct from checked (pointers actually re-resolved), so a checked: 0 result is never ambiguous about whether anything was examined.
  • Unbound and ambiguous edges are excluded from normal reads. Every existing two-hop traversal, search, and MCP read path requires both hops to exist — a pointer that hasn’t resolved is invisible the same way any other incomplete edge would be, no special-casing needed. Aggregate counts are visible via knowledge_status’s cross_group_pointers: {bound, unbound, ambiguous} field, so a refresh in progress is observable without a dedicated inspection endpoint.

Ingestion results (process_chunk)

knowledge_process_chunk reports what it could not write, not only how much. Two additive result fields exist for that, both introduced in 0.13.2.

  • dropped_edges reports the edges behind edges_dropped_unresolvable. An extracted edge whose source or target endpoint resolves to no entity — neither in the current extraction batch nor in the persisted graph — is dropped rather than written (ADR-0051). edges_dropped_unresolvable counts those drops; dropped_edges describes them, with one entry per counted edge, in extraction order:

    {
    "edges_dropped_unresolvable": 1,
    "dropped_edges": [
    {
    "source_name": "Ada Lovelace",
    "target_name": "Analytical Engine",
    "relation_type": "WROTE_NOTES_ON",
    "fact": "Ada Lovelace wrote the first published algorithm for the Analytical Engine.",
    "unresolved_endpoint": "target"
    }
    ]
    }

    unresolved_endpoint is "source", "target", or "both". relation_type may be null, mirroring the fact that it is already optional on an extracted edge before resolution is attempted. The dropped edge’s content is not persisted anywhere, so this result is the only place it appears — a consumer that wants to tell a user which fact was lost must read it here rather than recover it later. dropped_edges is always present, an empty list when nothing was dropped, so it can be iterated unconditionally. edges_dropped_unresolvable’s existing meaning is unchanged, and a caller reading only the count is unaffected (issue #411).

  • edges_dropped_self_loop counts edges dropped because both endpoints resolved to one entity. Two differently named endpoints can resolve to the same entity UUID (through dedup merges or salvage); inserting that edge would create a self-loop, so Phase C drops it. The count is a separate top-level field, always present, and these edges are not in dropped_edges or edges_dropped_unresolvable. (Edges whose endpoint names are identical are filtered earlier and never counted.) Separately, an off-list edge endpoint that already exists exactly (same group, case-insensitive) in the stored graph resolves to that entity and is never salvaged onto a similar entity of the current chunk; a name stored under more than one eligible kind is dropped as ambiguous rather than salvaged (ADR-0666).

  • warning reports oversized input. A chunk_text longer than the advisory threshold (LCG_CHUNK_TEXT_ADVISORY_MAX_CHARS, default 8,000 characters — see Configuration) adds a warning field naming the actual and recommended character counts, and emits a chunk_text_oversized telemetry event. This is visibility only: nothing is truncated, split, or rejected, and the call succeeds exactly as it did before. Splitting oversized input is the caller’s responsibility (issue #407).

  • dedup_paths reports how this chunk’s entities were resolved. An additive object of per-path tallies for the chunk’s Phase B entity resolution (ADR-0650), always present:

    {
    "dedup_paths": {
    "exact_name": 3,
    "embedding_merge": 1,
    "vetoed": 2,
    "adapter_rejected": 0,
    "llm_confirmed": 0,
    "llm_rejected": 0,
    "llm_unavailable": 0,
    "salvage_vetoed": 0
    }
    }

    exact_name counts entities merged by the exact case-insensitive name match; embedding_merge counts embedding-path merges where a candidate survived the identifier veto and the dedup adapter confirmed it; vetoed counts entities that had above-threshold embedding candidates, all of which were rejected by the identifier-mismatch veto (names differing by a number or identifier, such as ADR 2018 / ADR 2019, or by a distinct short all-caps code such as ACDS / ACDM), so a new entity was inserted; adapter_rejected counts candidates that survived the veto but were rejected by a non-LLM (custom) dedup adapter; llm_confirmed, llm_rejected and llm_unavailable count candidates judged by the LLM dedup check (LCG_DEDUP_LLM, ADR-0652) as duplicate (merged), distinct (inserted) and unanswerable — error, timeout or malformed answer — (inserted); when the check is off, a surviving candidate merges and counts as embedding_merge; salvage_vetoed counts off-list edge endpoints whose salvage candidates were all vetoed. The tallies are per chunk; sum them across chunks for a run total.

Progress notifications

The six long-running operations — knowledge_rebuild_from_wal, knowledge_canonicalize_relations, knowledge_backfill_relation_types, knowledge_backfill_summary_embeddings, knowledge_reprocess_relation_types, and knowledge_reprocess_entity_types — bridge to MCP progress notifications when the client attaches a progress token to the tools/call request (_meta.progressToken), in both standalone and attached mode. Without a progress token, these calls simply block until they complete, same as over the socket protocol. In attached mode, each progress notification also re-arms LCG_ATTACHED_CALL_TIMEOUT_MS’s per-read-line idle timer (see DB-access modes above), so a progress-tracked call isn’t falsely reported as timed out just because it runs longer than that timeout in total.

Example MCP client config

{
"mcpServers": {
"liminis-context-graph": {
"command": "liminis-context-graph",
"args": ["--mcp-stdio", "--scope=read,write"],
"cwd": "/path/to/your/workspace"
}
}
}

To attach to an already-running socket service instead of opening the DB directly:

{
"mcpServers": {
"liminis-context-graph": {
"command": "liminis-context-graph",
"args": ["--mcp-stdio", "--connect", "/path/to/your/workspace/.lcg/service.sock", "--scope=read"]
}
}
}

See ADR 0035 for the transport’s internal architecture.

To point the embedder used by an MCP-launched instance at a different backend (a local OpenAI-compatible server, a hosted provider with an API key, or a non-default sidecar socket), see MCP client config recipes in Configuration.

Documents liminis-context-graph v0.16.4.