Local Wiki RAG: LightRAG Graph Stack
The wiki uses a two-retrieval-path architecture: qmd for fast lexical+vector search inside Claude Code sessions, and LightRAG for graph-aware synthesis in the TUI and MCP server. qmd runs fully locally at zero cost. LightRAG retrieval (embeddings via ollamanomic-embed-text + graph traversal) is local, but extraction and synthesis call a hosted OpenCode Go model (OpenAI-compatible API); a self-hosted llama.cpp server is planned for the LLM (slot reserved, not wired) and later for embeddings.
Architecture
Why two paths
The agentic-search-vs-rag experiment validated the LightRAG path: graph search achieved 2× retrieval IoU with 99% fewer tokens vs flat RAG. See Agentic Search Vs Rag.
Tools
wiki-chat — interactive TUI
OPENCODE_GO_API_KEY_LIGHTRAG, otherwise queries return an error). Modes match LightRAG’s query modes (local/global/hybrid/naive).
TUI prompt commands:
/mode local|global|hybrid|naive— switch mid-session/reindex— trigger wiki-index for new pages/status— show manifest stats
wiki-index — graph indexer
.env or the shell):
LLAMACPP_BASE_URLset → self-hosted llama.cpp. Reserved, not wired yet: setting it raises an error. The slot exists so a llama-server instance (OpenAI-compatible/v1) can be added later.OPENCODE_GO_API_KEY_LIGHTRAGset → OpenCode Go via its OpenAI-compatible endpoint (https://opencode.ai/zen/go/v1); default modeldeepseek-v4.1-flash, overridable withOPENCODE_LIGHTRAG_MODEL/OPENCODE_LIGHTRAG_BASE_URL. That model reasons by default and, probed 2026-10-04, spent the whole 4096-token budget on thinking for extraction prompts (empty content,finish_reason=length), so requests sendthinking: {"type": "disabled"}(a sample prompt dropped from 2774 to 185 completion tokens); setOPENCODE_LIGHTRAG_DISABLE_THINKING=0for models that reject the field. OpenCode also rejects requests without anx-opencode-sessionheader.- neither → no backend:
wiki-index/--testand queries exit with a clear error;wiki-index --statusandwiki_statusstill work. The check runs before--fullwipes the index.
--full therefore requires --yes; incremental updates (1–3 new pages per ingest) are cheap.
Incremental by default: a manifest.json tracks {path: mtime}. Only changed/new pages are re-extracted. The manifest is saved after each page so partial runs resume automatically.
The post-commit hook triggers wiki-index in the background after any commit touching wiki/. Progress: tail -f .lightrag/last-index.log.
wiki-mcp — MCP server
Graph-aware wiki queries from Claude Code or OpenCode (synthesis counts against the OpenCode Go plan limits). Exposes two tools:wiki_query(question, mode="hybrid")— graph-aware synthesiswiki_status()— show index stats
OPENCODE_GO_API_KEY_LIGHTRAG is set; otherwise wiki_query returns an error). LightRAG graph is initialized once as a singleton; retrieval is always local (nomic-embed-text + graph traversal).
Wire into OpenCode (~/.config/opencode/opencode.json):
Setup
install.sh handles: copying binaries to ~/.local/bin, setting up the post-commit hook, pulling the nomic-embed-text embedding model via ollama. uv handles Python deps via PEP 723 inline metadata — no pip or venv needed.
One-time graph build (required before wiki-chat or wiki-mcp):
Design choices
Graph over flat RAG
Per Agentic Search Vs Rag: graph search wins on cross-concept queries (99% fewer tokens, 2× IoU). Flat RAG only wins on explicit dependency recall. The wiki’s primary use case — “how do X and Y relate?”, “what patterns apply to problem Z?” — is exactly where graph search wins.One concept per page = natural graph nodes
The wiki rule “one thing per page” (CLAUDE.md) makes each page a clean entity for LightRAG to extract. Entities extracted fromconcepts/context-degradation naturally link to concepts/context-compression, concepts/ralph-loop, etc. Cross-links become graph edges.
Hosted LLM, local embeddings
Extraction and synthesis use OpenCode Go (billed against the plan’s usage limits, not per page); embeddings stay local on ollama (nomic-embed-text) because OpenCode has no embedding offering: probed 2026-10-04, neither the Go (36 models) nor the Zen (50 models) catalog lists an embedding model, and POST /embeddings returns 404 on both endpoints. The earlier local-LLM path (qwen2.5 via ollama) was removed. A self-hosted llama.cpp server is the planned replacement for the LLM (env slot LLAMACPP_BASE_URL, reserved but not wired) and for embeddings (not started). When embeddings move, rebuild the index if the embedding model or dimension changes (training data — verify: LightRAG vector stores are tied to the embedding dimension).
Manifest-based incremental indexing
Building the full graph from scratch is the expensive operation (3 extraction phases per page). The manifest approach means each new ingest only costs extraction time for the new pages (typically 1–3 pages). Post-commit automation makes this transparent.Performance
Related
- Agentic Search Vs Rag — experiment validating graph search for this wiki
- Local Rag Elasticsearch — stack comparison; retrieval latency benchmarks
- Contextual Retrieval — chunk context technique; wiki pages are pre-contextualized (one concept per page)
- BM25 — lexical retrieval used by qmd (wiki-context path)
- Reranking — post-retrieval filtering; not yet applied here
- qmd — BM25 + vector engine for the wiki-context skill path
- Wikilink Graph Extraction: Reducing LightRAG Indexing Cost — Obsidian wikilink hints injected at chunk time to reduce LightRAG extraction cost ~40–55%
- Wiki Indexing Pipeline: qmd and LightRAG — how both indexes are built and kept current, step by step, with the failure modes found in practice