Wiki Indexing Pipeline: qmd and LightRAG
This repo is indexed two ways, for two different jobs. qmd finds pages; LightRAG connects them. Architecture and tool usage are in Local Wiki RAG: LightRAG Graph Stack; this page records how each index is built, kept current, and what goes wrong. Facts are as of 2026-10-04.
qmd’s own embedding and rerank models are downloaded by
qmd pull; I have not verified which ones.
When indexing runs
Thepost-commit hook runs after every commit: qmd update and qmd embed synchronously (incremental, seconds), then wiki-index in the background (incremental; progress in .lightrag/last-index.log). wiki-index --full --yes is the manual full rebuild.
LightRAG pipeline, step by step
- Page selection.
wiki-indexlistswiki/**/*.mdand compares each file’s mtime with.lightrag/manifest.json; only new or changed pages go through. - Chunking. A custom chunker prepends a graph-structure header (the page title, type, tags and up to 25 confirmed wikilinks) so the LLM only extracts what the links do not already state (Wikilink Graph Extraction: Reducing LightRAG Indexing Cost), then cuts ~800-token chunks with 64 tokens of overlap.
- Extraction. Each chunk goes to the LLM (default
deepseek-v4.1-flashvia OpenCode Go’s OpenAI-compatible endpoint,thinkingdisabled) to produce entities and relations. LightRAG then runs relation processing and merges duplicates into the graph (the 1.5.7 logs show extraction, relation processing, merging). - Embedding. Entity, relation and chunk texts are embedded locally by ollama
nomic-embed-text(768 dimensions) in small batches. - Persist. Written to
graph_chunk_entity_relation.graphml,vdb_{entities,relationships,chunks}.json(nano-vectordb) and thekv_store_*.jsonfiles. The manifest entry is saved after each page, so an interrupted run resumes.
wiki-chat, wiki-mcp): embed the question, retrieve from the graph and vectors, then have the LLM write the answer. Modes: local (specific entities), global (community-level, cross-concept), hybrid (both, default), naive (flat vector search).
Keeping it correct
- Failed pages are not recorded. LightRAG’s
ainsertlogs pipeline errors instead of raising, sowiki-indexchecks the document status afterwards; a failed page stays out of the manifest, is retried next run, and the run exits 1. - Deleted or moved pages linger. Incremental runs never remove their entities. The 2026-10-04 audit found 67 of 209 manifest entries pointing at missing files; a
--fullrebuild cleared them. Back up.lightrag/first (--fullwipes it before rebuilding). - No backend, no wipe. With no LLM backend configured,
wiki-indexexits before--fullwould wipe the index. - Full rebuilds are heavy. Many LLM calls per page; on OpenCode Go they count against 5-hour, weekly and monthly usage limits, hence
--yes. The 2026-10-04 rebuild of 175 pages took roughly 1.5 hours (estimate from the log).
Gotchas found in practice
- A long
sources:frontmatter line can stall a page. One page with 31 source filenames timed out on chunk 0 every time (3 attempts, 2 models) while the same text indexed in 28 s withsourcesemptied, and each section indexed alone. Fix: keepsourcesshort; the page was split into a hub plus three topical pages (see OWASP Security Checklist). - Reasoning models eat
max_tokens. Withdeepseek-v4.1-flash’s default thinking, extraction prompts spent the whole 4096 budget on reasoning and returned empty content. Sendingthinking: {"type": "disabled"}cut a sample prompt from 2774 to 185 completion tokens.OPENCODE_LIGHTRAG_DISABLE_THINKING=0opts out for models that reject the field. - OpenCode requires an
x-opencode-sessionheader (HTTP 400MissingSessionIDwithout it). The scripts send it per request viaextra_headers, because LightRAG overwrites client-leveldefault_headers. - LightRAG’s API moves. 1.5.7 dropped
max_extract_input_tokens, renamedcosine_thresholdtocosine_better_than_threshold, and movedchunking_by_token_sizetolightrag.chunker. The scripts pinlightrag-hku>=1.5.7,<1.6. - OpenCode has no embeddings. Neither the Go nor the Zen catalog lists an embedding model and
POST /embeddingsis 404 on both, so embeddings stay on local ollama until a self-hosted llama.cpp server takes over. Changing the embedding model or dimension probably needs a full rebuild (training data — verify: LightRAG vector stores are tied to the embedding dimension). install.shcopies from the checkout. It copiestemplates/wiki-*out of the main checkout into~/.local/bin, so pull before installing or you reinstall stale scripts.
Related
- Local Wiki RAG: LightRAG Graph Stack: the two-path architecture, tool usage and design choices
- Wikilink Graph Extraction: Reducing LightRAG Indexing Cost: why the chunker prepends the wikilink header
- qmd: the BM25 + vector engine
- OpenCode Go: the extraction LLM provider and its usage limits
- BM25: lexical retrieval used by qmd
- Linux Machine Setup Guide: installing the indexing tools on a fresh machine