Skip to main content

Wiki Indexing Pipeline: qmd and LightRAG

This repo is indexed two ways, for two different jobs. qmd finds pages; LightRAG connects them. Architecture and tool usage are in Local Wiki RAG: LightRAG Graph Stack; this page records how each index is built, kept current, and what goes wrong. Facts are as of 2026-10-04. qmd’s own embedding and rerank models are downloaded by qmd pull; I have not verified which ones.

When indexing runs

The post-commit hook runs after every commit: qmd update and qmd embed synchronously (incremental, seconds), then wiki-index in the background (incremental; progress in .lightrag/last-index.log). wiki-index --full --yes is the manual full rebuild.

LightRAG pipeline, step by step

  1. Page selection. wiki-index lists wiki/**/*.md and compares each file’s mtime with .lightrag/manifest.json; only new or changed pages go through.
  2. Chunking. A custom chunker prepends a graph-structure header (the page title, type, tags and up to 25 confirmed wikilinks) so the LLM only extracts what the links do not already state (Wikilink Graph Extraction: Reducing LightRAG Indexing Cost), then cuts ~800-token chunks with 64 tokens of overlap.
  3. Extraction. Each chunk goes to the LLM (default deepseek-v4.1-flash via OpenCode Go’s OpenAI-compatible endpoint, thinking disabled) to produce entities and relations. LightRAG then runs relation processing and merges duplicates into the graph (the 1.5.7 logs show extraction, relation processing, merging).
  4. Embedding. Entity, relation and chunk texts are embedded locally by ollama nomic-embed-text (768 dimensions) in small batches.
  5. Persist. Written to graph_chunk_entity_relation.graphml, vdb_{entities,relationships,chunks}.json (nano-vectordb) and the kv_store_*.json files. The manifest entry is saved after each page, so an interrupted run resumes.
Query (wiki-chat, wiki-mcp): embed the question, retrieve from the graph and vectors, then have the LLM write the answer. Modes: local (specific entities), global (community-level, cross-concept), hybrid (both, default), naive (flat vector search).

Keeping it correct

  • Failed pages are not recorded. LightRAG’s ainsert logs pipeline errors instead of raising, so wiki-index checks the document status afterwards; a failed page stays out of the manifest, is retried next run, and the run exits 1.
  • Deleted or moved pages linger. Incremental runs never remove their entities. The 2026-10-04 audit found 67 of 209 manifest entries pointing at missing files; a --full rebuild cleared them. Back up .lightrag/ first (--full wipes it before rebuilding).
  • No backend, no wipe. With no LLM backend configured, wiki-index exits before --full would wipe the index.
  • Full rebuilds are heavy. Many LLM calls per page; on OpenCode Go they count against 5-hour, weekly and monthly usage limits, hence --yes. The 2026-10-04 rebuild of 175 pages took roughly 1.5 hours (estimate from the log).

Gotchas found in practice

  • A long sources: frontmatter line can stall a page. One page with 31 source filenames timed out on chunk 0 every time (3 attempts, 2 models) while the same text indexed in 28 s with sources emptied, and each section indexed alone. Fix: keep sources short; the page was split into a hub plus three topical pages (see OWASP Security Checklist).
  • Reasoning models eat max_tokens. With deepseek-v4.1-flash’s default thinking, extraction prompts spent the whole 4096 budget on reasoning and returned empty content. Sending thinking: {"type": "disabled"} cut a sample prompt from 2774 to 185 completion tokens. OPENCODE_LIGHTRAG_DISABLE_THINKING=0 opts out for models that reject the field.
  • OpenCode requires an x-opencode-session header (HTTP 400 MissingSessionID without it). The scripts send it per request via extra_headers, because LightRAG overwrites client-level default_headers.
  • LightRAG’s API moves. 1.5.7 dropped max_extract_input_tokens, renamed cosine_threshold to cosine_better_than_threshold, and moved chunking_by_token_size to lightrag.chunker. The scripts pin lightrag-hku>=1.5.7,<1.6.
  • OpenCode has no embeddings. Neither the Go nor the Zen catalog lists an embedding model and POST /embeddings is 404 on both, so embeddings stay on local ollama until a self-hosted llama.cpp server takes over. Changing the embedding model or dimension probably needs a full rebuild (training data — verify: LightRAG vector stores are tied to the embedding dimension).
  • install.sh copies from the checkout. It copies templates/wiki-* out of the main checkout into ~/.local/bin, so pull before installing or you reinstall stale scripts.