ADR-0030: Late-interaction retrieval for semantic search
- Status: Accepted (not yet implemented — tracked in #412)
- Date: 2026-08-26
- Related: 📝 ADR-0002 (vector backend), 📝 ADR-0013 (pure-Go driver), SPEC-0002 (Search)
Context and Problem Statement
Semantic search today stores one pooled float32 vector per message
(embeddings table, schema v3) and ranks by brute-force cosine in Go. That is a
"no-interaction" retriever: everything the model knows about a message is
compressed into a single vector before it is ever compared to a query. On short,
jargon-heavy texts — exactly what message archives are — pooling loses token-level
nuance (names, code-ish fragments, slang), and there is no way to explain why
something matched.
Late-interaction models (ColBERT and descendants) keep one vector per token per document and score at query time with MaxSim (for each query token, take its best-matching document token, then sum). They match bi-encoder speed with cross-encoder-like precision and give token-level explainability for free. Should msgbrowse switch?
Decision Drivers
- Retrieval quality on short texts is the whole point of the Overview/Journal; better recall of "that one phrase" queries is the top user-visible win.
- Local-first ethos (📝 ADR-0004): the embedding path must run against Joe's own hardware; no new always-on service.
- One SQLite file stays non-negotiable (📝 ADR-0002, 📝 ADR-0013).
- The existing LLM/embedding gateway is OpenAI-compatible and returns one vector per input — it cannot produce token-level multi-vectors.
Considered Options
- Re-run embeddings through the current endpoint — not possible: pooled single-vector APIs cannot emit multi-vectors; re-running changes nothing.
- Self-hosted ColBERT-style encoder + multi-vector storage + MaxSim in Go — schema change, new local model dependency, full re-index.
- Add a cross-encoder reranker over existing single-vector results — keeps the pipeline, adds precision at query time only.
- Stay put.
Decision Outcome
Chosen option: (2) a self-hosted late-interaction encoder, because it is the only option that improves recall (not just reranking) and yields token-level explainability, while staying within the local-first, one-file constraints at personal-archive scale. Storage is the classic objection, and the honest numbers are better than the objection assumes. At int8 a token vector costs ~130 B (128 dimensions plus per-vector norm bookkeeping), so a message costs that times its token count — against today's flat 6 KiB pooled 1536-dim float32 vector, paid on every message regardless of length. Break-even is around 47 tokens: on a message archive, where most messages are far shorter, the multi-vector index is smaller than the pooled one it replaces, and only long messages cost more (capped at ~32 KiB by the 256-token limit). Binary 1-bit packing would cut this eightfold to 16 B/token and was considered, but rejected: retrieval quality is the entire point of this change, and storage was never the binding constraint. Cross-encoder reranking (option 3) is explicitly deferred as the fallback if the indexing cost proves annoying — it is a strictly smaller change and composes with either backend.
Consequences
- Good, because MaxSim ranking materially beats pooled-cosine on short, entity-dense messages (the archive's dominant shape).
- Good, because matches become explainable: which query token hit which message token can be surfaced in the UI/MCP later.
- Bad, because embeddings must come from a locally-run model (ColBERTv2 ≈ 110M params: ~250 MB fp16 VRAM or CPU-only at query time); the hosted embedding endpoint is no longer on the semantic-search path.
- Bad, because it is a full-corpus re-embed plus a schema migration, coordinated
via
modelkeying (the coexistence rule from schema v3 already anticipates this). - Neutral: ColPali/ColQwen (image patches over PDFs/screenshots) is out of scope until attachments are indexed; the storage layout chosen here does not block it.
Architecture Diagram
More Information
- ColBERT paper: https://arxiv.org/abs/2004.12832 · ColBERTv2: https://arxiv.org/abs/2112.01488
- Weaviate overview that prompted this: https://weaviate.io/blog/late-interaction-overview
- Formalized in SPEC-0019.
- Accepted 2026-09-04. Implementation is tracked by epic #412: #413 (store
migration + blob codec), #414 (encoder + backfill), #415 (MaxSim + cutover),
#424 (Settings → LLM encoder picker). Note for #414: the current gateway
serves
bge-m3only through the OpenAI-compatible/v1/embeddingsroute, which returns pooled vectors — the encoder endpoint has to be a native late-interaction API, not that route.