Skip to content

Retrieval & RAG

  • RAG = retrieve, then generate. Retrieval quality is the ceiling on answer quality — a model can’t ground on what you didn’t fetch. Most “the LLM is wrong” bugs are retrieval bugs.
  • Hybrid beats either half. Dense (semantic) retrieval and lexical (BM25) retrieval fail in opposite ways; fuse them — Reciprocal Rank Fusion is the simple, strong default — and add graph retrieval when relationships matter.
  • Chunking is the highest-leverage knob. How you split documents (size, overlap, recursive vs semantic boundaries) decides what can be retrieved. Get this wrong and no reranker saves you.
  • Encoders map meaning to geometry. A bi-encoder embeds query and document separately (fast, indexable); a cross-encoder scores a pair jointly (accurate, slow) — so you retrieve with the first and rerank with the second.
  • Metadata is a first-class retrieval signal. Filtering by source, time, type, and permissions before/with vector search is what makes retrieval correct and secure in production (ties to Authorization).
  • Multimodal retrieval reduces to text + a unifying record. Normalize heterogeneous media (image, table, audio, PDF) into a common searchable representation so one index serves them all.
  • Build the minimal RAG loop: chunk → embed → store (pgvector) → retrieve top-k → stuff context → generate. Then measure how answer quality moves as you change only the retriever.
  • Implement BM25 from the formula and compare it to dense retrieval on the same queries — find the queries where each wins (exact terms/IDs vs paraphrases).
  • Add a cross-encoder reranker over the fused candidates; measure NDCG/MRR before and after.
  • Re-chunk the same corpus three ways (fixed, recursive, semantic) and watch retrieval recall change.

A language model knows what was in its training data, frozen and unattributed. RAG turns it into an open-book system: at query time you retrieve the relevant evidence and put it in the context window, so the model reasons over fresh, private, citable facts instead of its parametric memory. That reframes most of “AI engineering” as an information retrieval problem wearing a neural coat — and IR has 50 years of patterns. The pipeline is always the same shape (ingest → chunk → encode → index → retrieve → rerank → assemble → generate → evaluate); the skill is knowing which knob to turn, and retrieval — not the LLM — is almost always the bottleneck.

ingest → chunk → encode (embed) → index ──┐
├─ retrieve → rerank → assemble context → generate → cite
query ──────────── encode (embed) ────────┘

Every production RAG system is a specialization of this. The rest of this topic is the patterns at each stage.

Mapping text (and other media) to vectors

Key ideas:

  • Transformer encoder: a bidirectional stack (BERT-family) reads the whole input at once (no causal mask) and emits a contextual vector per token. Pooling (CLS token or mean over tokens) collapses them into one fixed-length embedding. (Foundations in Neural Architectures and Deep Learning.)
  • The geometry: training arranges the space so semantic similarity ≈ cosine similarity. Retrieval is then “nearest neighbors in this space.”
  • Bi-encoder vs cross-encoder — the central distinction of the topic:
    • Bi-encoder: encode query and document separately into vectors; compare by dot product. Documents can be embedded once and indexed → fast, scalable retrieval. Slightly less accurate (no query-document interaction).
    • Cross-encoder: feed query+document together through the model, output one relevance score. Sees the interaction → more accurate, but must run per candidate pair at query time → only feasible for reranking a short candidate list.
    • The production pattern: bi-encoder to retrieve hundreds, cross-encoder to rerank to the top few.
  • Matryoshka embeddings: train so a truncated prefix of the vector is still usable — store short vectors cheaply, expand when needed.

Splitting documents into retrievable units — the highest-leverage decision

StrategyHowWhen
Fixed-sizeN tokens with overlapBaseline; simple, ignores structure
Recursive / structuralSplit on a hierarchy of separators (¶ → sentence → word) to respect boundaries, with overlapThe strong default (saige’s recursive splitter)
SemanticSplit where embedding similarity between adjacent sentences dropsCoherent topical chunks; more compute
Document-awareSplit on headings/sections/code blocksStructured docs, markdown, code

Key ideas:

  • The tension: chunks too large dilute the embedding and waste context; too small lose the context needed to be meaningful. There’s a sweet spot per corpus (often a few hundred tokens).
  • Overlap carries context across boundaries so a fact split between chunks is still findable.
  • Parent/child (small-to-big): retrieve on small precise chunks, but feed the parent (larger surrounding) chunk to the model — precision in search, context in generation.
  • Carry metadata onto every chunk (source, section, timestamp, permissions) — you’ll filter and cite on it.

4. Indexing & Approximate Nearest Neighbor (ANN)

Section titled “4. Indexing & Approximate Nearest Neighbor (ANN)”

Key ideas:

  • Exact nearest-neighbor over millions of vectors is too slow; ANN trades a little recall for orders-of-magnitude speed.
  • HNSW (hierarchical navigable small-world graphs): the dominant ANN index — a multi-layer proximity graph you greedily descend. Great recall/latency; tunable via M and efSearch.
  • IVF / IVF-PQ: cluster vectors, search only nearby clusters; product quantization compresses vectors for memory.
  • Where it lives: pgvector (Postgres), Qdrant, Milvus, Weaviate, FAISS, Redis — see Distributed Data & Caching. The index is sharded, replicated, and cached like any other data.

5. Lexical Retrieval (BM25 & Bag-of-Words)

Section titled “5. Lexical Retrieval (BM25 & Bag-of-Words)”

The other half of hybrid — don’t skip it

Key ideas:

  • Bag-of-words: represent text as term counts, ignoring order. Sparse, high-dimensional, exact-term.
  • TF-IDF: weight a term by its frequency in the doc (TF) × its rarity across the corpus (IDF) — common words count less, distinctive words more.
  • BM25: the refined, dominant lexical ranker — TF-IDF with saturation (k1: extra occurrences matter less and less) and length normalization (b: don’t reward long docs for having more words). Decades old, still a brutally strong baseline.
  • Why keep lexical at all: dense retrieval misses exact matches — IDs, error codes, names, rare jargon, acronyms — that BM25 nails. They fail in opposite directions, which is exactly why you fuse them.

Combining dense + lexical (+ graph) into one ranked list

Key ideas:

  • Reciprocal Rank Fusion (RRF): combine multiple ranked lists by summing 1/(k + rank) across them. No score calibration needed, robust, embarrassingly simple — the default fusion (and what saige uses to fuse vector + BM25 + graph retrievers).
  • Score-based fusion: normalize and weight each retriever’s scores — more tunable, more fragile.
  • Graph retrieval / GraphRAG: pull a connected subgraph of facts from a knowledge graph, resolving entities to documents — adds multi-hop reasoning and relationships that flat chunks can’t express. saige fuses this as a third retriever and tracks temporal validity (ValidAt/InvalidAt) on relations so retrieval is point-in-time correct.
  • The pattern: retrieve broadly from several complementary sources, fuse, then rerank — breadth first, precision last.

What separates a demo from production

Key ideas:

  • Metadata filtering: constrain retrieval by structured fields — source, type, recency, language, tenant, and access permissions — combined with vector/lexical search (pre-filter or post-filter). Often the difference between a right and wrong answer.
  • Authorization-aware retrieval: a user must only retrieve chunks they’re allowed to see. Filter by permission at retrieval time (carry ACLs/labels on each chunk) — see Authorization & Access Control. This is “secure RAG,” and it’s a hard, mandatory requirement in the enterprise.
  • Context assembly: dedupe, order, and fit candidates into the context budget; add citations (which chunk supported which claim) and optionally compress context with an LLM (saige does both).
  • Freshness & invalidation: re-embed and re-index on document change; pin embedding-model versions (changing the encoder invalidates the whole index — a caching problem in disguise).
  • Query transforms: rewriting, expansion, HyDE (embed a hypothetical answer), and multi-query — improve recall before retrieval even runs.

One index over heterogeneous media

Key ideas:

  • The unifying pattern: model content as a hierarchy — Document → Sections → ContentVariants (saige’s data model) — where every variant, whatever its medium (image, table, audio, PDF), carries a .Text representation (a caption, transcript, OCR, or extracted text). Now one text/vector index searches across all of them uniformly.
  • Joint embedding spaces: models like CLIP embed images and text into the same space, so a text query can retrieve images directly (true multimodal retrieval) — see Foundation Models.
  • Ingestion plumbing: URI resolution (file://, s3://), content negotiation, and per-type extractors (OCR, ASR, table parsing) turn raw files into searchable variants.
  • Generation: feed retrieved text variants (and, for VLMs, the original media) into a multimodal model.

Retrieval has its own metrics, separate from generation (full treatment in LLM Evaluation):

  • Recall@k / Precision@k: did the relevant chunk make the top-k?
  • MRR (mean reciprocal rank): how high was the first relevant hit?
  • NDCG: rank-quality with graded relevance and position discounting.
  • Faithfulness / context-precision / context-recall (RAG-specific, à la Ragas / saige’s RAG scorers): is the answer grounded in the retrieved context, and was the right context retrieved?

NeedReach for
Semantic / paraphrase matchingDense bi-encoder + HNSW
Exact terms, IDs, rare jargonBM25 (lexical)
Best general retrievalHybrid (dense + BM25) fused with RRF
Multi-hop / relationship questions+ Graph retrieval (GraphRAG)
Squeeze accuracy from candidatesCross-encoder reranker
Split documents wellRecursive/structural chunking + overlap
Precise search, rich contextParent/child (small-to-big) chunks
Secure, tenant-correct resultsMetadata + permission filtering
Mixed mediaDocument→Section→Variant with a .Text per variant
  • Retrieval is the ceiling. Fix retrieval before touching the prompt or the model.
  • Fuse complementary retrievers, then rerank — breadth (dense + lexical + graph via RRF) first, precision (cross-encoder) last.
  • Bi-encode to find, cross-encode to rank — separate the fast indexable step from the accurate scoring step.
  • Chunk for what’s retrievable; carry metadata on every chunk — including permissions, timestamps, and source for filtering and citation.
  • Normalize every medium to a common searchable record — heterogeneous data, one index.
  • Changing the encoder invalidates the index — treat embeddings like a cache with a version key.
ConceptConnected TrackApplication
Encoders, embeddings, fine-tuning, CLIPTraining & Frameworks / Foundation ModelsThe models that embed
Vector stores, ANN, knowledge graphs, cachingDistributed Data & CachingWhere the index lives
Permission-filtered retrievalAuthorization & Access ControlSecure / multi-tenant RAG
Retrieval & faithfulness metricsLLM EvaluationMeasuring the pipeline
Token streaming of the answerStreaming & SSEDelivering the generation
CompanyThe pattern they lean onInstance
MicrosoftGraph + vector hybridGraphRAG over knowledge graphs
Perplexity / You.comHybrid retrieval + reranking + citationsAnswer engines
Elastic / VespaLexical + dense in one engineBM25 + ANN, fusion built in
Cohere / VoyageBi-encoder embeddings + cross-encoder rerankersRetrieval-as-a-service
Glean / enterprise searchPermission-filtered, metadata-rich RAGSecure multi-tenant retrieval
saige (reference)RRF over vector + BM25 + graph, multimodal variantsGo SDK for agents/KG/RAG
#ModuleChapterKindPass
1ag.06RAG ingest: crawler with SSRF allowlist, chunker, embedder, storebuild10
2ag.07Retrieval: BM25, flat and IVF vectors, RRF, MMRbuild10
3ag.08Context assembly, citations, search_docs toolbuild10