Retrievalcore7 min read

Retrieval-Augmented Generation (RAG)

Fetch relevant context from an external index at query time, then condition the model on it to ground its answer.

Preparing spaced-repetition scheduler…

RAG is a pattern, not a single algorithm. The shape is:

  1. Offline indexing. Documents are split into chunks, each chunk is embedded by an encoder model, and the resulting vectors are stored in a vector index (e.g. HNSW, IVF) along with the original text and metadata.
  1. Online query. The user query is embedded with the same encoder. The index returns the top-k chunks by cosine similarity (or another metric). Optional rerankers refine the ordering.
  1. Generation. The retrieved chunks are stuffed into the model's context window with the user query, and the model produces an answer that is conditioned on the retrieved evidence. Citations are typically included by referring to chunk IDs or document names.

Why RAG works:

  • Freshness: new documents become retrievable after ingestion and indexing complete — no retraining.
  • Attribution: the prompt contains the source, so the model can be instructed to cite it.
  • Cost: far cheaper than continual pretraining or frequent fine-tuning.
  • Privacy control: source documents can stay in your index and are not added to base-model training, but retrieved chunks still reach the inference endpoint unless the model runs inside your own security boundary.

Failure modes:

  • Retrieval miss: the right chunk is not retrieved; the model may answer incorrectly unless it can abstain.
  • Lost in the middle: LLMs attend less faithfully to mid-context information; placing the answer near the start or end of the prompt helps.
  • Stale or noisy chunks: garbage in, garbage out. Quality of the corpus and the chunking strategy matter.

RAG remains a standard production pattern for "chat with your data" features.

RAG pipeline. User query → embedding → vector search → top-k chunks → optional rerank → prompt with citations → LLM answer.

Key takeaways

  • RAG = offline index + online retrieval + grounded generation.
  • Attribution and freshness are the main wins over parametric knowledge.
  • Chunking strategy, embedding choice, and reranking dominate quality.

Self-check

  1. What does the retriever return to the LLM?

  2. What is the 'lost in the middle' problem?

References

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al. / arXiv

    Original RAG research; retrieval and parameter updates can be combined.

  2. Model optimization — OpenAI

    Evaluate prompting, retrieval, and fine-tuning against the task.

  3. Chroma — What is a Vector Database? — Chroma

    Open-source vector database reference.

Last verified 2026-09-04.