Fine-tuning vs RAG
Two complementary ways to specialize a model: bake knowledge into weights, or fetch it from an external index at query time.
Preparing spaced-repetition scheduler…
When you need a model to know something new, you have two fundamentally different levers:
Fine-tuning updates the model's weights using additional supervised or preference data. Training can change behavior and learned associations, but it does not reliably install a factual database. Serving cost depends on the model, adapters, and provider.
- ✅ Best for: style, tone, format, tool-use patterns, reasoning style.
- ❌ Worst for: fresh facts — every update requires another training run, and old knowledge can be forgotten (catastrophic forgetting).
- Cost: data labeling + GPU hours.
Retrieval-Augmented Generation (RAG) keeps the base model unchanged and instead provides external context at query time. A retriever finds relevant chunks from an index, and the model conditions on them.
- ✅ Best for: fresh, citable, fast-changing facts — docs, policies, internal wikis.
- ❌ Worst for: teaching a new reasoning style — the model already knows how to use the context, but RAG alone does not change how it reasons.
- Cost: an index and per-query retrieval latency.
Use them together. Fine-tune for behavior, RAG for knowledge. Whether both are needed depends on evaluation results and the application.
Key takeaways
- Fine-tuning changes weights; RAG adds context at inference time.
- Fine-tune for style, RAG for facts.
- Update latency depends on ingestion, indexing, training, and deployment choices.
Self-check
Which lever is best for teaching a model a new writing style?
Why is RAG preferable for fresh facts?
References
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Original RAG research; retrieval and parameter updates can be combined.
- Model optimization
Evaluate prompting, retrieval, and fine-tuning against the task.
Last verified 2026-08-19.