The RAG vs fine-tuning decision consumes more meeting time than almost any
other choice in an AI build, usually because the two options are compared as if
they solve the same problem. They do not.
Retrieval changes what the model knows right now. Fine-tuning changes how the
model behaves. Once you frame it that way, most cases answer themselves.
What each one actually fixes
| Symptom | Retrieval fixes it | Fine-tuning fixes it |
|---|
| Answers are out of date | Yes | No |
| Cannot cite a source | Yes | No |
| Ignores our house format | Partly, via prompt | Yes |
| Tone is wrong for our brand | Partly, via prompt | Yes |
| Too expensive per call at scale | No | Often |
| Latency too high | Sometimes, via smaller context | Often |
| Hallucinates about our product | Yes | No |
| Needs to follow a rigid decision protocol | Partly | Yes |
Two lines in that table do most of the work. If the answer depends on data that
changes — prices, policies, tickets, inventory, documentation — you need
retrieval, because a fine-tuned model's knowledge is frozen at training time and
cannot be corrected without retraining. And if you need to show where an answer
came from, only retrieval gives you that; it is a hard requirement in regulated
contexts and a trust requirement in most others.
Published decision frameworks land in the same place: retrieval is the correct
first choice for the large majority of enterprise LLM applications, with
fine-tuning reserved for behaviour and for cost distillation. Winder's
2026 decision framework
is a reasonable summary of the current consensus.
Start with the cheapest layer that could work
There is a ladder here, and teams routinely skip to the top of it:
- A better prompt. Explicit format, worked examples, stated constraints.
Hours of work. Fixes more than people expect.
- Better context. Getting the right passages in front of the model — which
is a retrieval-quality problem, not a model problem.
- Retrieval architecture. Chunking, hybrid search, reranking, metadata
filters.
- Fine-tuning. A training pipeline, a labelled dataset, an evaluation
suite, and a versioned artefact to maintain.
Each rung costs roughly an order of magnitude more than the one below it in
engineering time and ongoing maintenance. A surprising share of "we need to
fine-tune" conclusions are really "our retrieval returns the wrong chunks" — and
you cannot tell which without measurement.
Where retrieval actually goes wrong
When RAG underperforms, the model is rarely the cause. The usual suspects:
- Chunking that splits meaning. A clause separated from its condition
retrieves as a confident half-truth.
- Pure vector search on keyword-shaped queries. Product codes, error codes
and names need lexical matching; hybrid search plus a reranker fixes more
quality complaints than any model upgrade.
- A corpus with contradictions. Three refund policies from different years,
all indexed. The model cannot know which is current — nothing can.
- No metadata filtering. Retrieving across all tenants, regions or product
lines when the query concerns one.
Fix those before considering a training run. They are cheaper, they are testable,
and they compound.
When fine-tuning genuinely earns its place
Three cases, reliably:
Output structure that must hold every time. If every response must be valid
JSON in a fixed schema, or follow a strict clinical or legal format, fine-tuning
holds it more reliably than instructions do — and it lets you drop the long
formatting prompt from every call.
Cost and latency at volume. This is the strongest quantitative case.
A smaller fine-tuned model with a short prompt can be dramatically cheaper per
call than a frontier model with a long system prompt and retrieved context. The
arithmetic is simple: if monthly call volume times per-call saving exceeds the
one-off project cost within a horizon you care about, fine-tune.
A decision protocol with many implicit rules. When the correct behaviour is
"what our best analyst would do" and the rules resist being written down,
examples teach it better than instructions.
Note what is not on that list: adding knowledge. That is retrieval's job.
The combination most production systems land on
Mature systems typically use both, with a clean division:
- Fine-tuning owns behaviour — format, tone, decision protocol, refusal
boundaries.
- Retrieval owns facts — current, attributable, updatable without retraining.
That split also keeps you portable. Behaviour lives in an artefact you control;
facts live in a store you control; the base model becomes a component you can
swap when something better ships. Committing your product's knowledge to a
fine-tuned model does the opposite — it welds you to a snapshot.
The decision in five questions
- Does the answer depend on data that changes? Yes → retrieval, always.
- Do you need citations? Yes → retrieval.
- Have you fixed prompting and retrieval quality, with measurements? No →
do that first.
- Is per-call cost or latency the binding constraint at real volume? Yes →
evaluate fine-tuning a smaller model.
- Is the required behaviour hard to state but easy to demonstrate? Yes →
fine-tuning, on top of retrieval.
If you cannot answer question 3 with numbers, that is the actual next task. An
evaluation set is a prerequisite for this decision, not a follow-up to it —
the same discipline that keeps
agents from failing in production,
and a meaningful share of the effort in
what an AI MVP actually costs.
When you want the architecture designed around the constraint that actually
binds, that is what we do on AI SaaS builds.