RAG or fine-tuning: which one does your product need?

Wooden library card catalogue drawers

RAG vs fine-tuning explained by engineers: when each fits, the data each needs, freshness, cost drivers, hybrid setups and a step-by-step way to decide.

Bhaskar Bhatt7 min read

The short answer to RAG vs fine-tuning: use retrieval-augmented generation (RAG) when the model needs to know things, such as your documents, prices or policies, and fine-tuning when it needs to behave differently, such as following a strict format, tone or classification scheme. Most products need RAG first. Fine-tuning is a later optimisation for specific, measurable problems.

On this page
  1. What is the difference between RAG and fine-tuning?
  2. When is RAG the better choice?
  3. When does fine-tuning make more sense?
  4. How do the two compare on data, freshness and cost?
  5. Can you combine RAG and fine-tuning?
  6. How do you decide, step by step?
  7. When isn’t either approach the right one?
  8. Frequently asked questions

Teams often frame this as a choice between two competing technologies. It isn’t really. The two methods change different things about a system, and many production setups use both. The useful question is which problem you have right now, and what evidence you’d need to justify the more expensive option.

What is the difference between RAG and fine-tuning?

RAG leaves the model unchanged and gives it relevant information at query time, by searching your content and putting the results into the prompt. Fine-tuning changes the model itself by training it further on examples of the inputs and outputs you want. RAG changes what the model can see. Fine-tuning changes how it responds.

The term RAG comes from a 2020 paper by Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, which combined a language model with a searchable index of documents. Today a RAG system is usually a pipeline: split documents into chunks, store them in a search index (often vector search, sometimes combined with keyword search), retrieve the best matches for each question, and ask the model to answer using only those passages.

Fine-tuning ranges from full retraining, which is rarely sensible outside large labs, to parameter-efficient methods such as LoRA (Hu et al., 2021), which train a small set of added weights while the base model stays frozen. Hosted model providers also offer managed fine-tuning on some of their models. Either way, you’re teaching the model patterns from examples, not giving it a library to look things up in.

When is RAG the better choice?

RAG is the better choice when answers depend on facts that change, are specific to your organisation, or need to be traceable to a source. Support assistants, internal knowledge search, policy questions and document Q&A all fit. You update the content, not the model, and every answer can cite where it came from.

That last point carries more weight than it first appears to. When a user or an auditor asks “where did this answer come from?”, a RAG system can show the passage. A fine-tuned model can’t point to the training example that shaped an answer. For regulated industries, or anywhere a wrong answer has consequences, that traceability is often the deciding factor.

RAG also handles access control more naturally. You can filter retrieval by the user’s permissions, so a junior employee’s question never pulls in board papers. Knowledge baked into model weights through fine-tuning has no such filter; whoever can query the model can, in principle, get at whatever it learned.

When does fine-tuning make more sense?

Fine-tuning makes sense when the problem is behaviour, not knowledge. Examples include producing a fixed output structure every time, matching a house style, classifying text into your own categories, or getting a smaller, cheaper model to do one narrow task as well as a larger one. You need hundreds or more good examples of the target behaviour.

Before fine-tuning, try harder with prompting. Clear instructions, a few worked examples in the prompt and structured output features solve a lot of formatting and tone problems. If you’ve done that, measured the result on a proper evaluation set, and the model still misses in a consistent way, fine-tuning is a reasonable next step.

The other honest case for fine-tuning is cost and speed at volume. A small fine-tuned model running a narrow task can be faster and cheaper per request than a large general model with a long prompt. That only pays off once the task is stable and the traffic is real. Fine-tuning a model for a product that hasn’t found its shape yet means retraining every time the requirements move.

How do the two compare on data, freshness and cost?

RAG needs well-organised source content and good search; fine-tuning needs a large set of high-quality labelled examples. RAG updates as soon as you re-index a document, while fine-tuned behaviour is fixed until you retrain. Cost drivers differ too: RAG adds tokens and retrieval work to every request, fine-tuning adds training and maintenance work upfront.

FactorRAGFine-tuning
ChangesWhat the model sees at query timeHow the model behaves
Data you needClean, current source documents with useful metadataHundreds to thousands of reviewed input/output examples
FreshnessUpdates when content is re-indexedFixed until the next training run
TraceabilityCan cite the source passageNo direct link from answer to source
Access controlFilter retrieval by user permissionsHard to restrict what the model learned
Main cost driversLonger prompts, embedding and search infrastructure, keeping the index in syncPreparing training data, training runs, hosting a custom model, retraining when the base model changes
Typical failureRetrieves the wrong passages, so the answer is confidently wrongLearns the wrong pattern, or forgets general ability on other tasks

Note the data row. People underestimate how much work it is to get good training examples for fine-tuning. Labelling, review and agreement between reviewers are real effort; our post on data labeling for AI covers what that involves. RAG has its own data work, but it’s mostly cleaning and structuring documents you already have, which is often the job of a data engineering team rather than an ML team.

Can you combine RAG and fine-tuning?

Yes, and mature systems often do. A common pattern fine-tunes a model to follow the output format, cite sources properly or stay within policy, then uses RAG to supply the facts. Another fine-tunes only a supporting component, such as the embedding model or a reranker, so retrieval works better on your domain’s vocabulary.

The order matters. We build RAG first, with an evaluation set, because it gives a working product sooner and shows where the real failures are. Sometimes those failures are retrieval problems: the right document exists but isn’t found. That’s fixed with better chunking, metadata, hybrid keyword and vector search, or a reranker, not with fine-tuning. Sometimes the right passages are retrieved and the model still misuses them. That’s where a fine-tuned generator, or a stronger base model, earns its place.

Don’t overlook the simplest option. Context windows are now large enough that small, stable knowledge bases can sometimes go straight into the prompt with no retrieval at all. That has limits; research by Liu et al. found models use information in the middle of long contexts less reliably than information at the start or end. Test it on your own data before relying on it.

How do you decide, step by step?

Start by naming the failure you’re trying to fix, then try the cheapest fix that addresses it, and measure. Most teams can decide in a few weeks by building a basic RAG prototype, testing it against real questions, and only then looking at fine-tuning for whatever consistent problems remain.

  1. Write the job down. One sentence on what the feature does, and a list of what a wrong answer looks like.
  2. Build an evaluation set. 50 to 200 real questions with expected answers, reviewed by someone who knows the domain. Our post on testing an LLM feature before it ships explains how.
  3. Try prompting alone. A good base model with clear instructions sets the baseline.
  4. Add retrieval if answers need your facts. Measure retrieval quality separately from answer quality.
  5. Fix retrieval before anything else. Chunking, metadata, hybrid search and reranking usually help more than a model change.
  6. Consider fine-tuning for what’s left. Only if failures are consistent, behavioural and you can collect enough good examples.
  7. Re-run the same evaluation. Keep the change only if it beats the baseline on the metrics you agreed.

When isn’t either approach the right one?

When the task doesn’t need a language model’s judgement at all. If the answer is a lookup in a database, a rules engine or a plain search box may serve users better, faster and more cheaply. Neither RAG nor fine-tuning fixes source content that is wrong, contradictory or out of date.

We’d also hold off when nobody can say what a correct answer looks like. Without that, you can’t build an evaluation set, and without an evaluation set you’ll end up choosing between RAG and fine-tuning based on a demo. Get the source content and the success criteria in order first. Our GenAI solutions work usually starts there, and when a project does need custom training, our AI and ML engineering team handles the model side.

Frequently asked questions

Does fine-tuning stop a model from making things up?

Not reliably. Fine-tuning can make a model sound more confident in your domain without making it more accurate. Grounding answers in retrieved sources, and checking that each claim is supported by them, does more to reduce made-up answers than training alone.

How much data do you need to fine-tune a model?

It depends on the task and the method. Narrow tasks like classification or fixed formatting can work with a few hundred good examples; broader behaviour changes need more. Quality and consistency of the examples matter more than the count.

Is RAG secure enough for confidential documents?

It can be, if retrieval respects the same permissions as the source systems and the index is stored and hosted under your data rules. The common mistake is indexing everything into one store with no per-user filtering, which lets anyone with access to the assistant read anything in it.

Do we need to retrain when the model provider releases a new version?

For fine-tuning, usually yes. A fine-tuned model is tied to its base model, so moving to a newer base means preparing and running training again. A RAG system can typically switch models with a configuration change and a re-run of the evaluation set.

Scroll to Top