Blog
4 min read

Fine-Tuning vs RAG: How to Teach an AI About Your Data

You want an AI model to know about your products, documents or customers. RAG looks things up and adds them to the prompt; fine-tuning retrains the model. What each actually does, what each is good at, costs and pitfalls, and why most apps should start with neither.

A general AI model doesn't know your refund policy, your product catalogue, or last week's support tickets. There are two famous ways to fix that — RAG and fine-tuning — and a third, simpler way most people should try first.

Option 0: Just put it in the prompt

Modern models have large context windows. If your knowledge fits in a few thousand words — a policy document, a product list, an FAQ — paste it into the system prompt. With prompt caching, repeating it on every request is cheap.

This is the right answer surprisingly often. Try it before building anything.

RAG: look it up, then answer

RAG (retrieval-augmented generation) is for when your knowledge is too big for the prompt — thousands of documents, a help centre, a database.

  1. Split your documents into chunks and store them in a searchable index (often with embeddings in a vector database like pgvector).
  2. When a question comes in, search for the most relevant chunks.
  3. Add those chunks to the prompt and ask the model to answer using them.

It's an open-book exam: the model doesn't memorise your data, it reads the relevant pages each time. (What is RAG?)

Good at: facts, documents, anything that changes, answers with citations, data that differs per user.

Fine-tuning: change the model itself

Fine-tuning means training a model further on your own examples — usually hundreds or thousands of input/output pairs — so its behaviour changes permanently.

It's teaching, not looking things up. The model learns patterns: a style, a format, a classification, a specialised way of responding.

Good at: consistent tone or format, narrow repetitive tasks (classify this ticket, extract these fields), getting a smaller, cheaper model to do one job as well as a bigger one.

Bad at: teaching facts. Fine-tuned models still make things up about details, and you can't easily update or remove what they "learned." Retraining every time your docs change isn't practical.

Side by side

Prompt RAG Fine-tuning
Best for Small, stable knowledge Large or changing knowledge Behaviour, style, format
Updating knowledge Edit the prompt Re-index documents Retrain
Shows sources Can Yes, naturally No
Per-user data Yes Yes No
Setup effort Minutes Days Days to weeks, plus data
Hallucination on facts Low if info is present Low if retrieval works Still happens
Cost profile Tokens per request Tokens + search infrastructure Training cost + often cheaper inference

Common mistakes

  • Fine-tuning to teach facts. The most common misconception. Use RAG.
  • RAG with bad retrieval. If search returns the wrong chunks, the model answers confidently from the wrong text. Most RAG quality problems are search problems — test what's retrieved, not just the final answer.
  • Building RAG for ten pages of docs. Just put them in the prompt.
  • Skipping evaluation. Whatever you choose, keep a set of real questions with known good answers and check them after every change.

Can you combine them?

Yes. A fine-tuned model can answer in your house style while RAG supplies the facts. But it's rare for a small app to need both. Start simple and add complexity when tests show you need it.

The decision, simply

  1. Does the knowledge fit in the prompt? → Put it there.
  2. Too big or changes often? → RAG.
  3. The model knows enough but behaves wrong (format, tone, narrow task) and prompting doesn't fix it? → Consider fine-tuning.

The summary

  • Try the prompt first; it's often enough.
  • RAG looks things up at question time — best for facts and changing data.
  • Fine-tuning changes behaviour — best for style, format and narrow tasks.
  • Don't fine-tune to teach facts.

EasySpawn servers come with PostgreSQL ready for pgvector, so your app, its documents and its RAG index can live in one database on one server. See how it works or join the waitlist.

Related: What Is RAG? · What Are Embeddings? · pgvector Tutorial · What Is a Context Window?

Keep reading