RAG or fine-tuning: choosing for a corporate LLM
Two approaches to deploying an LLM on corporate data. When RAG wins on cost and accuracy, when fine-tuning is justified, and why we pick RAG in four out of five B2B projects.
Short answer: corporate tasks almost always need RAG, not fine-tuning. RAG is cheaper to deploy and maintain, it copes with documents that change, and it gives answers with a link to the source. Fine-tuning solves a different problem — the style and format of answers, not knowledge of facts — and it is needed considerably less often than people ask about it.
Every second consultation about an AI assistant arrives with the same opening line:
“We want to train our own LLM on our documents.”
It sounds good, but what stands behind it is nearly always RAG. Here is how the two differ and how to pick.
What a local RAG system is
A local RAG system is three parts deployed entirely inside the company perimeter: a store of indexed documents (a vector database), a model that searches by meaning, and a language model that assembles an answer from the retrieved fragments and cites the source. “Local” means exactly one thing: neither the documents nor the employees’ questions leave your infrastructure. No request goes to an outside service.
It differs from a hosted assistant not in answer quality but in where the data sits and who is responsible for it. It differs from ordinary enterprise search in that it returns a formulated answer with citations rather than a list of documents.
The minimum working stack: a vector database (Qdrant), an embedding model, a local LLM (Mistral, Qwen and similar) served through vLLM, and a layer that splits documents into fragments and assembles context. The GPU is needed for generation; retrieval often manages without one.
What fine-tuning does
Fine-tuning is additional training of a base model on your data. The weights change, and the model starts to “know” your domain, your phrasing and your terminology.
What it takes for fine-tuning to work:
- Somewhere between 5,000 and 100,000 question-answer pairs, labelled by hand or semi-automatically.
- GPU infrastructure — for 7B–13B models, high-end accelerators for several hours at minimum.
- Retraining whenever the data changes. A new batch of documents means a new fine-tuning cycle.
Fine-tuning works well when you need to generate repetitive documents (legal opinions, insurance assessments), or adapt the model to a specific language register (medical terminology, industrial jargon).
What RAG does
RAG is a two-step pipeline:
- Retrieval — the user’s question is turned into a vector and relevant context is found in your document store.
- Generation — a base LLM answers the question using the retrieved context.
The model itself does not change. Your data stays in the vector database.
The key difference is where your data ends up, and whether you can check where an answer came from.
The main difference is where your data ends up, and whether you can check where an answer came from.
Comparison on real parameters
| Criterion | RAG | Fine-tuning |
|---|---|---|
| Initial cost | Lower, by a factor of three to five | Higher |
| Updating data | Reindexing, hours | A new training cycle, days |
| Verifiable answers | Citations from sources | The model simply “knows” |
| GPU infrastructure | Optional | Mandatory |
| Hallucinations | Controlled through context | Possible with no warning |
| Data residency | Data stays in your store | Data absorbed into the weights |
When we choose RAG (four times out of five)
- The knowledge base changes more often than once a quarter.
- Answers need to cite their sources — legal teams, support, sales.
- You would rather not spend the budget on GPUs and MLOps.
- The client is not prepared to wait two weeks for each model update.
A typical example: an AI assistant for B2B sales with twelve thousand items in the catalogue, where managers get answers in twelve minutes instead of two hours. RAG on Qdrant plus a local model, with the catalogue reindexed automatically once a week.
When we choose fine-tuning
- Mass generation of repetitive documents: standard contracts, invoices, notices.
- Stylistic adaptation: legal language, industrial jargon, a formal register with no exceptions.
- The data is stable: medical protocols, safety regulations, standards.
The hybrid
Sometimes we use both: fine-tuning adapts the base model to tone and format, RAG pulls in the current data. This shows up in projects with strict output-format requirements and rapidly changing content — automatically generated commercial proposals, for instance.
The practical test
One question that removes most of the doubt:
“How often does the data the model must rely on change?”
- More often than monthly → RAG.
- Less often than quarterly, and you have 5,000 or more labelled examples → fine-tuning may pay back.
- Not sure → start with RAG and add fine-tuning later if you need it.
The mistake in most projects that arrive with fine-tuning already written into the specification: the client sees “train our own model” as the only route to quality. In practice RAG wins on deployment cost by a factor of three to five, and on answer quality it is comparable or better for nine tasks out of ten, because hallucinations are held in check by citing sources.
What next
If you are unsure which fits your task, describe your knowledge base. We will show you the architecture for either approach and give you a budget range. The companion pieces are what a RAG system costs and why parsing is most of the work.
Facing a similar problem?
Tell us what you are building. We will walk through the architecture and give you a budget range in one call.