Skip to content
AI and LLM

What a RAG system costs a company: budget, timeline and where the money goes

Real numbers for deploying RAG on corporate data: from a pilot to a production system with roles, auditing and index updates. What is in the budget and where you can genuinely save.

By Nikolay Mazur · · 7 min read · Originally published in Russian: читать оригинал

Short answer: a RAG pilot on a limited document corpus takes three to six weeks. A production system — with access roles, answer auditing and automatic index updates — takes two to four months. GPU infrastructure is a separate line: renting is a monthly cost, buying a server is a one-off capital expense that only pays back under steady load.

We quote in roubles, so the figures below are given as proportions and timelines rather than converted amounts — convert at the rate of the day if you need an absolute number, or ask us for a quote.

Now, where those numbers come from and why the range is so wide.

What a RAG budget is made of

RAG is not “a neural network”, it is a pipeline: document preparation → vector index → retrieval → answer generation with citations. The budget splits roughly like this:

LineShare of budgetWhat is inside
Data preparation25–35%Parsing PDF, spreadsheets and internal systems, cleaning, chunking, metadata
Retrieval and index20–25%Vector database, hybrid search, reranking
Generation and prompts15–20%Model selection, prompt engineering, source citation
Integrations15–20%CRM, messengers, corporate portal, SSO
Quality and security15–20%Answer evaluation, access roles, logging, human fallback

The main surprise for clients: the most expensive part is not the model, it is preparing the data. If your knowledge base is ten thousand PDFs with scans, tables and three versions of the same regulation, a third of the budget goes there. We wrote about why in the piece on parsing.

How a RAG project budget splits: data preparation 25–35 percent, retrieval and index 20–25, generation and prompts 15–20, integrations 15–20, quality and security 15–20 percent — the most expensive part is not the model but preparing the documents

Work on the model is the most discussed and one of the cheapest lines in the budget. The money goes before it.

What moves the price most

1. The state of your data. A clean wiki and a dumping ground of scans in a network folder are projects of very different cost, even when the task is described identically.

2. Hosted or local model. A hosted API is cheaper at the start and billed per token. A local LLM costs more to deploy, but the data never leaves your perimeter and the cost does not grow with query volume. For sectors with sensitive data — legal, medical, defence — a local deployment is effectively the only option.

3. Accuracy requirements. “An assistant for internal staff questions” and “a bot that answers customers on behalf of the company” are different levels of responsibility. For the second, budget for an answer evaluation system and a human in the loop.

4. Number of integrations. Every source system is a separate connector and separate testing.

How to save without losing quality

  • Start with a pilot on one scenario. Two or three typical questions, a limited corpus, twenty to fifty real queries to check against. It gives you an honest answer to “does this work on our data” before you spend the full budget.
  • Do not fine-tune the model. In four out of five B2B tasks fine-tuning is unnecessary — RAG on a good index gives a better result for less money. More in RAG versus fine-tuning.
  • Rent the GPU until your load is stable. Buying a server pays back under a constant stream of requests, but during a pilot it is capital sitting idle.

Timeline by stage

  1. Discovery (one to two weeks). Data audit, scenario selection, risk assessment. This is where the real scope becomes visible.
  2. Pilot (three to six weeks). A working prototype on real documents with quality metrics.
  3. Production (two to four months from the start). Roles, monitoring, index updates, staff training.

If a vendor promises production RAG “turnkey in a month”, either they have not done it before, or by “RAG” they mean a wrapper around a hosted chatbot that never touches your data.

In short

  • Pilot: three to six weeks, an honest test on your own data.
  • Production: two to four months, with security, roles and quality metrics.
  • GPU: rental or a purchased server, always a separate line.

Want an estimate for your own corpus? Describe it to us and we will walk through the architecture, the risks and what a pilot could prove.

Facing a similar problem?

Tell us what you are building. We will walk through the architecture and give you a budget range in one call.