Skip to content
AI and LLM

A local LLM for B2B: when you need one, what it costs, how long it takes

The economics of on-premise LLMs: when local deployment is genuinely justified, what it really costs to run, what infrastructure it needs, and how to estimate the timeline.

By Nikolay Mazur · · 6 min read · Originally published in Russian: читать оригинал

Short answer: not everyone needs a local LLM. It is justified when the data cannot go to an outside service — trade secrets, personal data, security requirements. In every other case a hosted API comes out cheaper and more accurate, and the budget you would have spent on your own infrastructure is freed for the things that genuinely need it.

Every month we hear the same sentence:

“We need our own local LLM. We do not want to send our data to a third party.”

And almost always what stands behind it is fear of hosted models rather than an actual requirement. Here is when a local model really is the answer, and what it costs in practice.

When a local LLM is justified

Three scenarios where we always go local:

  1. Critical infrastructure rules. Where regulation says data does not leave the perimeter, there is nothing to negotiate.
  2. Defence and government supply chains. Even a domestically hosted model may fail the customer’s own regulations.
  3. Medical, financial or legal data under strict compliance. Here a local model is cheaper than justifying a hosted one to an auditor every time.

When a local LLM is overkill

Five scenarios where a hosted API is cheaper and more accurate:

  1. Ordinary B2B with no classified data. A hosted provider with a data processing agreement and a stated data residency usually satisfies the requirement.
  2. Tasks that need a genuinely strong model. Local 7B–13B models lose to frontier hosted models, especially on long-form answers.
  3. Low query volume. Under roughly ten thousand requests a day the GPU economics do not work: hosted billing is several times cheaper.
  4. You need to start quickly. A local model takes four to eight weeks to an MVP; a hosted one, two to three.
  5. No engineer for GPU DevOps. Running that infrastructure costs money and requires expertise you may not have.

Often the answer is not either-or but a two-track setup: sensitive requests go to the local model, public ones to the hosted API. The retrieval layer is shared, so the decision stays reversible during the project instead of being baked into the architecture.

A two-track setup: a query enters a routing layer that checks for sensitive data; such queries go to a local model inside the perimeter next to the vector store of corporate documents, while public queries go to a hosted model

The two-track setup turns the choice of model into a routing setting rather than an architectural decision you cannot reverse.

Infrastructure: what it actually takes

For an MVP with a local LLM:

  • A GPU server: one high-memory accelerator, or two consumer cards for a smaller model. Rented rather than bought at this stage.
  • A vector database: Qdrant self-hosted or pgvector. No separate server needed.
  • The model: Mistral 7B, Qwen 2.5 7B, or a model tuned for your working language.
  • Plumbing: Docker, a reverse proxy, monitoring.

For production with thousands of requests a day:

  • Two to four GPU servers with load balancing and failover.
  • A dedicated server for the vector database, with backups.
  • CI/CD for regular index updates.
  • An on-call engineer, in house or on retainer.

Cost: the shape of it

Development is a one-off: a pilot of three to six weeks, then a full production system over three to four months. Infrastructure is monthly and does not stop: GPU rental plus support and DevOps.

For comparison, the same functionality on a hosted API costs roughly half as much to build and a fraction of that monthly, billed per token.

A local LLM pays back when it saves enough employee time to exceed its monthly running cost, or when compliance makes the hosted option impossible. Those are the only two honest reasons.

Choosing between models

Selection runs on three criteria:

  1. The task: conversation versus document generation versus retrieval.
  2. GPU budget: a 7B model fits one large accelerator; 13B needs two, or quantisation.
  3. Latency requirements: 7B answers in two to four seconds, 13B in five to ten.

Language matters more than benchmark tables suggest. If your working language is not English, test the candidates on your own material before committing: general benchmarks will mislead you.

Ask these instead of “we need our own LLM”

  1. Is there an actual compliance prohibition on hosted models? If not, start hosted.
  2. How many requests a day will there realistically be? A few thousand — hosted. Tens of thousands — local starts to pay.
  3. Are you ready for the DevOps of local GPU infrastructure? If not, budget for a retainer.
  4. Do you already have three to five tasks where an LLM would genuinely save time? If not, pilot one first and decide about local afterwards.

A real example

One of our assistants runs a local model on the client’s own infrastructure with twelve thousand catalogue items, because the client works with defence contractors and the compliance requirement was real. The project paid back in three months.

Without that requirement we would have used a hosted API, and the development budget would have been half the size. That is worth saying plainly: the local deployment was the right call there and would have been the wrong call in most other places.

Next

If you are unsure, tell us about the constraint. We will work out where you have a genuine compliance prohibition and where you have a fear, and show you the option that costs least for the quality you need. The companion pieces are what a RAG system costs and RAG versus fine-tuning.

Facing a similar problem?

Tell us what you are building. We will walk through the architecture and give you a budget range in one call.