Quote preparation cut from two hours to twelve minutes
An AI assistant with search across a 12,000-SKU catalogue, quote generation and automatic client history in the CRM. A local LLM on the client's own infrastructure, so data never leaves the perimeter.
This is a representative build based on our expertise and stack. The architecture and the approaches are real. The metrics and the context are given as a reference point for projects of similar complexity, not as a delivered result.
About the client
A manufacturer and supplier of industrial equipment. Forty sales reps work with factories, machine-building plants and defence contractors. Average deal size is in the millions of roubles, and the sales cycle runs two to four months. The catalogue holds 12,000 SKUs, each with dozens of technical parameters.
This is textbook B2B with a complex product and a long sales cycle. There are several hundred companies like it in the market, and their problems are identical.
The problem
The symptoms a rep saw every day
A request arrived in free form — for example, “we need flanges and valves for steam at 40 atmospheres, up to 250 °C, twenty units, two weeks.” Then the routine began:
- Thirty minutes or more finding suitable items. The catalogue existed as fourteen PDF files plus tables in the accounting system. People searched with
Ctrl+Fand by scrolling. - Two hours assembling a quote. A Word template, prices copied by hand, technical parameters, payment terms. Half the templates were out of date, but everyone used them, because nobody had time to refresh them.
- Total loss of context when a rep changed. The conversation history lived in personal email, in a messenger and in someone’s head. A new rep started from zero, even with a client of five years.
What it cost
Out of a forty-hour week, a rep spent fourteen hours on routine — 35% of the department’s time on work that produces no revenue. Against the payroll of the sales department that is a seven-figure monthly loss in roubles.
On top of the direct loss came the deals that got away: while a rep spent two hours on a quote, the client had already heard back from a faster competitor. By the sales director’s estimate, one deal in six entering the funnel was lost this way.
Why off-the-shelf tools did not work
Before commissioning custom development the client tried three simpler routes.
A chatbot in the browser. Reps did not use it. You have to log in, copy data out of the catalogue, and the answer cannot go into the client record. Two weeks in they abandoned it: the workload went up, not down.
Ready-made AI extensions for the CRM. They answer in generalities with no knowledge of this specific catalogue. Useless where the exact SKU and parameter compatibility are the whole point.
Hosted LLM APIs directly. Blocked on security. Part of the documentation contains technical data on equipment for defence enterprises, and sending it to an external API is prohibited under the NDAs with the client’s own customers.
By the time they came to us the requirement was clear: a local LLM on their own infrastructure, tightly integrated into the tools the reps already used.
The approach
Discovery: the first two weeks
We spoke with six reps of differing seniority, with the sales director, with two of their clients and with the IT director. We recorded real conversations: how a request becomes a quote. Three things emerged:
- Reps spend most of their time not on the search itself but on checking technical compatibility — a flange against a valve, pressure against temperature.
- Half the quote templates were out of date, and no rep knew current ones existed.
- The CRM had been abandoned: 60% of records were filled in perfunctorily, because “nobody reads them anyway”.
That changed the focus considerably: what was needed was not a chatbot but an assistant embedded across the whole deal cycle.
Key architectural decisions
A local LLM, not an external API. The client’s documentation contains specifications for defence enterprises, and sending it to a hosted provider is inadmissible under NDA. We deployed a 7B open-weights model, fine-tuned on the client’s corpus, on their own GPU server. With a hybrid fallback: general queries — etiquette, common terminology — route to a hosted API, so the local model is not loaded unnecessarily.
RAG over Qdrant rather than fine-tuning the main model. The catalogue updates weekly; with fine-tuning it would need retraining every time. RAG means a simple overnight reindex. We vectorise the PDFs, the accounting tables and the conversation history using bge-m3 embeddings.
A messenger bot plus a CRM extension, not a separate application. The reps already live in the messenger. A mobile app would have put a wall between them and the tool. A bot plus an iframe extension inside the CRM record: two entry points, one backend.
The solution in detail
Catalogue search
The 12,000 SKUs were broken into semantic chunks: each item is one record in Qdrant with an embedding of its name, key parameters and application area. On each query the assistant:
- runs a vector search for the top 50 candidates;
- applies hard filters — pressure, temperature, diameter — extracted from the query through function calling;
- passes the result through a cross-encoder reranker;
- returns the top five, with links to the source page in the PDF catalogue and the record in the accounting system.
The key feature: the assistant always shows the source of its quote. That removes the rep’s anxiety that the model may have invented it.
Quote generation
A quote is assembled in thirty seconds from a description of the client’s need:
- the rep writes in the messenger: “quote for company X, twenty flanges, valves, delivery through agent Y, fourteen days”;
- the assistant extracts the structure: recipient, line items, counterparty, deadline;
- it fills in current prices from the accounting system over REST, technical parameters from RAG, and payment terms from the admin panel;
- it generates a DOCX from the template using
python-docx. Templates are maintained by the sales director in the admin panel, so reps always work from the latest version; - the file arrives in the messenger and is attached to the CRM record automatically.
Conversation summaries and follow-up
Every morning at eight the assistant processes:
- email, over IMAP;
- messenger conversations with clients;
- call recordings from the phone system, over webhook.
For each active account it produces a digest: what was discussed, what was promised, what the next step is. The digest arrives before the working day starts. If a record has seen no action for 24 hours past a promised follow-up, the assistant messages the rep directly: “Client X was expecting the quote today — what are we doing?”
CRM integration
Not another tab, but embedded in the existing process:
- a webhook on stage change triggers the assistant to check whether the required artefact exists — quote, contract, shipping documents;
- when a new lead appears, the assistant asks the rep whether a preliminary quote should be prepared;
- the record is filled automatically with conversation summaries, links to generated quotes and the next follow-up date.
This solved the “60% of records filled in perfunctorily” problem: the data started accumulating by itself.
Admin panel and observability
A separate view for the sales director showing:
- how many queries each rep made;
- which SKUs are searched most often, which helps plan stock;
- which queries the assistant could not handle, for further tuning;
- response time, errors, GPU load.
Metrics go to Grafana, logs are structured through structlog, errors to Sentry.
The hardest parts
Technical compatibility between items
The problem. RAG returned relevant items but ignored technical compatibility — a flange rated for 50 atmospheres with a valve rated for 25 is wrong. Reps lost trust after two or three bad recommendations.
What we tried first. Adding compatibility rules to the system prompt. It did not work: the model forgot them on longer queries.
What we did. A validation layer after generation — plain Python code with compatibility rules extracted from the datasheets. If the model proposes an incompatible pair, the assistant says so itself: “Note: this valve is rated for 25 atmospheres but the request specifies 40 — a different valve is needed. Here are alternatives.” After that change, complaints from reps dropped to zero.
Long technical documentation in RAG
The problem. PDF catalogues of 200 to 400 pages contain tables, diagrams and footnotes. Standard chunking at 512 tokens broke the context: half the parameters were separated from the item name.
What we did. We wrote a parser on pdfplumber and unstructured.io that keeps a table as a single chunk and glues text descriptions to their section headings. Every chunk in Qdrant carries metadata: page, type (table, text, diagram), parent section. Recall on a test set of 200 typical queries went from 71% to 94%.
Working under NDAs with defence customers
The problem. Part of the documentation physically cannot go to an external API. But reps write queries from their phones, sometimes outside the company network.
What we did. A two-track architecture: queries on public data go to the external track with a hosted model; queries touching restricted data go to a separate on-premise GPU track with no internet access, reachable only over VPN. Routing is driven by tags in Qdrant and the client category in the CRM. It took about a week of discovery on its own, separate from the main development.
Results
Ninety days after full launch
92% less time finding information. From thirty minutes to two per item. Across a department of forty that is 280 hours a week — the equivalent of seven people freed from routine and returned to talking to clients.
84% less time on quotes. From two hours to twelve minutes. A rep now sends the quote the same day, where it used to take one or two working days.
14 percentage points better lead-to-deal conversion. Not from the AI directly, but because reps started responding faster. The client no longer had time to leave for a faster competitor.
38% better CRM completeness. From 60% to 98%, because the data is written automatically rather than “when someone gets round to it”.
Three-month payback, against the monthly saving on the sales department’s time.
“We tried a chatbot, extensions, bolting a frontier model on. None of it worked, because the reps did not want to change their process. This worked, because the assistant sits inside the messenger and the CRM instead of existing separately. Two weeks after launch three reps came to us with their own ideas for improving it. That was the first time a tool became theirs rather than something imposed from above.”
— Sales director, the client
What changes for each role
Different people in the company see the effect of a project like this differently. Here is what changes day to day in each seat.
The sales rep. Selecting compatible items stops being a half-hour job: the assistant proposes suitable flanges and valves with pressure and temperature checked, and shows the source in the catalogue for every recommendation. That source link matters more than the recommendation itself — without it the rep would not forward the answer to a client. A quote is assembled from one message and attaches to the CRM record immediately, so the client gets the offer the same day rather than two days later.
The sales director. A morning digest per client — what was discussed, what was promised, what comes next — changes the department’s discipline more than any written process would. The assistant chases overdue follow-ups on its own, and when a client moves to another rep the context is not lost: the history sits in one place instead of in a departed employee’s inbox.
The IT director. A local model on the client’s own infrastructure was not a preference but a condition of entry: technical documentation does not leave the perimeter and the NDA requirements hold. The two-track setup, with public queries going to a hosted model and sensitive ones handled on-premise, removes the need to choose between answer quality and security. RAG on Qdrant with a weekly catalogue reindex is an architecture an in-house team can take over.
The finance director. The saving here is not headcount reduction but released time: the routine of searching and preparing documents goes away and reps return to deals. Measure the effect by time to produce a quote and by lead-to-deal conversion — those are the two metrics that move first.
What comes next
We are working on a second version:
- voice commands for reps on site visits;
- extension to the service department — diagnosing equipment problems from a client’s description and selecting spare parts;
- deal probability forecasting from conversation history and client attributes, as an ML model over three years of data.
Support and development continue on a monthly retainer.
Stack
- Backend: Python 3.12, FastAPI, Celery with Redis, PostgreSQL, structlog
- AI: a 7B open-weights model locally, a hosted API as fallback, bge-m3 embeddings, a bge reranker, Qdrant as the vector store
- Integrations: CRM REST and webhooks, Telegram Bot API, IMAP, phone system, accounting over HTTP services
- Frontend: React 18 with TypeScript for the admin panel and the CRM iframe extension
- Infrastructure: on-premise Ubuntu 22.04, Docker Compose, an NVIDIA A100 40GB, Grafana and Sentry
When this approach fits
It fits if:
- you are B2B with a complex product and a catalogue of a thousand items or more;
- reps spend more than 30% of their time on routine — searching, quoting, filling in the CRM;
- ready-made AI extensions did not stick and deeper integration is needed;
- there are compliance requirements and data must not leave the perimeter.
It does not fit if:
- the catalogue is under 200 items — hiring a second rep is cheaper;
- the business is mass-market B2C with simple transactions, where off-the-shelf SaaS works;
- the expectation is magic without changing the process — without the sales director’s backing, deployment fails.
A similar problem in your business?
In one call we will work out what this would be worth to you and which architecture fits.
Services used in this project.
- AI for sales
AI assistant for a sales team
Building an AI assistant for sales: RAG over catalogues, quote drafting, CRM integration, messengers, local LLMs and data security.
- Local LLMs
Local LLM and RAG for business
Deploying local LLMs and RAG systems on corporate data: retrieval, answers with sources, on-premise, security and integrations.
- Messenger bots
Messenger bot development for business
Telegram bot development for B2B: order intake, CRM and accounting integration, payments, a user portal inside the messenger, AI answers, support and sales bots.