A general-purpose language model knows a lot about the world but nothing about your prices, contracts or internal rules. Retrieval-augmented generation, or RAG, closes that gap without training a new model.
How it works
RAG adds a search step before the answer. Your documents - knowledge base articles, policies, product descriptions, past answers to customers - are split into small fragments and indexed by meaning. When a question arrives, the system finds the most relevant fragments and passes them to the model together with the question. The model then writes an answer based on that material, ideally with links to the sources it used.
When RAG is the right choice
RAG works well when the information changes often and must stay accurate: product catalogues, support knowledge, internal procedures, legal or technical documentation. Updating the assistant is as simple as updating the documents - no retraining is required. It is also easier to control: you can see which fragments were used and check the answer against them.
What usually goes wrong
Most problems come from the data, not the model. Outdated or contradictory documents produce contradictory answers. Huge unstructured files are split badly and the search returns irrelevant pieces. Access rights are often forgotten: an assistant for clients must not see internal documents. Finally, without a set of test questions with known correct answers, nobody can tell whether a change made the system better or worse.
How to start
Pick one clear area, for example answers to the hundred most common customer questions. Clean up the documents for that area, define who may see what, prepare twenty or thirty test questions, and measure how often the assistant answers correctly and cites the right source. Only then expand to new topics.
The short version
RAG turns a general model into an assistant that speaks from your documents. Its quality depends on clean data, careful access rules and regular testing far more than on the choice of model.