Retrieval-Augmented Generation (RAG)
Retrieval-augmented generation (RAG) connects a language model with your own sources: for each question, matching passages of text are looked up and handed to the model as the basis for its answer.
Also known as: RAG
How it works
RAG combines search and text generation. Your documents (manuals, knowledge base, product information) are split into small sections beforehand and stored as embeddings in a vector database. When someone asks a question, the system looks up the best-matching sections and gives them to the language model as a basis. The model writes the answer from these sources.
A practical example
A chatbot answers questions about delivery times and returns — based on the company’s current help pages, not on the model’s general knowledge.
What you should know
- RAG brings current and internal knowledge into the model without retraining it.
- It lowers the risk of hallucinations but does not rule it out. Quality depends heavily on how well your sources are maintained and prepared.
- Source references in the answer make it verifiable.
- You decide which documents the system may see — important for data protection and access rights.
Matching tools