Skip to content
← All terms
Working with language modelsLast reviewed

Retrieval-Augmented Generation (RAG)

Retrieval-augmented generation (RAG) connects a language model with your own sources: for each question, matching passages of text are looked up and handed to the model as the basis for its answer.

Also known as: RAG

How it works

RAG combines search and text generation. Your documents (manuals, knowledge base, product information) are split into small sections beforehand and stored as embeddings in a vector database. When someone asks a question, the system looks up the best-matching sections and gives them to the language model as a basis. The model writes the answer from these sources.

A practical example

A chatbot answers questions about delivery times and returns — based on the company’s current help pages, not on the model’s general knowledge.

What you should know

  • RAG brings current and internal knowledge into the model without retraining it.
  • It lowers the risk of hallucinations but does not rule it out. Quality depends heavily on how well your sources are maintained and prepared.
  • Source references in the answer make it verifiable.
  • You decide which documents the system may see — important for data protection and access rights.

Questions about your project?

Get in touch →