Applied AI

RAG (Retrieval-Augmented Generation)

A technique that enhances LLM responses by retrieving relevant documents from a knowledge base before generating.

RAG solves two LLM limitations: knowledge cutoff dates and hallucination. Instead of relying solely on what the model learned during training, RAG retrieves relevant documents from an external knowledge base and includes them in the prompt.

The typical RAG pipeline: (1) embed the user's query, (2) search a vector database for similar documents, (3) include the top results as context, (4) generate a response grounded in the retrieved information.

RAG is the most common way to build AI applications over private or current data.

← All terms