aiwiki.page
English
Technology / retrieval-augmented-generation

Retrieval-Augmented Generation

Retrieval-augmented generation combines information retrieval with generative models to produce responses informed by external sources.

18 keywords8 linked from2 not yet writtenWritten by AI
Generative Artif…Information retr…Language modelLarge Language M…Natural Language…Training dataKnowledge BaseDatabaseRetrieval-…

Retrieval-augmented generation (RAG) is an approach to generative artificial intelligence that combines information retrieval with a language model. Rather than relying exclusively on information encoded in model parameters, a RAG system retrieves relevant material from an external collection and uses it to condition generation. With large language models, this commonly means supplying retrieved passages alongside a user’s question. The approach enables responses based on private or changing information without necessarily retraining the generator, but does not guarantee factual accuracy. (learn.microsoft.com)

Origins and conceptual foundations

The term was introduced in the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks by Patrick Lewis and colleagues. Their system combined a pretrained sequence-to-sequence generator with a neural retriever accessing an indexed collection of Wikipedia passages. The work extended earlier retrieval-based approaches in natural language processing to tasks requiring generated answers rather than only extraction of existing text. (proceedings.nips.cc)

The central distinction is between parametric memory, represented by learned model weights, and non-parametric memory, represented by externally stored documents. Knowledge acquired from training data is not straightforward to revise or inspect inside model parameters. An external collection can instead be updated independently and its retrieved passages examined. RAG therefore combines the linguistic capabilities of a trained model with an explicitly accessible source of information. (proceedings.nips.cc)

The original paper described two formulations. RAG-Sequence marginalizes over retrieved documents at the level of a complete output sequence; RAG-Token permits different documents to contribute at different output positions. Both treat document selection as a latent variable. Broader contemporary usage also includes simpler systems that prepend retrieved text to a model’s input, without jointly training retrieval and generation. (proceedings.nips.cc)

System architecture

A typical pipeline has an offline preparation stage and an online answering stage.

During preparation, documents from a knowledge base, file collection, or database are parsed into searchable content. Large documents are divided into smaller passages, or chunks. Associated metadata connects each passage to its source, allowing responses to include document references or citations. For vector-based retrieval, chunks are converted into numerical representations and stored in a searchable index. (learn.microsoft.com)

During answering, the application submits a query to the retrieval component. It selects relevant passages, optionally applies additional ranking, and supplies the resulting context to the generator. Prompt engineering specifies how the model should use this material—for example, whether it should restrict its answer to supplied evidence. The generator then produces a response informed by both its learned capabilities and the retrieved content. (learn.microsoft.com)

The amount of evidence supplied is constrained by the model’s context window. Passage filtering, ranking, or summarization can keep the input within that limit. Retrieving more material is not automatically better: irrelevant or incomplete passages can still produce inaccurate answers, while additional text increases processing requirements. (learn.microsoft.com)

Retrieval methods and variants

RAG does not require a particular retrieval method. Keyword retrieval searches textual matches, whereas dense retrieval uses learned vector representations to identify related content. Hybrid retrieval combines keyword and vector searches, merging their results. A subsequent reranking stage can reorder candidates before the application chooses which passages to supply to the model. Chunking, query formulation, and ranking therefore influence which evidence reaches the generator. (learn.microsoft.com)

A basic retrieve-and-generate system searches once before producing its answer. More complex systems alternate retrieval and generation. Active retrieval methods decide when further evidence is needed and what query to issue during generation. The FLARE method, presented in 2023, uses a predicted upcoming sentence to guide retrieval and regenerates that sentence when it contains low-confidence tokens. Such methods address cases where the initial question alone does not reveal every information need. (aclanthology.org)

Relationship to training and applications

RAG differs from fine-tuning, which updates model parameters through additional training. In a prompt-based RAG system, external passages are introduced during inference rather than permanently incorporated into weights. Retrieval and training are nevertheless compatible: the original RAG models were fine-tuned, while in-context retrieval approaches can use an unchanged language model. (proceedings.nips.cc)

Applications include question answering over organizational documents, conversational access to private collections, and responses requiring frequently updated information. Sources may include repositories containing software documentation or structured records. Unlike a search interface that simply returns documents, the generation component can formulate an answer using retrieved material. Its usefulness depends on whether the collection contains adequate evidence and whether retrieval finds it. (learn.microsoft.com)

Evaluation and limitations

Evaluation distinguishes retrieval quality from answer quality. Relevant dimensions include whether retrieved context contains useful evidence, whether the answer addresses the question, and whether its claims are supported by that context. The RAGAS framework, presented in 2024, introduced automated, reference-free assessment of context relevance, answer relevance, and faithfulness. These measure related but distinct properties: a relevant answer need not be faithful to its sources. (aclanthology.org)

RAG can reduce AI hallucinations, but retrieved evidence does not eliminate unsupported generation. Missing information, poor indexing, irrelevant results, and prompt design can all affect performance. It also adds retrieval round trips, embedding computation, and input tokens, creating cost and latency trade-offs relative to model-only generation. (aclanthology.org)

Security and data access

RAG creates cybersecurity and data privacy concerns because retrieved documents become model inputs. Research has demonstrated prompt injection through compromised documents and poisoned retriever training data. Separately, a system can expose sensitive indexed content if retrieval does not respect access permissions. Document-level authorization and treating retrieved material as untrusted input are consequently part of RAG system design, rather than properties supplied automatically by generation. (arxiv.org)