What is Agentic RAG?
Agentic RAG is retrieval-augmented generation in which an AI agent decides what to retrieve, from which source, and whether the results are good enough. When they are not, it retrieves again, instead of running one fixed lookup before answering. A 2025 survey of the field describes it as "embedding autonomous AI agents into the RAG pipeline" so that retrieval strategies are managed dynamically and contextual understanding is refined iteratively.
Classic RAG, where the term comes from
The term RAG comes from a 2020 paper by Patrick Lewis and colleagues, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks". Their models combined a pre-trained seq2seq model (parametric memory) with a dense vector index of Wikipedia (non-parametric memory), accessed through a pre-trained neural retriever. The paper reports state-of-the-art results on three open-domain question-answering tasks, and more specific, diverse and factual generation than a parametric-only baseline.
Two details matter for the comparison. The original RAG was fine-tuned end to end; a common production pattern (embed the question, fetch the top-k chunks, paste them into a prompt) is a looser descendant of it. And in both, retrieval happens the same way for every question: one query, one pass, no judgement about whether the passages were relevant.
Classic RAG: question → retrieve top-k → generate answer
Agentic RAG: question → plan → choose source/tool → retrieve → judge results
↳ (rewrite query / retrieve again / split into sub-questions)
→ generate → check answer against evidence → answer
How agentic RAG works
The survey (Singh et al., arXiv 2501.09136) lists four agentic design patterns that these systems draw on: reflection, planning, tool use and multi-agent collaboration. In practice they show up as a loop:
- Plan. Decide whether the question needs retrieval at all, and whether it should be broken into sub-questions.
- Choose a source or tool. A vector index, a SQL database, a web search or an API can each be exposed to the agent as a tool. LangChain's LangGraph tutorial builds exactly this: the retriever is wrapped as a tool, and the agent "can decide when to use" it.
- Retrieve.
- Judge the results. Grade the retrieved passages for relevance before they reach the answer.
- Retrieve again if needed. Rewrite the query, switch source, or search for a missing piece. The LangGraph tutorial's steps are create a retriever tool, generate a query or respond, grade documents, rewrite the question, generate an answer.
- Self-check. Compare the draft answer with the evidence before returning it.
The survey classifies architectures by agent cardinality (one agent or several), control structure, autonomy and knowledge representation.
Named patterns with published papers
| Pattern | What it adds | Paper |
|---|---|---|
| Self-RAG (Asai et al., 2023) | A single LM trained to retrieve on demand and to critique retrieved passages and its own output with special "reflection tokens" | arXiv 2310.11511 |
| Corrective RAG, CRAG (Yan et al., 2024) | A lightweight retrieval evaluator scores the retrieved documents and returns a confidence degree that triggers different retrieval actions, including web search when the static corpus falls short, plus a decompose-then-recompose step that filters irrelevant text | arXiv 2401.15884 |
Both papers start from the same diagnosis. Self-RAG's authors say that retrieving a fixed number of passages "regardless of whether retrieval is necessary, or passages are relevant" can lead to unhelpful answers; CRAG's authors say RAG "relies heavily on the relevance of retrieved documents". Note the difference in approach: Self-RAG trains the retrieve-and-critique behaviour into the model, while CRAG is described by its authors as plug-and-play with existing RAG pipelines. Self-RAG shows that agentic behaviour does not have to come from a separate orchestration framework.
Trade-offs
| Classic RAG | Agentic RAG | |
|---|---|---|
| Model calls per question | One generation after one retrieval | Several (planning, grading, rewriting, checking) |
| Latency and token cost | Lower, predictable | Higher, varies with how many loops a question triggers |
| Points of failure | Retriever and generator | Also the planner, the grader, tool calls and loop-stopping logic |
| Handling a bad first retrieval | Answers from it anyway | Can detect it and retry |
| Evaluation | Retrieval quality plus answer quality | Also has to judge each decision along a variable path |
Every row of the table is an inference from the design (each extra step is another model or tool call), not a measured figure; none of the sources here gives a general number. The survey itself lists evaluation, coordination, memory management, efficiency and governance as open research challenges.
When to use it, and when not to
Use it when questions are multi-part, when the answer may sit in different sources (documents, databases, the web), or when a wrong answer drawn from irrelevant passages is costly.
Do not reach for it when a single retrieval already answers most questions well, when latency or per-query cost is tightly bounded, or when you cannot yet measure retrieval quality. An agent that loops over a weak index retrieves more weak passages; fix chunking, embeddings and the index first. A longer context window does not by itself replace retrieval: the model still has to be given the right text.
In the index
The index lists frameworks that can be used to build these loops: LangChain (whose LangGraph tutorial is cited above), LlamaIndex, PydanticAI, smolagents and Mastra, all compared under agent frameworks. For the vector-search layer underneath, see HNSW and Pinecone. A hands-on walkthrough is in how to connect local RAG with MCP. This index does not test or benchmark these frameworks.
Sources
- Patrick Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", 2020 (origin of the term RAG; seq2seq parametric memory plus dense Wikipedia index with a neural retriever; state of the art on three open-domain QA tasks; more specific, diverse and factual generation): https://arxiv.org/abs/2005.11401
- Aditi Singh et al., "Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG", 2025 (definition quoted; reflection, planning, tool use, multi-agent collaboration; taxonomy by agent cardinality, control structure, autonomy, knowledge representation; open challenges incl. evaluation and efficiency): https://arxiv.org/abs/2501.09136
- Akari Asai et al., "Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection", 2023 (on-demand retrieval, reflection tokens, critique of fixed-passage retrieval): https://arxiv.org/abs/2310.11511
- Shi-Qi Yan et al., "Corrective Retrieval Augmented Generation", 2024 (lightweight retrieval evaluator, confidence degree, web-search extension, decompose-then-recompose, plug-and-play): https://arxiv.org/abs/2401.15884
- LangChain, "Build a custom RAG agent with LangGraph" (retriever as a tool the agent decides when to use; grade documents and rewrite-question steps): https://docs.langchain.com/oss/python/langgraph/agentic-rag
