Five RAG architectures (and how to pick one)
Plugging a vector DB into an LLM is not a RAG architecture, it's the starting point. Five patterns hybrid, graph, corrective, agentic, multimodal and what each one actually costs.
Most RAG implementations stop at "plug a vector DB into an LLM" and then wonder why it still hallucinates in production. That's not a RAG architecture it's the starting point for one.
There isn't a single RAG architecture. There are at least five, each fixing a different failure mode, and picking the wrong one for your data shape is usually the actual bug, not the model.
The five patterns
Hybrid RAG
Fuses vector search with keyword search, typically combined with Reciprocal Rank Fusion (RRF). It closes the gap pure semantic search leaves open exact SKUs, error codes, clause numbers, part numbers strings where "similar meaning" is the wrong thing to search for, because there is exactly one right answer and it has to match exactly. It's the cheapest upgrade over naive RAG, and for most systems, the single biggest jump in accuracy.
Graph RAG
Retrieves from a knowledge graph instead of flat chunks, walking relationships across entities rather than ranking documents by similarity. Reach for it when the answer isn't "similar to the query," it's "connected to it through three other things" org charts, regulatory dependencies, bill of materials questions. The cost is on the other side of the ledger: building the graph and keeping it in sync with the source data is real, ongoing work.
Corrective RAG (CRAG)
Scores retrieved documents before generation. Low-confidence retrieval triggers a fallback re-query, broaden the search, or say "I don't know" instead of handing the model bad context and letting it generate on top of it anyway.
This is the difference between "wrong" and "confidently wrong." In a regulated domain, that difference is non-negotiable.
Agentic RAG
Wraps retrieval in a planning loop: decompose the question, query multiple sources, evaluate whether you have enough to answer, iterate if not. It's the right shape for multi-source research assistants, where a single retrieval pass was never going to be enough. It's also the easiest of the five to turn into a latency and token-budget problem an unbounded loop will eventually find its own way to burn both.
Multimodal RAG
Retrieves across text, images, audio, and video using multimodal embeddings, instead of assuming everything can be flattened to clean text first. The moment your knowledge actually lives in scanned PDFs, charts, or diagrams not the text extracted from them this stops being optional.
What each one actually costs
| Architecture | Fixes | Costs you |
|---|---|---|
| Hybrid | Exact-match misses — SKUs, error codes, clause numbers | A fusion step (RRF); the cheapest of the five |
| Graph | Multi-hop, relationship-shaped questions | Building and maintaining the graph |
| Corrective (CRAG) | Confident answers built on bad context | An extra scoring pass before generation |
| Agentic | Questions that need several sources and iteration | Latency and token spend, unless the loop is bounded |
| Multimodal | Knowledge trapped in scans, charts, diagrams, audio | A multimodal embedding pipeline |
Stack them, don't pick one
None of these are mutually exclusive, and most serious production systems don't run just one. A common shape:
- Hybrid as the retrieval base
- Corrective as the safety layer over it
- Agentic or Graph bolted on for the specific queries that need them
The actual question
Not "which RAG is best." It's "what does my data actually look like" flat text or a graph of relationships, clean documents or scanned images, single-hop lookups or multi-source research. The architecture follows from that answer; it doesn't precede it.