← All posts

3 min read

Five RAG architectures (and how to pick one)

Plugging a vector DB into an LLM is not a RAG architecture, it's the starting point. Five patterns hybrid, graph, corrective, agentic, multimodal and what each one actually costs.

rag llm genai architecture

View image

Most RAG implementations stop at "plug a vector DB into an LLM" and then wonder why it still hallucinates in production. That's not a RAG architecture it's the starting point for one.

There isn't a single RAG architecture. There are at least five, each fixing a different failure mode, and picking the wrong one for your data shape is usually the actual bug, not the model.

The five patterns

Hybrid RAG

Fuses vector search with keyword search, typically combined with Reciprocal Rank Fusion (RRF). It closes the gap pure semantic search leaves open exact SKUs, error codes, clause numbers, part numbers strings where "similar meaning" is the wrong thing to search for, because there is exactly one right answer and it has to match exactly. It's the cheapest upgrade over naive RAG, and for most systems, the single biggest jump in accuracy.

Graph RAG

Retrieves from a knowledge graph instead of flat chunks, walking relationships across entities rather than ranking documents by similarity. Reach for it when the answer isn't "similar to the query," it's "connected to it through three other things" org charts, regulatory dependencies, bill of materials questions. The cost is on the other side of the ledger: building the graph and keeping it in sync with the source data is real, ongoing work.

Corrective RAG (CRAG)

Scores retrieved documents before generation. Low-confidence retrieval triggers a fallback re-query, broaden the search, or say "I don't know" instead of handing the model bad context and letting it generate on top of it anyway.

This is the difference between "wrong" and "confidently wrong." In a regulated domain, that difference is non-negotiable.

Agentic RAG

Wraps retrieval in a planning loop: decompose the question, query multiple sources, evaluate whether you have enough to answer, iterate if not. It's the right shape for multi-source research assistants, where a single retrieval pass was never going to be enough. It's also the easiest of the five to turn into a latency and token-budget problem an unbounded loop will eventually find its own way to burn both.

Multimodal RAG

Retrieves across text, images, audio, and video using multimodal embeddings, instead of assuming everything can be flattened to clean text first. The moment your knowledge actually lives in scanned PDFs, charts, or diagrams not the text extracted from them this stops being optional.

What each one actually costs

Architecture Fixes Costs you
Hybrid Exact-match misses — SKUs, error codes, clause numbers A fusion step (RRF); the cheapest of the five
Graph Multi-hop, relationship-shaped questions Building and maintaining the graph
Corrective (CRAG) Confident answers built on bad context An extra scoring pass before generation
Agentic Questions that need several sources and iteration Latency and token spend, unless the loop is bounded
Multimodal Knowledge trapped in scans, charts, diagrams, audio A multimodal embedding pipeline

Stack them, don't pick one

None of these are mutually exclusive, and most serious production systems don't run just one. A common shape:

  • Hybrid as the retrieval base
  • Corrective as the safety layer over it
  • Agentic or Graph bolted on for the specific queries that need them

The actual question

Not "which RAG is best." It's "what does my data actually look like" flat text or a graph of relationships, clean documents or scanned images, single-hop lookups or multi-source research. The architecture follows from that answer; it doesn't precede it.