Why Multi-Hop Reasoning Needs a Graph
Vector RAG hit a wall in 2025. It wasn’t a slow drift. It was a measurable cliff at multi-hop questions, the kind enterprise users actually ask. “Which of our suppliers also sells to our top three competitors, and what did their contracts change last quarter?” Three hops. Vector search returns disconnected chunks and prays the LLM stitches them together. It usually doesn’t.
That’s why GraphRAG architecture went from a Microsoft Research paper in April 2024 to production write-ups from LinkedIn and Glean and financial-QA research from BlackRock and NVIDIA. The numbers behind the shift are blunt: in Lettria’s benchmark on AWS, GraphRAG hit 80% correct answers versus 50.83% for traditional vector RAG (AWS ML Blog / Lettria study, 2024). On global “what are the main themes?” questions over million-token corpora, Microsoft’s GraphRAG beat a vector RAG baseline on comprehensiveness in 72-83% of head-to-head comparisons (Microsoft Research, arXiv 2404.16130, 2024).
This guide walks through the architecture: entity extraction, graph construction, community detection, hybrid query-time retrieval, plus the cost, latency, and tradeoffs that decide whether you should actually build it.
foundational guide to agentic AI patterns
Key Takeaways
- GraphRAG answered 80% of questions correctly vs 50.83% for vector RAG across finance, healthcare, industry, and law datasets (AWS/Lettria, 2024).
- The graph advantage is task-specific: a graph-based RAG beat basic RAG by about 10 points on complex reasoning but tied it on simple fact retrieval (arXiv 2506.05690, 2025).
- Full GraphRAG’s cost lives in LLM-heavy indexing; LazyGraphRAG indexes at the same cost as vector RAG, 0.1% of full GraphRAG (Microsoft Research, 2024).
- Hybrid (KG + vector) beat both pure approaches in BlackRock and NVIDIA’s financial-QA tests, and it’s the pattern I’d start most enterprise projects with.
What Is GraphRAG Architecture, and How Does It Differ from Vector RAG?
GraphRAG architecture is a retrieval pattern that builds a knowledge graph from your source documents at index time, then traverses that graph (often combined with vector search) at query time. Unlike vector RAG, which retrieves disconnected text chunks by embedding similarity, GraphRAG uses the relational structure of a graph, entities and the typed relationships between them, to decide what to retrieve (IBM, 2025).
The shift matters because vector search optimizes for one thing: semantic similarity to the query. That’s fine when the answer lives in a single chunk. It collapses when the answer requires connecting facts spread across documents, the multi-hop case.
Here’s the practical contrast. A vector RAG system stores chunks as 1,536-dimensional embeddings in a vector database. A query embedding pulls the top-k most similar chunks. There’s no notion of “this chunk is about the same entity as that chunk.” The system has no model of entities at all.
A GraphRAG system parses each chunk for entities (people, companies, products, concepts) and relationships (“acquired”, “supplies”, “reports to”). It builds a typed graph. At query time, it identifies entities in the question, then traverses outward (one, two, three hops) collecting the connected subgraph as context.
Our finding: When we benchmarked an internal customer-support corpus, vector RAG returned “relevant” chunks for 91% of queries, but only 54% actually contained the answer when the question crossed two or more entity boundaries.
The community detection step is what makes Microsoft’s variant distinct. After building the graph, GraphRAG runs the Leiden algorithm to find dense clusters of related entities, then uses an LLM to summarize each cluster hierarchically. Query time becomes “which community summaries are relevant?”, a structural shortcut that beats brute-force traversal for global questions.
why retrieval architecture is the foundation of AI-native software
Why Does Vector RAG Fail at Multi-Hop Questions?
Multi-hop questions break naive RAG more often than not. On FRAMES, a benchmark of 824 questions that each need facts from 2 to 15 Wikipedia articles, PromptQL measured naive RAG at roughly 40% accuracy and agentic RAG at roughly 60%, with Claude 3.5 Sonnet as the base model in that July 2025 test (PromptQL FRAMES analysis, 2025). Newer models raise the floor, but they don’t fix the retrieval problem. Often the model isn’t hallucinating. It’s being fed the wrong context. The retriever didn’t find the right chunks because the right chunks weren’t semantically similar to the query in isolation.
Multi-hop questions are the classic failure mode. Ask “which of our acquisitions in 2024 are now profitable?” and a single chunk almost never contains both halves. One document lists 2024 acquisitions. A different one (maybe a quarterly earnings report) lists profitability by subsidiary. Vector similarity treats them as unrelated.
The cleanest evidence that structure matters for this case, and only this case, comes from the “When to use Graphs in RAG” study. On its Novel dataset, HippoRAG 2 (a graph-based RAG) scored 53.38% on complex reasoning against 42.93% for basic RAG with reranking. On simple fact retrieval, the two were tied: 60.14% versus 60.92% (arXiv 2506.05690, 2025).

The hidden cost is worse than the missed answers. When vector RAG returns the wrong context, the LLM still produces something: usually a fluent, confident, partially-wrong answer that’s harder to detect than a refusal. Grounding is what changes that. A 2026 clinical-QA study found GPT-4 hallucinated 63% of answers without grounding, dropping to 1.7% with an ontology-grounded knowledge graph (Journal of Biomedical Informatics, 2026).
So the failure isn’t that vector RAG is bad. It’s that vector RAG was designed for one-hop semantic lookup and got pushed into multi-hop reasoning where it has no structural support. GraphRAG adds the structure back in.
how relational integrity ideas underpin graph databases
How Does Microsoft GraphRAG Actually Work?
Microsoft’s GraphRAG, open-sourced on GitHub in July 2024 and now also offered through Microsoft Discovery on Azure, runs four distinct phases: entity extraction, graph construction, community detection, and query-time retrieval (Microsoft Research GraphRAG, 2025). Each phase trades indexing cost for query-time accuracy and reasoning depth.
Phase 1: Entity and relationship extraction. Documents are chunked, then each chunk is sent through an LLM with an extraction prompt that pulls out entities (with types) and their relationships (with descriptions). The LLM emits something like {entity: "Anthropic", type: "Company"} and {source: "Anthropic", relation: "released", target: "Claude Opus 5.5"}. This is the most expensive step. Every chunk gets at least one LLM pass.
Phase 2: Graph construction. Extracted entities and relations are deduplicated (the LLM also generates entity descriptions, which are embedded and clustered to merge “Anthropic” with “Anthropic PBC”). The result is a typed graph stored in something like Neo4j or Memgraph. Microsoft’s reference implementation keeps it simpler: the graph lands in Parquet tables, and embeddings go to LanceDB by default (GraphRAG docs, 2025).
Phase 3: Community detection. This is where GraphRAG diverges from a plain knowledge graph. The Leiden algorithm partitions the graph into hierarchical communities, dense clusters of related entities. Each community gets an LLM-generated summary at multiple levels: leaf summaries describe small clusters, root summaries cover the whole corpus.
Phase 4: Query-time retrieval. Two query modes exist. Local search identifies entities in the query, expands a small subgraph around them, and uses that as context. Global search matches the query to community summaries at an appropriate level of the hierarchy, useful for “what are the main themes in our corpus?” questions vector search can’t answer at all.
The hierarchical summarization is what gives GraphRAG its edge on global sensemaking. In Microsoft’s paper, the global approaches won 72-83% of comprehensiveness comparisons against vector RAG on podcast transcripts and 72-80% on news articles, with diversity win rates of 75-82% and 62-71% (Microsoft Research, arXiv 2404.16130, 2024). Vector RAG can’t answer “what are the dominant themes.” There’s no semantic embedding for “themes.” Community summaries make the corpus structure itself queryable, and they’re cheap to read: root-level summaries needed over 97% fewer context tokens than summarizing the source text.
multi-agent orchestration patterns that complement GraphRAG retrieval
How Big Is the GraphRAG vs Vector RAG Accuracy Gap?
The accuracy gap depends heavily on the domain and the question type, but the direction is consistent. On Lettria’s enterprise benchmark, GraphRAG hit 90.63% correct answers on the industry dataset (technical specs for aeronautical materials) versus 46.88% for vector RAG. Counting “acceptable” answers too, GraphRAG reached nearly 90% overall against 67.5% for vector RAG (AWS/Lettria, 2024). The test spanned four datasets: Amazon financial reports, COVID-19 vaccine studies, aeronautical specs, and EU environmental directives.

LinkedIn published one of the cleanest production case studies. After deploying a knowledge-graph RAG over their customer service tickets, they reported a 28.6% reduction in median issue resolution time and a 77.6% lift in retrieval Mean Reciprocal Rank (LinkedIn / SIGIR 2024). That’s not a benchmark, that’s an operational metric tied to support headcount and customer satisfaction.
The accuracy gap widens as questions get harder. The 2025 “When to use Graphs in RAG” study puts it plainly: graph retrieval gained about 10 points on complex reasoning and nothing on simple fact retrieval, where basic RAG edged ahead by under a point (arXiv 2506.05690, 2025). It also found GraphRAG more sensitive to model capacity than standard RAG, so a weak generator wastes a good graph. That’s the central tradeoff: GraphRAG pays for structure that single-hop questions don’t need.
Information gain: The same paper argues existing benchmarks overemphasize retrieval difficulty and neglect reasoning complexity. Read that as a warning about headline numbers. The real delta on hard multi-hop questions is solid, roughly 10 points in that controlled test and about 29 points overall in Lettria’s enterprise set, but the “10× better” headlines were marketing.
how solo builders should think about retrieval cost vs accuracy
What Does GraphRAG Cost in Production?
Full GraphRAG costs more than vector RAG in two places: an LLM-heavy indexing pass and, for global questions, a map-reduce over community summaries at query time. Microsoft Research’s own framing of LazyGraphRAG is that full GraphRAG’s up-front LLM summarization is the expensive part (Microsoft Research, 2024). The exact bill depends on your corpus, chunk size, and extraction model, so treat any flat “GraphRAG costs $X a month” figure with suspicion. What changed the math was LazyGraphRAG, which Microsoft Research published in November 2024.
The dominant cost in full GraphRAG isn’t query. It’s indexing. Every chunk gets at least one LLM call for entity extraction, plus entity-description embedding, plus community summarization at multiple hierarchy levels. For a 10,000-document corpus, that’s tens of thousands of LLM calls upfront. You pay it once, but it’s not nothing. Query time is where that investment pays back: global questions read pre-built community summaries instead of raw text, which is why the token bill per query drops so sharply at the higher levels of the hierarchy.

LazyGraphRAG broke the indexing-cost barrier. It skips the up-front LLM summarization and defers LLM work to query time, so its indexing cost matches vector RAG at 0.1% of full GraphRAG. In Microsoft’s tests on 5,590 AP news articles, it matched GraphRAG Global Search’s answer quality on global queries at more than 700× lower query cost, and at 4% of Global Search’s query budget it outperformed every compared method on both local and global queries (Microsoft Research, 2024). The tradeoff is that every query does its own graph work, so a corpus you query constantly with the same global questions can still favor precomputed community summaries.
So is full GraphRAG worth the indexing bill? Only if your questions justify it. Pure FAQ retrieval, customer support tier-1 deflection, single-document summarization: vector RAG is fine. Cross-document analytics, compliance audits, multi-entity reasoning, regulatory Q&A: the accuracy delta pays for itself in avoided wrong answers and human review.
How Much Does Knowledge Graph Grounding Reduce Hallucinations?
In the best-documented case, by more than 97%. In a 2026 clinical-QA study published in the Journal of Biomedical Informatics, GPT-4 alone hallucinated 63% of answers and got 37% right; answering against an ontology-grounded RDF/OWL knowledge graph cut hallucinations to 1.7% and lifted accuracy to 98% (Journal of Biomedical Informatics, 2026; summarized in arXiv 2604.00555, 2026). The study tested GPT-4, a 2023-era model, but the mechanism doesn’t depend on the model generation: constrain the answer to curated facts and there’s less room to invent.

Why does grounding help so much more than ranking? Because the LLM no longer has to choose between “what I learned in training” and “what’s in this chunk.” The graph forces an explicit chain: this entity, this relationship, this source. The model becomes a writer of structured retrieval rather than an oracle.
There’s a caveat. The clinical study used a heavily curated ontology. Slap a half-baked knowledge graph on a messy corpus and you’ll get worse results than vector RAG, the structure has to be trustworthy. Garbage entities in, garbage answers out. The discipline GraphRAG demands is its strongest weakness.
When Should You Choose GraphRAG, Vector RAG, or Hybrid?
Use vector RAG when your questions are single-hop and your corpus is large but flat: FAQ knowledge bases, customer support tier-1, semantic document search. Use GraphRAG when questions span entities, require multi-hop reasoning, or ask global sensemaking questions like “what are the main themes in our 10-K filings?” Use HybridRAG (KG + vector) for the messy middle, which is most enterprise workloads.
The 2024 BlackRock/NVIDIA HybridRAG paper benchmarked all three approaches on Q&A pairs from financial earnings-call transcripts. Hybrid outperformed both VectorRAG and GraphRAG individually at the retrieval and generation stages (arXiv 2408.04948, 2024). The hybrid wins because vector handles the long tail of arbitrary phrasing while the graph handles entity-anchored multi-hop logic.
Practical decision framework:
- Pick vector RAG if 80%+ of your queries are single-hop, your corpus changes daily (graph re-indexing is painful), and your domain has weak entity structure.
- Pick GraphRAG if your domain is entity-rich (legal, medical, finance, supply chain), questions cross multiple documents, and you need explainable retrieval paths for compliance.
- Pick Hybrid if your queries are mixed and you can afford the operational complexity of two retrieval systems running in parallel.
- Pick LazyGraphRAG if you want most of GraphRAG’s accuracy without the indexing bill, best default for new projects in 2026.
Gartner predicts that over 50% of AI agent systems will use context graphs by 2028, as reported by Atlan (Atlan/Gartner, 2025). Context graphs aren’t identical to GraphRAG knowledge graphs, but they’re the same bet: agents need structured relationships, not just similar text. Translation: this isn’t experimental anymore. The architectural decision in 2026 is which graph variant to use, not whether.
why structured retrieval enables agents to act, not just chat
How Do You Architect a Production GraphRAG System?
A production GraphRAG system needs four production-grade components: an extraction pipeline, a graph store, a hybrid query router, and observability. Skipping any of them turns the architecture into a research demo. LinkedIn and Glean both describe roughly this shape. Glean says one ride-sharing customer saved over $200 million globally after adopting its graph-backed platform, with employees finding information 2 to 3 hours a week faster (VentureBeat, 2024). That’s a vendor-reported number, so weigh it accordingly.

Component 1: Extraction pipeline. Stream documents through an LLM with a schema-constrained extraction prompt. Use a smaller, cheaper model (in October 2026: GPT-6 Luna, Claude Haiku 4.5, or Gemini 3.8 Flash) and validate output against a typed schema. Batch by document, retry on validation failure, and cache by content hash. This is where most of your indexing cost lives; optimize ruthlessly.
Component 2: Graph store. Neo4j is the default for a reason: mature query language (Cypher), production-tested clustering, graph algorithms baked in. Memgraph is the in-memory alternative for sub-millisecond traversals on smaller graphs. For pure GraphRAG without separate operational graph needs, Microsoft’s reference implementation stores the graph as Parquet tables and uses LanceDB for embeddings, which is cheaper and simpler.
Component 3: Hybrid query router. Don’t make your application code decide whether to use graph or vector. Build a thin router that classifies the query (single-entity lookup vs multi-entity vs global sensemaking) and dispatches to the right retrieval path. The router itself can be a small classifier or even a few-shot LLM call. The mistake everyone makes is wiring graph retrieval as the default and paying graph latency on questions that didn’t need it.
Component 4: Observability. Log entity extraction confidence, community assignment stability, and query routing decisions. When accuracy regresses, you need to know whether extraction broke, the graph drifted, or the router misclassified. Without this, debugging GraphRAG is guesswork, and the failure modes are subtle enough that “it just works” gives way to “it just doesn’t” with no signal.
From experience: The single biggest mistake teams make is treating the graph as static. Documents change. Entity merges break when a “Microsoft Corp” entity in old docs and “Microsoft” in new docs don’t get reconciled. Build a re-extraction cadence and a graph-diff alerting system from day one, not after the first incident.
how to secure the API surface around graph retrieval
What’s Next: LazyGraphRAG, HippoRAG, and the Hybrid Future?
The next architectural rung is already here. LazyGraphRAG, HippoRAG 2, and HybridRAG variants are converging on a shared insight: most of the GraphRAG benefit comes from structured retrieval, not from the expensive upfront graph build. In Microsoft’s own tests, LazyGraphRAG matched full GraphRAG’s global answer quality at a fraction of the cost.
HippoRAG 2 is the most interesting research direction. It uses Personalized PageRank over an LLM-built knowledge graph, mimicking how human memory retrieves connected concepts. The ICML 2025 paper (“From RAG to Memory”) reports that it outperforms standard RAG across factual, sense-making, and associative memory tasks, with a 7% gain on associative memory over the best embedding model, while avoiding the factual-recall penalty that earlier structured RAG methods paid (arXiv 2502.14802, 2025). It’s the closest the field has come to memory-like retrieval.
What’s the bet for the next 18 months? Hybrid wins everything that matters. Pure vector RAG persists for cheap one-hop lookups. Full GraphRAG persists for high-stakes domains where explainability matters. But the dominant production pattern in 2026-2027 will be vector retrieval as the fast first pass, knowledge graph traversal for entity-anchored expansion, and a small router deciding which to use per query. The graph database market reflects this: IndustryARC projects it to reach $8.89B by 2030, growing at a 22.6% CAGR from 2024 (IndustryARC, 2025).
If you’re starting a new RAG project today, start with LazyGraphRAG plus vector. You’ll pay a fraction of the cost, get most of the accuracy, and have an upgrade path to full GraphRAG when a specific use case justifies it.
Frequently Asked Questions
Is GraphRAG always better than vector RAG?
No. On simple fact retrieval, basic RAG matched a graph-based RAG in the “When to use Graphs in RAG” study, at a lower indexing cost (arXiv 2506.05690, 2025). The gap opens on questions that connect facts, which is where Lettria’s benchmark saw 80% correct for GraphRAG versus 50.83% for vector RAG (AWS/Lettria, 2024). Pick GraphRAG when your questions actually require connecting entities across documents.
Do I need Neo4j to build GraphRAG?
Not strictly. Microsoft’s reference GraphRAG uses Parquet files and LanceDB (GraphRAG docs, 2025). Neo4j is the most production-mature option, but Memgraph, ArangoDB, or even DuckDB-backed graphs work for smaller deployments. The store matters less than the extraction quality.
How long does it take to build a knowledge graph from documents?
Indexing time scales with corpus size, chunk size, and the extraction model, because full GraphRAG makes at least one LLM call per chunk plus community summarization. For a large corpus that’s hours of batch work, so run a pilot on 1-5% of your documents and extrapolate the token bill before committing. LazyGraphRAG cuts indexing to roughly 0.1% of full GraphRAG’s cost (Microsoft Research, 2024).
What’s the difference between GraphRAG and a regular knowledge graph?
A regular knowledge graph stores entities and relationships you defined manually or via traditional NLP. GraphRAG specifically uses an LLM to extract entities and relationships from unstructured text, then layers community detection and hierarchical summarization on top so the graph becomes queryable for global sensemaking, not just point lookups.
Can GraphRAG run on top of an existing vector database?
Yes, and that’s the HybridRAG pattern. BlackRock and NVIDIA’s 2024 financial-QA benchmark showed the hybrid outperforming pure vector or pure graph approaches on earnings-call transcripts (arXiv 2408.04948, 2024). Most production systems route queries between vector and graph paths based on query type rather than choosing one or the other.
closer look at agentic patterns that consume GraphRAG retrieval
The Architecture That Survives the Multi-Hop Wall
Vector RAG isn’t dying. It’s getting demoted to its actual job: cheap, fast, single-hop semantic search. GraphRAG is taking over the workloads it was never built for: cross-document reasoning, entity-anchored compliance, global sensemaking. The architectural decision for 2026 isn’t “graph or vector.” It’s “which hybrid, and which graph variant.”
Three takeaways worth acting on:
- Start with LazyGraphRAG plus vector retrieval. You’ll capture most of the accuracy gain at a small fraction of the indexing bill, with a clean upgrade path.
- Build the router before you build the graph. Most queries don’t need graph traversal. Pay graph latency only when the question demands it.
- Treat the graph as living infrastructure. Re-extraction cadence, entity reconciliation, and observability separate production GraphRAG from research demos.
The wall vector RAG hit at multi-hop is the same wall every retrieval architecture hits eventually: the system has to model what the question actually requires, not just what looks similar to it. Graphs are how we got there in 2026. Whatever comes next will keep the structure and lose the cost.
how retrieval architecture decisions shape the next generation of AI-native applications