/init.news
Agent systems / FIELD NOTE

Before you add GraphRAG, test the questions your memory misses.

A decision worksheet for teams choosing between a simpler retrieval baseline and a graph-based memory system.

THE ANSWER

Evaluate GraphRAG when your important questions need relationships across documents or themes across a collection. Keep a retrieval baseline and compare both approaches on the same questions, corpus, model, and scoring rubric. Measure answer support and operating cost as well as answer quality. A graph is useful only if it resolves a failure you can demonstrate.

Match the method to the failure

Microsoft's GraphRAG documentation describes a pipeline that extracts entities and relationships, forms communities, and summarizes them. Global search uses community summaries for collection-wide questions; local search focuses on entities and nearby context. The documentation identifies connecting information across sources and summarizing broad themes as weaknesses of a plain similarity-search baseline. These are the project's stated motivations, not evidence that every workload improves. [1]

Make a question set before a migration plan

Our proposed evaluation begins with a small, versioned set of real work questions. Separate exact lookup, relationship traversal, collection-wide synthesis, and questions the system should decline to answer. Write expected evidence before comparing outputs. If you choose questions after seeing the graph's answers, it becomes too easy to select a flattering test.

Question classExampleWhat to inspect
Exact lookupWhich document records the approved release date?Correct source and date, without extra inference
RelationshipsWhich decisions depended on the same assumption?Explicit supporting links across source documents
SynthesisWhich recurring constraints explain delayed releases?Coverage and support across the collection
AbstentionWhat did the customer approve when no approval is recorded?An explicit evidence gap instead of a confident invention

Keep the comparison honest

  • Freeze the document snapshot and record its identifier. Keep prompts, model versions, and retrieval settings with the run.
  • Start with a credible baseline: appropriate chunking, useful metadata filters, and a retrieval method that suits the corpus. Do not use a deliberately weak baseline.
  • Score supported claims, missing evidence, unsupported assertions, and useful abstentions separately. Have a reviewer inspect the source passages without knowing which system produced the answer when practical.
  • Record indexing time, refresh cost, query latency, token use, and failure recovery. A quality gain can still be the wrong operational choice.
  • Repeat with one update to the corpus. Test whether changed facts, deleted sources, and contradictory records are handled in a way your application can explain.

Define the adoption rule in advance

Pick the question classes that matter to the product, then set acceptable quality and cost thresholds before running the comparison. Keep the simpler system if the graph does not clear those thresholds. If only synthesis improves, consider routing that class of question to a different retrieval path instead of migrating every request.

That routing recommendation is an implementation hypothesis. This article has not run a GraphRAG benchmark, measured a speedup, or established a universal threshold. The useful artifact is the test design: a way to reject an unnecessary migration as confidently as you accept a justified one.

Sources & checks

  1. Microsoft GraphRAG: overview and query modes ↗

    [1] Graph construction, community summaries, global/local query modes, and the stated limitations of baseline retrieval.

    CHECKED 2026-10-07

Found an error? Use the contact route on Metanomicon and include this page's URL and the correction.