Graph RAG in the Wild: Insights and Best Practices from Real-World Applications

Tracking #: 4027-5241

Authors: 
Diego Collarana Vargas
Christopher Ingo Pack
Yan-Ying Liao
Marlena Flüh
Jonathan Lehmkuhl
Abhishek Nageri
Alexander Graß
Moritz Busch
Prinon Das
Lara Dingels
Stefan Decker
Christian Beecks

Responsible editor: 
Harald Sack

Submission type: 
Application Report
Abstract: 
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by enabling access to external knowledge without retraining. While effective, traditional RAG methods—typically reliant on vector-based retrieval—face limitations in understanding complex semantics, connecting dispersed information, and supporting user-centric search workflows. Graph Retrieval-Augmented Generation (Graph RAG) addresses these challenges by incorporating knowledge graphs into the retrieval process, enabling semantically enriched and structured query handling. This paper presents the application of Graph RAG across seven real-world applications, including legal compliance, customer support, enterprise knowledge management, finance, education, data protection enforcement, and time series analytics. For each application, we outline the distinct challenges, solutions, and design decisions made. In addition, we introduce a modular Graph RAG Engine, named FRAG-KEDA, to support ingestion, graph construction, hybrid retrieval, and LLM orchestration. We present empirical evidence demonstrating improvements in accuracy, latency, and user trust. Additionally, we address cross-domain challenges, including graph drift and evaluation strategies.
Full PDF Version: 
Tags: 
Reviewed

Decision/Status: 
Minor Revision

Solicited Reviews:
Click to Expand/Collapse
Review #1
Anonymous submitted on 26/Mar/2026
Suggestion:
Accept
Review Comment:

This paper provides an overview of Graph RAG applied in seven scenarios from Fraunhofer partners. Given the number of authors and scenarios, the paper is written consistently and shows how Graph RAG is applied in each of the presented use cases.

The paper is now submitted as an application report, which fits the presented content much better. Based on my first review, all points were addressed very well. There are minor points where the authors can still slightly improve the paper:

Section 2 covers LLMs and their limitations. The authors state that LLMs are not retrained and thus only have information up to a certain point in time (page 3, line 37). This is true in general, but a few sentences on web search capabilities (in essence, standard RAG) would help here, because the statement is quite strong. Directly afterwards, the authors mention that most of the training data is in English and that "this impairs LLMs’ ability to process other languages". It might be useful to highlight that low-resource languages remain challenging for LLMs, while other languages are still handled quite well.

Section 3 introduces the FRAG-KEDA engine and the possible configuration options. In the following section, all use cases are presented: first the schemaless applications, then the schema-first applications. These are all well motivated and described in detail. One smaller improvement would be to write "FRAG-KEDA Configuration" instead of just "Engine Configuration", which would make the reference clearer. These configurations also show that for each application, the software requires a significant amount of manual work to get it right. On the one hand, this demonstrates that a framework is useful and enables such configurations; on the other hand, it also shows that there is still manual effort involved and no one-size-fits-all solution is available. It might be interesting in the future to evaluate how good the results are with fewer manual configurations.

In Figure 2a, it would be helpful to further describe what the image shows. Is the connection between the entities an aggregated measure of the number of connections between each instance of these classes, or is there another interpretation?

Since the authors introduce the FRAG-KEDA engine, it would also be very interesting to make the implementation publicly available, or to explain why this is not possible, so that it can serve as a reference for future readers. The same applies to the benchmarks, although I understand that the data might not be released due to licensing issues.

Smaller comments: On page 15, line 26, the authors could also add a reference to the schema in Figure 3. It might also be beneficial to move the schemas into the sections where they belong.

Section 6 highlights best practices and summarizes which approaches are used in each use case. In the following section, open challenges are highlighted and explained. This is an exhaustive list, and I do not think that major points are missing.

Overall, I think the paper can be accepted as is, and I hope that the very small improvements are addressed by the authors for the camera-ready version.

Review #2
Anonymous submitted on 24/Jun/2026
Suggestion:
Minor Revision
Review Comment:

I thank the authors for considering my previous comments, and the manuscript has been improved in some of the aspects.

First, each application has a much more thorough description. The authors give more precise explanations of how the knowledge graphs are constructed, how retrieval is done, how the systems are implemented, and what beneficial impacts they have in real-life scenarios.

Second, a clearer explanation of the differences between schemaless and schema-first Graph RAG techniques is now provided. This is one of the paper's strongest points and gives readers a greater understanding of the various ways Graph RAG systems can be constructed.

Finally, the authors are more transparent about the limitations of their evaluation methodology. In particular, they acknowledge the shortcomings of using LLM-as-a-Judge. This discussion improves the credibility of the paper and shows a more balanced assessment of the reported findings.

However, the paper still needs improvement in the following aspects:

As it was pointed out earlier, the paper does not show any ablation studies. It remains unclear which components contribute most to the observed improvements, e.g., graph-based indexing versus graph-guided retrieval.

I would also recommend that the authors look into the HHEM scoring model or semantic similarity instead of solely relying on LLM-as-a-judge.

The revised paper provides some additional discussion on retrieval strategies and computational considerations. For example, the authors describe replacing semantic clustering with HyDE in one application to reduce the cost of traversing large graph structures. These additions help readers better understand some of the practical design decisions behind the system.
However, the paper still lacks a systematic analysis of scalability and efficiency. There is no information about retrieval latency, response time, memory consumption, graph construction costs, or how the system performs as the size of the knowledge graph increases. While the paper mentions that graph construction can incur "significant upfront indexing costs," it lacks a concrete numerical comparison of the token costs or latency between traditional vector indexing and graph-based indexing. Since one of the main challenges in Graph RAG systems is maintaining efficient retrieval over large graphs, such information would provide valuable insights for practitioners considering real-world deployment. Overall, this concern has been only partially addressed.

The paper uses different evaluation frameworks for different application RAGAS for the CRR and ABEL apps and DeepEval for the GDPR Verifier. This inconsistency makes it difficult to compare the relative effectiveness of the FRAG-KEDA engine across different domains. Also, Table 7 proves how easily these results can be swayed by prompt engineering.

The revised version introduces additional methodological detail compared to the earlier draft. However, several key components required for full reproducibility are still missing. The authors have not released any of the following: source code, datasets, prompting templates, graph schema definitions, or evaluation scripts.

The revised manuscript broadens the discussion by presenting applications from multiple domains rather than focusing exclusively on a single scenario. Nevertheless, the settings remain tied to Fraunhofer collaborations. The paper would benefit from a more explicit discussion of which findings generalize across domains, which lessons are application-specific, and what prerequisites are required for successful Graph RAG deployment. The current manuscript remains primarily a collection of case studies rather than a systematic analysis of general Graph RAG principles.

The graph statistics are missing: number of entities, relations, graph density, ontology size, centrality, etc. Graph quality is a central issue in Graph RAG: an analysis of entity extraction errors, entity linking quality, relation extraction quality, and propagation of KG errors into generated answers would be beneficial.

Review #3
Anonymous submitted on 19/Aug/2026
Suggestion:
Accept
Review Comment:

# Summary
The paper presents seven applications that use Graph-RAG methods across heterogeneous domains (legal, finance, customer support, education, agriculture, and enterprise knowledge management). Specifically, each application is built using a shared, custom GraphRAG framework. A brief overview of this framework is provided, followed by introductions of each application. The work focuses on the knowledge organisation/KG construction step of these Graph-RAG systems and covers applications with and without prior schemas. A minimal evaluation is provided for most applications using private, synthetic data. Based on the implementation and evaluation insights and open challenges are discussed.

# Strengths
* Many heterogeneous applications are presented in sufficient detail to give a good overview of KG-RAG-based applications, which are especially interesting for practitioners.
* Overall, the paper is well written, follows a clear structure and motivates the presented use-cases of Graph-RAG well.
* The presented general insights and the comparision between applications could be a valuable resource for the development of further Graph-RAG applications.
# Weaknesses
* The framework as well as each application description only provide a top-level view and many details are missing. Consequently, the impact of the individual applications and the framework is likely to be low.
* No serious evaluation for any of these applications is provided. Due to this, the generated insights might not be generalizable and, at best, could count as anecdotal evidence in a scientific setting.

# Comments
* Is the RAGAS-generated dataset publicly available? What are the basic statistics of this data, e.g. number of samples, typical length, and size of the retrieved data? Confidence intervals might be helpful as well. That is relevant for all evaluations of this work. An evaluation based on unknown data cannot yield reproducible scientific results.
* Was the LLM-As-A-Judge framework reliable here? Was a subset manually evaluated?
* Tables 5 and 6 present an average response time, but this is only comparable if the same hardware is used for all models. For GPT4o, Mixtral and Gema2 27b, this is not the case.
* 5.3 mentions the demonstration of this application at Hannover Messe 2025 at the start of the evaluation. A demonstration is not an evaluation. I suggest removing this part or integrating it more similarly to the mention in 5.4 (as a factor in the impact of this application, not as a replacement for an evaluation).

## Minor Comments
* In table 3, Entity Retrieval and EF Retrieval seem to refer to same component. I suggest using one term consistently. What is "Clus" in this table?
* "4.3 Wenn Fraunhofer Wüsste -Wass Fraunhofer Weiß" (page 10, line 39) is not a very descriptive title, especially for non-German speakers.
* 2.1 provides unecessary details about LLMs. Most of the information is not relevant to the presented applications.

# Score
The resubmission as an application report addresses many of my initial concerns. Overall, the authors improved the structure and removed redundant parts from the previous version. Only the evaluations remain a major drawback of this work from my perspective. However, given the useful overview of a range of heterogeneous applications, I recommend publishing this paper.