Review Comment:
I thank the authors for considering my previous comments, and the manuscript has been improved in some of the aspects.
First, each application has a much more thorough description. The authors give more precise explanations of how the knowledge graphs are constructed, how retrieval is done, how the systems are implemented, and what beneficial impacts they have in real-life scenarios.
Second, a clearer explanation of the differences between schemaless and schema-first Graph RAG techniques is now provided. This is one of the paper's strongest points and gives readers a greater understanding of the various ways Graph RAG systems can be constructed.
Finally, the authors are more transparent about the limitations of their evaluation methodology. In particular, they acknowledge the shortcomings of using LLM-as-a-Judge. This discussion improves the credibility of the paper and shows a more balanced assessment of the reported findings.
However, the paper still needs improvement in the following aspects:
As it was pointed out earlier, the paper does not show any ablation studies. It remains unclear which components contribute most to the observed improvements, e.g., graph-based indexing versus graph-guided retrieval.
I would also recommend that the authors look into the HHEM scoring model or semantic similarity instead of solely relying on LLM-as-a-judge.
The revised paper provides some additional discussion on retrieval strategies and computational considerations. For example, the authors describe replacing semantic clustering with HyDE in one application to reduce the cost of traversing large graph structures. These additions help readers better understand some of the practical design decisions behind the system.
However, the paper still lacks a systematic analysis of scalability and efficiency. There is no information about retrieval latency, response time, memory consumption, graph construction costs, or how the system performs as the size of the knowledge graph increases. While the paper mentions that graph construction can incur "significant upfront indexing costs," it lacks a concrete numerical comparison of the token costs or latency between traditional vector indexing and graph-based indexing. Since one of the main challenges in Graph RAG systems is maintaining efficient retrieval over large graphs, such information would provide valuable insights for practitioners considering real-world deployment. Overall, this concern has been only partially addressed.
The paper uses different evaluation frameworks for different application RAGAS for the CRR and ABEL apps and DeepEval for the GDPR Verifier. This inconsistency makes it difficult to compare the relative effectiveness of the FRAG-KEDA engine across different domains. Also, Table 7 proves how easily these results can be swayed by prompt engineering.
The revised version introduces additional methodological detail compared to the earlier draft. However, several key components required for full reproducibility are still missing. The authors have not released any of the following: source code, datasets, prompting templates, graph schema definitions, or evaluation scripts.
The revised manuscript broadens the discussion by presenting applications from multiple domains rather than focusing exclusively on a single scenario. Nevertheless, the settings remain tied to Fraunhofer collaborations. The paper would benefit from a more explicit discussion of which findings generalize across domains, which lessons are application-specific, and what prerequisites are required for successful Graph RAG deployment. The current manuscript remains primarily a collection of case studies rather than a systematic analysis of general Graph RAG principles.
The graph statistics are missing: number of entities, relations, graph density, ontology size, centrality, etc. Graph quality is a central issue in Graph RAG: an analysis of entity extraction errors, entity linking quality, relation extraction quality, and propagation of KG errors into generated answers would be beneficial.
|