Review Comment:
This paper presents a systematic literature review (SLR) of methods for extracting metadata from tables and charts in scientific publications. Following the Kitchenham et al. methodology for systematic reviews, the authors searched across five scholarly databases and four aggregator sources, retrieved 8,398 records, and arrived at a final corpus of 68 peer-reviewed papers after title, abstract, and full-text screening. The review is structured around seven research questions covering extraction approaches and tools, model architectures, PDF-to-table conversion pipelines, chart data extraction, chart-to-table reconstruction, variable interpretation, and cross-modal alignment. The authors propose an evaluation framework that annotates each paper across six characteristics: modality coverage, task coverage, model type, dataset and benchmark use, evaluation metrics, artifact availability, and level of extracted information. The paper concludes that structural detection is well addressed by current systems, while variable interpretation and multimodal alignment between tables and charts remain substantially open problems.
* The decision to analyze table and chart extraction within a single review is the most distinctive feature of this manuscript.
* The paper consistently situates its findings within the context of knowledge graph generation from scientific publications. This contextualization is valuable and positions the SLR as more than a catalog of extraction methods: it argues that the limitations identified have direct consequences for the reproducibility and completeness of scientific knowledge bases. This framing is appropriate for the SWJ audience.
* The seven research questions are specific, well-motivated, and cover the extraction pipeline from low-level detection through to multimodal alignment. The six-dimensional evaluation framework provides a systematic basis for comparison that goes beyond listing methods and their scores.
* The paper reports the full set of keyword combinations, the databases queried, and the number of records retrieved at each filtering stage. The deposit of the complete annotation table on Zenodo and the availability of processing scripts on GitHub are a good thing nd make the review more reproducible than many published SLRs.
The main concerns are methodological.
* The annotation of 68 papers across six characteristics was performed manually, but the paper does not state how many annotators were involved, does not report any inter-rater reliability measures , and does not describe a procedure for resolving disagreements. The Kitchenham methodology, which the paper explicitly claims to follow, requires at least two independent reviewers for screening and a reliability assessment for annotation. This is not a minor procedural point: it affects whether the findings of the review can be considered systematic or whether they reflect the judgments of a single analyst.
* Any SLR is expected to follow the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) or an equivalent reporting standard. The paper does not include a PRISMA flow diagram showing the number of records identified, screened, assessed for eligibility, and included, with reasons for exclusion at each stage. The exclusion criteria are defined, but the paper does not report how many papers were excluded at the full-text stage and for which reasons. This makes it difficult to assess whether the corpus is well-defined and replicable. A PRISMA-style flow diagram should be added as a figure. The number of papers excluded at full-text review and the primary reasons for exclusion should be tabulated.
* Several tables present counts that require clarification. Table 10 (techniques for table conversion from PDF) states that visual layout analysis appears in 45 studies, but the note at the bottom says a single study may contribute to multiple techniques. Given that the full corpus is 68 papers, of which only 28 focus on tables and 8 on both modalities, 45 seems too large unless studies focused on charts are also counted for table conversion techniques. The authors should clarify the denominator for each count in this table. A similar issue arises with Table 11 for chart extraction tasks.
* Given that this is a submission to the Semantic Web Journal, the treatment of semantic table interpretation and ontology-based approaches is surprisingly thin. The paper acknowledges that prior work on lifting tabular data to RDF, semantic annotation of tables, and knowledge graph population exists, citing relevant references, but these are discussed only briefly in the context of some research questions. None of these semantic approaches appears to be included in the 68 reviewed papers, which suggests either that they were excluded because they do not focus on the extraction step, or that the search strategy missed them. Clarification about these issues is needed.
* The discussion section identifies several important themes but does not systematically connect findings back to the individual research questions. For a systematic review, the norm is to provide a synthesized answer to each RQ and then discuss broader implications. The current structure of the discussion, organized around thematic challenges (multimodal alignment, reproducibility, interpretability, knowledge integration), is reasonable but makes it harder to see whether each RQ has been adequately addressed. A brief table summarizing the answer to each RQ, with references to the supporting evidence, would improve the structure.
* The recommendations for future research, including unified multimodal architectures, new benchmarks, and better evaluation practices, are reasonable but not specific enough to be actionable. The paper does not articulate what a new benchmark should measure that existing ones do not, or what kind of evaluation metric would capture correctness at another level
Minor issue
* Reference list entries should be checked for completeness and conformity to journal format. In particular, check the entries for Cedeno and Garijo 2025a and 2025b, which do not seem to conform to the journal citation format.
|