Metadata Extraction from Tables and Charts in Scientific Publications: A Systematic Literature Review

Tracking #: 4000-5214

Authors: 
Erick Cedeño
Daniel Garijo
Oscar Corcho

Responsible editor: 
Guest Editors ML and KR 2025

Submission type: 
Survey Article
Abstract: 
Extracting metadata from tables and charts in scientific publications is essential for enabling structured knowledge representation, automated retrieval of experimental results, and large-scale evidence synthesis. Such metadata includes descriptive information that goes beyond structural parsing and captures the roles and relationships of elements such as headers, units, variables, legends, and axis labels. Although numerous methods have been proposed for table extraction and chart understanding, prior work remains fragmented, meaning that most approaches are developed and evaluated independently for a single modality and therefore lack mechanisms for connecting information across tables and charts. This specific approach to each modality has led to heterogeneous processes and inconsistent assessment practices, limiting comparability across studies. In this systematic literature review, we analyze 68 peer-reviewed studies using a unified evaluation framework specifically designed to examine metadata extraction capabilities in both tables and charts. Guided by explicit research questions, we compare these systems in terms of their task definitions, model architectures, metadata outputs, and reporting practices. Rather than reproducing implementations, our analysis evaluates the extent to which each method supports metadata identification, variable interpretation, and multimodal alignment. The findings highlight unresolved challenges in linking related information across modalities (e.g., associating table headers with chart axes), interpreting variables beyond their superficial textual labels, and establishing standardized benchmarks that measure correctness at the metadata level.
Full PDF Version: 
Tags: 
Reviewed

Decision/Status: 
Major Revision

Solicited Reviews:
Click to Expand/Collapse
Review #1
By ISMAILOVA Nigar submitted on 04/May/2026
Suggestion:
Minor Revision
Review Comment:

The paper addresses a relevant and timely topic and provides a useful overview of existing work. The structure based on research questions is clear and existence of summary tables are useful for readers. The manuscript includes a substantial number of recent scientific works, indicating good coverage of the literature. The paper also highlights important challenges, which points to meaningful future research directions. However, a deeper analytical comparison of methods besides the task-level descriptions and an improved taxonomy which supports systematic evaluation and visual representation of taxonomy would improve the quality of the paper
6

Review #2
Anonymous submitted on 20/May/2026
Suggestion:
Minor Revision
Review Comment:

The objective of this paper is to provide a broad literature review of metadata extraction methods from tables and charts. The authors analyze and compare recent approaches, tools , datasets , and model architectures used for extracting structured information from systematic and graphical content. The paper provides a discussion on current limitations and forthcoming directions for multimodal and explicable technological document analysis .

Here I detail pro and con that need to be addressed in the revised version of the paper

Pros:

1. The paper provides a review of both table and chart extraction methods rather of treating them separately.
2. The evaluation framework is very well detailed allowing a comparison across 68 studies.
3. The paper clearly identifies important research challenges, especially in multimodal alignment and benchmark standardization.

Cons:

1. The definition of “metadata” is sometimes broad and would benefit from a clearer formal definition.
2. The multimodal analysis could be improved, despite being presented as a central motivation of the paper. Could the authors further develop and compare present strategies for table-chart alignment and multimodal integration?
3. Several sections are insistent and too detailed, which affects readability and conciseness . For example , discussions about datasets, benchmarks , and extraction pipelines are sometimes repeated across dual sections with modest added value . Some methodological descriptions are overly described and could be summarized in relative tables or synthesis figures .

Additional comments

Table 9 could also be improved by organizing the approaches into clearer families of models (e.g., CNN-based, Transformer-based, graph-based, LLM-based). In particular, why do no Vision-Language Model (VLM) approaches appear in the comparison, despite their growing importance in multimodal document understanding?

Review #3
Anonymous submitted on 20/May/2026
Suggestion:
Major Revision
Review Comment:

This paper presents a systematic literature review (SLR) of methods for extracting metadata from tables and charts in scientific publications. Following the Kitchenham et al. methodology for systematic reviews, the authors searched across five scholarly databases and four aggregator sources, retrieved 8,398 records, and arrived at a final corpus of 68 peer-reviewed papers after title, abstract, and full-text screening. The review is structured around seven research questions covering extraction approaches and tools, model architectures, PDF-to-table conversion pipelines, chart data extraction, chart-to-table reconstruction, variable interpretation, and cross-modal alignment. The authors propose an evaluation framework that annotates each paper across six characteristics: modality coverage, task coverage, model type, dataset and benchmark use, evaluation metrics, artifact availability, and level of extracted information. The paper concludes that structural detection is well addressed by current systems, while variable interpretation and multimodal alignment between tables and charts remain substantially open problems.
* The decision to analyze table and chart extraction within a single review is the most distinctive feature of this manuscript.
* The paper consistently situates its findings within the context of knowledge graph generation from scientific publications. This contextualization is valuable and positions the SLR as more than a catalog of extraction methods: it argues that the limitations identified have direct consequences for the reproducibility and completeness of scientific knowledge bases. This framing is appropriate for the SWJ audience.
* The seven research questions are specific, well-motivated, and cover the extraction pipeline from low-level detection through to multimodal alignment. The six-dimensional evaluation framework provides a systematic basis for comparison that goes beyond listing methods and their scores.
* The paper reports the full set of keyword combinations, the databases queried, and the number of records retrieved at each filtering stage. The deposit of the complete annotation table on Zenodo and the availability of processing scripts on GitHub are a good thing nd make the review more reproducible than many published SLRs.

The main concerns are methodological.
* The annotation of 68 papers across six characteristics was performed manually, but the paper does not state how many annotators were involved, does not report any inter-rater reliability measures , and does not describe a procedure for resolving disagreements. The Kitchenham methodology, which the paper explicitly claims to follow, requires at least two independent reviewers for screening and a reliability assessment for annotation. This is not a minor procedural point: it affects whether the findings of the review can be considered systematic or whether they reflect the judgments of a single analyst.
* Any SLR is expected to follow the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) or an equivalent reporting standard. The paper does not include a PRISMA flow diagram showing the number of records identified, screened, assessed for eligibility, and included, with reasons for exclusion at each stage. The exclusion criteria are defined, but the paper does not report how many papers were excluded at the full-text stage and for which reasons. This makes it difficult to assess whether the corpus is well-defined and replicable. A PRISMA-style flow diagram should be added as a figure. The number of papers excluded at full-text review and the primary reasons for exclusion should be tabulated.
* Several tables present counts that require clarification. Table 10 (techniques for table conversion from PDF) states that visual layout analysis appears in 45 studies, but the note at the bottom says a single study may contribute to multiple techniques. Given that the full corpus is 68 papers, of which only 28 focus on tables and 8 on both modalities, 45 seems too large unless studies focused on charts are also counted for table conversion techniques. The authors should clarify the denominator for each count in this table. A similar issue arises with Table 11 for chart extraction tasks.
* Given that this is a submission to the Semantic Web Journal, the treatment of semantic table interpretation and ontology-based approaches is surprisingly thin. The paper acknowledges that prior work on lifting tabular data to RDF, semantic annotation of tables, and knowledge graph population exists, citing relevant references, but these are discussed only briefly in the context of some research questions. None of these semantic approaches appears to be included in the 68 reviewed papers, which suggests either that they were excluded because they do not focus on the extraction step, or that the search strategy missed them. Clarification about these issues is needed.
* The discussion section identifies several important themes but does not systematically connect findings back to the individual research questions. For a systematic review, the norm is to provide a synthesized answer to each RQ and then discuss broader implications. The current structure of the discussion, organized around thematic challenges (multimodal alignment, reproducibility, interpretability, knowledge integration), is reasonable but makes it harder to see whether each RQ has been adequately addressed. A brief table summarizing the answer to each RQ, with references to the supporting evidence, would improve the structure.
* The recommendations for future research, including unified multimodal architectures, new benchmarks, and better evaluation practices, are reasonable but not specific enough to be actionable. The paper does not articulate what a new benchmark should measure that existing ones do not, or what kind of evaluation metric would capture correctness at another level

Minor issue
* Reference list entries should be checked for completeness and conformity to journal format. In particular, check the entries for Cedeno and Garijo 2025a and 2025b, which do not seem to conform to the journal citation format.