Metadata Extraction from Tables and Charts in Scientific Publications: A Systematic Literature Review

Tracking #: 4118-5332

Authors: 
Erick Cedeño
Daniel Garijo
Oscar Corcho

Responsible editor: 
Guest Editors ML and KR 2025

Submission type: 
Survey Article
Abstract: 
Extracting metadata from tables and charts in scientific publications is essential for enabling structured knowledge representation, automated retrieval of experimental results, and large-scale evidence synthesis. Such metadata includes descriptive information that goes beyond structural parsing and captures the roles and relationships of elements such as headers, units, variables, legends, and axis labels. Although numerous methods have been proposed for table extraction and chart understanding, prior work remains fragmented, meaning that most approaches are developed and evaluated independently for a single modality and therefore lack mechanisms for connecting information across tables and charts. This specific approach to each modality has led to heterogeneous processes and inconsistent assessment practices, limiting comparability across studies. In this systematic literature review, we analyze 68 peer-reviewed studies using a unified evaluation framework specifically designed to examine metadata extraction capabilities in both tables and charts. Guided by explicit research questions, we compare these systems in terms of their task definitions, model architectures, metadata outputs, and reporting practices. Rather than reproducing implementations, our analysis evaluates the extent to which each method supports metadata identification, variable interpretation, and multimodal alignment. The findings highlight unresolved challenges in linking related information across modalities (e.g., associating table headers with chart axes), interpreting variables beyond their superficial textual labels, and establishing standardized benchmarks that measure correctness at the metadata level.
Full PDF Version: 
Tags: 
Reviewed

Decision/Status: 
Accept

Solicited Reviews:
Click to Expand/Collapse
Review #1
By ISMAILOVA Nigar submitted on 16/Sep/2026
Suggestion:
Accept
Review Comment:

This manuscript was submitted as 'Survey Article' and should be reviewed along the following dimensions: (1) Suitability as introductory text, targeted at researchers, PhD students, or practitioners, to get started on the covered topic. (2) How comprehensive and how balanced is the presentation and coverage. (3) Readability and clarity of the presentation. (4) Importance of the covered material to the broader Semantic Web community. Please also assess the data file provided by the authors under “Long-term stable URL for resources”. In particular, assess (A) whether the data file is well organized and in particular contains a README file which makes it easy for you to assess the data, (B) whether the provided resources appear to be complete for replication of experiments, and if not, why, (C) whether the chosen repository, if it is not GitHub, Figshare or Zenodo, is appropriate for long-term repository discoverability, and (4) whether the provided data artifacts are complete. Please refer to the reviewer instructions and the FAQ for further information.

Review #2
Anonymous submitted on 19/Sep/2026
Suggestion:
Accept
Review Comment:

Originality. The review's contribution, a joint systematic treatment of table and chart metadata extraction rather than the usual single-modality survey, remains the paper's main value. The revision strengthens this by adding a five-dimension categorization (input modality, extraction level, model family, output representation, integration capability, Table 6), which gives the comparison more analytical structure than the original task-level listing.

Significance. The response to Reviewer 3 on quantitative errors is the most consequential change. The authors re-examined the annotations and corrected the denominators and counts in the PDF table-conversion techniques table (previously reporting 45 studies against a base that could not support that count; now correctly scoped to n=17, with revised per-technique counts) and clarified the chart-extraction task table's denominator (n=40). Reporting an error of this kind, rather than only adjusting the surrounding prose, is a meaningful correction that affects the credibility of the quantitative findings. The added PRISMA-style flow diagram and the explicit accounting of two-reviewer annotation and disagreement resolution (Reviewer 3's methodological concern) bring the review closer to standard SLR reporting practice. The new RQ-summary table (Table 15) and the expanded, more concrete benchmark/metric recommendations (variable-value mapping accuracy, provenance accuracy, semantic alignment accuracy) directly answer the reviewers' calls for actionable specificity.

Quality of writing. Clear and well organized; the response letter's transparent, itemized handling of every comment, including an explicit acknowledgment that no formal inter-rater reliability metric (e.g., Cohen's kappa) was computed, is exactly the kind of forthright methodological reporting expected in a resubmission. The scope clarification separating primary extraction systems from downstream Semantic Web work (entity linking, ontology alignment, KG population) resolves what had been a genuine ambiguity for a Semantic Web Journal submission.