Review Comment:
Summary
------
The paper proposes an integrated semantic framework for analyzing the scoliosis research literature. Starting from 31,836 English-language abstracts indexed in Scopus (2000–2025), the authors apply BERTopic to identify latent research themes, construct a topic-aware property graph through automated head-relation-tail extraction, quantify inter-topic semantic proximity via cosine similarity on centroid embeddings, and trace temporal trends in knowledge intensity across year-binned periods. A stratified validation of 250 extracted triplets (50 per predicate across five relation types) yields precision estimates ranging from 0.78 (Treats) to 0.92 (Associated_with). The integration of implicit topic structure with explicit relational representation is presented as the paper's primary methodological contribution, with scoliosis serving as the application domain. The motivation is sound and the corpus scale is a genuine asset. Overall, I found the paper interesting and well organized. However, my principal reservations concern the systematic omission of quantitative details that make reproducibility impossible, a validation design that does not establish the credibility of its precision claims, the absence of any comparative baseline, a purely qualitative temporal analysis, and the near-total absence of the resource metadata required by the journal's long-term stable resource criteria. None of these are insurmountable, but they must be addressed before the paper can be accepted.
Strengths
------
The corpus scale - 31 836 abstracts spanning 25 years - is a genuine asset and positions the work well as a large-scale semantic analysis rather than a proof of concept. The decision to combine topic modeling with explicit relational extraction is conceptually well motivated: the paper correctly identifies that distributional representations alone do not encode typed semantic links, and that knowledge graphs alone do not capture latent thematic organization. The validation design distinguishes explicitly between incorrect and ambiguous relations rather than conflating uncertainty with error, which is methodologically honest and aligned with best practices in information extraction evaluation. The temporal perspective - examining how thematic intensity and diversity evolved across decades - adds a longitudinal dimension that is absent from most prior semantic bibliometric work. The Sankey diagram (Figure 7) is an effective visualization device that integrates time, topic, and journal dimensions in a single view.
Major limits
------
1. Several methodological parameters that are essential for reproducibility are either missing or insufficiently emphasized in the manuscript. In particular, the final number of retained topics produced by BERTopic is not explicitly stated in the main text, the sentence embedding model is described only generically ("a sentence-level transformer embedding model"), the total size of the resulting knowledge graph is not reported, and the cosine similarity threshold used in Section 2.4 is discussed qualitatively without providing its numerical value. Likewise, the temporal aggregation used in Section 2.5 can only be inferred from Figure 7 rather than being explicitly described. Reporting these parameters explicitly would greatly improve the reproducibility of the proposed framework.
2. Section 2.6 mentions "annotators" in the plural but never states how many annotators evaluated the triplets, what their domain expertise was, or whether inter-annotator agreement was measured. For a paper presenting precision values as evidence of moderate-to-high reliability, this is a fundamental gap. A single judge with high familiarity with the source material would inflate precision on precisely the relation types that are most ambiguous. Authors should report the annotation design (number of judges, qualification, instructions), and provide at least one agreement metric (Cohen’s κ or Krippendorff’s α) on a common subset. If the validation was effectively single-annotator, this must be acknowledged as a limitation bearing directly on the precision figures.
3. The reported accuracy values are difficult to interpret without further context. Although relation-specific accuracy values are provided (0.78-0.92), the article does not explain how these results compare with those reported in previous biomedical relation extraction studies or in comparable semantic exploration frameworks from the literature. Such a comparison would allow readers to understand whether the reported performance reflects the expected level of accuracy for this type of task or represents a significant improvement. Even if a direct experimental comparison is not possible, a discussion situating these results within the existing literature would greatly strengthen the evaluation.
4. Section 3.4 reports observations such as "a gradual but consistent increase" and "accelerated growth trajectories" without any statistical test (e.g., Mann-Kendall, Spearman rank correlation, or change-point detection). The "intensity proxy" defined as the count of extracted triplets per bin is novel but not validated against any external measure of knowledge richness. Without statistical backing, the temporal claims reduce to visual inspection of Figure 6. Authors should apply appropriate trend-detection methods to distinguish genuine directional trends from noise, and should validate or at least discuss the assumptions underlying the triplet-count proxy.
5. Section 2.3 describes the extraction technique as "rule-based and dependency-aware natural language processing technique" without naming the NLP framework (spaCy, CoreNLP, etc.), the dependency patterns or rules used to identify each relation type, or any error analysis at the extraction stage prior to validation. The five relation types (Treats, Used_for, Associated_with, Measured_by, Predicts) are introduced without operational definitions. Readers cannot assess whether the patterns are specific enough to avoid spurious triplets, whether they are appropriate for abstract-level text, or whether the schema is grounded in any biomedical ontology. A dedicated subsection describing the extraction rules and any ontological grounding (UMLS, SNOMED-CT, MeSH) is needed.
6. The integration between BERTopic and the knowledge graph is well illustrated throughout the manuscript and constitutes the main conceptual contribution of the work. However, its added value over applying topic modelling and knowledge graph construction independently is never quantitatively demonstrated. For example, the manuscript does not investigate whether the topic-aware graph improves the interpretability of the extracted relations, facilitates exploration, or provides better analytical capabilities than a conventional knowledge graph. A dedicated experiment or, at least, a more explicit discussion of the practical benefits brought by the topic-aware representation would considerably strengthen the central contribution of the paper.
7. Figures 2 and 4 presumably display topic labels and IDs, but the text of the paper never states how many topics BERTopic produced, what labels were assigned to them, or what the five largest topics are. Figure 7 references labels T0-T10, implying at most 11 topics, which is inconsistent with the "top-20 topics" used for the similarity analysis in Section 3.3. Authors should include a table listing all topics (ID, label, document count, coherence score) and resolve the apparent inconsistency between the top-20 mention and the T0-T10 labeling in Figure 7.
Minor limits
------
The sentence opening Section 3.4 reads "o examine how scoliosis research has evolved" - the initial "T" is missing. This must be corrected.
The abstract states that validation "demonstrated moderate-to-high precision across relation categories" without giving the numerical range. Given that the abstract is the only part of the paper many readers will see, the range (0.78-0.92) should be reported there.
Section 2.4 says the threshold was chosen based on "empirical inspection of network density, connectedness, and the emergence of visually and analytically coherent communities." This is a reasonable heuristic, but the specific threshold value and the resulting network density (edges retained/total possible edges) should be stated so that other researchers can reproduce the network.
The outlier category (topic −1) is mentioned briefly: "outlier documents were excluded from subsequent analyses." The number of outlier documents should be reported alongside the 31,836 total so readers know what fraction of the corpus was excluded.
The co-authorship network analysis (Section 3.5 and Figure 9) is introduced and briefly described but then plays no role in the discussion or conclusion. If it is not analytically central, it should either be developed into a substantive analysis with network statistics (degree distribution, clustering coefficient, identified communities) or removed to sharpen the paper's focus.
Assessment of Long-term Stable Resources
------
A. Availability. I could not identify any persistent URL or repository associated with the proposed resource in either the manuscript or the supplementary material. Section 2.6 notes that preprocessing and validation scripts were implemented in Python, but no code repository is mentioned. Without a publicly accessible location, the primary resource claimed by the paper is inaccessible to readers. A citable, stable repository (e.g., Zenodo, GitHub with a tagged release, or an institutional SPARQL endpoint) must be provided before the resource can be assessed.
B. Licensing. No license is specified for any component — the knowledge graph, the code, or any derived dataset. Given that the underlying abstracts are Scopus-indexed and may be subject to commercial copyright, authors must clarify what they are actually permitted to distribute and under which license.
C. Standards compliance and format. Section 2.3 states that the property graph was designed to be "compatible with Resource Description Framework (RDF) principles," but no actual RDF serialization is provided, no ontology or vocabulary alignment is described (no SKOS for topics, no UMLS for relation types), and no SPARQL endpoint or VOID descriptor is mentioned. For a paper submitted to the Semantic Web Journal, the gap between "designed to be compatible" and "serialized and queryable as RDF" is significant. Authors should either provide a concrete RDF dump with documented prefixes and schema, or revise the resource claims accordingly.
Typos and Editorial Suggestions
------
"o examine how scoliosis research has evolved" "To examine how scoliosis research has evolved" (Section 3.4, first sentence)
The reference list mixes DOI formats: some entries use "doi: 10.XXX" and others "Available: https://arxiv.org/..." — standardize to one format per entry type.
Reference [9] (Angelov, Top2Vec) is cited in Section 2.2 to support BERTopic's use of contextual embeddings but Top2Vec is a competing method, not supporting evidence; replace with a more appropriate citation for transformer-based topic modeling advantages.
Overall Impression
------
The paper addresses a real need - scalable semantic analysis of domain-specific biomedical literature - and the combination of BERTopic with knowledge graph construction is conceptually appealing. The corpus scale and the explicit handling of ambiguity in validation are genuine positives. However, the paper currently presents the framework at a level of description that cannot support the reproducibility, comparability, or resource-contribution claims that a Semantic Web Journal submission requires. The omission of basic quantitative parameters (topic count, embedding model, KG size, threshold) across every component is systematic rather than incidental, and it must be addressed comprehensively. The validation, as described, is insufficient to establish the reliability of the extraction pipeline to the standard expected for a resource paper. The paper would benefit substantially from a revision that adds one comparative experiment, tightens the resource description to meet FAIR standards, and completes the missing quantitative reporting.
|