Review Comment:
The paper presents WikiBias, a framework intended to explore bias in the Wikimedia ecosystem by combining semantic modeling, Wikidata/Wikipedia data, biographical event extraction, and network analysis. The paper focuses on writers and studies how gender and origin interact in the representation and visibility of different groups, especially women and Transnational writers. The topic is important and timely, since Wikimedia resources are widely used and can propagate existing biases.
Overall, the paper is generally well written and addresses a relevant problem. The authors have also carried out an extensive set of experiments in several directions: ontology modeling, biographical event extraction, entity linking, knowledge graph augmentation, representational bias analysis, and network analysis. The results are interesting and relevant, especially those showing that extracted biographical predicates can reveal gendered patterns in the representation of writers. In particular, the paper gives evidence that women writers are more often associated with private-life events such as marriage and parenthood, while men are more often associated with public or career-related events.
The paper presents many components and experiments, but several central methodological details are not sufficiently clear. In particular, the framework itself is not described as a clear reusable framework, the ontology is under-specified, the event detection model is not described in enough detail, and the external resources do not seem sufficient to reproduce the experiments. Please carefully read the following concerns:
1- The title already says a "framework" as the main contribution, but the paper does not clearly describe the framework as a structured system. The components are discussed across different sections, but there is no overall architecture showing the modules, inputs, outputs, intermediate resources, and expected use. A framework paper should make it clear how the different parts fit together (source data, ontology, knowledge graph, extraction pipeline, bias analysis, and network analysis). A diagram and a clearer description of the workflow/achitecture would substantially improve the paper.
2- The ontology description is under-specified. The paper says that the People in the Media Ontology is based on DOLCE and Ontology Design Patterns, but the ontology is introduced mostly through examples around pim:BiographicalSituation. It is not clear how many new classes and properties were created, which axioms are used, and how roles are modeled across different situations. For example, the current explanation suggests that a person has a role, but a person can have different roles in different biographical situations. The authors should clarify the ontology design, provide more details about the classes and properties, and explain the modeling choices more explicitly.
3- The event detection model is not sufficiently described. The paper spends considerable space describing the corpus, training sets, and evaluation results, but the actual model architecture and training setup are not clear enough. The authors should explain which model was used, how it was trained, what the input and output format is, which hyperparameters were used, and how the development set was used. It is also unclear why there is no cross-validation experiment, given the relatively small size of the annotated data used for training, development, and testing.
4- The choice to work only with English Wikipedia pages is a major limitation for a paper about bias and underrepresentation. The paper starts from a much larger set of writers, but then keeps only writers with an English Wikipedia page. This removes a large part of the original set (61%) and may introduce an important language and cultural bias. Many Transnational writers may have pages only in their native language, and these writers are excluded from the extraction pipeline. This limitation should be quantified and discussed in more depth. If possible, discuss whether multilingual Wikipedia extraction could reduce this bias.
5- The definition of "Transnational" and the operationalization of underrepresentation need more analysis. The classification depends on several things (former-colony status, HDI, passport mobility, ethnic minority status, and so on). However, the paper does not analyze how much each parameter affects the final classification. A sensitivity analysis would be useful to show whether the results are stable when these parameters change. The criterion about belonging to an ethnic minority in a Western country also needs clarification: which Wikidata properties are used, and how are missing or noisy ethnicity data handled?
6- The entity linking pipeline should be justified in more detail. The paper says that the pipeline combines neural methods and heuristics, and it describes the main steps of the process, including PERSON detection, Wikipedia API search, string-similarity filtering, and Wikidata ID retrieval. However, the choices behind this pipeline are not sufficiently justified or compared with alternatives. Since the extraction is based on Wikipedia pages, the authors should also justify why they used Wikipedia API search and string similarity instead of comparing with existing entity linking tools such as DBpedia Spotlight, which is built using resources derived from Wikipedia and its entities can often be mapped to Wikidata entities. DBpedia Spotlight could be considered as a baseline or alternative for the entity-linking step, especially for future multilingual extensions, as there are versions trained in different languages (https://demo.dbpedia-spotlight.org/). This would not solve the full multilingual extraction problem, but it could help with the entity-linking part.
7- The analysis of results would benefit from more qualitative validation of the augmented triples and their usefulness. The quantitative increase in triples is large, but it is not always clear how meaningful the new triples are. Although the paper includes some qualitative examples, more systematic validation would be useful. For example, the authors could analyze whether some extracted relations correspond to knowledge that previously existed in Wikidata history, whether they represent genuinely missing knowledge, or whether they are relations outside the current Wikidata taxonomy. This would help assess the practical value of the knowledge graph augmentation.
8- The representational bias analysis needs clearer explanation of normalization. In Table 5, it is not clear whether the most frequent predicates are ranked by raw counts or normalized by the size of each socio-demographic group. Since the groups have very different sizes, raw counts may not reflect the actual representativity of each predicate within each group. Table 6 partially addresses this through Jensen-Shannon Divergence, but the paper should make the methodology clearer.
9- The external resources and code need substantial improvement. The paper presents many experiments and many derived datasets (subsets of writers, training and test configurations, extracted triples, socio-demographic labels, manual evaluation samples, and network analysis data). However, the linked GitHub repository does not appear to contain all the resources, code, intermediate datasets, and outputs needed to reproduce the paper. This is a major reproducibility issue. The authors should make available all resources and scripts used in the different sections of the paper, or clearly explain how they can be regenerated. At minimum, the repository should include the exact data splits, writer lists and labels, the extracted triples, manual evaluation data, scripts for producing the reported tables, configuration files, dependency versions, and clear instructions.
10- The related work section should better position the contribution of the paper. At the moment, it mostly lists related papers one after another. It would be helpful to include a clearer comparison showing what each related work studies, which type of bias it addresses, which dataset it uses, and how the proposed work differs from it. A comparison table would improve readability and make the contribution clearer.
11- The paper would benefit from stronger comparison with existing methods in the literature. Many of the reported results are computed using the authors’ own pipeline, datasets, subsets, and experimental variants. This is understandable for a framework paper, but some components could still be compared with existing solutions or baselines. For example, the entity linking step could be compared with established entity linking systems.
List of minor issues:
- Page 2, line 2: The word "demonstrated" may suggest mathematical proof. A softer wording such as "shown", "reported", or "provided evidence" would be better.
- Page 2, line 24: "from all the English Wikidata pages" -> "from all the English Wikipedia pages".
- Page 3, lines 8–10: The related work discussion does not mention the relevant results of the cited work. Please add the main finding or explain why the cited work is relevant.
- Related work section: Citation style is inconsistent. Sometimes the text uses "Author et al. [34]", while in other places the sentence starts directly with "[34]". Please make this consistent.
- Page 4, lines 36–37: The sentence "They were not born before..." is hard to read. Consider rewriting as "They were born after..." or "They were born in or after...".
- Page 4, lines 49–51: URLs should include the date of last access.
- Section 3.3: At the time of review, the SPARQL endpoint (https://kgccc.di.unito.it/sparql/wikibias) was not available for testing and returned "Service Unavailable".
- Section 4: Please clarify the role of the development set.
- Section 5.2, page 12, line 21: "To do si" -> "To do so".
- Section 5.3: More detail is needed on how the "knowledge graph based on Wikidata relationships" was extracted.
- Page 14, line 50: "connections within a each group" -> "connections within each group".
- Table 8: To facilitate comparison, all numbers should be written using the same scientific notation format.
- Conclusion, page 16, lines 46–49: The sentence beginning "Augmenting the triples about them..." seems to lack a connector, please check it.
Overall, I found the paper very interesting and with promising results. The main problems I see are: (1) the framework as an artifact should be the central part of the paper; (2) the methods, algorithms, and approaches used in each experiment should be better justified, detailed, and compared with existing solutions where possible; and (3) reproducibility should be substantially improved, as the GitHub repository lacks most of the resources and code needed to generate the reported results.
|