Pith. sign in

REVIEW 5 major objections 4 minor 78 references

Wikidata from a Research Perspective -- A Systematic Mapping Study of Wikidata

T0 review · 5 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A systematic review of 67 peer-reviewed papers argues that Wikidata research is concentrated on data quality and in Europe, with multilingualism, usability, and knowledge diversity standing out as the clearest gaps.

desk verdict Useful first map of Wikidata research, with a real but manageable sampling caveat; worth refereeing after the promised data links and flowchart arithmetic are fixed. read the letter →

arxiv 1908.11153 v2 pith:XATYOCTJ submitted 2019-08-29 cs.DL cs.CYcs.SI

classification cs.DLcs.CYcs.SI
keywords Wikidatasystematicmappingstudyresearchclassificationknowledgegraphdataqualitymultilingualismpeerproduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to establish that a systematic review of 67 peer-reviewed papers can serve as a reliable map of the first six years of research on Wikidata, the collaborative knowledge base behind Wikipedia. The authors gathered papers using a single search term in four academic search engines, filtered out non-English work, short papers, and publications outside journals and conference proceedings, then sorted what remained into five categories. On that basis they argue that Wikidata research is growing year by year, that it is concentrated in Europe, and that its most mature topic is data quality, while multilingualism, knowledge diversity, user-interface usability, and use in non-technical disciplines remain underexplored. A reliable map matters because researchers need to know which parts of Wikidata have already been studied before choosing a new question.

What carries the argument

The machinery that carries the argument is the classification scheme built from a systematic reading of the 67 papers. The authors collected candidates with the single search term “Wikidata” in four academic search engines, removed duplicates, non-English results, non-papers, work shorter than five pages, and publications outside journals or conference proceedings, then labeled the remaining papers and grouped them by constant comparison into five categories: community-oriented, engineering-oriented, application use cases, knowledge-graph-oriented, and data-oriented research. This taxonomy is what turns a list of references into a map, because a paper's placement determines where the authors report density and where they see clear gaps.

What would settle it

Re-run the same procedure for October 2012 to June 2018 with a broader set of related search terms, without the five-page minimum, and including non-English publications; if the five categories change materially or new high-volume topics appear, the paper's claim that its 67-paper sample maps the field would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that the 67 selected papers, read and classified carefully, show the state of Wikidata research between October 2012 and June 2018. Descriptively, conference papers outnumber journal articles, and most first authors are based in Europe, with 39 of the 67 papers coming from 2017 and the first half of 2018. Topically, the data-oriented category is the largest, with 22 papers, followed by knowledge-graph-oriented research with 15 and community-oriented research with 14, while engineering-oriented research and application use cases account for 9 and 7 papers respectively. From that distribution the authors identify the field's white spots: uneven language coverage, little study of what plurality and contradictory claims do to trust, few usability studies, and applications concentrated mainly in biomedicine and linguistics. They present these gaps as directions for future work rather than as failures of Wikidata itself.

Load-bearing premise

The load-bearing premise is that the 67 filtered papers fairly represent all high-quality Wikidata research; if substantial work was published under other search terms, in other languages, in short-paper form, or in venues outside the four search engines used, the map and its reported gaps would shift.

Editorial extensions

If this is right

  • If the map is accurate, Wikidata research is expanding quickly: 39 of the 67 papers appeared in 2017 or the first half of 2018, and the 12 papers logged by June 2018 point to a projected 40–50 full papers for that year.
  • Data quality is the field's most developed strand, covering completeness, references, provenance, and vandalism, so new researchers entering this area have a firm baseline rather than an open field.
  • The clearest white spots are language coverage, the effect of contradictory claims on trustworthiness, user interface usability, and use cases outside biomedicine and linguistics, which the paper names as recommended future directions.
  • Because most first authors are based in Europe, the current research perspective is geographically concentrated, and the paper suggests this Western perspective may limit how knowledge diversity in Wikidata is studied.
  • The five-category classification offers a reusable skeleton for later mapping studies, since it labels both what has been done and what has not.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The use of a single search term means a follow-up search with related software and data-service names would likely catch additional papers; my expectation is that it would enrich rather than overturn the five categories.
  • The exclusion of short papers may have removed early exploratory results, such as system demonstrations, which are often where new community and engineering topics first appear.
  • The paper's projection of 40–50 full papers for 2018 is directly testable by counting papers in late-2018 indexes and comparing the result with the growth curve it reports.
  • Comparing this Europe-heavy profile with the more global research community around Wikipedia would clarify whether the concentration is a property of Wikidata or of peer-production research more generally.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This paper reports a systematic mapping study of research on Wikidata, covering publications from October 2012 to June 2018. The authors searched ACM Digital Library, Springer Link, DBLP, and Google Scholar using the single keyword "Wikidata," screened 1,497 initial results down to 67 peer-reviewed full papers, and manually classified the final set into five categories: community-oriented, engineering-oriented, application use cases, knowledge-graph-oriented, and data-oriented research. The paper describes publication frequency, venues, and the geographic/institutional origin of authors, and it identifies current research foci (most prominently data quality) and "white spots" such as multilingualism, usability, and non-Western perspectives. The central claim is that the 67-paper sample provides a reliable map of Wikidata research and that the identified gaps are meaningful directions for future work.

Significance. If the map is taken at face value, this is a useful early systematic overview of a young and rapidly growing research area. The study follows established guidelines by Petersen et al., documents the search and screening steps in detail, and makes the final paper set accessible through a Zotero group, which are notable strengths. The descriptive findings on publication growth, conference dominance, and the predominance of European authors are likely correct for the sampled literature and provide a baseline for later studies. However, the paper's value rests on the representativeness of the sample and the transparency of the classification; both currently have weaknesses that need to be addressed before the map can be considered reliable. The claim about future research directions is particularly sensitive to the search and exclusion protocol.

major comments (5)
  1. [Section 3.1 (Frequency of Publication)] The extrapolation that the number of research articles "are expected to reach between 40-50 by the end of 2018" is not statistically justified. The paper reports 12 papers in the first half of 2018, but no model, confidence interval, or discussion of publication delay is given to support this range. This statement should be removed or explicitly labeled as a rough, non-committal guess rather than a finding.
  2. [Figure 1 and Section 2.3] The arithmetic in the selection process is inconsistent between the text and Figure 1. Section 2.3 states that 833 papers were excluded after reading abstracts and 17 were excluded as theses and short papers, yielding 67 from 1,125. Figure 1 shows 86 and 19 at the corresponding steps, and the intermediate count after abstract screening appears as 86 rather than 833 - 1125 = 84. These discrepancies undermine the reported repeatability of the screening process and must be corrected so that the figure and text agree.
  3. [Sections 2.2-2.3 (Search Process and Exclusion Criteria)] The search protocol relies on the single exact keyword "Wikidata," four digital libraries, English-only results, and a minimum length of five pages. The five-page rule in particular excludes short and companion papers that the authors themselves mark as not part of the study, such as the WSDM Cup submissions in references [6] and [77]. This exclusion is not sampling-neutral: short papers often describe tools, position statements, or work from less dominant communities, and their removal could change the category distribution and the list of under-studied topics. Non-English papers are the most likely to address multilingualism and knowledge diversity, so the reported "white spots" for these topics may be artifacts of the search rather than genuine gaps in the literature. The paper should either provide a sensitivity analysis or substantially hedge the RQ4 conclusions.
  4. [Section 2.5 (Research Paper Classification)] The classification of the 67 papers into five categories and subcategories is performed manually by the authors without inter-rater reliability or a second independent coder. Because the map and the reported "research foci" are entirely based on this subjective coding, the paper should report the coding process in more detail, ideally providing the category assignment for each of the 67 papers, or at least a random subset coded by both authors with a measure of agreement.
  5. [Section 6 (Limitations of Research)] The paper promises that the search log and the final article sample are available on GitHub and Zenodo, respectively, but no links are given in this version and the text says "This information will be included in the final version." For a systematic mapping study, the availability of the full search log and the complete paper list is a core requirement for repeatability. The final version must include working links and, if possible, the full screening decisions.
minor comments (4)
  1. [Section 2 and footnotes] There are several typographical errors, including "abbrviation" in footnote 5, "Formarly" in footnotes 17 and 18, and "Sarbadani" instead of "Sarabadani" in Section 4.2.2. These should be corrected.
  2. [Section 3.2 (Publishers and Publication Types)] The text contains a stray URL fragment ("https://www.overleaf.com/project/5c3f2935235d8259ff21db4e") embedded in the sentence about journal articles. This appears to be an editing artifact and must be removed.
  3. [Section 2.3] The sentence "In a first step, we already excluded 160 non-English search results ... second, 132 duplicates ... third, 80 non-papers" is clear, but the subsequent sentence "After applying the aforementioned criteria on the remaining 1,125 articles, 208 articles were excluded by reading the titles and, another 833 papers were excluded after reading the abstracts" does not match Figure 1. Please align the figure and text.
  4. [References] The use of asterisks to mark excluded references is helpful, but the paper should state explicitly in the text whether a reference marked with an asterisk is a paper that was excluded from the analysis or a non-paper source. Currently the reader must infer this from context.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the mapping study's descriptive findings are derived from its stated search and classification protocol, not from the authors' prior work or from an equation-level reduction.

full rationale

This paper is a systematic mapping study: it searches bibliographic databases for the keyword 'Wikidata', applies inclusion/exclusion criteria, and manually classifies 67 resulting papers into categories. There is no mathematical derivation, no fitted parameter, and no prediction that is statistically forced by construction. The classification categories were induced from the papers themselves (Section 2.5), which is inherent to the mapping-study method rather than a circular reduction of the kind this analysis targets. The paper cites two works involving the second author (Müller-Birn et al. [43] and Cuong and Müller-Birn [13]), but these are simply included as objects of classification among the 67 sampled papers; the central claims about publication frequency, venue, geographic origin, and topical focus do not depend on the correctness of those cited papers. The authors also cite their own workshop experience in the discussion, but that is anecdotal context, not load-bearing evidence. The main validity concern is whether the sample is representative, given the single-keyword search, English-only criterion, and five-page minimum (Sections 2.2-2.3), and the paper itself acknowledges repeatability limitations in Section 6 while noting that the search log and Zenodo sample would be supplied in the final version. These are correctness and transparency concerns, not circularity. No step in the paper's chain reduces to its own inputs, and no load-bearing premise is justified only by a self-referential citation. The honest finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on search recall, inclusion criteria, and manual coding. These are standard assumptions for a mapping study, not ad hoc inventions. No free parameters or invented entities are introduced.

assumptions (3)
  • domain assumption The four search engines, queried with the single keyword 'Wikidata', retrieve all relevant research on Wikidata.
    The completeness of the map depends on this search recall; the paper does not validate recall against an external benchmark. Section 2.2.
  • domain assumption Excluding papers with fewer than five pages removes only short papers, posters, and demonstration papers, not substantive research.
    This filter removes 17 items and shapes the sample; some full results might appear as short papers. Section 2.3.
  • domain assumption The manual classification of papers into the five reported categories is a valid representation of the research topics.
    The map and the 'white spots' depend on the authors' coding scheme, which is not independently validated. Section 2.5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Wikidata from a Research Perspective -- A Systematic Mapping Study of Wikidata." pith.science (2026). https://pith.science/paper/XATYOCTJ

@misc{pith2026190811153,
  author       = {Pith},
  title        = {Pith review of: Wikidata from a Research Perspective -- A Systematic Mapping Study of Wikidata},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XATYOCTJ}},
  note         = {Machine review of arXiv:1908.11153}
}
read the original abstract

Wikidata is one of the most edited knowledge bases which contains structured data. It serves as the data source for many projects in the Wikimedia sphere and beyond. Since its inception in October 2012, it has been increasingly growing in term of both its community and its content. This growth is reflected by an expanding number of research focusing on Wikidata. Our study aims to provide a general overview of the research performed on Wikidata through a systematic mapping study in order to identify the current topical coverage of existing research as well as the white spots which need further investigation. In this study, 67 peer-reviewed research from journals and conference proceedings were selected, and classified into meaningful categories. We describe this data set descriptively by showing the publication frequency, the publication venue and the origin of the authors and reveal current research focuses. These especially include aspects concerning data quality, including questions related to language coverage and data integrity. These results indicate a number of future research directions, such as, multilingualism and overcoming language gaps, the impact of plurality on the quality of Wikidata's data, Wikidata's potential in various disciplines, and usability of user interface.

Figures

Figures reproduced from arXiv: 1908.11153 by the authors.

Figure 1
Figure 1. Article selection process and the number of in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Frequency of publications per year. more insights about each research, and therefore, we read the arti￾cle in more detail by focusing on the findings. The tools used for data extraction and analysis are Zotero and Microsoft Excel. 2.5 Research Paper Classification At this stage, we read the abstract, introduction and conclusion parts of all articles to get more insights about each research for cat￾egorization. In a … view at source ↗
Figure 3
Figure 3. Research contributions from countries and conti [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 68 canonical work pages

  1. [6]

    *Tomoya Yamazaki, Mei Sasaki, Naoya Murakami, Takuya Makabe, and Hiroki Iwasawa. 2017. Ensemble Models for Detecting Wikidata Vandalism with Stacking - Team Honeyberry Vandalism Detector at WSDM Cup 2017. In WSDM Cup 2017 Notebook Papers. Cambridge, UK. 31This information will be included in the final version of this paper. Wikidata from a Research Perspe...

  2. [77]

    Denny Vrandecic. 2013. The rise of Wikidata. IEEE Intelligent Systems 28, 4 (2013), 90–95

  3. [1]

    Abián, F

    D. Abián, F. Guerra, J. Martínez-Romanos, and Raquel Trillo-Lado. 2018. Wikidata and DBpedia: A Comparative Study. In Semantic Keyword-Based Search on Structured Data Sources, Julian Szymański and Yannis Velegrakis (Eds.). Vol. 10546. Springer International Publishing, Cham, 142–154. https://doi.org/10.1007/ 978-3-319-74497-1_14

  4. [2]

    Albin Ahmeti, Simon Razniewski, and Axel Polleres. 2014. Assessing the Completeness of Entities in Knowledge Bases. In The Semantic Web: ESWC 2017 Satellite Events (Lecture Notes in Computer Science) . Springer, Cham, 7–11. https://doi.org/10.1007/978-3-319-70407-4_2

  5. [3]

    Paulo Dias Almeida, Jorge Gustavo Rocha, Andrea Ballatore, and Alexander Zipf

  6. [4]

    Isabelle Augenstein. 2014. Joint Information Extraction from the Web Using Linked Data. InThe Semantic Web – ISWC 2014 (Lecture Notes in Computer Science). Springer, Cham, 505–512. https://doi.org/10.1007/978-3-319-11915-1_32

  7. [7]

    *Tuo Yu, Yiran Zhao, Xiaoxiao Wang, Yiwen Xu, Huajie Shao, Yuhang Wang, Xin Ma, and Dipannita Dey. 2017. Vandalism Detection Midpoint Report—The Riberry Vandalism Detector at WSDM Cup 2017. (2017). University of Illinois at Urbana?Champaign Student Report, not published

  8. [8]

    Adrian Bielefeldt, Julius Gonsior, and Markus Krötzsch. 2018. Practical Linked Data Access via SPARQL: The Case of Wikidata. Lyon, France, 10

Show all 78 references
  1. [9]

    Almeida, Victorio A

    Freddy Brasileiro, João Paulo A. Almeida, Victorio A. Carvalho, and Giancarlo Guizzardi. 2016. Applying a multi-level modeling theory to assess taxonomic hierarchies in Wikidata. In Proceedings of the 25th International Conference Com- panion on World Wide Web. International W...

  2. [10]

    Sebastian Burgstaller-Muehlbacher, Andra Waagmeester, Elvira Mitraka, Julia Turner, Tim Putman, Justin Leong, Chinmay Naik, Paul Pavlidis, Lynn Schriml, and Benjamin M. Good. 2016. Wikidata as a semantic framework for the Gene Wiki initiative. Database 2016 (2016)

  3. [12]

    Garcia Calabria, Pablo Albani, Diego Tauziet, Adriana Baravalle, and Andrés Sebastián D’Ambrosio

    Rafael Crescenzi, Marcelo Fernandez, Federico A. Garcia Calabria, Pablo Albani, Diego Tauziet, Adriana Baravalle, and Andrés Sebastián D’Ambrosio. 2017. A Production Oriented Approach for Vandalism Detection in Wikidata. In WSDM Cup 2017 Notebook Papers . arxiv.org, Cambridge, UK

  4. [13]

    To Tu Cuong and Claudia Müller-Birn. 2016. Applicability of Sequence Anal- ysis Methods in Analyzing Peer-Production Systems: A Case Study in Wiki- data. In Social Informatics. Springer, Cham, 142–156. https://doi.org/10.1007/ 978-3-319-47874-6_11

  5. [14]

    Dennis Diefenbach, Kamal Singh, and Pierre Maret. 2017. WDAqua-core0: A Question Answering Component for the Research Community. In Semantic Web Challenges (Communications in Computer and Information Science) . Springer, Cham, 84–89. https://doi.org/10.1007/978-3-319-69146-6_8

  6. [15]

    Dennis Diefenbach, Kamal Singh, and Pierre Maret. 2018. WDAqua-core1: A Question Answering Service for RDF Knowledge Bases. InCompanion Proceedings of the The Web Conference 2018 (WWW ’18) . International World Wide Web Conferences Steering Committee, Republic and Canton of Ge...

  7. [16]

    Putman, Sebastien Lelong, Sebastian Burgstaller, Andra Waagmeester, Colin Diesh, Nathan Dunn, Monica Munoz-Torres, Gregory S

    Tim E. Putman, Sebastien Lelong, Sebastian Burgstaller, Andra Waagmeester, Colin Diesh, Nathan Dunn, Monica Munoz-Torres, Gregory S. Stupp, Chunlei Wu, Andrew I. Su, and Benjamin Good. 2017. WikiGenomes: An open web application for community consumption and curation of gene an...

  8. [17]

    Fredo Erxleben, Michael Günther, Markus Krötzsch, Julian Mendez, and Denny Vrandečić. 2014. Introducing Wikidata to the linked data web. In International Semantic Web Conference. Springer, 50–65

  9. [18]

    *Michael Felderer and Jeffrey C. Carver. 2018. Guidelines for Systematic Mapping Studies in Security Engineering. arXiv:1801.06810 [cs] (Jan. 2018). http://arxiv. org/abs/1801.06810 arXiv: 1801.06810

  10. [20]

    Michael Färber, Frederic Bartscherer, Carsten Menne, and Achim Rettinger. 2016. Linked data quality of dbpedia, freebase, opencyc, wikidata, and yago. Semantic Web Preprint (2016), 1–53

  11. [21]

    Michael Färber, Basil Ell, Carsten Menne, and Achim Rettinger. 2015. A Com- parative Survey of DBpedia, Freebase, OpenCyc, Wikidata, and YAGO. Semantic Web 1 (2015) (2015), 26

  12. [22]

    Gad-Elrab, Daria Stepanova, Jacopo Urbani, and Gerhard Weikum

    Mohamed H. Gad-Elrab, Daria Stepanova, Jacopo Urbani, and Gerhard Weikum

  13. [23]

    Suchanek

    Luis Galárraga, Simon Razniewski, Antoine Amarilli, and Fabian M. Suchanek

  14. [24]

    InThe Semantic Web – ISWC 2016 (Lecture Notes in Computer Science)

    Exception-Enriched Rule Learning from Knowledge Graphs. InThe Semantic Web – ISWC 2016 (Lecture Notes in Computer Science) . Springer, Cham, 234–251. https://doi.org/10.1007/978-3-319-46523-4_15

  15. [25]

    Alexey Grigorev, Association for Computing Machinery, and Special Interest Group on Information Retrieval (Eds.). 2017. Large-Scale Vandalism Detection with Linear Classifiers The Conkerberry Vandalism Detector at WSDM Cup 2017 . Association for Computing Machinery, New York, ...

  16. [26]

    Ben Hachey, Will Radford, and Andrew Chisholm. 2017. Learning to generate one-sentence biographies from Wikidata. In Proceeding of the 15th Conference of the European Chapter of the Association for Computational Linguistics , Vol. 1. LongPress, valencia, Spain, 633–642

  17. [27]

    Johanna Geiß, Andreas Spitz, and Michael Gertz. 2017. NECKAr: A Named Entity Classifier for Wikidata. In Language Technologies for the Challenges of the Digital Age (Lecture Notes in Computer Science) . Springer, Cham, 115–129. https://doi.org/10.1007/978-3-319-73706-5_10

  18. [28]

    Stefan Heindorf, Martin Potthast, Gregor Engels, and Benno Stein. 2017. Overview of the Wikidata Vandalism Detection Task at WSDM Cup 2017. arXiv:1712.05956 [cs] abs/1712.05956 (Dec. 2017). http://arxiv.org/abs/1712.05956 arXiv: 1712.05956

  19. [29]

    Stefan Heindorf, Martin Potthast, Benno Stein, and Gregor Engels. 2016. Van- dalism detection in wikidata. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management . ACM, 327–336

  20. [30]

    *Stefan Heindorf, Martin Potthast, Hannah Bast, Björn Buchhold, and Elmar Haussmann. 2017. WSDM Cup 2017: Vandalism Detection and Triple Scoring. In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining (WSDM ’17) . ACM, New York, NY, USA, 827–828...

  21. [31]

    Daniel Hernández, Aidan Hogan, Cristian Riveros, Carlos Rojas, and Enzo Zerega

  22. [32]

    Ali Ismayilov, Dimitris Kontokostas, Sören Auer, Jens Lehmann, and Sebastian Hellmann. 2015. Wikidata through the Eyes of DBpedia. arXiv:1507.04180 [cs] (July 2015). http://arxiv.org/abs/1507.04180 arXiv: 1507.04180

  23. [33]

    Daniel Hernández, Aidan Hogan, and Markus Krötzsch. 2015. Reifying RDF: What works well with wikidata? SSWS@ ISWC 1457 (2015), 32–47

  24. [34]

    Lucie-Aimée Kaffee, Hady Elsahar, Pavlos Vougiouklis, Christophe Gravier, Frédérique Laforest, Jonathon Hare, and Elena Simperl. 2018. Mind the (Lan- guage) Gap: Generation of Multilingual Wikipedia Summaries from Wikidata for ArticlePlaceholders. In The Semantic Web (Lecture ...

  25. [35]

    In The Semantic Web – ISWC 2016

    Querying Wikidata: Comparing SPARQL, Relational and Graph Databases. In The Semantic Web – ISWC 2016 . Springer, Cham, 88–103. https://doi.org/10. 1007/978-3-319-46547-0_10

  26. [36]

    *Barbara Kitchenham, Pearl Brereton, and David Budgen. 2010. The Educational Value of Mapping Studies of Software Engineering Literature. InProceedings of the 32Nd ACM/IEEE International Conference on Software Engineering - Volume 1 (ICSE ’10). ACM, New York, NY, USA, 589–598....

  27. [37]

    Lucie-Aimée Kaffee, Hady Elsahar, Pavlos Vougiouklis, Christophe Gravier, Fred- erique Laforest, Jonathon Hare, and Elena Simperl. 2018. Learning to Generate Wikipedia Summaries for Underserved Languages from Wikidata. Proceedings of the 2018 Conference of the North American C...

  28. [38]

    Maximilian Klein, Harsh Gupta, Vivek Rai, Piotr Konieczny, and Haiyi Zhu

  29. [40]

    Werner Leyh and Homero Fonseca Filho. 2017. Interlinking Standardized Open- StreetMap Data and Citizen Science Data in the OpenData Cloud. In Advances in Human Factors and Systems Interaction (Advances in Intelligent Systems and Com- puting). Springer, Cham, 85–96. https://doi...

  30. [41]

    Kitchenham, David Budgen, and O

    *Barbara A. Kitchenham, David Budgen, and O. Pearl Brereton. 2011. Using Mapping Studies As the Basis for Further Research - A Participant-observer Case Study. Inf. Softw. Technol. 53, 6 (2011), 638–651. https://doi.org/10.1016/j.infsof. 2010.12.011

  31. [42]

    Hatem Mousselly Sergieh and Iryna Gurevych. 2016. Enriching Wikidata with Frame Semantics. In Proceedings of the 5th Workshop on Automated Knowledge Base Construction. https://doi.org/10.18653/v1/W16-1306

  32. [43]

    In Proceedings of the 12th International Symposium on Open Collaboration (OpenSym ’16)

    Monitoring the Gender Gap with Wikidata Human Gender Indicators. In Proceedings of the 12th International Symposium on Open Collaboration (OpenSym ’16). ACM, New York, NY, USA, 16:1–16:9. https://doi.org/10.1145/2957792. 2957798

  33. [44]

    Markus Krötzsch. 2017. Ontologies for Knowledge Graphs?. In Proceedings of the 30th International Workshop on Description Logics (CEUR Workshop Proceedings) , Vol. Vol-1879. CEUR-WS. org, France

  34. [45]

    Finn Årup Nielsen and Lars Kai Hansen. 2018. Inferring visual semantic similarity with deep learning and Wikidata: Introducing imagesim-353. InProceedings of the First Workshop on Deep Learning for Knowledge Graphs and Semantic Technologies (DL4KGS), Vol. Vol-2106. Department ...

  35. [46]

    Elvira Mitraka, Andra Waagmeester, Sebastian Burgstaller-Muehlbacher, Lynn M Schriml, Andrew I Su, and Benjamin M Good. 2015. Wikidata: A platform for data integration and dissemination for the life sciences and beyond. bioRxiv (Jan. 2015). https://doi.org/10.1101/031971

  36. [47]

    *Chitu Okoli, Mohamad Mehdi, Mostafa Mesgari, Finn Årup Nielsen, and Arto Lanamäki. 2012. The People’s Encyclopedia Under the Gaze of the Sages: A Systematic Review of Scholarly Research on Wikipedia. SSRN Electronic Journal (2012). https://doi.org/10.2139/ssrn.2021326

  37. [48]

    Claudia Müller-Birn, Benjamin Karran, Janette Lehmann, and Markus Luczak- Rösch. 2015. Peer-production system or collaborative ontology engineering effort: What is Wikidata?. In Proceedings of the 11th International Symposium on Open Collaboration. ACM, 20

  38. [50]

    *Kai Petersen, Sairam Vakkalanka, and Ludwik Kuzniarz. 2015. Guidelines for conducting systematic mapping studies in software engineering: An update. Information and Software Technology 64 (Aug. 2015), 1–18. https://doi.org/10. 1016/j.infsof.2015.03.007

  39. [51]

    Finn Årup Nielsen, Daniel Mietchen, and Egon Willighagen. 2017. Scholia, Scientometrics and Wikidata. In The Semantic Web: ESWC 2017 Satellite Events (Lecture Notes in Computer Science) . Springer, Cham, 237–259. https://doi.org/10. 1007/978-3-319-70407-4_36

  40. [52]

    Alessandro Piscopo, Lucie-Aimée Kaffee, Chris Phethean, and Elena Simperl. 2017. Provenance Information in a Collaborative Knowledge Graph: An Evaluation of Wikidata External References. In The Semantic Web – ISWC 2017 (Lecture Notes in Computer Science) . Springer, Cham, 542–...

  41. [53]

    Thomas Pellissier Tanon, Denny Vrandečić, Sebastian Schaffert, Thomas Steiner, and Lydia Pintscher. 2016. From freebase to wikidata: The great migration. In Proceedings of the 25th international conference on world wide web . International World Wide Web Conferences Steering C...

  42. [54]

    *Kai Petersen, Robert Feldt, Shahid Mujtaba, and Michael Mattsson. 2008. Sys- tematic Mapping Studies in Software Engineering. In Proceedings of the 12th International Conference on Evaluation and Assessment in Software Engineer- ing (EASE’08). BCS Learning & Development Ltd.,...

  43. [55]

    Alessandro Piscopo, Pavlos Vougiouklis, Lucie-Aimée Kaffee, Christopher Phethean, Jonathon Hare, and Elena Simperl. 2017. What Do Wikidata and Wikipedia Have in Common?: An Analysis of Their Use of External References. In Proceedings of the 13th International Symposium on Open...

  44. [56]

    Boyce, and Matthias Samwald

    Alexander Pfundner, Tobias Schönberg, John Horn, Richard D. Boyce, and Matthias Samwald. 2015. Utilizing the Wikidata system to improve the quality of medical content in Wikipedia in diverse languages: a pilot study. Journal of medical Internet research 17, 5 (2015)

  45. [57]

    Simon Razniewski, Vevake Balaraman, and Werner Nutt. 2017. Doctoral Advisor or Medical Condition: Towards Entity-Specific Rankings of Knowledge Base Properties. In Advanced Data Mining and Applications (Lecture Notes in Computer Science). Springer, Cham, 526–540. https://doi.o...

  46. [58]

    Alessandro Piscopo, Chris Phethean, and Elena Simperl. 2017. What Makes a Good Collaborative Knowledge Graph: Group Composition and Quality in Wikidata. In Social Informatics (Lecture Notes in Computer Science) . Springer, Cham, 305–322. https://doi.org/10.1007/978-3-319-67217-5_19

  47. [59]

    Alessandro Piscopo, Christopher Phethean, and Elena Simperl. 2017. Wikidatians are born: paths to full participation in a collaborative structured knowledge base. In Proceedings of the 50th Hawaii International Conference on System Sciences . University of Hawaii, 4354–4363. h...

  48. [60]

    John Samuel. 2017. Collaborative Approach to Developing a Multilingual On- tology: A Case Study of Wikidata. In Metadata and Semantic Research (Com- munications in Computer and Information Science) . Springer, Cham, 167–172. https://doi.org/10.1007/978-3-319-70863-8_16

  49. [61]

    Radityo Eko Prasojo, Fariz Darari, Simon Razniewski, and Werner Nutt. 2016. Managing and consuming completeness information for wikidata using COOL- WD. CEUR-WS. org

  50. [62]

    Andreas Spitz, Johanna Geiß, and Michael Gertz. 2016. So far away and yet so close: augmenting toponym disambiguation and similarity with text-based networks. ACM Press, 1–6. https://doi.org/10.1145/2948649.2948651

  51. [63]

    Simon Razniewski, Fabian Suchanek, and Werner Nutt. 2016. But What Do We Actually Know? Association for Computational Linguistics, 40–44. https: //doi.org/10.18653/v1/W16-1308

  52. [64]

    Daniel Ringler and Heiko Paulheim. 2017. One Knowledge Graph to Rule Them All? Analyzing the Differences Between DBpedia, YAGO, Wikidata & co.. In KI 2017: Advances in Artificial Intelligence (Lecture Notes in Computer Science) . Springer, Cham, 366–372. https://doi.org/10.100...

  53. [65]

    Thang Hoang Ta and Chutiporn Anutariya. 2014. A Model for Enriching Multilingual Wikipedias Using Infobox and Wikidata Property Alignment. In Semantic Technology . Springer, Cham, 335–350. https://doi.org/10.1007/ 978-3-319-15615-6_25

  54. [66]

    Amir Sarabadani, Aaron Halfaker, and Dario Taraborelli. 2017. Building Auto- mated Vandalism Detection Tools for Wikidata. InProceedings of the 26th Interna- tional Conference on World Wide Web Companion (WWW ’17 Companion). Interna- tional World Wide Web Conferences Steering ...

  55. [67]

    Katherine Thornton, Euan Cochrane, Thomas Ledoux, Bertrand Caron, and Carl Wilson. 2017. Modeling the Domain of Digital Preservation in Wikidata. In In Proceedings of ACM International Conference on Digital Preservation . iPres’17, Kyoto, Japan

  56. [68]

    Thomas Steiner. 2014. Bots vs. Wikipedians, Anons vs. Logged-Ins (Redux): A Global Study of Edit Activity on Wikipedia and Wikidata. In Proceedings of The International Symposium on Open Collaboration (OpenSym ’14) . ACM, New York, NY, USA, 25:1–25:7. https://doi.org/10.1145/2...

  57. [69]

    Tomás Sáez and Aidan Hogan. 2018. Automatically Generating Wikipedia Info- boxes from Wikidata. In Companion Proceedings of the The Web Conference 2018 (WWW ’18). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, Switzerland, 1823–1830...

  58. [70]

    Jakob Voß. 2016. Classification of Knowledge Organization Systems with Wikidata. In Proceedings of the 15th European Networked Knowledge Organi- zation Systems Workshop (NKOS 2016), Vol. Vol-1676. CEUR-WS. org, Hannover. https://doi.org/10.5281/zenodo.61767

  59. [71]

    Endris, Jose M

    Harsh Thakkar, Kemele M. Endris, Jose M. Gimenez-Garcia, Jeremy Debattista, Christoph Lange, and Sören Auer. 2016. Are Linked Datasets Fit for Open- domain Question Answering? A Quality Assessment. In Proceedings of the 6th International Conference on Web Intelligence, Mining ...

  60. [72]

    Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM 57, 10 (2014), 78–85

  61. [73]

    Turki, D

    H. Turki, D. Vrandecic, H. Hamdi, and I. Adel. 2017. Using WikiData as a Multi- lingual Multi-dialectal Dictionary for Arabic Dialects. In 2017 IEEE/ACS 14th International Conference on Computer Systems and Applications (AICCSA) . 437–

  62. [74]

    Xi Yang, Shiya Ren, Yuan Li, Ke Shen, Zhixing Li, and Guoyin Wang. 2017. Relation Linking for Wikidata Using Bag of Distribution Representation. In Natural Language Processing and Chinese Computing . Springer, Cham, 652–661. https://doi.org/10.1007/978-3-319-73618-1_55

  63. [75]

    Theo van Veen, Juliette Lonij, and Willem Jan Faber. 2016. Linking Named Entities in Dutch Historical Newspapers. In Metadata and Semantics Research (Communications in Computer and Information Science) . Springer, Cham, 205–210. https://doi.org/10.1007/978-3-319-49157-8_18

  64. [76]

    Eva Zangerle, Wolfgang Gassler, Martin Pichl, Stefan Steinhauser, and Günther Specht. 2016. An Empirical Evaluation of Property Recommender Systems for Wikidata and Collaborative Knowledge Bases. In Proceedings of the 12th Inter- national Symposium on Open Collaboration (OpenS...

  65. [79]

    Gerhard Wohlgenannt, Nikolay Klimov, Dmitry Mouromtsev, Daniil Razdyakonov, Dmitry Pavlov, and Yury Emelyanov. 2017. Using Word Embeddings for Visual Data Exploration with Ontodia and Wikidata. In Joint Proceedings of BLINK2017: 2nd International Workshop on Benchmarking Linke...

  66. [81]

    Xue-lu Yu and Lin Qiao. 2017. Meronymy Relation Extraction Based on 3-Motif in Wikidata. DEStech Transactions on Computer Science and Engineering 0, cnsce (2017). https://doi.org/10.12783/dtcse/cnsce2017/8915

  67. [83]

    *Qi Zhu, Hongwei Ng, Liyuan Liu, Ziwei Ji, Bingjie Jiang, Jiaming Shen, and Huan Gui. 2017. Wikidata Vandalism Detection - The Loganberry Vandalism Detector at WSDM Cup 2017. In Proceedings of the WSDM Cup 2017: Vandalism Detection and Triple Scoring , Vol. abs/1712.06922. arx...

  68. [442]

    https://doi.org/10.1109/AICCSA.2017.115

  69. [2016]

    In Computational Science and Its Applications – ICCSA 2016 (Lecture Notes in Computer Science)

    Where the Streets Have Known Names. In Computational Science and Its Applications – ICCSA 2016 (Lecture Notes in Computer Science) . Springer, Cham, 1–12. https://doi.org/10.1007/978-3-319-42089-9_1

  70. [2017]

    In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining (WSDM ’17)

    Predicting Completeness in Knowledge Bases. In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining (WSDM ’17) . ACM, New York, NY, USA, 375–383. https://doi.org/10.1145/3018661.3018739

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.