{"id":"3264abd5-114e-416e-8c91-e247b613af22","arxiv_id":"1908.11153","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic mapping study classifies 67 peer-reviewed Wikidata research papers into five topic areas and identifies gaps such as multilingualism and usability.","lead":"This paper maps 67 peer-reviewed studies about Wikidata, the open knowledge base that feeds structured data to Wikipedia, and sorts them into topics like data quality, community behavior, and knowledge graph engineering. It shows that Wikidata research is growing steadily, is concentrated in Europe, and has unexplored areas such as multilingualism and user interface design.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The map's conclusions are only as sound as the 67-paper sample; the English-only, full-paper-only, single-keyword search could systematically omit exactly the short-paper and non-English research that the claimed white spots predict.","rationale":"The reader's weakest assumption—that the 67-paper sample is representative—is the same load-bearing concern I identify. The central claim is descriptive and depends on the sample being a fair map of high-quality Wikidata research. My concern is not that the authors were careless; their process follows Petersen et al. and includes Google Scholar to widen coverage. Rather, the inclusion criteria are not neutral with respect to the very conclusions drawn. Short papers and non-English work are not marginal in this community, and the paper itself marks some excluded items in the references. The concrete test would settle whether the protocol changes the map: if the category and country distributions remain stable under a broader search, the concern does not land; if they move, the conclusions are search artifacts. Because this potential bias is not demonstrated by the paper, the appropriate verdict remains CONDITIONAL, matching the reader's assessment; no change to the verdict is needed, but the promised artifacts should be linked and the search protocol should be independently reproducible.","tokens_in":18341,"tokens_out":7454,"duration_ms":71487,"concrete_test":"Re-run the mapping with an expanded protocol: search Scopus, OpenAlex, and Web of Science in addition to the original four sources; use the query 'Wikidata' OR 'Wikibase'; do not apply the English or five-page filters before relevance screening; include peer-reviewed workshop and companion papers meeting the authors' content criteria. Code the resulting set with the same five categories and compute category proportions and first-author country distributions, then compare them with Table 2 and Figure 3 using percentage-point differences or a chi-square test. If category or continental percentages shift by more than about 10 percentage points, the original map is protocol-dependent; if the distributions are stable, the central descriptions are robust to sampling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claims—data quality as the dominant focus, Europe as the leading region, and the identification of white spots such as multilingualism, usability, and non-Western perspectives—are statements about Wikidata research as a whole, not merely about the 67 sampled papers. The load-bearing link is the search protocol in Sections 2.2–2.3: it uses the single exact keyword 'Wikidata', four bibliographic sources, English-only results, and a five-page minimum. The five-page rule is especially consequential in this field because a substantial share of Wikidata-related work appears as short/companion/workshop papers, several of which the authors themselves mark as excluded in the reference list. Excluding those papers is not sampling-neutral: applied and position work could plausibly change the category distribution and the list of under-studied topics. Similarly, non-English studies are the most likely to address the multilingual and knowledge-diversity gaps that the paper presents as future directions, so filtering them out may manufacture those very gaps. If the population is mis-sampled, RQ4's 'aspects still to be studied' are artifacts of the search rather than findings about Wikidata research. The paper's own limitation section acknowledges repeatability concerns, but the promised search log and Zenodo sample are not actually linked in this version, which prevents an independent check of this core assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic mapping study of research on Wikidata, covering publications from October 2012 to June 2018. The authors searched ACM Digital Library, Springer Link, DBLP, and Google Scholar using the single keyword \"Wikidata,\" screened 1,497 initial results down to 67 peer-reviewed full papers, and manually classified the final set into five categories: community-oriented, engineering-oriented, application use cases, knowledge-graph-oriented, and data-oriented research. The paper describes publication frequency, venues, and the geographic/institutional origin of authors, and it identifies current research foci (most prominently data quality) and \"white spots\" such as multilingualism, usability, and non-Western perspectives. The central claim is that the 67-paper sample provides a reliable map of Wikidata research and that the identified gaps are meaningful directions for future work.","tokens_in":18553,"tokens_out":3569,"duration_ms":34931,"significance":"If the map is taken at face value, this is a useful early systematic overview of a young and rapidly growing research area. The study follows established guidelines by Petersen et al., documents the search and screening steps in detail, and makes the final paper set accessible through a Zotero group, which are notable strengths. The descriptive findings on publication growth, conference dominance, and the predominance of European authors are likely correct for the sampled literature and provide a baseline for later studies. However, the paper's value rests on the representativeness of the sample and the transparency of the classification; both currently have weaknesses that need to be addressed before the map can be considered reliable. The claim about future research directions is particularly sensitive to the search and exclusion protocol.","major_comments":[{"comment":"The extrapolation that the number of research articles \"are expected to reach between 40-50 by the end of 2018\" is not statistically justified. The paper reports 12 papers in the first half of 2018, but no model, confidence interval, or discussion of publication delay is given to support this range. This statement should be removed or explicitly labeled as a rough, non-committal guess rather than a finding.","section":"Section 3.1 (Frequency of Publication)"},{"comment":"The arithmetic in the selection process is inconsistent between the text and Figure 1. Section 2.3 states that 833 papers were excluded after reading abstracts and 17 were excluded as theses and short papers, yielding 67 from 1,125. Figure 1 shows 86 and 19 at the corresponding steps, and the intermediate count after abstract screening appears as 86 rather than 833 - 1125 = 84. These discrepancies undermine the reported repeatability of the screening process and must be corrected so that the figure and text agree.","section":"Figure 1 and Section 2.3"},{"comment":"The search protocol relies on the single exact keyword \"Wikidata,\" four digital libraries, English-only results, and a minimum length of five pages. The five-page rule in particular excludes short and companion papers that the authors themselves mark as not part of the study, such as the WSDM Cup submissions in references [6] and [77]. This exclusion is not sampling-neutral: short papers often describe tools, position statements, or work from less dominant communities, and their removal could change the category distribution and the list of under-studied topics. Non-English papers are the most likely to address multilingualism and knowledge diversity, so the reported \"white spots\" for these topics may be artifacts of the search rather than genuine gaps in the literature. The paper should either provide a sensitivity analysis or substantially hedge the RQ4 conclusions.","section":"Sections 2.2-2.3 (Search Process and Exclusion Criteria)"},{"comment":"The classification of the 67 papers into five categories and subcategories is performed manually by the authors without inter-rater reliability or a second independent coder. Because the map and the reported \"research foci\" are entirely based on this subjective coding, the paper should report the coding process in more detail, ideally providing the category assignment for each of the 67 papers, or at least a random subset coded by both authors with a measure of agreement.","section":"Section 2.5 (Research Paper Classification)"},{"comment":"The paper promises that the search log and the final article sample are available on GitHub and Zenodo, respectively, but no links are given in this version and the text says \"This information will be included in the final version.\" For a systematic mapping study, the availability of the full search log and the complete paper list is a core requirement for repeatability. The final version must include working links and, if possible, the full screening decisions.","section":"Section 6 (Limitations of Research)"}],"minor_comments":[{"comment":"There are several typographical errors, including \"abbrviation\" in footnote 5, \"Formarly\" in footnotes 17 and 18, and \"Sarbadani\" instead of \"Sarabadani\" in Section 4.2.2. These should be corrected.","section":"Section 2 and footnotes"},{"comment":"The text contains a stray URL fragment (\"https://www.overleaf.com/project/5c3f2935235d8259ff21db4e\") embedded in the sentence about journal articles. This appears to be an editing artifact and must be removed.","section":"Section 3.2 (Publishers and Publication Types)"},{"comment":"The sentence \"In a first step, we already excluded 160 non-English search results ... second, 132 duplicates ... third, 80 non-papers\" is clear, but the subsequent sentence \"After applying the aforementioned criteria on the remaining 1,125 articles, 208 articles were excluded by reading the titles and, another 833 papers were excluded after reading the abstracts\" does not match Figure 1. Please align the figure and text.","section":"Section 2.3"},{"comment":"The use of asterisks to mark excluded references is helpful, but the paper should state explicitly in the text whether a reference marked with an asterisk is a paper that was excluded from the analysis or a non-paper source. Currently the reader must infer this from context.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a preprint (arXiv:1908.11153v2) and appears to be an earlier version of a planned journal submission. The missing data links and the arithmetic inconsistencies in Figure 1 suggest the authors have not yet finalized the reproducibility artifacts. For a mapping study, the search protocol and screening log are the equivalent of source code and data; without them the central claim cannot be independently verified. The paper is salvageable with revision, but the author should be asked to either provide the artifacts or soften the repeatability claims. The geographic and topical findings are interesting but should be framed as hypotheses about the field rather than definitive conclusions given the sampling limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first systematic map of Wikidata research, and it does what a mapping study should—clear protocol, sensible five-category classification, and a readable descriptive picture. The contribution is organizational, not deep: no new claims about Wikidata itself, just a structured survey of 67 full-length English peer-reviewed papers up to mid-2018. I think the reader's conditional verdict is about right, and the stress-test worry is fair but not fatal.\n\nWhat's good: the method follows Petersen et al., the screening steps are mostly transparent, and the category labels (community, engineering, applications, knowledge graphs, data) feel natural. The finding that data quality dominates and Europe (Germany especially) leads in output is credible for the included sample. The paper also flags its own limitations and promises the search log and sample dataset—good practice, even if the promises are not yet fulfilled in this version.\n\nSoft spots, in rough order of importance. First, the sampling frame is narrower than the language of the conclusions. Single keyword 'Wikidata', four bibliographic sources, English only, five-page minimum. Short/workshop papers are exactly where applied and position work often lives, and non-English work is exactly where multilingualism gaps might be studied. So RQ4's 'white spots' should be read as white spots within this frame, not demonstrated gaps in Wikidata research as a whole. The authors mostly stay close to their data, but the abstract's 'white spots which need further investigation' overreaches.\n\nSecond, reproducibility is promised but not delivered in this version: the search log and Zenodo link are announced as future additions, and there is a stray Overleaf URL in Section 3.2. Those need to actually exist before I'd call the study repeatable.\n\nThird, the selection flowchart arithmetic doesn't reconcile with the text (the 80/208/833/17 bookkeeping is off somewhere), and the extrapolation from 12 papers in the first half of 2018 to 40-50 by year-end is a linear guess with no error bar; call it a rough projection, not a finding. Fourth, no inter-rater reliability is reported for the classification; the second-author check is something but not much. Minor, given the scale.\n\nCitation pattern: fine. The two second-author papers in the corpus are legitimately part of the literature, and citing them is not inflation.\n\nBottom line: this is a useful reference for anyone entering Wikidata research and a reasonable baseline for future mapping studies. It deserves peer review, not desk rejection. I'd send it out with requests to fix the arithmetic, link the actual data, and soften or re-scope the white-spot claims.","headline":"Useful first map of Wikidata research, with a real but manageable sampling caveat; worth refereeing after the promised data links and flowchart arithmetic are fixed.","tokens_in":19067,"tokens_out":3591,"would_cite":true,"duration_ms":33791,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 67 peer-reviewed papers argues that Wikidata research is concentrated on data quality and in Europe, with multilingualism, usability, and knowledge diversity standing out as the clearest gaps.","keywords":["Wikidata","systematic mapping study","research classification","knowledge graph","data quality","multilingualism","peer production"],"falsifier":"Re-run the same procedure for October 2012 to June 2018 with a broader set of related search terms, without the five-page minimum, and including non-English publications; if the five categories change materially or new high-volume topics appear, the paper's claim that its 67-paper sample maps the field would be refuted.","tokens_in":70,"feed_emoji":"🗺️","tokens_out":12711,"duration_ms":172174,"temperature":0.7,"pith_summary":"This paper sets out to establish that a systematic review of 67 peer-reviewed papers can serve as a reliable map of the first six years of research on Wikidata, the collaborative knowledge base behind Wikipedia. The authors gathered papers using a single search term in four academic search engines, filtered out non-English work, short papers, and publications outside journals and conference proceedings, then sorted what remained into five categories. On that basis they argue that Wikidata research is growing year by year, that it is concentrated in Europe, and that its most mature topic is data quality, while multilingualism, knowledge diversity, user-interface usability, and use in non-technical disciplines remain underexplored. A reliable map matters because researchers need to know which parts of Wikidata have already been studied before choosing a new question.","feed_headline":"Data quality dominates Wikidata research, 67-paper map shows","feed_subtitle":"A 2012–2018 systematic review finds data quality, completeness, and vandalism are the field's mature topics.","key_machinery":"The machinery that carries the argument is the classification scheme built from a systematic reading of the 67 papers. The authors collected candidates with the single search term “Wikidata” in four academic search engines, removed duplicates, non-English results, non-papers, work shorter than five pages, and publications outside journals or conference proceedings, then labeled the remaining papers and grouped them by constant comparison into five categories: community-oriented, engineering-oriented, application use cases, knowledge-graph-oriented, and data-oriented research. This taxonomy is what turns a list of references into a map, because a paper's placement determines where the authors report density and where they see clear gaps.","core_discovery":"The paper's central claim is that the 67 selected papers, read and classified carefully, show the state of Wikidata research between October 2012 and June 2018. Descriptively, conference papers outnumber journal articles, and most first authors are based in Europe, with 39 of the 67 papers coming from 2017 and the first half of 2018. Topically, the data-oriented category is the largest, with 22 papers, followed by knowledge-graph-oriented research with 15 and community-oriented research with 14, while engineering-oriented research and application use cases account for 9 and 7 papers respectively. From that distribution the authors identify the field's white spots: uneven language coverage, little study of what plurality and contradictory claims do to trust, few usability studies, and applications concentrated mainly in biomedicine and linguistics. They present these gaps as directions for future work rather than as failures of Wikidata itself.","pith_inferences":["The use of a single search term means a follow-up search with related software and data-service names would likely catch additional papers; my expectation is that it would enrich rather than overturn the five categories.","The exclusion of short papers may have removed early exploratory results, such as system demonstrations, which are often where new community and engineering topics first appear.","The paper's projection of 40–50 full papers for 2018 is directly testable by counting papers in late-2018 indexes and comparing the result with the growth curve it reports.","Comparing this Europe-heavy profile with the more global research community around Wikipedia would clarify whether the concentration is a property of Wikidata or of peer-production research more generally."],"forward_implications":["If the map is accurate, Wikidata research is expanding quickly: 39 of the 67 papers appeared in 2017 or the first half of 2018, and the 12 papers logged by June 2018 point to a projected 40–50 full papers for that year.","Data quality is the field's most developed strand, covering completeness, references, provenance, and vandalism, so new researchers entering this area have a firm baseline rather than an open field.","The clearest white spots are language coverage, the effect of contradictory claims on trustworthiness, user interface usability, and use cases outside biomedicine and linguistics, which the paper names as recommended future directions.","Because most first authors are based in Europe, the current research perspective is geographically concentrated, and the paper suggests this Western perspective may limit how knowledge diversity in Wikidata is studied.","The five-category classification offers a reusable skeleton for later mapping studies, since it labels both what has been done and what has not."],"supporting_citations":[{"why":"It supplies the systematic mapping study guidelines that define the paper's search, screening, and classification steps.","marker":"[50]"},{"why":"It frames a mapping study as a way to reveal existing topics and white spots, which is the purpose of this study.","marker":"[18]"},{"why":"It supports the premise that mapping studies provide a baseline that assists new research efforts.","marker":"[37]"},{"why":"It distinguishes a mapping study from a systematic literature review, setting the boundary for what a 67-paper sample can claim.","marker":"[49]"},{"why":"It defines Wikidata's design principles, which the paper uses to decide which topics fall into which research category.","marker":"[72]"},{"why":"It gives the early description of Wikidata's purpose that anchors the community-oriented research category.","marker":"[71]"}],"fun_headline_variants":["67-paper map: Wikidata research led by data quality","Wikidata studies cluster around data quality, gaps remain","Systematic review shows Wikidata research's blind spots","Wikidata research map: quality focus, language gaps"],"cache_read_input_tokens":21248,"weakest_assumption_plain":"The load-bearing premise is that the 67 filtered papers fairly represent all high-quality Wikidata research; if substantial work was published under other search terms, in other languages, in short-paper form, or in venues outside the four search engines used, the map and its reported gaps would shift.","fun_headline_variants_meta":{"raw":{"variants":["67-paper map: Wikidata research led by data quality","Wikidata studies cluster around data quality, gaps remain","Systematic review shows Wikidata research's blind spots","Wikidata research map: quality focus, language gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1576,"prompt_tokens":927,"completion_tokens":649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":586}},"tokens_in":543,"tokens_out":649,"duration_ms":6202,"temperature":1.0,"reasoning_tokens":586,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:22:30.307181+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same procedure for October 2012 to June 2018 with a broader set of related search terms, without the five-page minimum, and including non-English publications; if the five categories change materially or new high-volume topics appear, the paper's claim that its 67-paper sample maps the field would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the systematic mapping study guidelines that define the paper's search, screening, and classification steps."},{"cited_title":"Guidelines for Systematic Mapping Studies in Security Engineering","cited_arxiv_id":"1801.06810","evidence_quote":"It frames a mapping study as a way to reveal existing topics and white spots, which is the purpose of this study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supports the premise that mapping studies provide a baseline that assists new research efforts."},{"cited_title":"Endris, Jose M","cited_arxiv_id":null,"evidence_quote":"It gives the early description of Wikidata's purpose that anchors the community-oriented research category."}],"review_version":1}