{"id":"d84a6206-33fa-445a-90ab-230bc43cf0bc","arxiv_id":"1909.01109","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using edit-history mentions as capture-recapture samples, species-richness estimators can estimate Wikidata class completeness, with Jack1 and SOR most accurate and a new stability metric flagging incomplete classes.","lead":"Wikidata is incomplete in ways users cannot easily see. This paper uses edit history and species-counting statistics to estimate how many total items a category should have, testing the method on eight known classes and applying it across Wikidata.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing sampling assumptions in Section 3.2 are asserted, not validated; bot-driven and bulk-import edit patterns visibly violate them, so the estimator gains in Table 1 may not generalize.","rationale":"The reader's weakest assumption is the same one I consider load-bearing: the capture-recapture model depends on Section 3.2's independence, with-replacement, and stationarity assumptions. The eight-class evaluation with external ground truth is genuine evidence, and the released code is useful, but it only tests the estimators under whatever sampling process actually generated those classes' edit histories. Since the estimators' theoretical guarantees come from species-richness methods that require the stated assumptions, the fact that Jack1 and SOR beat the distinct-count baseline in Table 1 is not yet evidence that the method works generally. The concrete bot-filtering test would settle whether the performance is driven by a valid sampling process or by the particular mix of bot and bulk-import activity. I therefore keep the reader's CONDITIONAL verdict: the idea is plausible and the empirical setup is real, but the sampling model must be validated or its failure modes characterized before the completeness-measurement claim is accepted as stated.","tokens_in":12249,"tokens_out":7781,"duration_ms":85032,"concrete_test":"Using the released code and the per-edit user field in the Wikibase XML dump, recompute Table 1 and the rho splits of Table 2 under three observation streams: (a) all edits, as reported; (b) only non-bot human edits; (c) human edits with all edits by one user in one 30-day period collapsed into a single sample-period mention per entity. If Jack1 and SOR no longer consistently beat the distinct-count lower bound, or if the <0.001 versus >0.1 rho separation degrades under (b) or (c), then the reported accuracy is an artifact of bot and bulk-import sampling, and Assumptions 2-4 are violated in the target setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2's model is the foundation of every estimator in Eqs. (1)-(9): observations are independent, with-replacement, random draws from a closed class with time-invariant per-instance mention probabilities (Assumptions 2-4). The paper itself concedes that bursts of inserts cause estimators to overestimate (Fig. 3(h), Section 4.3), and Wikidata editing is in large part driven by bots and bulk imports that touch many entities in one period, producing dependent, non-random singletons. These are not rare edge cases; they are a dominant mode of growth for the very classes shown in Fig. 3. If Assumptions 2-4 fail, the frequency statistics f1 and f2 - the only inputs to Jack1, SOR, and the coverage terms - are systematically inflated, and every estimator is biased. The paper does not measure the degree of violation: the Section 3.2 claim that 'we have not observed any significant correlations in the edits' is not supported by any statistic, and no confidence intervals are reported anywhere, so the point estimates in Table 1 cannot separate real signal from sampling noise. Consequently, the central claim that edit-history frequencies yield a measurement of class completeness is only demonstrated for the particular mixture of human and bot activity in eight selected classes; it is not established as a property of collaborative KGs generally. This is not an objection to the approach, but an objection to accepting the empirical headline before the sampling model is validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a capture-recapture methodology to estimate the true cardinality of classes in Wikidata from edit histories. It formalizes class completeness estimation, introduces four non-parametric estimators (Jack1, N1-UNIF, SOR, Chao92) built on frequency-of-frequency counts, and defines an error metric φ (against external ground truth) and a convergence metric ρ (against the observed distinct count over a trailing window). The experimental study covers eight Wikidata classes with authoritative ground truth and reports per-estimator φ and ρ values, a discussion of burst-insertion effects, and large-scale examples of classes with low and high ρ. The authors conclude that Jack1 and SOR are the most accurate estimators and that ρ can be used to flag incomplete classes.","tokens_in":12515,"tokens_out":9361,"duration_ms":92847,"significance":"The paper addresses an important practical problem—measuring class completeness in collaboratively curated KGs—and the design has several strengths: the estimators are imported from independent ecological statistics rather than fitted to the data; the evaluation uses external authoritative counts for eight classes; the distinct-count lower bound is included as a baseline; and the authors release code, data, and a public dashboard. If the sampling model is valid, the approach provides a scalable way to prioritize editing effort in Wikidata. However, the current evidence for the model is indirect; the paper asserts independence and stationarity of edit mentions without diagnostics, and reports only point estimates. The contribution is therefore promising but not yet sufficiently validated for its stated claims.","major_comments":[{"comment":"Assumptions 2–4 (independence, with-replacement sampling, and time-invariant mention probabilities) are asserted rather than tested. The paper states 'we have not observed any significant correlations in the edits' without reporting any statistic, and Section 4.3 and Fig. 3(h) document batch insertions that violate these assumptions. Since f1 and f2 are the only inputs to Jack1, SOR, sample coverage, and Chao92, a systematic violation directly biases every estimator. Please provide diagnostics such as autocorrelation of class-level mention counts, comparison of estimates computed from human-only versus bot and bulk-import edit subsets, or a sensitivity analysis that removes burst periods, to bound the resulting bias.","section":"Section 3.2; Eqs. (1)–(9)"},{"comment":"No uncertainty is reported for any estimate. Variance estimators are available for the jackknife (Burnham and Overton) and for Chao92, and a nonparametric bootstrap is straightforward from the frequency counts; without them, the differences in φ between estimators (e.g., Jack1 27.4 versus SOR 36.0 for Video Game Consoles, or Jack1 1538 versus SOR 2663 for Hospitals) cannot be distinguished from sampling noise. Please add confidence intervals or standard errors to Table 1 and to the ρ comparisons.","section":"Section 3.4; Table 1"},{"comment":"The claim that 'Jack1 and SOR consistently achieve the lowest error rate across all classes' is not supported by the table. SOR has the lowest φ in only one class (Skyscrapers, 650.4); Jack1 has the lowest φ in six classes, but for Municipalities of the CZ the best estimator is N1-UNIF (22.2), and Jack1 (86.3) and SOR (31.3) are both worse than the distinct lower bound (26.6). Please either rephrase the conclusion to reflect the per-class pattern or provide a statistical test of the claimed dominance.","section":"Section 4.2; Table 1"},{"comment":"The convergence metric ρ measures the distance between an estimator and the observed distinct count D_i, not between the estimator and the true class size. If an estimator tracks D_i closely, ρ can be near zero even for an incomplete class; conversely, an incomplete class can have an intermediate ρ (Cathedrals of Mexico has N=93, D=63, and SOR ρ=0.0162, which is neither below 0.001 nor above 0.1). The binary thresholds therefore need to be validated against ground truth rather than illustrated by random examples.","section":"Section 3.4, Eq. (11); Section 4.3, Table 2"},{"comment":"The user-specified parameters—sample period length (30 days), convergence window (w=4), and the SOR singleton cap (2σ+μ)—are fixed without a sensitivity analysis. Because all estimates and ρ values depend on these choices, the reported thresholds and estimator rankings may change under different settings; please report how Table 1 and Table 2 vary with the sample period and w.","section":"Section 4.1; Eq. (11)"}],"minor_comments":[{"comment":"The definition of φ is garbled in the manuscript; please rewrite it with explicit summation limits and indices so that Table 1 is reproducible.","section":"Section 3.4, Eq. (10)"},{"comment":"The definitions of μ and σ are typeset incorrectly ('F∑ ... |F|−1', 'vuu√'); please clarify the summation index and whether the denominator is |F|−1 or |F|−2.","section":"Section 3.3, Eq. (7)"},{"comment":"The class label 'Paintings by Vincent van Gogh (T = 864)' should use N for the class size, and the composite class notation in Fig. 3(g)–(h) is hard to parse.","section":"Section 4.2"},{"comment":"Listing 1.1 relies on an 'edit:' prefix that is not available on the public Wikidata endpoint; a reproducible description of how mentions were extracted from the Wikibase XML dump would strengthen the paper.","section":"Listing 1.1"},{"comment":"For Mountains, the ground truth is acknowledged to be 'suggestive'; given that this class contributes the largest errors in Table 1, the ranking for that row should be interpreted with caution.","section":"Section 4.2, Fig. 3(e)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a data-management or knowledge-graph venue, and the public release of code and data is a strength. The main revision requirement should focus on validating Assumptions 2–4 and adding uncertainty quantification; without those, the headline empirical claims remain conditional. I would not require new theory, but the revision should be substantial rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper takes established species-richness estimators (Jackknife, Good-Turing, Chao92, SOR) and applies them to Wikidata edit mentions to estimate class cardinality and completeness. That application is new, and the rho convergence metric is a sensible addition. The empirical work is honest: eight classes with independent ground truth, and Jack1 and SOR beat the distinct-count lower bound on most of them. They also ship code, preprocessed data, and per-class results, which makes the work reproducible.\n\nThe main problem is the load-bearing sampling model in Section 3.2. The estimators and rho all assume independent, stationary, with-replacement observations on a closed class. The paper acknowledges that dependencies exist but says \"we have not observed any significant correlations\" without giving any statistic. More importantly, Wikidata editing is heavily driven by bots and bulk imports, and the paper itself shows a batch-import burst causing overshooting in Fig. 3(h). That is exactly what a violation of the independence/stationarity assumptions predicts. So the gains in Table 1 are only demonstrated for the particular mix of human and bot activity in eight selected classes, not for collaborative KGs generally.\n\nThe other soft spots are minor by comparison. There are no confidence intervals anywhere, so the point estimates in Table 1 cannot be separated from sampling noise. The rho metric inherits whatever bias the underlying estimators have, so calling a class \"converged\" is only as good as the estimator. And while the related work is cited well, no direct comparison is made to existing KG completeness methods (ReCoin, Galarraga et al., Soulet et al.), though those target different notions of completeness and the omission is not fatal.\n\nThe paper is clearly written, the math is standard, and the authors are upfront about several limitations, including burst sensitivity and the closed-class restriction. The central idea is plausible and the empirical evidence is suggestive, but the sampling model needs validation or at least a robustness check before the headline claim is fully supported. The right fix would be to measure the degree of assumption violation, report confidence intervals, and compare against a simple baseline (e.g., distinct count plus a perturbation). This is a solid candidate for peer review, not a desk reject, because it is a legitimate new application with reproducible artifacts and a useful practical tool. I would engage with it, and I would cite the rho metric and the empirical comparison if I worked on KG data quality.","headline":"A useful and reproducible application of species richness estimators to Wikidata edit logs, with independent ground truth on eight classes; the main soft spot is the unvalidated sampling model, not the estimator math.","tokens_in":768,"tokens_out":886,"would_cite":true,"duration_ms":27705,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that a knowledge graph's edit history is enough to estimate class completeness.","keywords":["knowledge graph completeness","class cardinality estimation","Wikidata","species richness estimation","capture-recapture","non-parametric estimators","edit history","collaborative knowledge graphs"],"falsifier":"Take a class with a known true size and a known history of bulk imports, then recompute Jack1 and SOR estimates from the edit log: if the estimators overshoot the true size whenever a bot inserts many instances at once, and the $\\rho$ metric nonetheless drops below 0.001, then the stationarity and independence assumptions do not hold.","tokens_in":12060,"feed_emoji":"🗃️","tokens_out":5306,"duration_ms":52726,"temperature":0.7,"pith_summary":"This paper tries to establish that a knowledge graph's edit history can be used to estimate how many instances a class ought to have, and therefore how complete the class is. It treats every edit that mentions an entity of a class as a capture event, borrows species-richness estimators from ecology, and applies them to monthly samples of Wikidata edits. The authors argue that the resulting estimates of true class size are reliable enough to distinguish classes that are complete from classes that still have missing instances. If true, this gives Wikidata editors and consumers a practical way to find knowledge gaps without manually checking candidate entities.","feed_headline":"Edit logs can reveal how many entries a Wikidata class is missing","feed_subtitle":"Species-richness estimators and a convergence score turn monthly edit mentions into completeness estimates.","key_machinery":"The machinery is the frequency-of-frequencies vector computed from mentions: $f_1$ counts instances mentioned exactly once, $f_2$ counts instances mentioned exactly twice, and so on, with $f_0$ the unobserved instances that the estimators try to recover. Four non-parametric estimators use this vector: Jack1 and Jack2 leave one sample period out; N1-UNIF uses the sample-coverage estimate $\\hat{S} = 1 - f_1/n$; SOR caps $f_1$ at two standard deviations above the mean to reduce singleton bursts; and the coverage-based estimator adds a coefficient-of-variation correction. The convergence metric $\\rho$ averages the relative distance between the estimate and the observed distinct count over the last $w$ sample periods, providing a completeness signal that does not require ground truth.","core_discovery":"The paper's central claim is that the true size $N$ of a finite class in a collaborative knowledge graph can be estimated from the edit history alone, by treating each month's mentions of class instances as a capture-recapture sample. Given the currently observed count $D$, the class is complete when $D = N$, so an estimate of $N$ becomes a completeness estimate. On Wikidata classes whose true sizes are known from external sources, the jackknife estimator Jack1 and the singleton-outlier-reduction estimator SOR consistently give the lowest error, while the convergence metric $\\rho$ separates complete classes ($\\rho < 0.001$) from incomplete ones ($\\rho > 0.1$).","pith_inferences":["Editorial inference: the $\\rho$ threshold could be turned into a monitoring signal for bulk-import workflows, because the paper itself notes that bursts of inserts cause some estimators to overestimate; a class whose $\\rho$ suddenly drops after a mass import should be re-checked rather than trusted.","Editorial inference: classes maintained mostly by one bot or a small group of editors should be expected to produce less reliable estimates than classes edited by many independent volunteers, since independence of mentions is the load-bearing assumption.","Editorial inference: a natural testable extension is to weight mentions by page-view attention, which the paper lists as future work; if popularity drives mention probability, the stationary assumption is violated and estimates should degrade on celebrity-heavy classes.","Editorial inference: the same frequency-of-frequencies machinery could be applied to detect systematic misclassification in the ontology, since a wave of reclassifications would look like a burst of mentions that changes the apparent class size."],"forward_implications":["A Wikidata consumer can ask whether a class is complete by checking whether the convergence metric $\\rho$ stays below 0.001 over the last few monthly samples, without needing an external ground-truth count.","Editors and projects can direct effort toward classes with high $\\rho$ values, since those are the classes the estimators judge to be far from complete.","The method generalizes to any collaborative knowledge graph that keeps an action log with timestamps, not just Wikidata.","For classes that are already complete, the estimators slowly converge toward the true size from below, so a low error does not require waiting for the class to finish growing."],"supporting_citations":[{"why":"Supplies the review of species-richness estimators that motivates the choice of non-parametric capture-recapture techniques.","marker":"[2]"},{"why":"Provides the closed-form first- and second-order jackknife estimators used for Jack1 and Jack2.","marker":"[3]"},{"why":"Introduces the sample-coverage approach and the coverage-based estimator that the paper adapts to class cardinality estimation.","marker":"[4]"},{"why":"Gives the Good-Turing sample-coverage estimator that underlies N1-UNIF and the coverage-based estimator.","marker":"[10]"},{"why":"Contributes the crowdsourced enumeration setting, the error metric, and the singleton-threshold idea that becomes SOR.","marker":"[19]"}],"fun_headline_variants":["Edit logs reveal missing Wikidata class entries","Species estimation methods show Wikidata class completeness","How many Wikidata items are missing? Edit history estimates","Jackknife and convergence score gauge Wikidata completeness","From edit bursts to completeness: Wikidata class estimation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method works only if the mentions of a class's instances in the edit log are like random draws from a fixed list, with each instance having a constant chance of being mentioned in any month and edits made independently of one another.","fun_headline_variants_meta":{"raw":{"variants":["Edit logs reveal missing Wikidata class entries","Species estimation methods show Wikidata class completeness","How many Wikidata items are missing? Edit history estimates","Jackknife and convergence score gauge Wikidata completeness","From edit bursts to completeness: Wikidata class estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1421,"prompt_tokens":930,"completion_tokens":491,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":423}},"tokens_in":546,"tokens_out":491,"duration_ms":6044,"temperature":1.0,"reasoning_tokens":423,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:26:16.817714+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a class with a known true size and a known history of bulk imports, then recompute Jack1 and SOR estimates from the edit log: if the estimators overshoot the true size whenever a bot inserts many instances at once, and the $\\rho$ metric nonetheless drops below 0.001, then the stationarity and independence assumptions do not hold.","supporting_citations":[{"cited_title":"Journal of the American Statistical Association 88(421), 364–373 (1993)","cited_arxiv_id":null,"evidence_quote":"Supplies the review of species-richness estimators that motivates the choice of non-parametric capture-recapture techniques."},{"cited_title":"Ecology 60(5), 927–936 (1979)","cited_arxiv_id":null,"evidence_quote":"Provides the closed-form first- and second-order jackknife estimators used for Jack1 and Jack2."},{"cited_title":"Journal of the American statistical Association 87(417), 210–217 (1992)","cited_arxiv_id":null,"evidence_quote":"Introduces the sample-coverage approach and the coverage-based estimator that the paper adapts to class cardinality estimation."},{"cited_title":"Biometrika 40(3-4), 237–264 (1953)","cited_arxiv_id":null,"evidence_quote":"Gives the Good-Turing sample-coverage estimator that underlies N1-UNIF and the coverage-based estimator."},{"cited_title":"In: ICDE","cited_arxiv_id":null,"evidence_quote":"Contributes the crowdsourced enumeration setting, the error metric, and the singleton-threshold idea that becomes SOR."}],"review_version":1}