{"id":"47b94ef3-3c22-4c85-b23c-c3f515dbc98b","arxiv_id":"2412.14351","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Early citations forecast later citations far better than publication venue does across ACL, PubMed, and arXiv papers.","lead":"This paper asks whether where a paper is published or how quickly it is cited is the better crystal ball for its future importance. Across ACL, PubMed, and arXiv papers, early citations win, and the authors suggest trimming peer review in favor of citation-based triage.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'predictive' claim is never evaluated out of sample; all reported evidence is in-sample, so the forecasting conclusion is not actually demonstrated.","rationale":"The reader's weakest_assumption pins the argument on the normative equation of citations with value. That is a legitimate interpretive concern, but the paper's central claim, as stated in the abstract and by the reader, is empirical: early returns are more predictive than venue. The most load-bearing problem for that claim is the absence of any out-of-sample evaluation. The paper repeatedly uses 'predictive' and 'forecasting' language, yet every result—correlations, group summaries, regression coefficients, ANOVA—is computed on the same data used to fit the model. Figure 3 even omits actual outcomes, so it cannot establish predictive accuracy. The paper does have strengths: the descriptive correlations are large, the data are public, and the result is replicated across three corpora and multiple years. Those make the direction of the effect likely robust. But the magnitude of the advantage over venue, and whether it holds prospectively, remains unquantified. A conditional acceptance should require the simple train/test split described above. This is not a rejection: the effect is large enough that out-of-sample validation will almost certainly confirm early citations are strongly predictive. The concern is that the paper's central claim is stated more strongly than its evidence. We therefore partially agree with the reader: the normative assumption matters for the policy conclusion, but the missing out-of-sample validation directly undermines the empirical headline.","tokens_in":14336,"tokens_out":8588,"duration_ms":78551,"concrete_test":"Split the data by publication year: train the Eq. 1 regression on papers published in 2016-2018 and test on 2019 papers. For three models—(a) early-citation factors only, (b) venue only, and (c) both—compute out-of-sample Spearman correlation and RMSE between predicted and observed fourth-year citation percentiles. Report the difference (a)-(b) with bootstrap confidence intervals. If the out-of-sample advantage of early citations over venue is small or not significantly different from zero, the headline claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that early citations are more predictive of future citations than venue. Every piece of evidence is in-sample. Table 2 reports autocorrelations among citation counts for the same papers; Table 3 reports correlations with a venue dummy, which is not directly comparable to a continuous count. Tables 4 and 5 compare group means after the fact, not forecasts. The regression in Eq. 1 is fit and described on the same data; Figure 3 plots model predictions, not predicted-versus-actual outcomes, and no out-of-sample R², rank correlation, or forecast-accuracy metric is reported. The ANOVA cited in Section 3.3 decomposes variance in the fitted sample; it does not measure predictive skill. Because the DDI proposal in Section 4.2 explicitly proposes using early citations as a selection mechanism, the paper needs prospective evidence that early citations would select future highly cited papers better than venue. Without out-of-sample validation, the statement 'early returns are more predictive than venue' is an in-sample association, not a demonstrated forecasting result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper frames the question of whether peer review identifies important papers as a forecasting problem. Using Semantic Scholar citation data for papers published in 2016-2019 in ACL Anthology, PubMed, and arXiv (with detailed analyses for 2016-2017), it compares two predictors of fourth-year citation standing: early citations (counts in the first year after publication) and publication venue. Evidence is presented in three forms: lagged correlation matrices of citation counts (Table 2) versus point-biserial correlations with venue dummies (Table 3, Figure 1); group summaries (h-index, impact factor mu, N, sigma) for papers grouped by early-citation thresholds versus by venue (Tables 4-5, Figure 2); and a regression model (Eq. 1) with a percentile outcome and factor terms for venue and early citations (Table 6, Figure 3). The authors conclude that early returns are more predictive and more robust than venue, propose posting papers on arXiv and using early citations plus nominations to triage reviewing (the DDI proposal in Section 4.2), and explicitly state the normative assumption that assessments of value should be leading indicators of future citations (Section 2.2).","tokens_in":14512,"tokens_out":15997,"duration_ms":129310,"significance":"If the forecasting claim is substantiated, the paper offers a broadly useful, policy-relevant result: a simple, transparent signal (early citations) that dominates venue labels for anticipating future citation standing, with implications for reading and reviewing effort at scale. The paper is strongest in its reproducibility and scale: data and code are posted on GitHub, samples include roughly one million papers per year for PubMed, and the main comparisons are replicated across three corpora and multiple publication years. The central descriptive finding is internally consistent and agrees with prior bibliometric literature (e.g., Abramo et al. 2019; Wang et al. 2013), and the normative assumption is stated transparently rather than hidden. However, two issues currently cap the significance. First, all reported evidence is in-sample; a prospective evaluation is needed before the word 'predictive' can support the DDI selection mechanism. Second, the evaluative conclusion inherits an untested assumption that citations measure value.","major_comments":[{"comment":"The forecasting claim is only demonstrated in-sample, and this gap is load-bearing because the abstract frames the task as prediction and Section 4.2 proposes early citations as a selection mechanism. The regression described by Eq. (1) is fit once per publication year, and the boxplots in Figure 3 show fitted values for the same papers used to estimate the coefficients; no train/test split, out-of-sample R², rank correlation, or forecast-error metric is reported. The ANOVA invoked in Section 3.3 is stated without any test statistics, and Section 6 lists limitations about gaming, language coverage, and equity risks but does not acknowledge that the evaluation is entirely in-sample. This is not a circularity problem (early citations are lagged observations of the same citation stream), but it is a missing prospective test. Given the sample sizes, the shrinkage one would expect out of sample is probably modest, but the manuscript provides no way to assess it. I recommend fitting on 2016-2018 and evaluating on 2019 (or leave-one-year-out), reporting out-of-sample forecast accuracy for venue-only, early-citations-only, and combined models.","section":"§3.3, Eq. (1), Figure 3"},{"comment":"The headline correlation comparison is not metric-comparable as presented. Table 2 reports Pearson correlations between two continuous citation counts (e.g., 0.80 for 2016 vs 2017 ACL papers), while Table 3 reports point-biserial correlations of individual venue indicator variables with future citation counts (e.g., 0.14 for ACL Conf). Comparing these magnitudes is apples-to-oranges: a continuous variable carries more information than a single binary dummy, so the gap in Figure 1 partly reflects variable coding rather than predictive merit. The regression in Eq. (1), which treats venue as a full factor and early citations as a factor, is the appropriate symmetric design; reporting the variance explained or multiple correlation of each factor from that model, together with the out-of-sample evaluation requested above, would put both predictors on a common footing.","section":"§3.1, Tables 2-3, Figure 1"},{"comment":"The evaluative conclusion is conditional on an untested assumption, and this condition should be built into the conclusions and the DDI proposal. The paper states in Section 2.2 that 'we assume reviews and other assessments of value should be leading indicators of future citations,' and the recommendation to use early citations instead of reviewing inherits that assumption. If citations track visibility, author reputation, or trend rather than scholarly value, the demonstrated association advantage of early citations does not by itself establish that early citations should replace review. This is a scope concern rather than an internal inconsistency: the paper is transparent about the assumption, but the title question 'Is peer-reviewing worth the effort?' overreaches the evidence unless the conclusions read as conditional on the citation-impact definition. I suggest either explicitly conditioning the conclusions (e.g., 'for the purpose of predicting citation impact') or adding a small validation against an independent signal of value such as expert ratings, awards, or follow-on usage data.","section":"§2.2, §4.2"}],"minor_comments":[{"comment":"In the PubMed row for 2017, the sample size appears as '107,7437,' which is likely a typo for 1,077,437; please correct it.","section":"Table 5"},{"comment":"The text contains a few editing artifacts: 'Table 4 does this The main observation is...' is missing punctuation, and 'as evidence by the large σ' should read 'as evidenced by.'","section":"§3.2"},{"comment":"The percentile outcome is not fully defined: percentile within which reference set (per source, per year, pooled over sources)? Since the coefficients in Table 6 are interpreted as percentile changes, this should be specified precisely.","section":"§3.3"},{"comment":"The h-index comparisons between early-citation groups and venue groups conflate group size, since h grows with N; the µ comparisons are size-normalized, but the bullets in §3.2 and §4.1 that cite h should be read alongside N or a size-normalized variant such as h/N.","section":"§3.2, Tables 4-5"},{"comment":"Semantic Scholar's venue field mixes journals, conferences, and non-peer-reviewed outlets, so venue correlations are an attenuated proxy for peer-review outcomes; a brief caveat near the headline figure would help readers connect the results to the title question.","section":"Figure 1 caption, §1.1"},{"comment":"The motivation for the citation-as-value assumption cites only the authors' prior essays (Church 2005, 2020); citing the broader literature on citation-based research evaluation would strengthen the grounding.","section":"§2.2"},{"comment":"The header 'V enue Id' contains a stray space; Table 1 is illustrative only and would benefit from a sentence saying so.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper's empirical core is sound but largely confirms established bibliometric findings; its distinctive contribution is the framing and the DDI/nomination proposal, which the paper explicitly offers as a discussion starter. The main fixable gap is the absence of out-of-sample validation for the forecasting claim, which I judge to be within the authors' reach. I also noted that the motivating assumption in Section 2.2 leans on the authors' own prior essays; the external literature is cited elsewhere, so this is a minor issue. The scope question worth considering is whether the journal's audience expects archival empirical results or accepts a position-paper contribution with a small empirical core."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does one thing well and one thing prematurely. What it does well: it replays the Abramo et al. (2019) result on much larger data—ACL, PubMed, arXiv, multiple years—and shows over and over that early citations correlate with future citations far more than venue does. The GitHub data is a real asset, the methods are simple enough to check, and the paper is honest that the central finding is not new. Credit where due: this is a useful replication and an interpretable demonstration for a policy audience.\n\nThe premature part is the forecasting frame. The abstract says the question is forecasting, and the conclusion says early returns are 'more predictive' than venue. But the regression evidence is entirely in-sample. Figure 3 plots fitted values, not predicted-versus-actual; the ANOVA decomposes variance in the training data; and there are no confidence intervals or out-of-sample metrics anywhere. The correlations in Table 2 are lagged autocorrelations, which are legitimate evidence of association, but they do not by themselves establish that a selection rule based on early citations would prospectively outperform one based on venue. For the DDI proposal in Section 4.2—which is explicitly about using early citations to decide what gets reviewed—prospective evidence is exactly what is needed.\n\nThere are two other soft spots worth naming, though neither sinks the empirical comparison. First, the venue correlations in Table 3 use binary dummy variables against continuous citation counts, so the correlation sizes are not directly comparable; the h/µ tables in Section 3.2 are the more compelling evidence. Second, the normative assumption that citations measure value is stated (Section 2.2) but not defended. The entire 'is peer review worth the effort' framing depends on that assumption, and the paper would be stronger if it acknowledged more clearly that it is evaluating peer review against a citation-based yardstick, not against any independent notion of quality.\n\nI agree with the reader's conditional verdict. The central empirical claim is probably true and is worth taking seriously, but the paper overreaches when it turns an in-sample association into a forecasting result and then uses that result to recommend changing how conferences operate. The flaws are addressable: add a held-out year or a proper temporal split, report error bars, and temper the policy language. I would send this to a serious referee, ask for those changes, and expect a publishable paper after revision. I would also bring it to our reading group; it is a good conversation starter on peer review and bibliometrics.","headline":"A clean, large-scale replication that early citations beat venue, but the forecasting claim needs out-of-sample validation before the DDI proposal carries weight.","tokens_in":15023,"tokens_out":1564,"would_cite":true,"duration_ms":16605,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Citations in the first year after publication predict a paper's future impact better than the venue that published it.","keywords":["peer review","citation prediction","early citations","venue prestige","h-index","impact factor","bibliometrics","forecasting"],"falsifier":"Find a corpus where year-one citation counts are decoupled from later impact: for example, a field with many 'sleeping beauties' (papers ignored for years that later become highly cited) or a venue where papers with zero early citations but strong reviewer endorsements consistently beat 20+ early-citation papers in fourth-year percentiles. Showing such a pattern in a large sample would refute the paper's claim that early returns dominate venue as a forecast.","tokens_in":14128,"feed_emoji":"📈","tokens_out":13619,"duration_ms":98906,"temperature":0.7,"pith_summary":"This paper asks whether peer review is worth the effort by turning review quality into a forecasting problem: can we predict which papers will be highly cited later? The authors find, across three large publication corpora, that a paper's citation count in its first year after publication (early returns) is a much stronger predictor of its fourth-year citation percentile than the venue that published it. In their tables, papers with 20+ early citations have higher impact and h-index than every venue they examined, and papers in less selective venues with a few early citations outrank papers in more selective venues with none. The authors conclude that selecting papers by early citations would be more selective, more inclusive, and more stable than selecting by venue, and they sketch a 'Don't Do It' (DDI) alternative in which papers are posted on a preprint server and program committees focus on papers with impressive early citations plus nominations. The paper frames this as evidence that a simple citation-based rule could do much of the work peer review is currently asked to do, and it offers the DDI sketch mainly to start discussion.","feed_headline":"Early citations beat venue as impact forecast","feed_subtitle":"In three large publication corpora, year-one citations outpredict the publication venue.","key_machinery":"The central objects are 'early returns' and 'venue' as competing grouping variables, measured by three summary statistics: correlation $\\rho$ between year-by-year citation counts, the h-index $h$, and impact factor $\\mu$ (average citations). The paper's main mechanical move is to compare groups of papers conditioned on early citations (1+, 2+, 3+, 10+, 20+ citations in year one) against groups conditioned on venue, using fourth-year citations as the outcome, and then to fit an interpretable linear regression $percentile_{year+4} \\sim venue + factor(\\min(T, citations_{year+1}))$ so that each early-citation count receives its own coefficient. The h-index and impact factor are usually computed per author or per venue; here they are repurposed as evaluation metrics for the two grouping schemes, which is what lets the paper compare exclusivity, inclusivity, and stability in one analysis.","core_discovery":"The paper's central claim is that early returns are more predictive than venue. Concretely, when papers are grouped by how many citations they received in the first year after publication, those groups have larger correlations with fourth-year citations ($\\rho \\approx 0.8$ between consecutive years, versus venue correlations mostly below $0.15$), higher h-index $h$ and impact $\\mu$, and larger counts $N$ than grouping by venue. A regression predicting fourth-year citation percentiles from early citations plus venue shows the early-citation coefficients dominate: a paper with 10+ early citations is predicted to fall in the 75th percentile or better, while venue coefficients are small and unstable across years. The authors frame this as evidence that peer review, as currently practiced, is a costly way to achieve what a simple citation threshold already does.","pith_inferences":["The paper does not say this, but its logic implies that review scorecards could be benchmarked against early-citation forecasts: a review is useful exactly to the extent it improves prediction of later impact beyond what year-one citations already give.","A direct testable extension would be to run the same early-returns-versus-venue comparison on fields with long citation lags, where 'sleeping beauties' (papers ignored for years that later become highly cited) are more common; the authors note the phenomenon but do not quantify how much it erodes the rule.","If committees adopted the early-citation rule, author incentives would shift toward quick visibility and self-citation; the paper acknowledges that citations can be gamed but does not model how thresholds would change behavior.","One implicit consequence is that the 'Don't Do It' proposal treats the venue label as almost redundant for impact forecasting, which would reallocate prestige from editorial selection to community citation behavior; that change would take years to validate."],"forward_implications":["If early returns really are more predictive than venue, readers, authors, and committees can rank candidate papers by first-year citations rather than by journal or conference prestige.","Conference programs that shift effort toward papers with early citations plus nominations could cut the number of full reviews while keeping or improving the quality of accepted papers, since the marginal value of a venue label is small.","The regression results imply that a paper with 10+ citations in year one is expected to land above the 75th percentile in year four, so early-citation thresholds provide a concrete and cheap triage rule.","Because the finding replicates across three large corpora and across 2016 and 2017 publication cohorts, the authors expect the advantage of early citations over venue to generalize across fields and time periods.","Venue effects, while statistically significant, are so small and unstable from year to year that they have little practical consequence for prioritization."],"supporting_citations":[{"why":"Supplies the citation counts and venue labels used in every table and figure; the entire empirical comparison depends on this dataset.","marker":"Wade (2022)"},{"why":"Establishes early citations as a predictive feature for long-term scientific impact, the feature the paper shows dominates venue.","marker":"Wang et al. (2013)"},{"why":"Prior result that early citations make journal impact factor negligible after two years; the paper extends this result from impact factor to venue.","marker":"Abramo et al. (2019)"},{"why":"Treats the identification of high-impact papers as a forecasting task, the framing this paper adopts.","marker":"Davletov et al. (2014)"},{"why":"Defines the h-index used as one of the two summary statistics for comparing early-citation groups with venue groups.","marker":"Hirsch (2005)"},{"why":"Defines the impact factor used as the other summary statistic for the same comparison.","marker":"Garfield (2006)"},{"why":"Justifies using percentile ranks of citation counts, the transform applied to the outcome in the regression model.","marker":"Bornmann et al. (2012)"},{"why":"Supports improving early citation prediction with percentile-based measures, informing the paper's forecasting setup.","marker":"Bornmann et al. (2014)"},{"why":"Provides a conference review experiment showing review decisions correlate only weakly with future citations, motivating the comparison.","marker":"Cortes and Lawrence (2021)"},{"why":"Systematic review concluding peer review rests on faith rather than facts, which frames the paper's question.","marker":"Jefferson et al. (2002)"}],"fun_headline_variants":["Peer review loses to early citation counts","Year-one citations beat venue as impact signal","Is peer review worth it? Early citations say no","Citation forecast beats journal prestige"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is stated in Section 2.2: reviews and other assessments of value should be leading indicators of future citations, meaning the paper equates 'important paper' with 'highly cited paper'; if citations track visibility, fashion, or manipulation rather than scholarly value, the comparison says nothing about whether peer review is worth the effort.","fun_headline_variants_meta":{"raw":{"variants":["Peer review loses to early citation counts","Year-one citations beat venue as impact signal","Is peer review worth it? Early citations say no","Citation forecast beats journal prestige"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1315,"prompt_tokens":752,"completion_tokens":563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":368,"completion_tokens_details":{"reasoning_tokens":510}},"tokens_in":368,"tokens_out":563,"duration_ms":5098,"temperature":1.0,"reasoning_tokens":510,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:17:55.405149+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a corpus where year-one citation counts are decoupled from later impact: for example, a field with many 'sleeping beauties' (papers ignored for years that later become highly cited) or a venue where papers with zero early citations but strong reviewer endorsements consistently beat 20+ early-citation papers in fourth-year percentiles. Showing such a pattern in a large sample would refute the paper's claim that early returns dominate venue as a forecast.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the citation counts and venue labels used in every table and figure; the entire empirical comparison depends on this dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes early citations as a predictive feature for long-term scientific impact, the feature the paper shows dominates venue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior result that early citations make journal impact factor negligible after two years; the paper extends this result from impact factor to venue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the h-index used as one of the two summary statistics for comparing early-citation groups with venue groups."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the impact factor used as the other summary statistic for the same comparison."},{"cited_title":"The use of percentiles and percentile rank classes in the analysis of bibliometric data: Opportunities and limits","cited_arxiv_id":"1211.0381","evidence_quote":"Justifies using percentile ranks of citation counts, the transform applied to the outcome in the regression model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports improving early citation prediction with percentile-based measures, informing the paper's forecasting setup."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic review concluding peer review rests on faith rather than facts, which frames the paper's question."}],"review_version":1}