{"id":"273d7f5e-8235-4433-9760-6d4446233a9b","arxiv_id":"2411.09768","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"On 117,650 articles across 26 fields, ChatGPT 4o-mini gave systematically higher quality scores to newer articles, with field and country differences also present.","lead":"Researchers sent titles and abstracts of over 117,000 journal articles to ChatGPT and found that scores rise with publication year and differ a lot by field and country. The result matters because ChatGPT is being proposed as a support tool for research evaluation, and any such tool needs known biases before it is used.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's year-bias conclusion depends on an unproven constant-quality assumption; the 786-character abstract cutoff may exacerbate it.","rationale":"Reader's verdict is already CONDITIONAL and identifies the same weakest assumption. My stress-test agrees: the descriptive and regression results are clear, the 26-field robustness is strong, and the practical field/year normalization is a reasonable safeguard. The soft spot is not statistical but inferential: 'year-dependent' is established; 'biased by year' requires a quality baseline. I add the abstract-threshold interaction as a specific mechanism that could distort the year comparison even if journal identity is controlled. The proposed expert re-scoring is the decisive check because it supplies the missing counterfactual. If such a baseline is infeasible, the paper should soften the conclusion to 'year-dependent scores' and present normalization as cautionary rather than as correction of a proven bias. This does not change the reader's conditional verdict.","tokens_in":8798,"tokens_out":4792,"duration_ms":52622,"concrete_test":"Take a random subsample of, say, 50 articles per field from 2003 and 2023 within the same sampled journals (N≈2,600), redact the publication year, and have two REF-style expert reviewers score each abstract with the same four guidelines used for ChatGPT. Compare the mean expert year gap with the ChatGPT year gap for the same articles. If the expert gap is near zero while ChatGPT's gap remains large, the bias interpretation is supported; if experts also score 2023 substantially higher, the year trend may be correct quality assessment rather than bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The observed year trend is statistically solid, but the inferential step from trend to bias requires the counterfactual that the average intrinsic quality of articles drawn from the same journals was roughly constant across 2003-2023. The paper concedes: 'the quality of journals seems to be relatively stable. This is unproven' (Methods, Data). Journal identity is held fixed, but editorial standards, journal scope, and Scopus field coverage can drift over 20 years. Moreover, the 786-character abstract threshold discards 25% of articles; if abstract-length norms have increased over time, the threshold removes different fractions per year and can create or mask a year gradient. Since there is no human expert baseline for the same articles, the year coefficient could reflect genuine improvements in research quality rather than ChatGPT bias. The normalization recommendation would then remove true quality differences. This is not an internal inconsistency, but it makes the central 'bias' claim conditional on an untested assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale measurement study in which 117,650 articles from 26 Scopus broad fields and five publication years (2003, 2008, 2013, 2018, 2023) were scored by ChatGPT 4o-mini using REF-style evaluation guidelines. The central descriptive findings are that ChatGPT scores increase with publication year in all 26 fields, that the increase survives controls for first-author country and title/abstract length, and that average scores differ substantially between fields. The paper also reports associations between scores and abstract length, first-author country, and citation counts. The authors interpret the year and field differences as biases and recommend normalizing ChatGPT scores by field and year before use in research evaluation.","tokens_in":8946,"tokens_out":4386,"duration_ms":47526,"significance":"If the year and field differences are genuine biases, the paper makes an important practical contribution: it would show that ChatGPT-based research quality scores cannot be compared across time or fields without normalization, and it would extend the known bias literature for LLM-based evaluation. The study has notable strengths: a large and carefully constructed journal-balanced sample, transparent random sampling, explicit control variables in regression analyses, and prior validation of the REF-based prompt in other work. The positive correlations with citation counts provide useful context. The main weakness is that the central inferential step—from a monotone time trend to the claim that the trend is a bias—depends on the unproven assumption that average intrinsic quality in the sampled journals was roughly constant over the period. The paper itself acknowledges this assumption is unproven, which makes the normalization recommendation conditional on an untested premise.","major_comments":[{"comment":"The inference that the positive year coefficient is a bias rather than a reflection of genuine quality change requires that the average intrinsic quality of articles in the sampled journals was approximately constant between 2003 and 2023. The paper explicitly states: \"the quality of journals seems to be relatively stable. This is unproven.\" No human expert baseline is available for the same articles, so the positive year coefficient in all 26 field regressions is equally compatible with a real improvement in research quality over time. This assumption is load-bearing because it converts a descriptive trend into the recommendation that ChatGPT scores be normalized for year. I would like to see either (a) evidence that average expert-assessed quality was stable in these journals over the period, (b) a subsample with human scores on the same articles, or (c) a clear reframing of the conclusion as documenting a year association rather than a demonstrated bias.","section":"Methods, Data and Conclusions"},{"comment":"The minimum abstract length threshold of 786 characters discards 25% of Scopus articles. If abstract-length norms have changed over the 20-year period, the truncation removes different fractions of articles in different years and can create, amplify, or mask a year gradient. The paper does not report the fraction of excluded articles by year and field, nor does it provide a sensitivity analysis at lower thresholds. Because both the year effect and the abstract-length effect are central to the paper (RQ1 and RQ3), this design choice is not merely a detail. Please report the distribution of excluded articles and re-estimate the main regressions with a lower threshold or an explicit selection correction.","section":"Methods, Data"},{"comment":"The regression analysis treats ChatGPT scores as an interval scale and fits publication year as a single linear term. This imposes a linear trend across the five discrete years, even though the paper's headline pattern is that each successive year scores higher than the previous one in 101 of 104 cases. Fitting year as a categorical variable (or using an ordinal model) would show whether the increase is uniform or concentrated in particular years, which matters for the recommended field-year normalization. If the trend is nonlinear, the normalization procedure may need to be adjusted accordingly.","section":"Regression"}],"minor_comments":[{"comment":"The word \"dependant\" should be \"dependent\".","section":"Regression"},{"comment":"The citation \"Thelwall, 2024ab\" combines two distinct references; the text should cite \"Thelwall, 2024a, 2024b\" individually.","section":"ChatGPT procedure"},{"comment":"Figure 4's error bars are described as the minimum and maximum values from the 26 field estimates, but it is unclear whether the plotted coefficients are standardized; please clarify because year and log-length coefficients are not on the same scale as binary country indicators.","section":"RQ1-4: Regression results"},{"comment":"The choice of 786 characters as the abstract-length threshold is justified heuristically; a sensitivity check around this threshold would strengthen the paper, as noted in the major comments.","section":"Methods, Data"},{"comment":"The discussion of the abstract-length association states that \"the second and third options play a role\" but does not quantify the relative contribution of journal-level quality effects versus short-form content; this remains an interpretation rather than a formal test.","section":"RQ3"}],"recommendation":"major_revision","confidential_remarks":"This is a useful descriptive study with a strong dataset, but the causal language of \"bias\" is not fully supported without a credible identification strategy for the constant-quality assumption. The authors' own admission that the assumption is unproven should be taken seriously, since the recommended normalization would be harmful if the year trend reflects genuine quality improvement. The paper is likely fixable within its scope by adding sensitivity analyses and reframing the central claim more cautiously, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives the first large-scale evidence that ChatGPT 4o-mini scores increase with publication year across all 26 Scopus fields tested. The year coefficient is positive and statistically significant in every field-level regression, even after controlling for first-author country and abstract length. That descriptive result is solid and I would bet on it replicating. The open question is whether it is a bias. The authors infer bias from an assumption—that the average intrinsic quality of articles in the same journals was roughly constant from 2003 to 2023—and they admit in the Methods that this is unproven. Without a human expert baseline on the same articles, the trend could reflect genuine improvements in research quality. This is not a hidden flaw; they flag it, but it is load-bearing.\n\nWhat is genuinely new: I know of no other study that probes year, country, length, and field bias in one design with 117,650 articles and regression controls. The sampling is careful (monodisciplinary journals, balanced by journal across years). The code and data pipeline are on GitHub, so the measurements are reproducible. The discussion of abstract length is also thoughtful—they distinguish a direct ChatGPT length preference from second-order effects like journal quality or short-form articles, and they conclude the bias interpretation is less likely there. The practical recommendation to normalize by field-year means is clearly motivated and easy to implement.\n\nSoft spots, in proportion. The constant-quality assumption is the main one. The 786-character abstract cutoff drops 25% of articles; if abstract norms drifted upward over time, the cutoff could select a different slice of the quality distribution per year. The log-length control reduces this inside the sample but does not test the selection effect. I'd want a robustness check with a lower threshold or an analysis of the excluded articles. Also, single submission per article adds noise but should not bias the year comparison. The R2 values are modest (mean 0.12), so year is a small influence—the authors acknowledge this.\n\nWho should read it: anyone building or using LLM-based research evaluation, and scientometricians interested in fair cross-year comparisons. It deserves a serious referee. I would send it out with a request for additional robustness on the cutoff and a softened claim about \"bias\" unless the constant-quality premise gets direct support.","headline":"Solid large-scale evidence that ChatGPT scores rise with publication year in all fields, but the 'bias' label rests on an untested constant-quality assumption; still worth refereeing.","tokens_in":9430,"tokens_out":2753,"would_cite":true,"duration_ms":28310,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that ChatGPT's research-quality scores rise systematically with publication year in all 26 fields tested, and argues the scores must be normalised for field and year before being used in research evaluation.","keywords":["ChatGPT","research evaluation","publication year","field normalization","abstract length","LLM bias","citation correlation","quality scoring"],"falsifier":"Have an expert panel blind-score a random subset of the same 117,650 articles; if expert ratings are flat across years while ChatGPT scores rise, the recency effect is a bias. Or submit the same abstracts with contemporary vocabulary swapped for neutral phrasing; if scores rise with the modern wording, the year effect is carried by textual cues rather than by true quality.","tokens_in":8596,"feed_emoji":"🤖","tokens_out":8034,"duration_ms":78247,"temperature":0.7,"pith_summary":"The paper tests ChatGPT's research-quality scores for biases by submitting 117,650 title-abstract pairs from 26 fields and five publication years to ChatGPT 4o-mini. It finds that scores rise with publication year in every field, differ widely by field, and increase with abstract length. It argues that the year and field effects are systematic biases because its sample held journals constant across years, so average underlying article quality should have been roughly stable. The practical consequence is that anyone using ChatGPT scores in evaluation should normalise them by field and year first.","feed_headline":"ChatGPT scores newer papers higher in every field tested","feed_subtitle":"A 117,650-article audit finds year and field effects, so ChatGPT quality scores need normalising before evaluation use.","key_machinery":"The machinery is a journal-stable, field-balanced sample combined with regression. Up to 1,000 articles were drawn per year and per field from journals that belonged exclusively to one of 26 subject categories and had published in all five sampled years, filtering out short or missing abstracts; this produced 117,650 articles. Each title and abstract was scored once by ChatGPT 4o-mini with an expert-review quality rubric, and ordinary least squares regression was run per field with publication year, log abstract length, and ten first-author country indicators as predictors. The sampling design is what makes a year trend interpretable as a candidate bias rather than as database drift.","core_discovery":"The central discovery is a universal recency effect: for all 26 broad fields, the average ChatGPT score for 2023 was higher than for 2003, and in 101 of 104 year-to-year comparisons each later year beat the previous one. In all 26 field-level regressions the year coefficient was positive and statistically significant even after controlling for abstract length and first-author country. Field averages differ substantially, with some fields scoring above all others in every year, and ChatGPT scores correlate positively with citation counts in all fields. The paper reads the year trend as bias because the sample was drawn from the same journals across years, so it assumes article quality was roughly constant over time.","pith_inferences":["One consequence the paper leaves implicit is that the same year trend may affect other language-model evaluators, so the normalisation advice likely extends beyond this one model and prompt.","The size of the year effect could be underestimated or overestimated if article quality in these journals genuinely improved over two decades; a direct expert-quality benchmark on the same sample would settle whether 'bias' is the right word.","The study's exclusion of multidisciplinary journals and of articles with short abstracts means the biases for high-profile venues and for letter-type outputs remain untested; those are exactly the outputs where peer review support is most contested.","Because ChatGPT was not explicitly told publication years, the recency effect must be carried by textual cues correlated with time; identifying those cues could allow targeted debiasing rather than blanket normalisation."],"forward_implications":["If the paper is right, any use of ChatGPT scores in research evaluation should first divide each article's score by the average score for its field and year, mirroring standard citation-normalisation practice.","The year bias is small but universal: it explains on average 3.6% of score variance, so it will distort comparisons of older and newer work unless corrected.","The abstract-length association is probably not a direct ChatGPT bias; it seems to reflect weaker short-form articles and national journals, but fields with strict abstract limits should be checked for anomalies.","Because ChatGPT scores correlate positively with citation counts in all fields, the paper strengthens the broader claim that LLM quality scores carry signal about research quality, even though the correlations are modest."],"supporting_citations":[{"why":"Supplies the evaluation prompt setup and the earlier evidence that ChatGPT scores correlate with expert quality scores, which this study extends.","marker":"Thelwall & Yaghi, 2024"},{"why":"Provides prior evidence that ChatGPT can evaluate research quality and that multiple submissions improve accuracy, informing the single-score design.","marker":"Thelwall, 2024ab"},{"why":"Earlier observational study of ChatGPT in peer review, used as the baseline for the claim that ChatGPT can estimate article quality.","marker":"Saad et al., 2024"},{"why":"Supplies the field-normalisation methodology that the paper recommends applying to ChatGPT scores.","marker":"Waltman & Van Eck, 2012"},{"why":"Shows how database coverage changes over time, motivating the journal-based sampling design rather than direct database sampling.","marker":"Moed et al., 2018"},{"why":"Provides the caveat that journals can change scope or editorial policy, supporting the paper's admission that journal quality stability is unproven.","marker":"Martin, 2016"},{"why":"Documents how editorial decisions can alter journal behaviour, reinforcing the same limitation.","marker":"Wilhite et al., 2019"},{"why":"Explains how citation practices have changed over time, giving context for the citation correlation results.","marker":"Larivière et al., 2008"},{"why":"Justifies excluding 2023 articles from the citation analysis because citations need time to accrue.","marker":"Wang, 2013"}],"fun_headline_variants":["ChatGPT scores newer papers higher in all 26 fields","Year and field bias found in ChatGPT quality scores","ChatGPT quality scores show universal recency bias","ChatGPT favors recent papers: 26-field audit","Newer papers get higher ChatGPT scores across fields"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the average true quality of articles in the same core journals stayed roughly constant from 2003 to 2023; the paper admits this is unproven, and if quality genuinely improved then the higher scores for newer articles would be accurate rather than biased.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT scores newer papers higher in all 26 fields","Year and field bias found in ChatGPT quality scores","ChatGPT quality scores show universal recency bias","ChatGPT favors recent papers: 26-field audit","Newer papers get higher ChatGPT scores across fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000148,"raw_usage":{"total_tokens":1165,"prompt_tokens":897,"completion_tokens":268,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":194}},"tokens_in":513,"tokens_out":268,"duration_ms":3090,"temperature":1.0,"reasoning_tokens":194,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:19:27.578509+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an expert panel blind-score a random subset of the same 117,650 articles; if expert ratings are flat across years while ChatGPT scores rise, the recency effect is a bias. Or submit the same abstracts with contemporary vocabulary swapped for neutral phrasing; if scores rise with the modern wording, the year effect is carried by textual cues rather than by true quality.","supporting_citations":[],"review_version":1}