{"id":"ac03554f-2a3d-4c98-ab46-b0c76768bdfa","arxiv_id":"2411.09675","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Citation sentiment in neuroscience tracks collaboration, h-index differences, wet/dry-lab culture, and country-level individualism and power distance.","lead":"Researchers used an AI language model to classify the tone of over 600,000 citation sentences in neuroscience papers. They found that citation sentiment reflects collaboration ties, scientific status, disciplinary culture, and country-level individualism and power hierarchies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fine-tuned sentiment labels have no reported validation; differential LLM error across countries or disciplines could manufacture the culture correlations, so the central claim needs a stratified held-out check.","rationale":"I agree with the reader that classifier validity is the weakest assumption. The paper has real strengths: a large corpus, a null model for content similarity and article type, bootstrapping, and an identity-blind prompting scheme that prevents the LLM from directly reading author country, gender, or h-index. Those features make pure attenuation from label noise an unlikely full explanation for the strong observed slopes. But the country and discipline analyses aggregate sentiment by strata that differ in exactly the linguistic registers an LLM is sensitive to, and the paper's own definition of 'critical' deliberately includes difference, distinction, and distance. A classifier keyed to contrastive syntax would therefore encode rhetorical style as sentiment, and the null model does not adjust for style. Because the manuscript reports no held-out evaluation for the fine-tuned GPT-3.5 model, there is no evidence that label quality is uniform across the strata being compared. The proposed stratified validation would settle whether differential measurement error exists and, if it does, whether the core slopes survive correction. Until that check is run, the conditional verdict is the right one; nothing in my stress test moves it.","tokens_in":24198,"tokens_out":6330,"duration_ms":65553,"concrete_test":"Construct a held-out set of about 600 citation sentences from the same corpus, stratified by country cluster (e.g., high vs low individualism), discipline benchwork/synthesis tertiles, and collaborator status. Have the five original annotators label them, then run the fine-tuned model with the published prompt. Report macro-F1 and per-class precision/recall, and fit a logistic regression of per-sentence label correctness on strata with interactions. If critical-sentiment precision is below about 0.70 or any interaction is significant at p<0.05, re-estimate the Section II.B–II.D regressions with stratum-calibrated labels (double sampling). If the individualism/power-distance slopes shrink by more than 30% or lose significance, the multiscale claim is unsupported; if no differential error appears, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—citation sentiment tracks country, discipline, and collaboration norms—rests entirely on the GPT-3.5-Turbo labels, yet Methods B ('Measuring Sentiment') reports only that fine-tuning on 300 sentences yielded 'an improvement of 28% over the baseline,' with no held-out accuracy, per-class precision/recall, or inter-annotator agreement. This matters not just because noise attenuates effects, but because LLM measurement error can be differential. The model sees only the sentence with the citation marker, so it cannot know author country, gender, or h-index; however, academic writing styles differ systematically by country and discipline (contrastive connectives, hedging, explicitness of critique). If the LLM labels such stylistic constructions as 'critical' at different rates across these strata, the country slopes in Section II.D and the discipline slopes in Section II.C partly reflect the classifier's stylistic bias rather than citer sentiment. The null model controls for title similarity, article type, and citation frequency, but not for stable register differences. With only 300 fine-tuning examples (from 5 volunteers, thresholded at ±0.4), differential error across the exact strata later compared is not ruled out, and no supplementary validation metrics are given in the manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript assembles a corpus of 627,108 citation sentences from neuroscience articles in the PubMed Central Open Access subset, labels each citation as favorable, neutral, or critical using a fine-tuned GPT-3.5-Turbo model, aggregates labels by citer-citee pair, and compares observed sentiment proportions against a null model that conditions on title similarity, citation frequency, and article type. It reports that researchers cite collaborators more favorably and less critically than non-collaborators, that high h-index authors critically cite lower h-index non-collaborators, that wetlab and high-synthesis disciplines are less critical, that countries high in individualism and low in power distance produce more critical citations, and that men use more sentiment overall while women show a stronger collaborator bias. The paper interprets these patterns as evidence that citation sentiment tracks sociocultural norms across individual, disciplinary, and national scales.","tokens_in":24392,"tokens_out":5050,"duration_ms":55552,"significance":"If the sentiment labels are valid, this is a substantial large-scale descriptive contribution to the science of science. The strengths are the size of the corpus, the explicit construction of a null model that removes some baseline effects of content similarity, citation frequency, and article type, and the clear separation of collaborator versus non-collaborator contrasts across scales. The paper also makes falsifiable predictions, such as the drop in critical sentiment after collaboration begins, that could be tested in other fields. However, the central claim rests entirely on the accuracy and measurement invariance of an LLM-based sentiment classifier whose validation is not reported. Because the aggregate effects are small relative to the base rates, differential measurement error across disciplines, countries, or author groups could produce or distort the main results. The manuscript would be acceptable only after the classifier validation and confound controls are added and shown to preserve the reported patterns.","major_comments":[{"comment":"The paper reports no validation of the sentiment classifier. It states only that fine-tuning on 300 manually annotated sentences produced \"an improvement of 28% over the baseline (see Supplement)\", but it does not define the baseline, report held-out accuracy, per-class precision/recall, inter-annotator agreement, or a confusion matrix. Since only 7.94% of the aggregated citations are critical, even small differential error in detecting critical language—for example, across national academic writing styles or disciplinary registers—can move the slopes in Figures 4 and 6 as much as the reported effects. The supplement is referenced repeatedly but is not part of the submitted text, so these metrics cannot be checked. Please add stratified held-out evaluation by discipline, country, gender, and collaborator status, and show how the main results change under alternative labeling thresholds and aggregation rules.","section":"V.B (Measuring Sentiment)"},{"comment":"The null model conditions only on title-similarity bins, citation frequency, and article type. It does not control for discipline, country, journal, h-index composition, or author seniority. The disciplinary and country analyses use aggregate means with n=27 and n=23, respectively; if high-individualism countries are overrepresented in drylab disciplines, in high-h-index author pools, or in particular journals, the country correlations in Figure 6 could reflect compositional confounding rather than a cultural effect. The manuscript should report the composition of the country and discipline aggregates and add control analyses using fixed effects, matched samples, or residualization.","section":"V.E (Departmental and Cultural Measures) with II.C-D"},{"comment":"Most significance tests are reported as one-sided F-tests, and many related correlations are tested across favorable/critical and collaborator/non-collaborator splits without any multiple-comparison correction. This inflates the number of nominally significant results. Please report two-sided p-values or explicit permutation tests, pre-specify directional hypotheses where possible, and apply a multiple-comparison adjustment or clearly justify why it is unnecessary. Several slopes in Figures 4D-F also have wide confidence intervals, so the strength of the claims should be calibrated to the precision actually achieved.","section":"II.C-D and Figures 4-8"}],"minor_comments":[{"comment":"\"Citation sentiment tracts disciplinary differences\" should be \"tracks disciplinary differences\".","section":"III.B (Discussion)"},{"comment":"The text and figure use \"woman\" where \"women\" is intended; please correct the grammar.","section":"II.E and Figure 8"},{"comment":"The phrase \"improvement of 28% over the baseline\" needs a precise definition of the baseline model, the evaluation split, and the metric being improved; the current sentence is not checkable.","section":"V.B"},{"comment":"The units and definitions of the quantities in panel A are unclear; please specify whether they are percentage differences from the null model and what the error bars represent.","section":"Figure 8"},{"comment":"No data or code availability statement is provided. Given the reliance on proprietary LLM APIs, the authors should at minimum release the prompts, the 300 annotated sentences, and the analysis code to allow independent replication.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the empirical questions are interesting, but the missing validation materials are a reproducibility concern. I would ask the editor to require the supplement, the fine-tuning data, the full prompt protocol, and the validation metrics before the paper is reconsidered. There is no indication of misconduct; the issue is reporting completeness and the load-bearing role of the unvalidated classifier."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper you'd want to know about: 627k citation sentences in neuroscience, classified by a fine-tuned GPT-3.5, with a null model, showing sentiment tracks collaboration, h-index, discipline, and country. The scale is the news; most prior work is an order of magnitude smaller. What's genuinely new is the multiscale framing and the attempt to separate content effects with a null model.\n\nWhat it does well: clear methods, explicit aggregation precedence, and the collaboration-distance analysis is a nice design because it uses within-dyad comparisons. The null model controlling title similarity, citation frequency, and article type is a real step beyond raw sentiment counts.\n\nSoft spots: the classifier validation is the load-bearing issue. '28% improvement over baseline' with no held-out accuracy, per-class precision/recall, or inter-annotator agreement is too thin, especially because only 300 sentences from neuroscience and physics were used for fine-tuning. If GPT-3.5's error patterns differ by discipline or country—plausible, given known stylistic differences in academic writing—the country and discipline slopes could partly reflect classifier bias. The null model doesn't control for those strata, so the stress-test concern is legitimate. Second, the country-level analyses use Hofstede scores of the last author's country as if authors represent national culture; that is a big ecological leap. Third, many one-sided tests without correction; at this scale, some of the weaker effects will be false positives. The selection (open-access, IF≥3, last authors) is a limitation but not a serious one for a first pass.\n\nThe collaborator and h-index findings are less exposed to the differential-error critique because they are within-discipline/within-country comparisons, so I'd trust those more. The country and discipline results should be treated as provisional until the classifier is validated stratified.\n\nWho it's for: science-of-science researchers and anyone using LLMs for citation analysis. It deserves peer review, but the referee should require a stratified validation and a multiple-comparison correction or a pre-registered analysis.","headline":"Plausible large-scale citation-sentiment study; missing classifier validation keeps the cultural claims provisional.","tokens_in":24945,"tokens_out":2416,"would_cite":true,"duration_ms":24156,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Neuroscience citations carry a measurable social fingerprint: collaborators get kinder, outsiders harsher.","keywords":["citation sentiment","sentiment analysis","large language models","sociocultural norms","collaboration networks","h-index","individualism","power distance"],"falsifier":"Take a stratified random sample of roughly one thousand citation sentences from the same corpus, have trained annotators label them as favorable, neutral, or critical, and compare with the model's labels. If agreement is poor or if the reported correlations with collaboration distance, h-index, discipline, and country disappear when using the human labels, the central claim fails.","tokens_in":23993,"feed_emoji":"🧠","tokens_out":5692,"duration_ms":54504,"temperature":0.7,"pith_summary":"Scientists rarely cite neutrally, and this paper argues that the warmth or sharpness of a citation is socially structured rather than random. Analyzing hundreds of thousands of citation sentences, the authors show that researchers cite their collaborators more favorably and less critically than they cite strangers, and that higher-status scholars are especially critical of lower-status outsiders. The same pattern reappears at larger scales: wet-lab and review-heavy subfields are gentler, while countries high in individualism and low in power distance produce more critical citations. The paper therefore tries to establish that citation sentiment is a multiscale record of ingroup/outgroup bias, status hierarchy, and cultural norms in science, not just a byproduct of the cited content.","feed_headline":"Citations get kinder for collaborators, harsher for outsiders","feed_subtitle":"Across 627,000 neuroscience citations, sentiment tracks status, discipline, and national culture.","key_machinery":"The central machinery is a sentiment-labeled citation network: each citation edge between two last authors is tagged favorable, neutral, or critical by a large language model fine-tuned on a small human-annotated set, with negative sentiment taking precedence when a paper cites another multiple times. To separate content effects from social effects, the analysis uses a null model that predicts expected sentiment from title similarity, citation frequency, and article type, so all reported results are residual deviations from chance. This machinery lets the authors compare sentiment across collaboration distances, h-index differences, disciplines scored by benchwork and synthesis, countries scored by individualism and power-distance, and author gender.","core_discovery":"The paper's central discovery is that the sentiment expressed in a citation—favorable, neutral, or critical—varies systematically with the social and cultural positions of the citing and cited authors. On a large sample of neuroscience papers, the authors find that collaboration distance predicts sentiment: collaborators receive more favorable and less critical citations, and pairs who will begin collaborating in the future show elevated critical sentiment just beforehand. Outside collaborations, the greater the citer's h-index relative to the citee's, the more critical the citation. At the discipline level, criticality decreases with benchwork (wet-lab) practice and with the proportion of review articles, and at the country level, critical sentiment rises with individualism and falls with power-distance. The authors conclude that citation sentiment tracks sociocultural norms across scales, with hierarchy and group membership playing a focal role.","pith_inferences":["If critical citations tend to precede first collaborations, as the temporal pattern suggests, citation sentiment could be mined as a forward-looking signal for team formation; the authors flag this possibility but do not test it prospectively.","Because the dataset is confined to neuroscience, the cleanest testable extension is to repeat the pipeline in older, more crystallized disciplines, where the multiscale effects should be attenuated if the mechanism is sociocultural maturity.","The same sentiment-labeled network could be turned into a bias-audit instrument: journals, funders, or institutions could monitor whether critical or favorable tone is systematically skewed by gender, nationality, or status, a practical use the paper does not develop."],"forward_implications":["Citation analyses that count only presence or absence will systematically miss the social signal; sentiment-aware counts should differ by collaboration and status.","Critical citations are concentrated among high-h-index authors citing lower-h-index non-collaborators, so scientific hierarchy shapes which work receives public challenge.","Disciplinary culture appears to modulate criticality: wet-lab disciplines and review-heavy fields cite more gently, suggesting that synthesis practices may cool polarization.","Country-level differences in individualism and power-distance predict citation tone, so cross-national bibliometric comparisons should adjust for cultural norms.","Men in this dataset cite with more sentiment overall while women show a stronger favoritism bias toward collaborators, indicating that gender shapes the emotional register of citations."],"supporting_citations":[{"why":"Supplies the definition of collaboration distance as the shortest path in the coauthorship network, the paper's main ingroup/outgroup variable.","marker":"[43]"},{"why":"Provides the individualism and power-distance country scores used in the country-level analyses.","marker":"[66]"},{"why":"Earlier deep-learning citation sentiment classifier whose scope this paper extends to large-scale sociocultural analysis.","marker":"[38]"},{"why":"Underpins the aggregation rule that negative sentiment overrides positive and neutral when a paper cites another multiple times.","marker":"[151]"},{"why":"Documents disciplinary differences in citation and attribution practices that motivate the discipline-level sentiment analyses.","marker":"[110]"},{"why":"Source of the open-access neuroscience article corpus from which citation sentences were extracted.","marker":"[144]"},{"why":"Provides evidence that critical sentiment rates differ between neuroscience and physics, guiding the fine-tuning sample selection.","marker":"[153]"}],"fun_headline_variants":["Collaborators get kinder citations, outsiders get harsher","Citation tone tracks culture, status, and collaboration","Science's citation sentiment mirrors societal norms","From lab to nation: citation sentiment follows culture","Citation sentiment reveals social norms across science"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the large language model's favorable/neutral/critical labels are accurate and consistent across disciplines, countries, genders, and time periods; the paper does not report held-out accuracy or inter-annotator agreement for the fine-tuned classifier.","fun_headline_variants_meta":{"raw":{"variants":["Collaborators get kinder citations, outsiders get harsher","Citation tone tracks culture, status, and collaboration","Science's citation sentiment mirrors societal norms","From lab to nation: citation sentiment follows culture","Citation sentiment reveals social norms across science"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000384,"raw_usage":{"total_tokens":2041,"prompt_tokens":965,"completion_tokens":1076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":1007}},"tokens_in":581,"tokens_out":1076,"duration_ms":8428,"temperature":1.0,"reasoning_tokens":1007,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:23:20.673473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of roughly one thousand citation sentences from the same corpus, have trained annotators label them as favorable, neutral, or critical, and compare with the model's labels. If agreement is poor or if the reported correlations with collaboration distance, h-index, discipline, and country disappear when using the human labels, the central claim fails.","supporting_citations":[{"cited_title":"Hyland, Academic attribution: citation and the con- struction of disciplinary knowledge, Applied Linguistics 20, 341 (1999)","cited_arxiv_id":null,"evidence_quote":"Documents disciplinary differences in citation and attribution practices that motivate the discipline-level sentiment analyses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the open-access neuroscience article corpus from which citation sentences were extracted."},{"cited_title":"Bordignon, Self-correction of science: A comparative study of negative citations and post-publication peer review, Scientometrics 124, 1225 (2020)","cited_arxiv_id":null,"evidence_quote":"Provides evidence that critical sentiment rates differ between neuroscience and physics, guiding the fine-tuning sample selection."}],"review_version":1}