Pith. sign in

REVIEW 4 major objections 6 minor 25 references

From Rapid Release to Reinforced Elite: Citation Inequality Is Stronger in Preprints than Journals

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Preprint citations are systematically more unequal than journal citations.

desk verdict Plausible and important question, but the preprint-journal Gini gap may be a sample-size artifact; needs a matched-sample-size robustness check before the elite-reinforcement story can be believed. read the letter →

arxiv 2506.07547 v2 pith:6JY7DZLZ submitted 2025-06-09 cs.DL

classification cs.DL
keywords preprintscitationinequalityGinicoefficientpreprintserverspreferentialattachmentauthorprestigescientometricsscholarlycommunication
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that citations to arXiv and bioRxiv preprints are distributed substantially more unequally than citations in comparable journals, and that the pattern holds in every one of the seventeen preprint categories examined. The authors match each preprint category to similar journals on field, age, and mean citation count, and still find that preprint Gini coefficients sit above the journal range, even after trimming extreme values and regardless of whether the preprint is later published in a journal. They trace the extra inequality to author prestige rather than to preferential attachment: scientists who are already highly cited in journals receive a disproportionate citation boost in preprints while producing fewer preprints. If this is right, preprint servers are not neutral rapid-release channels but systems that concentrate attention on established researchers, which carries stakes for who gets read, funded, and discovered.

What carries the argument

The Gini coefficient on five-year citation counts, justified by a lognormal citation model in which $G = \Phi(\sigma/\sqrt{2})$ depends only on the shape parameter, making the metric scale- and sample-size-independent; normalized relative inequality $z$ against matched journals; the preferential-attachment exponent $\alpha$ estimated from cumulative citation probability and compared with a standard preferential-attachment simulation; and an author-prestige measure splitting authors by the top 10% of journal Field-Weighted Citation Impact. The Gini and $z$ pair carries the headline result, while the exponent and prestige measures carry the attribution.

What would settle it

Resample each preprint category down to the size of its matched journal set (and bootstrap the journals up to preprint size), recompute the Gini gap $z$ for every category; if the positive $z$-scores collapse toward zero, the preprint-journal inequality is an artifact of sample size rather than a property of preprint culture. A direct check of the lognormal assumption by comparing empirical and theoretical Lorenz curves would invalidate the sample-size-independence proof if the curves diverge.

Watch

Extended reading notes

Core claim

The paper's central claim is that the five-year citation distributions of preprints are more concentrated than those of journals, measured by the Gini coefficient $G$, and that the excess is consistent across all categories. The claim is established by computing $z$-scores for each preprint category against matched journals; all $z>0$, with the largest excess in condensed matter, astrophysics, general physics, quantitative finance, and older preprint cultures, and a rank correlation of $0.59$ between category age and the gap. The paper further claims that the effect is not driven by the tails or by curation: trimming the top and bottom 1% leaves the pattern intact, and preprints that later appear in journals show the same elevated inequality. On mechanisms, the measured preferential-attachment exponent is below the value needed to produce the observed $G=0.88$, whereas top-decile journal authors have a higher relative preprint impact despite publishing fewer preprints; the authors therefore locate the cause in author journal prestige.

Load-bearing premise

The conclusion assumes that the Gini coefficient can be compared fairly across venues of very different sizes—that citation counts really follow the lognormal shape the paper assumes, and that the small-sample bias of the Gini estimator does not distort the comparison even though preprint categories are much larger than the matched journal sets.

Editorial extensions

If this is right

  • Research evaluation metrics that count preprint citations will inherit the extra prestige bias documented here, so preprint-based indicators should be calibrated against this baseline.
  • Simply curating preprints through journal publication will not reduce the inequality, since the gap is the same for preprints that later appear in journals.
  • Fields with older, more embedded preprint cultures show larger gaps, so the inequality may grow as preprint adoption spreads to other disciplines.
  • Because preferential attachment cannot explain the gap, interventions aimed at visibility or recommendation engines may be less effective than interventions that dampen author-prestige cues.
  • The finding implies a trade-off between the speed of preprint dissemination and the diversity of voices that gain traction in the literature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the mechanism should be observable in real time: citations to preprints from established names should spike immediately after release, before any quality signal exists, while equally new but less-known authors should lag; a release-date-resolved citation analysis would test that.
  • Because the sample-size-independence proof is load-bearing, a direct resampling experiment with equal-size subsamples of each preprint category and its matched journals would settle whether any of the measured gap is a finite-sample artifact.
  • A practical consequence the authors do not spell out: evaluation committees could adjust preprint citation counts by author-prestige baselines so that early-career research is not systematically drowned out.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper compares citation inequality, measured by the Gini coefficient over five-year citation counts, between preprint categories on arXiv and bioRxiv and matched journal sets. It reports that preprint categories consistently show higher inequality than their matched journals, even after trimming extremes and controlling for venue age and mean citation count. The paper further argues that this gap is not explained by preferential attachment but is associated with author prestige, and that the gap is larger in fields where preprints are more established.

Significance. If the descriptive claim survives scrutiny, the paper makes a valuable contribution: it documents a systematic inequality gap between preprint and journal citation distributions using a large bibliographic corpus, and it puts forward a concrete mechanism (author prestige rather than preferential attachment) that has policy implications for preprint evaluation and diversity in science. The manuscript also has useful strengths: it covers a wide range of fields, uses multiple robustness checks (trimming, curation status, field-level analyses), and explicitly discusses limitations. However, the central statistical comparison is currently vulnerable to a sample-size confound that the paper's theoretical justification does not resolve, and there is an algebraic error in the Gini derivation. The claimed universality of the preprint-journal inequality gap is therefore not yet established.

major comments (4)
  1. [Data collection, normalization, and matching; Table 1; Citation inequality metrics and its sample size independence] The central comparison is between Gini coefficients estimated from samples of very different sizes. The Gini estimator in Eq. (1) is a finite-sample estimator, and for skewed, zero-inflated citation distributions it is biased downward at small n. The lognormal derivation in 'Citation inequality metrics and its sample size independence' concerns the population parameter of a continuous distribution, not the finite-sample estimator actually used. Table 1 shows that the pooled matched-journal N is often two orders of magnitude smaller than the preprint category N (e.g., math: 168,088 preprints vs. 627 journal papers; astro-ph: 21,622 vs. 462), and each individual matched journal is smaller still. Because publication volume was deliberately excluded from the matching criteria, the observed universal z>0 pattern could be an artifact of comparing an essentially precise preprint Gini with noisy, downward-biased journal Ginis. The Discussion even concedes that 'outcomes may partially reflect underlying differences in the publication rate.' The authors need a matched-sample-size robustness analysis: for example, randomly downsample each preprint category to the sample size of its matched journal set, recompute Gini and z, or use a bias-corrected Gini estimator. Without such a check, the headline claim that preprints are consistently more unequal than journals is not supported.
  2. [Citation inequality metrics and its sample size independence, Eqs. (11)-(12)] The derivation of the lognormal Gini contains an algebraic error. From G = 1 - 2∫ L(p) dp and L(p) = Φ(Φ^{-1}(p) - σ), the correct result is G = 2Φ(σ/√2) - 1 (equivalently erf(σ/2)), not G = Φ(σ/√2) as stated in Eq. (12); Eq. (11) also incorrectly states G = Φ(-σ/√2), which would give G = 0.5 when σ = 0. The correct formula still depends only on σ, so the population-level scale independence claim survives, but the equations should be corrected, and the 'sample-size independence' wording should be restricted to the population parameter, not the empirical estimator.
  3. [Preferential attachment, Eqs. (3)-(4)] The estimation of the preferential attachment exponent is not clearly defined and the quantitative support is missing. If Π(c) ∝ c^α and π(c) = ∫_0^c Π(c) dc, then a log-log plot of π(c) versus c has slope α+1, not α; the text says the slope 'corresponds to the exponent α,' which would systematically misstate the fitted exponent. In addition, the key quantitative claim that α needs to be about 1.3 to produce G = 0.88 while the actual slope is α<1.0 contains an empty cross-reference '(see )' and no derivation or figure citation. The authors should clarify the fitting procedure and provide the missing reference to Figure 4 or the simulation details.
  4. [Author journal prestige] The author-prestige analysis divides authors into top and average groups based on their journal-based Field-Weighted Citation Impact, then compares the relative impact of these groups on preprints versus journals. This design shows an association between being a high-impact journal author and having high preprint impact, but it cannot separate 'prestige' from persistent author quality, field-specific citation norms, or selection effects in who posts preprints. The Discussion's language that 'researchers who are already influential in journals may exert even more substantial influence within preprint ecosystems' goes beyond what this observational comparison can establish. The authors should soften the causal claim and explicitly acknowledge that the FWCI-based prestige measure is confounded with author quality.
minor comments (6)
  1. [Abstract vs. Discussion] The abstract lists 'high-energy physics' as a field where the gap is pronounced, while the Discussion lists 'condensed matter physics'; these should be reconciled.
  2. [Eq. (1)] The estimator in Eq. (1) should explicitly define N as the number of papers in the venue and state whether it is the standard unbiased or the sample Gini estimator.
  3. [Eq. (5)] 'Lorentz curve' should be 'Lorenz curve.'
  4. [References] Reference [25] has a malformed author name ('family=Eck, p. u., given=Nees Jan'); it should be formatted properly as van Eck, N. J., and Waltman, L.
  5. [Data collection, normalization, and matching] There are typos and awkward phrasings, e.g., 'Gieger (2019)' should be 'Geiger (2019)', 'acconting' should be 'accounting', and 'diffrent' should be 'different.'
  6. [Preprint categories] The sentence about arXiv being 'taken top three major categories by publication' is unclear and should be rewritten.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the inference chain is externally grounded, though the sample-size invariance defense is statistically fragile.

full rationale

The paper's central comparison is computed directly from external citation data: Gini coefficients for preprint categories are contrasted with Gini coefficients of matched journals, and the z-score is a descriptive normalization rather than a fitted prediction. The sample-size-independence argument is based on a lognormal model and is mathematically vulnerable (finite-sample estimator bias and an algebraic slip in the lognormal Gini formula), but this is a statistical validity concern, not circularity: Eq. (12) is not derived from the empirical z > 0 result, and no parameter is fit to the claim it explains. The author-prestige analysis defines top authors by journal-based FWCI and then compares relative ratios for journal and preprint impact; it does not define prestige in terms of preprint impact, so the conclusion is not forced by construction. The only self-citation (ref [19]) appears in the Discussion as supporting evidence for delayed recognition and is not load-bearing for the headline result. No step in the derivation reduces to its own inputs.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The analysis relies on several fitted or hand-chosen parameters, mainly in the matching procedure, the preferential attachment estimation, and the author prestige split. No new physical or conceptual entities are introduced; the paper uses established bibliometric constructs.

free parameters (6)
  • Preferential attachment exponent alpha = alpha < 1.0 (preprints, exact value not reported)
    Estimated by fitting a line to the log-log cumulative citation probability in Figure 3.a; used to argue that preferential attachment does not explain the observed Gini of 0.88.
  • Journal age threshold N0 = 10
    Chosen to define the first year of a journal's existence; affects the age matching variable tau.
  • Winsorization and trimming threshold = 1% at both ends
    Applied to citation counts before matching and to Gini computation; the choice affects the reported z-scores and Gini values.
  • Citation window = 5 years (c5)
    Gini and preferential attachment analyses use five-year citation counts; different windows could change the results.
  • Top-author percentile = top 10% by mean Field-Weighted Citation Impact
    Used to split authors into top and average groups in the author prestige analysis.
  • Barabasi-Albert simulation parameters = m=4, 5000 iterations, 30 runs
    Used to estimate the alpha threshold needed to reach a given Gini coefficient; the threshold depends on these choices.
assumptions (5)
  • domain assumption Citation counts follow a lognormal distribution in both preprints and journals
    Invoked in the sample-size independence proof; based on references [23,24]. If the distribution deviates from lognormal, the theoretical Gini formula and the matching logic may fail.
  • domain assumption The empirical Gini estimator is approximately unbiased and comparable across sample sizes
    The paper uses the population formula for G to claim sample-size independence, but small-sample bias in the estimator is not accounted for. This is an unstated assumption that underlies the decision to exclude publication volume from matching.
  • domain assumption Preferential attachment follows the functional form Pi(c) proportional to c^alpha
    Used in Eqs. (3) and (4) to interpret the slope of the cumulative distribution as alpha. The paper's slope-to-alpha mapping appears inconsistent with Eq. (4).
  • domain assumption OpenAlex metadata correctly matches preprints to citations and authors
    All citation and author data comes from OpenAlex; matching rates were 66.8% for arXiv and 25.7% for bioRxiv, so unmatched preprints could bias the results.
  • ad hoc to paper The top three OpenAlex subfields represent a journal's scope
    Used in the matching procedure to identify journals similar to each preprint category; the choice of three subfields is arbitrary.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Rapid Release to Reinforced Elite: Citation Inequality Is Stronger in Preprints than Journals." pith.science (2026). https://pith.science/paper/6JY7DZLZ

@misc{pith2026250607547,
  author       = {Pith},
  title        = {Pith review of: From Rapid Release to Reinforced Elite: Citation Inequality Is Stronger in Preprints than Journals},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6JY7DZLZ}},
  note         = {Machine review of arXiv:2506.07547}
}
read the original abstract

Preprints have been considered primarily as a supplement to journal-based systems for the rapid dissemination of relevant scientific knowledge and have historically been supported by studies indicating that preprints and published reports have comparable authorship, references, and quality. However, as preprints increasingly serve as an independent medium for scholarly communication rather than precursors to the version of record, it remains uncertain how preprint usage is shaping scientific discourse. Our research revealed that the preprint citations exhibit significantly higher inequality than journal citations, consistently among categories. This trend persisted even when controlling for age and the mean citation count of the journal matched to each of the preprint categories. We also found that the citation inequality in preprints is not solely driven by a few highly cited papers or those with no impact, but rather reflects a broader systemic effect. Whether the preprint is subsequently published in a journal or not does not significantly affect the citation inequality. Further analyses of the structural factors show that preferential attachment does not significantly contribute to citation inequality in preprints, whereas author prestige plays a substantial role. Notably, the gap in citation inequality between the preprint category and the journal is more pronounced in fields where preprints are more established, such as mathematics, physics, and high-energy physics. This highlights a potential vulnerability in preprint ecosystems where reputation-driven citation may hinder scientific diversity.

Figures

Figures reproduced from arXiv: 2506.07547 by the authors.

Figure 1
Figure 1. Citation inequality G for each subcategory in arXiv and bioRxiv. The cool color indicates moderate inequality (blue≈ 0.6, green≈ 0.7), while warmer colors mean more substantial inequality (orange≈ 0.9, red means monopoly). Subcategories are positioned more closely when more journal articles cite the same preprint in each category. Here, Gpreprint is a Gini coefficient of a preprint category, Gjournals is a list of t… view at source ↗
Figure 2
Figure 2. Each point represents a z-score compared to its control group, matched on journal age [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. a) Preferential attachment for journals and preprints. The slope of the fitted line [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The simulation of preferential attachment. The slope of the fitted line corresponds to [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

25 extracted references · 18 canonical work pages

  1. [1]

    Communication Patterns in High-Energy Physics URL https: //cds.cern.ch/record/546422/files/sis-2002-163.html

    Goldschmidt-Clermont, L. Communication Patterns in High-Energy Physics URL https: //cds.cern.ch/record/546422/files/sis-2002-163.html

  2. [2]

    How swamped preprint servers are blocking bad coronavirus research 581, 130–131

    Kwon, D. How swamped preprint servers are blocking bad coronavirus research 581, 130–131. URL https://www.nature.com/articles/d41586-020-01394-6

  3. [3]

    Weissgerber, T. et al. Automated screening of COVID-19 preprints: Can we help authors to improve transparency and reproducibility? 27, 6–7. URL https://www.nature.com/ articles/s41591-020-01203-7

  4. [4]

    Eisen, M. B. et al. Implementing a ”publish, then review” model of publishing 9, e64910. URL https://elifesciences.org/articles/64910. 7

  5. [5]

    & Squazzoni, F

    Akbaritabar, A., Stephen, D. & Squazzoni, F. A study of referencing changes in preprint- publication pairs across multiple fields 16, 101258. URL https://www.sciencedirect.com/ science/article/pii/S1751157722000104

  6. [6]

    Nielsen, M. W. & Andersen, J. P. Global citation inequality is on the rise 118, e2012208118. URL https://www.pnas.org/doi/abs/10.1073/pnas.2012208118

  7. [7]

    Chu, J. S. G. & Evans, J. A. Slowed canonical progress in large fields of science 118, e2021636118. URL https://www.pnas.org/doi/abs/10.1073/pnas.2021636118

  8. [8]

    The expert game -- Cooperation in social communication

    Bendtsen, K. M., Uekermann, F. & Haerter, J. O. The expert game – Cooperation in social communication. URL http://arxiv.org/abs/1312.6715. 1312.6715

Show all 25 references
  1. [9]

    & Lycett, S

    Mesoudi, A. & Lycett, S. J. Random copying, frequency-dependent copying and cul- ture change 30, 41–48. URL https://www.sciencedirect.com/science/article/pii/ S1090513808000810

  2. [10]

    MacRoberts, M. H. & MacRoberts, B. R. Problems of citation analysis: A critical re- view 40, 342–349. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/%28SICI% 291097-4571%28198909%2940%3A5%3C342%3A%3AAID-ASI7%3E3.0.CO%3B2-U

  3. [11]

    URL https://apastyle.apa.org/about-apa-style

    About APA Style. URL https://apastyle.apa.org/about-apa-style

  4. [12]

    Price, D. J. D. S. Networks of Scientific Papers: The pattern of bibliographic refer- ences indicates the nature of the scientific research front. 149, 510–515. URL https: //www.science.org/doi/10.1126/science.149.3683.510

  5. [13]

    Abelson, P. H. Information Exchange Groups 154, 727–727. URL https://www.science. org/doi/10.1126/science.154.3750.727

  6. [14]

    The E-volution of preprints in the scholarly communication of physicists and astronomers 52, 187–200

    Brown, C. The E-volution of preprints in the scholarly communication of physicists and astronomers 52, 187–200. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/ 1097-4571%282000%299999%3A9999%3C%3A%3AAID-ASI1586%3E3.0.CO%3B2-D

  7. [15]

    arXiv E-prints and the journal of record: An analysis of roles and rela- tionships 65, 1157–1169

    Larivi` ere, V.et al. arXiv E-prints and the journal of record: An analysis of roles and rela- tionships 65, 1157–1169. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/asi. 23044

  8. [16]

    & Wang, K

    Xie, B., Shen, Z. & Wang, K. Is preprint the future of science? A thirty year journey of online preprint services. URL http://arxiv.org/abs/2102.09066. 2102.09066

  9. [17]

    & Barab´ asi, A

    Jeong, H., N´ eda, Z. & Barab´ asi, A. L. Measuring preferential attachment in evolving networks 61, 567. URL https://iopscience.iop.org/article/10.1209/epl/i2003-00166-9/meta

  10. [18]

    & Larivi` ere, V

    Golosovsky, M. & Larivi` ere, V. Uncited papers are not useless 2, 899–911. URL https: //doi.org/10.1162/qss_a_00142

  11. [19]

    & Sakata, I

    Miura, T., Asatani, K. & Sakata, I. Large-scale analysis of delayed recognition using sleeping beauty and the prince 6, 48. URL https://doi.org/10.1007/s41109-021-00389-0

  12. [20]

    ArXiV Archive

    Geiger, R. ArXiV Archive. URL https://zenodo.org/records/4990937

  13. [21]

    URL https://sharing.nih.gov/ public-access-policy/public-access-policy-overview?utm_source=chatgpt.com

    NIH Public Access Policy Overview — Data Sharing. URL https://sharing.nih.gov/ public-access-policy/public-access-policy-overview?utm_source=chatgpt.com

  14. [22]

    M., Pan, R

    Petersen, A. M., Pan, R. K., Pammolli, F. & Fortunato, S. Methods to account for cita- tion inflation in research evaluation 48, 1855–1865. URL https://www.sciencedirect.com/ science/article/pii/S0048733319301003

  15. [23]

    & Barab´ asi, A.-L

    Wang, D., Song, C. & Barab´ asi, A.-L. Quantifying Long-Term Scientific Impact342, 127–132. URL https://www.science.org/doi/10.1126/science.1237825. 8

  16. [24]

    & Peters, I

    Fraser, N., Momeni, F., Mayr, P. & Peters, I. The relationship between bioRxiv preprints, citations and altmetrics 1, 618–638. URL https://doi.org/10.1162/qss_a_00043

  17. [25]

    S4306400194

    Waltman, L. & family=Eck, p. u., given=Nees Jan. A systematic empirical comparison of different approaches for normalizing citation impact indicators 7, 833–849. URL https: //www.sciencedirect.com/science/article/pii/S1751157713000667. 9 Materials and Methods Data collection, ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.