REVIEW 4 major objections 5 minor 4 references
Effect of perceived preprint effectiveness and research intensity on posting behaviour
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that preprint posting is driven above all by how much a researcher already publishes in journals and books, and that valuing preprints for early access predicts more posting while valuing them for early feedback predicts…
desk verdict A clean reanalysis of an open survey with one genuinely new but fragile result — the negative early-feedback coefficient — and a robust productivity finding; worth refereeing with conditions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is one Ordinary Least Squares (OLS) linear regression model: the dependent variable is the number of preprints posted in the last three years (question Q6, level 5), and 53 coefficients are estimated across six blocks, namely five research-output counts (journal articles, conference proceedings, book chapters, books and monographs, peer reviews), five Likert-rated items on perceived preprint effectiveness (Q13), and then field, organisation type, region, research experience and gender as controls. The machinery does two things: the raw coefficients give ceteris paribus effects (each journal article adds $0.155$ expected preprints; each point of early-access effectiveness adds $0.465$, and each point of early-feedback effectiveness subtracts $0.371$), and the standardised coefficients rank the drivers by effect size. The analysis works because the survey questions on outputs, perceptions and demographics come from the same respondents, letting the model separate what is attributable to productivity, what to perception and what is simply demographic.
What would settle it
Go to the open survey file and recover the 5,272 respondents excluded for incomplete answers, then re-estimate the model with multiple imputation or compare the included and excluded groups' preprint counts and effectiveness ratings. If the excluded researchers post more preprints or rate early feedback more favourably than complete-case respondents, the OLS coefficients are biased; the early-feedback coefficient ($-0.371$) changing sign or the journal-article effect ($\beta = 0.280$) collapsing toward the perception effects would settle the paper's ranking of drivers differently.
Extended reading notes
Core claim
The paper's core claim is that preprint posting is a complement to traditional scholarly output, not a substitute for it. In the OLS model, the number of journal articles ($\beta = 0.280$, the largest standardized effect) and of books and monographs ($\beta = 0.163$, the third largest) are the top productivity predictors of how many preprints a researcher posted in the last three years, with the Physics and Astronomy field effect ($\beta = 0.191$) between them. Perceived effectiveness splits by function: each Likert point of perceived effectiveness for providing early access to new research raises expected preprint counts by $0.465$, while each point for receiving early feedback lowers them by $0.371$, a split the paper interprets as early dissemination being valued while the feedback function is perceived as risky or inadequate. Net of these, gender and organisation type show essentially no association with posting, whereas discipline, career stage and geography retain moderate effects, and the model as a whole explains $26.8\%$ of the variance (adjusted $R^2 = 0.268$). The authors take this as evidence against the assumption that preprints mainly help researchers who cannot get published in traditional venues.
Load-bearing premise
The model is estimated only on the 5,873 respondents who answered every question, a complete-case sample chosen without testing how the 5,272 dropped respondents differ; the load-bearing premise is that these complete cases behave like all 11,145 researchers surveyed, so if who finishes the survey depends on preprint activity or preprint perceptions, every coefficient—including the negative early-feedback effect—could be an artefact of selection rather than a genuine driver of posting.
Editorial extensions
If this is right
- Preprint promotion aimed at researchers who struggle to publish in traditional venues is likely aimed at the wrong group; the paper's corollary is that preprint incentives belong inside the reward structures that already track journal and book output.
- Because expecting useful early feedback currently predicts fewer preprints at every level of productivity, platforms that build more structured and supportive commenting mechanisms could remove a genuine barrier to posting.
- With gender and organisation type essentially null net of other factors, observed demographic gaps in preprint use are likely productivity, discipline and region effects in disguise rather than independent barriers.
- The persistent moderate effects of discipline, experience and region mean uniform open-science mandates will produce uneven uptake, so policy should be tailored to fields and regions with weaker preprint cultures.
- With adjusted $R^2 = 0.268$, most individual variation in posting is left unexplained, so institutional contexts, platform features and field-level norms that the survey does not measure still carry most of the weight.
Reading between the lines
- If preprint posting truly complements traditional output, a consequence the paper leaves implicit is a Matthew effect in scholarly communication: the same authors who already dominate journal visibility would also gain the earliest visibility through preprints, widening the gap with less prolific researchers.
- The regression is run on the 5,873 complete cases drawn from 11,145 survey respondents; re-estimating on the full open dataset with multiple imputation for missing answers is the natural next test of whether the negative early-feedback coefficient is a behavioural effect or a selection artefact.
- The five effectiveness items are strongly intercorrelated (early access and early feedback correlate $0.48$), so the opposite signs of their coefficients hint at a suppressor relationship; a partial-correlation or mediation decomposition would clarify whether the negative feedback effect is genuine or a by-product of collinearity.
- The effectiveness perception items may well interact with field: in disciplines where preprint cultures are mature, such as physics and mathematics, the early-feedback effect could differ from fields with younger commenting cultures, a pattern testable by adding interaction terms to a larger sample.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes open survey data from a global cOAlition S consultation (11,145 researcher responses, reduced to 5,873 complete cases) to explain researchers' preprint posting frequency over the last three years. Using OLS regression with the number of preprints as the outcome, the authors report that traditional productivity measures—especially journal articles and books/monographs—are the strongest predictors of preprint production, that perceived effectiveness for early access is positively associated with posting, while perceived effectiveness for early feedback is negatively associated, and that most demographic variables are not significant. The authors interpret these findings as evidence that preprints complement, rather than substitute for, traditional publishing, and that researchers' perceptions of early feedback present a nuanced barrier.
Significance. The study addresses a relevant and under-explored question in scholarly communication: how perceptions of preprint effectiveness and research intensity shape actual posting behavior. Its strengths include the use of a large, international, openly available survey dataset and the explicit focus on standardizing effect sizes to compare predictors. If the perception results were robust, the negative association between valuing early feedback and posting preprints would be a novel and policy-relevant contribution. The productivity finding—that journal articles and books are the strongest predictors—is plausible and consistent with prior work, and it is supported by simple correlations and a highly significant regression coefficient. However, the central perception claims rest on a cross-sectional design with substantial listwise deletion, and the manuscript does not provide the diagnostic checks needed to establish that these coefficients are not artifacts of selection or collinearity.
major comments (4)
- [Data] The manuscript reports that 11,145 researcher responses were reduced to 5,873 valid observations by listwise deletion, but it provides no analysis of missing-data patterns or comparison of included and excluded respondents. Because the regression estimates in Table 5 are computed only on complete cases, the coefficients—especially the perception coefficients—are unbiased only if missingness is completely at random or fully explained by the included covariates. The negative coefficient on 'receiving early feedback' (-0.371) is particularly fragile: its raw correlation with preprint production is only +0.03 (Table 3), so a modest selection-induced shift in the joint distribution of perceptions and preprints could flip its sign. The authors should report the missingness rate per variable, compare means and distributions on observed variables between complete and incomplete cases, and ideally present sensitivity analyses (e.g., multiple imputation or inverse-probability weighting).
- [Method and Table 5] The outcome variable—number of preprints in the last three years—is a bounded count (0–40, mean 2.2, SD 4.9) with a large mass at zero, yet it is modeled with OLS. This can produce inefficient estimates, incorrect standard errors, and predicted values outside the feasible range. The paper does not report residual diagnostics, heteroskedasticity-robust standard errors, or a count-model alternative (e.g., Poisson or negative binomial). Given that the main productivity coefficients are large and highly significant, they may survive a count-model re-estimation, but the smaller perception coefficients could change in significance or sign. The authors should at least re-estimate the model with robust standard errors and with a negative binomial or Poisson specification to confirm the key results.
- [Method and Table 3] The perception variables are five Likert items that are highly intercorrelated (Table 3 shows correlations up to 0.72 between 'enhancing accessibility' and 'accelerating discourse', and 0.65 between 'early feedback' and 'accelerating discourse'). The regression enters all five simultaneously, but the paper reports no multicollinearity diagnostics such as variance inflation factors. The negative coefficient on 'receiving early feedback' could partly reflect suppressor effects among collinear perception items; the manuscript's interpretation as a substantive negative perception is not warranted without evidence that the coefficient is stable across model specifications. I recommend reporting VIFs and showing a specification that enters each perception item separately or in subsets.
- [Method and Results] The perception measures are taken at the time of the survey, after respondents have accumulated years of experience with preprints, so the regression cannot distinguish between perceptions causing posting behavior and posting behavior shaping perceptions. The paper acknowledges self-selection and self-report bias in the limitations section, but it does not address reverse causality, which is especially relevant for the early-feedback coefficient. For instance, researchers who have posted preprints and received poor feedback may lower their rating of early feedback, producing a negative association without any causal effect of perception on posting. The authors should temper the causal language (e.g., 'hurts preprint production', 'decreases the likelihood') and discuss the direction-of-causality problem explicitly, or use instrumental-variable or longitudinal designs if available.
minor comments (5)
- [Table 3] The title says 'Linear correlations of Person'; this is a typo for 'Pearson'.
- [Table 5] The row for Sub-Saharan Africa reports a coefficient of 1.274 with a standard error of 4.241, which is implausibly large for a dummy variable in this model; please check the data entry or model specification.
- [General] The paper repeatedly uses the phrase 'likelihood of depositing a preprint' for a linear regression of the number of preprints; this wording is appropriate for logistic or count models, not for a linear probability interpretation of a count outcome. Please rephrase to 'predicted number of preprints' or similar.
- [Method] The Likert-scale perception items are treated as interval-level variables and entered linearly. While this is common practice, it should be stated explicitly as an assumption, and a robustness check with ordinal coding or with dummy variables for each Likert category would strengthen the inference.
- [Implications, limitations and recommendations] The limitations section does not mention the listwise deletion or the potential for collinearity among perception items; adding these to the limitations would give readers a more complete picture.
Circularity Check
No significant circularity: the regression outcome and predictors are distinct survey measures, and the paper's claims are empirical estimates rather than derivations from their own inputs.
full rationale
The paper is an empirical OLS regression using survey data. The dependent variable is the number of preprints posted in the last three years (Q6, level 5), while the key predictors are other self-reported research outputs from Q6 (journal articles, conference proceedings, book chapters, books and monographs, peer reviews) and perceptions of preprint effectiveness from Q13. These are distinct survey items, not transformations of the outcome, so the central coefficients (e.g., journal articles beta = 0.280, early access coef = 0.465, early feedback coef = -0.371) are estimated from data rather than forced by construction. The paper does not fit a parameter to a subset of the data and then rename it a prediction; it reports multivariate associations with clearly described variables. The two self-citations (Dorta-González and Dorta-González 2023; Dorta-González et al. 2024) appear only in the introduction as background on citation and impact effects and are not load-bearing for the preprint-posting analysis. No uniqueness theorem, ansatz, or imported modelling choice is invoked. The acknowledged limitations—self-selection bias, response bias, and the use of listwise deletion—are data-quality and inferential concerns, not circularity, because they do not make the outcome definitionally depend on the predictors. The derivation chain is therefore self-contained as a statistical analysis of an external survey dataset.
Assumptions & free parameters
assumptions (4)
- domain assumption The 5-point Likert items for perceived effectiveness are treated as interval-scale variables in an OLS regression.
- domain assumption Listwise deletion of records with any missing value yields a subsample that is unbiased for the regression.
- domain assumption Self-reported counts of preprints and other outputs correspond to actual behaviour.
- domain assumption The survey sample is representative of the global research community.
Cite this review
Pith. "Pith review of Effect of perceived preprint effectiveness and research intensity on posting behaviour." pith.science (2026). https://pith.science/paper/5WZ5FGXS
@misc{pith2026250418896,
author = {Pith},
title = {Pith review of: Effect of perceived preprint effectiveness and research intensity on posting behaviour},
year = {2026},
howpublished = {\url{https://pith.science/paper/5WZ5FGXS}},
note = {Machine review of arXiv:2504.18896}
}
read the original abstract
Open science is increasingly recognised worldwide, with preprint posting emerging as a key strategy. This study explores the factors influencing researchers' adoption of preprint publication, particularly the perceived effectiveness of this practice and research intensity indicators such as publication and review frequency. Using open data from a comprehensive survey with 5,873 valid responses, we conducted regression analyses to control for demographic variables. Researchers' productivity, particularly the number of journal articles and books published, greatly influences the frequency of preprint deposits. The perception of the effectiveness of preprints follows this. Preprints are viewed positively in terms of early access to new research, but negatively in terms of early feedback. Demographic variables, such as gender and the type of organisation conducting the research, do not have a significant impact on the production of preprints when other factors are controlled for. However, the researcher's discipline, years of experience and geographical region generally have a moderate effect on the production of preprints. These findings highlight the motivations and barriers associated with preprint publication and provide insights into how researchers perceive the benefits and challenges of this practice within the broader context of open science.
Reference graph
Works this paper leans on
-
[1]
Abdill, R. J., Adamowicz, E. M., & Blekhman, R. (2020). International authorship and collaboration across bioRxiv preprints. eLife, 9, e58496. https://doi.org/10.7554/eLife.58496 ASAPbio (2020). Preprint authors optimistic about benefits: Preliminary results from the #bioPreprints2020 survey. https://asapbio.org/biopreprints2020-survey-initial-results Bar...
-
[34]
https://doi.org/10.3390/publications7020034 Vale, R. D. (2015). Accelerating scientific publication in biology. Proceedings of the National Academy of Sciences of the United States of America, 112(44), 13439–13446. https://doi.org/10.1073/pnas.1511912112 Vale, R. D., & Hyman, A. A. (2016). Priority of discovery in the life sciences. eLife, 5, e16931. http...
-
[971]
Towards Responsible Publishing
https://doi.org/10.12688/f1000research.19619.2 Chiarelli, A., Cox, E., Johnson, R., Waltman, L., Kaltenbrunner, W., Brasil, A., Reyes Elizondo, A., & Pinfield, S. (2024). "Towards Responsible Publishing": Findings from a global stakeholder consultation. Zenodo. https://doi.org/10.5281/zenodo.11243942 Dorta-González, P., & Dorta-González, M. I. (2023). Cit...
arXiv 2024
-
[2019]
https://doi.org/10.11613/BM.2021.020201 Ng, J
Biochemia Medica, 31(2), 177–184. https://doi.org/10.11613/BM.2021.020201 Ng, J. Y ., Chow, V., Santoro, L. J., Armond, A. C. V., Pirshahid, S. E., Cobey, K. D., & Moher, D. (2023). An international, cross-sectional survey of preprinting attitudes among biomedical researchers. medRxiv. https://doi.org/10.1101/2023.09.17.23295682 Ni, R., & Waltman, L. (202...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.