{"id":"a2982e4c-bdc6-4143-9160-a9c3b0d51f89","arxiv_id":"2504.18896","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Researchers post more preprints when they are already productive in journal articles and books; perceptions that preprints provide early access encourage posting, while valuing early feedback is associated with less posting.","lead":"Analyzing survey responses from 5,873 researchers, this paper finds that traditional publishing productivity, especially journal articles and books, is the strongest predictor of how often researchers post preprints. It also reports that researchers who value preprints for early access post more, while those who value them for early feedback post fewer.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Listwise deletion of 47% of responses is an untested selection filter; key coefficients in Table 5 may be biased.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: listwise deletion without missing-data analysis. Since nearly half of the 11,145 researcher responses are excluded, the complete-case sample may not represent the survey population. The OLS estimates in Table 5, including the negative early-feedback coefficient, could be biased if missingness depends on preprint production, perceptions, or unobserved factors correlated with these. This is not a speculative threat: the paper itself acknowledges using listwise deletion for mathematical convenience, and the raw correlation between early-feedback perceptions and preprint counts is essentially zero. The recommended concrete test is feasible with the openly available dataset and would settle whether the concern lands. Because the reader already issued a conditional verdict based on this concern, I see no reason to move the verdict; the paper should remain conditional pending the missing-data robustness check and softened causal language.","tokens_in":14560,"tokens_out":9297,"duration_ms":95659,"concrete_test":"Download the open dataset (Chiarelli et al., 2024; Zenodo 10.5281/zenodo.11243942), reconstruct the full 11,145-researcher sample, and run two checks: (1) compare demographic and preprint-production distributions between the 5,873 complete cases and the 5,272 excluded records using available items; (2) re-estimate Table 5 with multiple imputation by chained equations (or inverse-probability weights) on the full sample. If the early-feedback coefficient changes by more than one standard error, loses significance, or changes sign, listwise deletion is the load-bearing problem. Also report the proportion of missingness per variable and a Little MCAR test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper reduces 11,145 researcher responses to 5,873 complete cases via listwise deletion (Data section) but reports no missing-data diagnostics. The central claim—that journal articles (beta=0.280) and books (beta=0.163) dominate and that perceived effectiveness for early feedback is negatively associated with preprint posting (coef=-0.371)—is only as trustworthy as the assumption that the 5,272 excluded records are missing completely at random or that missingness is fully explained by included covariates. If, say, high-prestige researchers or those with many preprints are more or less likely to complete all perception and demographic items, the OLS coefficients are biased. The negative early-feedback coefficient is particularly fragile: its raw correlation with preprints is only +0.03 (Table 3), so a selection-induced shift in the joint distribution of perceptions and preprints could flip it. The Data section explicitly invokes listwise deletion as a mathematical convenience without acknowledging that it alters the estimand if missingness is informative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes open survey data from a global cOAlition S consultation (11,145 researcher responses, reduced to 5,873 complete cases) to explain researchers' preprint posting frequency over the last three years. Using OLS regression with the number of preprints as the outcome, the authors report that traditional productivity measures—especially journal articles and books/monographs—are the strongest predictors of preprint production, that perceived effectiveness for early access is positively associated with posting, while perceived effectiveness for early feedback is negatively associated, and that most demographic variables are not significant. The authors interpret these findings as evidence that preprints complement, rather than substitute for, traditional publishing, and that researchers' perceptions of early feedback present a nuanced barrier.","tokens_in":14690,"tokens_out":2118,"duration_ms":24490,"significance":"The study addresses a relevant and under-explored question in scholarly communication: how perceptions of preprint effectiveness and research intensity shape actual posting behavior. Its strengths include the use of a large, international, openly available survey dataset and the explicit focus on standardizing effect sizes to compare predictors. If the perception results were robust, the negative association between valuing early feedback and posting preprints would be a novel and policy-relevant contribution. The productivity finding—that journal articles and books are the strongest predictors—is plausible and consistent with prior work, and it is supported by simple correlations and a highly significant regression coefficient. However, the central perception claims rest on a cross-sectional design with substantial listwise deletion, and the manuscript does not provide the diagnostic checks needed to establish that these coefficients are not artifacts of selection or collinearity.","major_comments":[{"comment":"The manuscript reports that 11,145 researcher responses were reduced to 5,873 valid observations by listwise deletion, but it provides no analysis of missing-data patterns or comparison of included and excluded respondents. Because the regression estimates in Table 5 are computed only on complete cases, the coefficients—especially the perception coefficients—are unbiased only if missingness is completely at random or fully explained by the included covariates. The negative coefficient on 'receiving early feedback' (-0.371) is particularly fragile: its raw correlation with preprint production is only +0.03 (Table 3), so a modest selection-induced shift in the joint distribution of perceptions and preprints could flip its sign. The authors should report the missingness rate per variable, compare means and distributions on observed variables between complete and incomplete cases, and ideally present sensitivity analyses (e.g., multiple imputation or inverse-probability weighting).","section":"Data"},{"comment":"The outcome variable—number of preprints in the last three years—is a bounded count (0–40, mean 2.2, SD 4.9) with a large mass at zero, yet it is modeled with OLS. This can produce inefficient estimates, incorrect standard errors, and predicted values outside the feasible range. The paper does not report residual diagnostics, heteroskedasticity-robust standard errors, or a count-model alternative (e.g., Poisson or negative binomial). Given that the main productivity coefficients are large and highly significant, they may survive a count-model re-estimation, but the smaller perception coefficients could change in significance or sign. The authors should at least re-estimate the model with robust standard errors and with a negative binomial or Poisson specification to confirm the key results.","section":"Method and Table 5"},{"comment":"The perception variables are five Likert items that are highly intercorrelated (Table 3 shows correlations up to 0.72 between 'enhancing accessibility' and 'accelerating discourse', and 0.65 between 'early feedback' and 'accelerating discourse'). The regression enters all five simultaneously, but the paper reports no multicollinearity diagnostics such as variance inflation factors. The negative coefficient on 'receiving early feedback' could partly reflect suppressor effects among collinear perception items; the manuscript's interpretation as a substantive negative perception is not warranted without evidence that the coefficient is stable across model specifications. I recommend reporting VIFs and showing a specification that enters each perception item separately or in subsets.","section":"Method and Table 3"},{"comment":"The perception measures are taken at the time of the survey, after respondents have accumulated years of experience with preprints, so the regression cannot distinguish between perceptions causing posting behavior and posting behavior shaping perceptions. The paper acknowledges self-selection and self-report bias in the limitations section, but it does not address reverse causality, which is especially relevant for the early-feedback coefficient. For instance, researchers who have posted preprints and received poor feedback may lower their rating of early feedback, producing a negative association without any causal effect of perception on posting. The authors should temper the causal language (e.g., 'hurts preprint production', 'decreases the likelihood') and discuss the direction-of-causality problem explicitly, or use instrumental-variable or longitudinal designs if available.","section":"Method and Results"}],"minor_comments":[{"comment":"The title says 'Linear correlations of Person'; this is a typo for 'Pearson'.","section":"Table 3"},{"comment":"The row for Sub-Saharan Africa reports a coefficient of 1.274 with a standard error of 4.241, which is implausibly large for a dummy variable in this model; please check the data entry or model specification.","section":"Table 5"},{"comment":"The paper repeatedly uses the phrase 'likelihood of depositing a preprint' for a linear regression of the number of preprints; this wording is appropriate for logistic or count models, not for a linear probability interpretation of a count outcome. Please rephrase to 'predicted number of preprints' or similar.","section":"General"},{"comment":"The Likert-scale perception items are treated as interval-level variables and entered linearly. While this is common practice, it should be stated explicitly as an assumption, and a robustness check with ordinal coding or with dummy variables for each Likert category would strengthen the inference.","section":"Method"},{"comment":"The limitations section does not mention the listwise deletion or the potential for collinearity among perception items; adding these to the limitations would give readers a more complete picture.","section":"Implications, limitations and recommendations"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a bibliometrics/scholarly communication journal and addresses a timely topic. The productivity finding is likely solid, but the perception-related claims require substantially stronger statistical support. Given that the weaknesses are fixable with additional analyses (missing-data checks, count models, collinearity diagnostics, and a more cautious causal interpretation), major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe only result here that will stay with you is the negative coefficient for perceived effectiveness of early feedback (-0.37, β=-0.09). It is new, and it is probably fragile. Everything else — productivity in journal articles and books predicts preprint posting — is robust and consistent with earlier work, but the paper frames it as a challenge to an assumption that few people actually hold.\n\nWhat the paper does well: it takes a large, open, global survey (11,145 researchers in the original consultation, 5,873 after complete-case cleaning), runs a clean OLS with demographic controls, reports standardized coefficients, and is transparent about the data source. That is real work. The complementarity finding — that preprint posters are already productive in traditional channels rather than using preprints as an alternative — is a useful confirmatory result, and the field/region patterns mirror Ni and Waltman.\n\nThe soft spots are mostly around the headline perception result. The raw correlation between 'receiving early feedback' and number of preprints is 0.03. In a multiple regression that includes four other perception items, all correlated with feedback at 0.45–0.65, the coefficient flips to −0.37. That is a textbook suppression/multicollinearity situation; the sign may be an artifact of the covariate set rather than something real about feedback. On top of that, the design is cross-sectional, so reverse causation is plausible: people who posted many preprints and got little or poor feedback may rate the feedback item lower. The paper's causal language ('hurts', 'decreases') is not justified.\n\nThe bigger technical concern is the listwise deletion. Dropping 5,272 of 11,145 respondents without any missing-data diagnostic means the estimates are only valid if missingness is completely at random. If researchers with more preprints (or stronger opinions) were less likely to answer all perception items, the coefficients, including the negative feedback one, could be biased. This is not a fatal flaw for the productivity results, but it matters for the fragile perception coefficients.\n\nAnother minor point: the outcome is a count (0–40 preprints) modeled with OLS. A Poisson or negative binomial would be worth a robustness check; the standard errors and possibly the significances could shift.\n\nSo: serious referee? Yes, but a conditional accept with specific requests: missing-data comparison, a multicollinearity/sensitivity analysis for the perception variables, and softened interpretation. The paper is a modest empirical contribution for scholarly communication and open science policy audiences, not a game-changer. I would bring it to a reading group if we were discussing methodological pitfalls in survey regression, but not for the substantive news.\n\nRecommendation: send to peer review, with those conditions in mind.","headline":"A clean reanalysis of an open survey with one genuinely new but fragile result — the negative early-feedback coefficient — and a robust productivity finding; worth refereeing with conditions.","tokens_in":15229,"tokens_out":2704,"would_cite":true,"duration_ms":27055,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that preprint posting is driven above all by how much a researcher already publishes in journals and books, and that valuing preprints for early access predicts more posting while valuing them for early feedback predicts…","keywords":["preprints","open access","open science","scholarly communication","researcher survey","OLS regression","perceived effectiveness","publication productivity"],"falsifier":"Go to the open survey file and recover the 5,272 respondents excluded for incomplete answers, then re-estimate the model with multiple imputation or compare the included and excluded groups' preprint counts and effectiveness ratings. If the excluded researchers post more preprints or rate early feedback more favourably than complete-case respondents, the OLS coefficients are biased; the early-feedback coefficient ($-0.371$) changing sign or the journal-article effect ($\\beta = 0.280$) collapsing toward the perception effects would settle the paper's ranking of drivers differently.","tokens_in":14338,"feed_emoji":"📄","tokens_out":13659,"duration_ms":123172,"temperature":0.7,"pith_summary":"The paper asks what actually makes researchers post preprints, and answers with an ordinary least-squares regression on 5,873 complete responses to a global survey of scholarly communication practices. It claims that the strongest predictor of preprint posting is how much a researcher already produces through traditional channels: journal articles ($\\beta = 0.280$) and books and monographs ($\\beta = 0.163$) carry the two largest standardized effects after the physics-and-astronomy field effect. Perceptions split by function: researchers who rate preprints effective for early access post more (coefficient $0.465$ per Likert point), while those who rate them effective for early feedback post fewer (coefficient $-0.371$). Net of these factors, gender and organisation type show essentially no association with posting, while discipline, career stage and region retain moderate effects. These findings matter because they undercut the idea that preprints are a workaround for researchers locked out of traditional venues, suggesting instead that preprints complement an already-active publication profile.","feed_headline":"Journal publishing is the strongest predictor of preprint posts","feed_subtitle":"A 5,873-researcher regression ranks journal output first; valuing early feedback predicts less posting.","key_machinery":"The central object is one Ordinary Least Squares (OLS) linear regression model: the dependent variable is the number of preprints posted in the last three years (question Q6, level 5), and 53 coefficients are estimated across six blocks, namely five research-output counts (journal articles, conference proceedings, book chapters, books and monographs, peer reviews), five Likert-rated items on perceived preprint effectiveness (Q13), and then field, organisation type, region, research experience and gender as controls. The machinery does two things: the raw coefficients give ceteris paribus effects (each journal article adds $0.155$ expected preprints; each point of early-access effectiveness adds $0.465$, and each point of early-feedback effectiveness subtracts $0.371$), and the standardised coefficients rank the drivers by effect size. The analysis works because the survey questions on outputs, perceptions and demographics come from the same respondents, letting the model separate what is attributable to productivity, what to perception and what is simply demographic.","core_discovery":"The paper's core claim is that preprint posting is a complement to traditional scholarly output, not a substitute for it. In the OLS model, the number of journal articles ($\\beta = 0.280$, the largest standardized effect) and of books and monographs ($\\beta = 0.163$, the third largest) are the top productivity predictors of how many preprints a researcher posted in the last three years, with the Physics and Astronomy field effect ($\\beta = 0.191$) between them. Perceived effectiveness splits by function: each Likert point of perceived effectiveness for providing early access to new research raises expected preprint counts by $0.465$, while each point for receiving early feedback lowers them by $0.371$, a split the paper interprets as early dissemination being valued while the feedback function is perceived as risky or inadequate. Net of these, gender and organisation type show essentially no association with posting, whereas discipline, career stage and geography retain moderate effects, and the model as a whole explains $26.8\\%$ of the variance (adjusted $R^2 = 0.268$). The authors take this as evidence against the assumption that preprints mainly help researchers who cannot get published in traditional venues.","pith_inferences":["If preprint posting truly complements traditional output, a consequence the paper leaves implicit is a Matthew effect in scholarly communication: the same authors who already dominate journal visibility would also gain the earliest visibility through preprints, widening the gap with less prolific researchers.","The regression is run on the 5,873 complete cases drawn from 11,145 survey respondents; re-estimating on the full open dataset with multiple imputation for missing answers is the natural next test of whether the negative early-feedback coefficient is a behavioural effect or a selection artefact.","The five effectiveness items are strongly intercorrelated (early access and early feedback correlate $0.48$), so the opposite signs of their coefficients hint at a suppressor relationship; a partial-correlation or mediation decomposition would clarify whether the negative feedback effect is genuine or a by-product of collinearity.","The effectiveness perception items may well interact with field: in disciplines where preprint cultures are mature, such as physics and mathematics, the early-feedback effect could differ from fields with younger commenting cultures, a pattern testable by adding interaction terms to a larger sample."],"forward_implications":["Preprint promotion aimed at researchers who struggle to publish in traditional venues is likely aimed at the wrong group; the paper's corollary is that preprint incentives belong inside the reward structures that already track journal and book output.","Because expecting useful early feedback currently predicts fewer preprints at every level of productivity, platforms that build more structured and supportive commenting mechanisms could remove a genuine barrier to posting.","With gender and organisation type essentially null net of other factors, observed demographic gaps in preprint use are likely productivity, discipline and region effects in disguise rather than independent barriers.","The persistent moderate effects of discipline, experience and region mean uniform open-science mandates will produce uneven uptake, so policy should be tailored to fields and regions with weaker preprint cultures.","With adjusted $R^2 = 0.268$, most individual variation in posting is left unexplained, so institutional contexts, platform features and field-level norms that the survey does not measure still carry most of the weight."],"supporting_citations":[{"why":"Supplies the dataset: the open global stakeholder consultation survey whose complete cases the regression analyses.","marker":"Chiarelli et al. (2024)"},{"why":"The prior global researcher survey that frames disciplinary and regional preprint adoption patterns and serves as the study's comparison baseline.","marker":"Ni & Waltman (2024)"},{"why":"Establishes the drivers and barriers of preprint adoption that motivate the productivity and perception variables.","marker":"Chiarelli et al. (2019)"},{"why":"The closest prior survey evidence on bioRxiv authors' motivations and concerns, against which the perception findings are positioned.","marker":"Fraser et al. (2022)"},{"why":"Supplies the catalogue of preprint benefits, including early access and early feedback, that the effectiveness items operationalize.","marker":"Puebla et al. (2021)"},{"why":"Documents how journal preprint policies vary sharply across disciplines, used to interpret the disciplinary coefficients.","marker":"Klebel et al. (2020)"}],"fun_headline_variants":["Preprints ride on publishing record, not substitute","Journal output fuels preprint posting, survey of 5,873 finds","Early access praised, early feedback feared in preprint adoption","Preprint posting tracks publication record, not demographics","Complement not substitute: preprint posts follow publishing output"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model is estimated only on the 5,873 respondents who answered every question, a complete-case sample chosen without testing how the 5,272 dropped respondents differ; the load-bearing premise is that these complete cases behave like all 11,145 researchers surveyed, so if who finishes the survey depends on preprint activity or preprint perceptions, every coefficient—including the negative early-feedback effect—could be an artefact of selection rather than a genuine driver of posting.","fun_headline_variants_meta":{"raw":{"variants":["Preprints ride on publishing record, not substitute","Journal output fuels preprint posting, survey of 5,873 finds","Early access praised, early feedback feared in preprint adoption","Preprint posting tracks publication record, not demographics","Complement not substitute: preprint posts follow publishing output"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3099,"prompt_tokens":973,"completion_tokens":2126,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":2050}},"tokens_in":589,"tokens_out":2126,"duration_ms":14695,"temperature":1.0,"reasoning_tokens":2050,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:05:53.895202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Go to the open survey file and recover the 5,272 respondents excluded for incomplete answers, then re-estimate the model with multiple imputation or compare the included and excluded groups' preprint counts and effectiveness ratings. If the excluded researchers post more preprints or rate early feedback more favourably than complete-case respondents, the OLS coefficients are biased; the early-feedback coefficient ($-0.371$) changing sign or the journal-article effect ($\\beta = 0.280$) collapsing toward the perception effects would settle the paper's ranking of drivers differently.","supporting_citations":[],"review_version":1}