{"id":"ac300445-1630-4cba-8cbb-450074efa880","arxiv_id":"2504.18971","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"An LLM-curated catalog of 18,247 scientific software repositories shows scientific projects outlive matched non-scientific ones, with infrastructure and government participation linked to longer survival.","lead":"Researchers built a catalog of 18,247 scientific open-source projects by using AI language models to read repository README files. They report that scientific software survives longer than similar non-scientific projects, with infrastructure, government involvement, and downstream users linked to longer lifespans.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3 comparison may be confounded by asymmetric sampling criteria: SciCat required recent activity after Nov 2018, while the matched non-scientific sample is not explicitly described as subject to the same recent-commit filter.","rationale":"I follow the reader's verdict in spirit. The paper is not fraudulent, and its transparency is commendable: the SciCat construction, LLM validation, and survival models are described in enough detail to reproduce or falsify the claims. The central, headline-grabbing claim is the RQ3 comparison, and its correctness hinges on the comparability of the matched non-scientific sample. The reader's weakest_assumption captures exactly this: the non-scientific sample is not documented as being subject to the same recent-commit/liveness filter as SciCat. My reading of Section 6 finds no explicit statement that the same Section 4.1 size/activity filters were applied to the non-scientific sample before matching; the bins for commits and authors could be satisfied by dead projects. Section 7's robustness check for stars does not address this temporal selection asymmetry, and the authors' own 'noninformative or ignorable' sampling requirement is not obviously met. Thus the load-bearing concern is sound. I do not go as far as REJECT because the concern is testable and likely repairable by re-matching or left-truncation analysis; the paper's contributions (SciCat dataset, RQ2 findings) are substantial and largely independent of this specific claim. CONDITIONAL means the headline should be reported as conditional on the matching being symmetric or on corrected survival analysis. I agree with the reader's weakest_assumption, so agreement_with_reader is agree.","tokens_in":25147,"tokens_out":1776,"duration_ms":16383,"concrete_test":"Re-run the RQ3 analysis after applying the identical Section 4.1 filters to the non-scientific sample: ≥10 files, >300 commits, ≥3 authors, >6 consecutive active months, and last commit after November 2018, then re-match on the same 48 bins. If the HR for Is Scientific Software remains 0.92 (or the confidence interval excludes 1.0 in the same direction), the concern is resolved. Additionally, fit the Cox model with delayed entry (left truncation) so that projects enter the risk set only once they satisfy the sampling filter, and compare the scientific indicator coefficient.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The RQ3 result (HR = 0.92, p < 0.001 for Is Scientific Software) compares SciCat, which was sampled with a last-commit-after-Nov-2018 filter (Section 4.1), against a non-scientific sample described only as matched on commit bin, author bin, and earliest-commit-year bin (Section 6, RQ3 Methods). The paper's own text states that for non-scientific projects they sampled from WoC using those bins, and does not state that the non-scientific sample was restricted to repositories with a commit after November 2018 or to those passing the same file/commit/author/activity filters. If the non-scientific sample includes many projects whose last activity was before the observation window (e.g., projects initiated before 2016 that died in 2012), those projects are 'known-dead' at time zero relative to the SciCat sample, whose members were all alive as of late 2018. Under a Cox model with earliest-commit-as-time-origin and no left-truncation/delayed-entry adjustment, the comparison mixes survival time with sample-selection artifact. The authors themselves assert in Section 6 that a random sample of non-scientific software would compare 'not longevity but precursors of longevity, such as activity,' and that the sampling must be 'noninformative or ignorable.' The burden is on the authors to show the matched sample was drawn under identical left-truncation and activity filters; the paper as written does not state this. The robustness check on stars (Section 7) does not address this temporal-selection asymmetry. If the concern lands, the 8% effect could be an artifact of conditioning on recent activity for one group only.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a two-part empirical study of scientific open-source software longevity. The authors build SciCat, a dataset of 18,247 repositories, by filtering World of Code with size/activity heuristics and then using a two-stage LLM pipeline (GPT-3.5 prescreening, GPT-4 classification) on READMEs to assign each repository a STEM field and a layer in Hinsen's software stack; the classification is validated with stratified human annotation and cross-checked against five external collections. Using SciCat, the paper estimates Kaplan-Meier survival curves and Cox proportional-hazards models to identify correlates of longevity (RQ2), and compares SciCat to a matched sample of 36,494 non-scientific repositories (RQ3). The headline findings are that infrastructure-layer position, downstream dependents, publication/funding mentions, and government participation are associated with lower abandonment hazard, that academic participation and recent start dates are associated with higher hazard, and that scientific software has an 8% lower abandonment hazard than matched non-scientific software.","tokens_in":25446,"tokens_out":10852,"duration_ms":109163,"significance":"The paper's dataset and classification methodology are solid and useful: the authors provide a replication package, report inter-rater reliability and per-class precision/recall, and transparently discuss classification errors. The RQ2 survival models are standard, with multiple operationalization checks, and the direction of effects is consistent with prior OSS ecosystem work. If the RQ3 comparison were valid, the finding that scientific software outlives comparable non-scientific software would be a notable and policy-relevant result. However, the current evidence for RQ3 does not support that conclusion because the non-scientific comparison sample is not described as being subject to the same eligibility filters as SciCat; the central claim therefore requires a substantial re-analysis. The dataset and RQ2 contributions remain valuable independently.","major_comments":[{"comment":"The paper's headline RQ3 result (HR = 0.92 for Is Scientific Software, Figure 7) is not supported by the sampling procedure as described. SciCat was restricted to repositories with at least 10 files, more than 300 commits, at least 3 authors, more than 6 consecutive active months, and a last commit after November 2018. The non-scientific sample is described only as matched on bins of commit count, author count, and earliest commit year, with the first commit bin being '<750 commits' and the first author bin being '<10 authors'. As written, the non-scientific sample can therefore include projects with very few commits or authors and, crucially, projects whose last commit predates the observation window. Because SciCat is selected on surviving to late 2018 while the non-scientific sample is not, the survival comparison confounds longevity with sample-selection criteria. The authors' assertion in Section 4.1 that 'we use the same filtering criteria' for the matched sample is not reflected in the RQ3 methods, and Section 6's own requirement that sampling be 'noninformative or ignorable' is not met. The star-count robustness check in Section 7 does not address this temporal-selection concern. The authors should either apply the identical eligibility filters, including the recent-commit filter, to the non-scientific sample and re-run the analysis, or adopt a left-truncation/delayed-entry Cox model and explain why the current comparison is valid.","section":"Section 6, RQ3 Methods; Section 4.1 sampling"},{"comment":"The Kaplan-Meier curves and restricted mean survival times (e.g., 6.44 years over 15 years) are estimated from time of earliest commit for a sample in which every project had a last commit after November 2018. This is length-biased sampling: repositories that died before November 2018 are excluded by construction, so the curves and RMST describe only projects that survived to that date, and they overstate typical longevity. The same conditioning affects the Cox models in RQ2, since they use the same sample without accounting for left truncation. Please report estimates with delayed entry (entry at the later of project start and November 2018) or explicitly reframe all RQ2 descriptive results as conditional on survival to the sample-entry date.","section":"Section 5.1, Figure 5; Section 5.2"}],"minor_comments":[{"comment":"The sentence 'which is has negative HR = 0.92' is ungrammatical; it should read 'which has an HR below 1'.","section":"Section 6, Results"},{"comment":"The censoring scheme should be described more precisely: a project with no commit in the last six months is treated as an event at its last commit, while an active project is censored; this is interval-censoring rather than pure right-censoring, and the exact censoring time used in the Cox model should be stated.","section":"Section 5, Methods"},{"comment":"What is called 'recall' in the cross-validation against external datasets is better described as overlap after filtering; consider renaming to avoid implying a well-defined ground-truth recall.","section":"Section 4.2"},{"comment":"The forest plot does not show confidence intervals; include them in the figure or a companion table to support the reported significance levels.","section":"Figure 7"},{"comment":"The November 2018 last-commit filter is motivated as avoiding 'less relevant repositories,' but this is precisely the condition that induces length-biased sampling; please discuss this explicitly in Threats to Validity.","section":"Section 4.1"},{"comment":"The word 'quantiative' should be 'quantitative'.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern from the reader's report lands: the RQ3 sampling asymmetry is real and is the central threat to the paper's headline. I would ask the editor to require the revised version to (a) state explicitly and verify that the non-scientific sample passes the identical eligibility filters, or re-sample it, and (b) address left-truncation in the RQ2 descriptive analysis. The SciCat dataset and RQ2 factor analysis are strong enough to warrant a revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I think the reader's rejection is too harsh. The central worry about RQ3—that the non-scientific sample wasn't subject to the same recent-commit filter as SciCat—is addressed by the paper itself. Section 4.1 states plainly that \"we use the same filtering criteria to identify the matching non-scientific OSS repositories we compare against.\" Section 6's methods description is terse, but it doesn't contradict that; it just emphasizes the stratified bins. So the stress-test's \"as written\" objection doesn't land.\n\nWhat's genuinely new: the LLM-based README classification for detecting scientific software across domains and stack layers, with human validation and cross-checks against JOSS, Papers with Code, and other external datasets. The SciCat dataset is a real reusable asset, and the authors are appropriately careful about its limitations. The RQ2 survival analysis is transparent, uses standard Cox models with sensible controls, applies Bonferroni correction, and includes robustness checks on the abandonment window. That's solid work.\n\nSoft spots, in proportion: the inter-rater agreement for the paper/funding mention label is moderate (0.467), and that variable feeds into the model. Matching is on coarse bins, not exact, so residual confounding is possible in RQ3. The headline effect (HR = 0.92) is small—statistically significant but practically modest. There are also a few copyediting slips (\"which which\", \"is has negative\"). None of these are load-bearing.\n\nWho gets value: empirical software engineering researchers, research software engineers, and science funders. The dataset alone will get cited. The paper should go to peer review; I'd ask for a one-sentence clarification in Section 6 that the non-scientific sample was drawn from the same filtered universe, and a methodologist look at the delayed-entry issue as a robustness check. But the correct verdict is \"revise,\" not \"reject.\"","headline":"The RQ3 matching concern that sank the reader's verdict is explicitly answered in Section 4.1; this is a solid empirical paper that deserves serious peer review, not a desk reject.","tokens_in":26009,"tokens_out":2267,"would_cite":true,"duration_ms":23737,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scientific open-source software projects are abandoned at an 8% lower rate than matched non-scientific open-source projects, according to survival models on a curated set of 18,247 repositories.","keywords":["scientific software","open source software","software longevity","software abandonment","survival analysis","large language models","software sustainability","software ecosystems"],"falsifier":"Re-run the RQ3 comparison after passing the non-scientific sample through the exact full eligibility filter used for the scientific catalog (at least 10 files, more than 300 commits, at least 3 authors, more than 6 active months, and a last commit after November 2018) before matching; if the 0.92 hazard ratio for 'is scientific software' moves to or above 1, the headline result is an artifact of how the non-scientific projects were sampled.","tokens_in":24942,"feed_emoji":"🔬","tokens_out":8342,"duration_ms":78144,"temperature":0.7,"pith_summary":"Using large language models to read repository README files, the authors assemble a catalog of 18,247 scientific open-source software projects spanning 13 STEM fields and three layers of the scientific software stack. They then track how long these projects stay active before going dormant (no commits for six months). Their central finding is that scientific projects outlive matched non-scientific open-source projects: in a Cox proportional-hazards model, being scientific software lowers the hazard of abandonment by roughly 8%. Within the catalog, scientific infrastructure, projects with many downstream users, projects that mention publications or funding, and government participants are all associated with longer lifespans, while newer start years and academic participants are associated with shorter lifespans. The paper offers the catalog itself and these survival results as a baseline for monitoring and sustaining scientific software.","feed_headline":"Scientific open-source software outlives comparable projects","feed_subtitle":"An 18,000-repo study finds science code has an 8% lower abandonment risk than matched non-scientific OSS.","key_machinery":"The load-bearing machinery is a two-stage pipeline that classifies repositories from README content, combined with survival models on the resulting labeled set. An initial large-language-model pass separates scientific from non-scientific software and detects mentions of publications or funding; a second pass assigns each scientific repository to one of three software-stack layers — scientific infrastructure, domain-specific code, or publication-specific code — and one of 13 STEM fields. Longevity is measured as time from earliest to latest commit, with abandonment defined as six months without commits, and the paper uses Kaplan-Meier curves and Cox proportional-hazards regression to estimate which factors raise or lower abandonment risk. The decisive comparison for the headline result is the binary 'is scientific software' indicator estimated on scientific repositories matched to non-scientific repositories by number of commits, number of authors, and earliest commit year.","core_discovery":"The paper's discovery is that scientific open-source software, far from being especially fragile, tends to be more durable than otherwise comparable non-scientific open-source software. Using a matched sample of 36,494 non-scientific repositories selected to resemble the 18,247 scientific ones in commit count, author count, and era, the authors fit a Cox proportional-hazards model in which the indicator 'is scientific software' has a hazard ratio of 0.92 (p < 0.001), corresponding to an 8% lower risk of abandonment. The same model shows scientific infrastructure outlasting domain-specific code, which in turn outlasts publication-specific code; projects with more downstream dependents, explicit mentions of publications or funding, and government participants also survive longer, whereas newer projects and projects with academic participants face higher abandonment risk. The authors interpret this as evidence that science-specific incentives such as publication credit and funding support, while short-term, give scientific projects a measurable durability advantage over generic open-source software.","pith_inferences":["If the longevity advantage is driven by external anchors such as grants and publication credit, then non-scientific OSS could borrow the practice of tying development to documented outcomes or institutional sponsorship — a transfer the paper mentions only as a possibility.","Because the classifier depends on READMEs, scientific software with sparse or outdated documentation may be underrepresented; applying the same method to code content or paper-mention mining could shift the estimated 8% advantage.","The matched comparison holds commit count, author count, and era constant but not popularity directly; a replication that also matches on stars or forks would test whether the scientific advantage survives once visibility is equalized.","The average scientific advantage hides real variation — computer science and data science projects show shorter lifespans than astronomy or mathematics — so policy responses aimed at the whole class may miss the projects most at risk."],"forward_implications":["The common assumption that scientific software is unusually prone to abandonment should be qualified: for large, collaboratively developed projects, scientific software appears to be the sturdier category.","Funding agencies concerned with sustainability can allocate attention by layer, since scientific infrastructure shows the lowest abandonment hazard and publication-specific code the highest.","Measurable project attributes — downstream dependents, documentation that mentions publications or funding, and government participation — are associated with longer lifespans and can serve as early indicators of durability.","The catalog of over 18,000 labeled repositories gives future studies a population from which to sample projects for qualitative or longitudinal work on sustainability."],"supporting_citations":[{"why":"supplies the universe of public repositories and the commit, author, and dependency data used for sampling, classification, and survival analysis","marker":"[66]"},{"why":"provides the scientific software stack taxonomy used to assign projects to layers and motivates the hypothesis that infrastructure should be longer-lived","marker":"[44]"},{"why":"introduces the proportional-hazards regression model used to estimate abandonment risk and the scientific-vs-non-scientific comparison","marker":"[25]"},{"why":"reports prior domain-level longevity measurements for research software that this study contrasts and extends","marker":"[42]"},{"why":"represents the prior publication-link-based identification approach that the README-based LLM classifier is designed to complement","marker":"[108]"},{"why":"shows maintainers sometimes return after long breaks, supporting the paper's robustness checks with extended inactivity thresholds","marker":"[15]"}],"fun_headline_variants":["Science open-source code outlasts non-science counterparts","8% lower abandonment risk for scientific open-source software","Study: Scientific software less likely to be abandoned than thought","18,000 scientific repos show greater longevity than typical OSS","Contrary to perception, science open-source projects prove durable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the non-scientific sample was selected under the same rules as the scientific one — in particular the requirement of a recent commit after November 2018 and the same minimum size and activity filters — so that the survival gap reflects genuine longevity rather than different selection windows.","fun_headline_variants_meta":{"raw":{"variants":["Science open-source code outlasts non-science counterparts","8% lower abandonment risk for scientific open-source software","Study: Scientific software less likely to be abandoned than thought","18,000 scientific repos show greater longevity than typical OSS","Contrary to perception, science open-source projects prove durable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000659,"raw_usage":{"total_tokens":3036,"prompt_tokens":991,"completion_tokens":2045,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1964}},"tokens_in":607,"tokens_out":2045,"duration_ms":13419,"temperature":1.0,"reasoning_tokens":1964,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:04:53.834422+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the RQ3 comparison after passing the non-scientific sample through the exact full eligibility filter used for the scientific catalog (at least 10 files, more than 300 commits, at least 3 authors, more than 6 active months, and a last commit after November 2018) before matching; if the 0.92 hazard ratio for 'is scientific software' moves to or above 1, the headline result is an artifact of how the non-scientific projects were sampled.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the universe of public repositories and the commit, author, and dependency data used for sampling, classification, and survival analysis"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the scientific software stack taxonomy used to assign projects to layers and motivates the hypothesis that infrastructure should be longer-lived"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"reports prior domain-level longevity measurements for research software that this study contrasts and extends"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"represents the prior publication-link-based identification approach that the README-based LLM classifier is designed to complement"}],"review_version":1}