{"id":"fcb72570-d7c0-46e2-9a33-6e330abbe2b1","arxiv_id":"2506.20217","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Ten practical rules to help principal investigators integrate Research Software Engineering into their research groups.","lead":"This paper gives principal investigators ten practical rules for bringing Research Software Engineering practices into their research groups. It argues that better software practices make research more reproducible and trustworthy, and explains each rule in plain language for non-experts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's causal promise that following the rules improves research outcomes is not supported by the cited evidence; the references are editorials and position papers, and no outcome data appear, so the central claim is an overclaim.","rationale":"The reader's weakest_assumption identifies exactly the concern that matters most: the paper asserts a causal benefit without supplying empirical evidence. My review of the full text and reference list confirms this. The rules themselves are internally coherent and practical, and the paper does not make formal mathematical or empirical claims that could be falsified from within. The absence of outcome data is not an internal inconsistency, but it does make the abstract's causal promise unsupported. Because the paper is explicitly a 'Ten Simple Rules' advisory piece, I would not reject it for lacking a randomized trial; however, the strength of the language should match the strength of the evidence. The acknowledgments' mention of surveys actually highlights the gap: the surveys are the closest thing to empirical grounding, yet their questions, results, and analysis are not provided. A structured reference-classification check would settle whether any cited source offers genuine outcome evidence; my reading suggests none does. This does not change the reader's conditional verdict: the paper is conditionally acceptable with a softened causal claim and, ideally, availability of the underlying survey materials. The concern is real but not a reason to reject the paper outright.","tokens_in":7560,"tokens_out":2591,"duration_ms":35235,"concrete_test":"Classify the study design of each reference invoked for the causal chain (refs [1], [2], [8], [9], [15], [17]-[19]). If every one is an editorial, position paper, descriptive case study, or perception survey with no control or comparison group measuring research outcomes, then the claim that following these rules 'ultimately leads to better research outcomes' is not established by the paper's own evidence and should be explicitly hedged as expert opinion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the abstract and repeated in the introduction, is that following the ten rules will improve the quality, reproducibility, and trustworthiness of research software and 'ultimately lead to better, more reproducible and more trustworthy research outcomes.' The load-bearing assumption is that a causal chain runs from PI adoption of these practices to measurable research improvements. The paper provides no empirical evidence for this chain. The citations offered in support are not empirical studies: ref. [1] is a position paper on the 'Four Pillars' of RSEng, ref. [2] is an editorial titled 'Better Software, Better Research,' and the other supporting references are guidelines, case descriptions, or systematic reviews of developer perceptions rather than outcome comparisons. The acknowledgments mention surveys that 'led to some of the content of our rules,' but the survey data, instruments, and analyses are not included, so the reader cannot assess whether the rules are grounded in evidence about outcomes or only in practitioner opinion. This is not an internal inconsistency: the individual rules are reasonable and mutually compatible heuristics, and the paper is clearly in the 'Ten Simple Rules' advisory genre. But the abstract's promise is stronger than the evidence presented, and the unsupported causal link is the key premise on which the paper's value to PIs rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents ten rules aimed at principal investigators for incorporating research software engineering (RSEng) practices into their research groups. The rules span early adoption of RSEng, requirements elicitation, software architecture, tool selection, version control, open-source development, documentation, quality assurance, and software publication and citation. The paper is a position piece in the 'Ten Simple Rules' genre; it contains no new empirical data and instead synthesizes existing literature and the authors' expertise. The abstract and introduction assert that following the rules will improve software quality, reproducibility, and trustworthiness, and ultimately lead to better research outcomes.","tokens_in":7844,"tokens_out":4766,"duration_ms":48854,"significance":"The paper addresses a real gap: most RSEng guidance is technical and aimed at developers, while PIs make the resource decisions. The rules are concrete, actionable, and largely consistent with established software engineering practice and prior 'Ten Simple Rules' literature. The authors correctly emphasize that RSEng involves management and knowledge transfer, not just coding. The principal weakness is that the paper's central promise—that adopting these rules causes better research outcomes—is asserted rather than demonstrated; the cited support is mostly editorial or position literature, and the surveys mentioned in the Acknowledgments are not provided. As a practical checklist, the paper is a useful contribution; as an empirical claim, it is unsubstantiated.","major_comments":[{"comment":"The abstract ('By following these rules, readers can improve ... ultimately leading to better, more reproducible and more trustworthy research outcomes') and the introduction ('better research software leads to better research' [2]) make a causal claim that is not supported by the evidence cited. References [1] and [2] are a position paper and an editorial, respectively, and the surveys mentioned in the Acknowledgments are not included in the manuscript. Because this causal promise is the paper's primary value proposition to PIs, please either (a) supply or cite empirical evidence (e.g., the survey instruments and results, or a synthesis of quantitative studies linking RSEng practices to research outcomes), or (b) temper the claims to state that these are expert recommendations whose benefits are widely acknowledged but not yet rigorously measured.","section":"Abstract and Introduction"}],"minor_comments":[{"comment":"The heading contains a typo: 'W e' should be 'We'.","section":"Rule 3"},{"comment":"The sentence 'Software isopen-source when its source code is publicly available' is missing a space; it should read 'Software is open-source'.","section":"Rule 7"},{"comment":"The phrase 'it also not to be disregarded lightly' is grammatically incomplete; it should read 'it also is not to be disregarded lightly' or 'it should not be disregarded lightly'.","section":"Rule 5"},{"comment":"Reference 14 has a typo ('Foundaton' should be 'Foundation') and Reference 20 has a typo ('Allaince' should be 'Alliance').","section":"References"},{"comment":"The informal aside 'oops, now I have made this a scary story you might tell PIs around a campfire' is out of place in a journal article and should be removed or rewritten.","section":"Rule 9"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's Ten Simple Rules series, and the rules themselves are sensible. The main concern is the overstated causal claim in the abstract, which is fixable by rewording or by including a brief evidence statement. The self-citation rate (e.g., refs. 1, 12, 20) is noticeable but not disqualifying. The informal aside in Rule 9 and several typos should be cleaned up."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the ten rules paper. It is exactly what it says on the tin: a clear, accessible set of recommendations for research group leaders who don't know RSE from a hole in the ground. The framing is genuinely useful—most of this material lives in technical tutorials or in papers aimed at developers, not at PIs who control budgets and hiring. The authors have done a good job translating version control, testing, documentation, and open source into management-level language, with concrete \"go do this\" guidance. Rules 3 and 5 are particularly well done: they address the actual decision context of a research group rather than just listing practices.\n\nThe new element here is audience, not content. Each rule maps to prior work, and the paper cites that work honestly. That's fine for the Ten Simple Rules genre, which is a synthesis-and-framing exercise, not a discovery vehicle.\n\nThe soft spot is the causal claim. The abstract and introduction say that following the rules will \"ultimately lead to better, more reproducible and more trustworthy research outcomes.\" That is a statement about efficacy, and the paper offers no evidence for it beyond plausibility and citations to other position papers. The acknowledgments mention surveys, but the instruments and data aren't included, so a reader can't check whether the rules are grounded in outcome data or just practitioner opinion. In fairness, the rules are all reasonable heuristics that line up with general software engineering evidence, so the overclaim is more a matter of tone than a fatal flaw. If the authors softened the language to \"can help\" or \"makes it more likely,\" the paper would be on solid ground.\n\nOtherwise: the citation pattern is fine, with some self-citation among the RSE community but proportional to the topic. No internal inconsistencies. The paper is exactly as deep as the genre requires, no more.\n\nWho's it for? Every PI who writes research software or manages people who do. This is a good thing to hand to a new group leader. It deserves a real peer review—the positioning, tone, and claims should be scrutinized by someone outside the immediate RSE circle. I'd send it to review with a request to soften the abstract and, ideally, to make the survey data available or at least describe how it shaped the rules.\n\nSerious thinker: yes.","headline":"A well-written, PI-facing synthesis of known RSE practices; the advice is sound but the abstract overpromises a causal payoff the paper can't back.","tokens_in":8340,"tokens_out":2238,"would_cite":true,"duration_ms":23714,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ten simple rules put research software engineering within a PI's reach.","keywords":["research software engineering","principal investigators","research group management","software quality assurance","reproducibility","open source software","software citation","ten simple rules"],"falsifier":"A controlled comparison of matched research groups—half adopting the ten rules, half continuing usual practice—measuring defect density, successful reproduction of results, or software citation rates over 12–24 months would settle the claim; if adopters show no improvement, the causal link fails. Alternatively, a survey-based study showing that groups reporting high rule adherence have no better reproducibility scores than low-adherence groups would falsify it.","tokens_in":7417,"feed_emoji":"💻","tokens_out":4404,"duration_ms":44736,"temperature":0.7,"pith_summary":"This paper argues that research leaders do not need a computer science degree to adopt Research Software Engineering (RSEng) in their groups. It offers ten practical rules—covering requirements, architecture, tool choice, version control, open source, documentation, quality assurance, and citation—that translate software-engineering practice into management decisions a principal investigator can make. The paper claims that applying these rules from the start of a project improves the quality, reproducibility, and trustworthiness of research software, and that better software in turn yields better research. The intended audience is PIs and group leaders who currently see RSEng as technically opaque and want a starting point.","feed_headline":"Ten simple rules help PIs build trustworthy research software","feed_subtitle":"A practical guide aims to make research software engineering accessible to group leaders, improving reproducibility and trust.","key_machinery":"The central object is the set of ten rules itself, framed as actionable guidance rather than a technical manual. The paper groups the work of RSEng into three activity categories—technical (code, architecture, testing), management (team, project, product), and knowledge transfer (documentation, training, community)—and each rule is a concrete practice a PI can mandate or resource. The devices that carry the argument are analogies between computational work and experimental science: version control and documentation are the computational scientist's lab notebook, architecture is the plan for a house, and untested software is uncalibrated lab equipment.","core_discovery":"The central claim is that Research Software Engineering is much more than writing code—it is a process built from technical, management, and knowledge-transfer activities—and that the barriers preventing PIs from engaging with it can be removed by following ten simple rules. The discovery, in effect, is an accessible framing of RSEng for group leaders: start the engineering practices at the beginning of a project, translate research needs into software requirements, think about architecture before and during coding, choose tools in context, keep everything under version control, develop in the open, document for real audiences, apply automated quality assurance with continuous integration, and treat software as a citable scholarly object. If the paper is right, a group that follows these rules produces software that is maintainable, reusable, and trustworthy, and the research built on that software inherits those properties.","pith_inferences":["If the rules work as claimed, they could serve as a lightweight audit rubric: funders and institutions could check a project's plan for version control, testing, and documentation before release, turning an advice document into an assessment checklist.","The causal claim invites a direct empirical test: compare matched research groups that follow the rules against those that do not on defect density, reproducibility success rates, or software citation counts after a fixed period.","The paper stops short of specifying how much effort each rule costs; an implicit testable extension is measuring time-to-first-release or maintenance burden under partial versus full rule adoption.","The 'lab notebook' framing implies that institutional research-integrity policies could recognise version control history and test logs as legitimate research records, a step that would give PIs a concrete incentive to adopt the practices."],"forward_implications":["Groups that adopt the rules will allocate time for requirements, architecture, version control, documentation, and testing from the first day of a project rather than as a cleanup phase.","Research software developed this way becomes easier to reuse, extend, and maintain, reducing waste and enabling collaborations that would otherwise stall.","Open-source development and published, citable software give the group scholarly credit and make results reproducible by third parties.","Automated quality assurance and code review reduce the chance of undetected errors, protecting the group from embarrassing retractions.","PIs can engage with RSEng at the level of practices and expectations, even when they do not personally write code."],"supporting_citations":[{"why":"Establishes the Four Pillars of Research Software Engineering, the basis for the claim that RSEng is a key success factor.","marker":"[1]"},{"why":"Supplies the 'better software, better research' premise that motivates the whole paper.","marker":"[2]"},{"why":"Prior ten simple rules on making research software robust, which this paper extends toward group management.","marker":"[4]"},{"why":"The reproducible-computing rules that ground the paper's emphasis on version control and documentation.","marker":"[6]"},{"why":"Systematic review of scientific-software testing used to justify the need for fine-grained quality assurance.","marker":"[18]"},{"why":"FAIR Principles for Research Software, which Rule 10 translates into publishing and citation guidance.","marker":"[20]"},{"why":"Software citation principles that define how users should credit research software.","marker":"[21]"},{"why":"Agile Development manifesto cited as a flexible methodology adaptable to research software teams.","marker":"[11]"}],"fun_headline_variants":["Ten rules for PIs to build better research software","A practical guide for group leaders on research software","Ten simple rules to make your lab's code trustworthy","Why PIs should care about research software engineering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a group that follows these ten rules will end up with higher-quality, more reproducible, more trustworthy software—and ultimately better research—but it presents no empirical evidence comparing adopters with non-adopters.","fun_headline_variants_meta":{"raw":{"variants":["Ten rules for PIs to build better research software","A practical guide for group leaders on research software","Ten simple rules to make your lab's code trustworthy","Why PIs should care about research software engineering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1340,"prompt_tokens":839,"completion_tokens":501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":440}},"tokens_in":455,"tokens_out":501,"duration_ms":5699,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:52:55.740539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison of matched research groups—half adopting the ten rules, half continuing usual practice—measuring defect density, successful reproduction of results, or software citation rates over 12–24 months would settle the claim; if adopters show no improvement, the causal link fails. Alternatively, a survey-based study showing that groups reporting high rule adherence have no better reproducibility scores than low-adherence groups would falsify it.","supporting_citations":[{"cited_title":"Ten Simple Rules for Making Research Software More Robust","cited_arxiv_id":null,"evidence_quote":"Prior ten simple rules on making research software robust, which this paper extends toward group management."},{"cited_title":"Manifesto for Agile Software Development; 2001","cited_arxiv_id":null,"evidence_quote":"Agile Development manifesto cited as a flexible methodology adaptable to research software teams."}],"review_version":1}