{"id":"974b79a5-7f57-4ec7-bbb2-9847e7682a3e","arxiv_id":"2501.10037","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 57 scientists finds most lack formal training in readable code, struggle with poor documentation and naming, and are increasingly using AI tools to improve code quality.","lead":"This paper surveys 57 research scientists about their coding habits and finds that most lack formal training in writing readable code and rely on comments and documentation. It matters because scientific software underlies much of modern research, and these findings point to where training and tools could improve reproducibility and collaboration.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'trend toward LLMs' claim in the abstract and contributions is a temporal inference from a cross-sectional survey that only measures current tool use; this overclaim should be removed or reframed.","rationale":"The paper is a honest, well-structured exploratory survey with an available artifact, and most descriptive claims are supported by the data. The reader's identified weakest assumption, the convenience/snowball sample, is real but explicitly acknowledged and does not undermine the modest 'preliminary insights' framing. However, the abstract and contributions advertise 'a trend towards the adoption of AI-based code generation tools' and the reader's strongest claim includes 'increasingly turning to large language models.' This temporal claim cannot be supported by a single cross-sectional survey that asks only about current tool use. The relevant result in RQ2 is a percentage of responses among current users of automated code quality tools; it says nothing about change over time. The Discussion itself uses 'popular' rather than 'trend,' suggesting the stronger wording in the abstract and contributions is an overreach. Since one of the paper's three advertised contributions depends on this unsupported temporal inference, the appropriate action is conditional acceptance with a required revision of the trend language, not rejection. If the authors reframe the finding as current reliance on LLMs rather than a trend, the paper's remaining claims stand as preliminary descriptive evidence.","tokens_in":9019,"tokens_out":5085,"duration_ms":54640,"concrete_test":"Download the archived survey instrument from Zenodo (doi:10.5281/zenodo.14676939) and inspect every item for temporal or change-based wording (e.g., 'have you started using,' 'compared to a year ago,' 'increase in'). If no such item exists, the 'trend toward LLMs' conclusion is not derivable from the instrument; the abstract and contributions should be revised to say 'current reported use' rather than 'trend' or 'increasingly.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing vulnerability is not sample size but the paper's temporal 'trend' claim about LLM adoption. Section III RQ2 reports that among the 50.88% of respondents who use automated code quality tools, 41.38% of responses mention LLMs (ChatGPT, Claude). This is a cross-sectional snapshot of current tool preference. A cross-sectional survey cannot establish that scientists are 'increasingly turning to' or showing 'a trend towards' LLMs; no question asks about past behavior or change over time. The abstract and the contribution list both state the trend language, while the Discussion softens to 'popular.' Because this is one of the paper's three advertised contributions, the overclaim matters more than the acknowledged convenience-sample limitation.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This exploratory survey paper reports on 57 research scientists at the University of Hawai'i at Manoa, recruited through convenience and snowball sampling, to understand their programming backgrounds, code comprehension practices, challenges, and tool usage. The main descriptive findings are that most respondents are self-taught or learned on the job, 57.9% report no formal training in writing readable code, Python and R dominate, comments and documentation are the most common readability practices, inadequate documentation and poor identifier names are the top challenges, traditional code quality tool adoption is low, and LLM-based tools (notably ChatGPT) are the most frequently mentioned tool category among those who use automated tools. The paper includes a Zenodo artifact with the dataset and thematic coding, and it explicitly acknowledges threats to validity from its single-institution convenience sample.","tokens_in":9152,"tokens_out":4420,"duration_ms":44022,"significance":"If taken as a preliminary descriptive study, this paper addresses a genuine gap: most program-comprehension research targets traditional software, not scientific programming. Its value lies in surfacing concrete candidate phenomena—self-taught programmers, reliance on comments, documentation deficits, cryptic names, and low traditional tool adoption—that can motivate follow-up work with larger, more representative samples. The authors provide the survey data and thematic coding in an artifact, which supports reproducibility and independent checking. The main advertised 'trend' claim about LLM adoption is not supported by the cross-sectional design and must be reframed, but the remaining descriptive findings are internally consistent and useful as an early-stage foundation.","major_comments":[{"comment":"The abstract and the contribution list state, respectively, that the findings show 'a trend towards utilizing large language models' and 'a trend towards the adoption of AI-based code generation tools.' Section IV-A similarly refers to 'the shift toward LLMs.' However, the survey is cross-sectional: Section III RQ2 only reports current tool use, namely that among the 50.88% of respondents who use automated code quality tools, 41.38% of responses mention LLMs such as ChatGPT or Claude. No survey question asks about past behavior, changes in tool use, or adoption over time, so the data cannot support a temporal trend or shift. This is a load-bearing claim because it appears as one of the paper's three advertised contributions. The 'trend' and 'shift' language should be removed or explicitly reframed as 'current use' or 'current popularity of LLM-based tools among users of automated tools'; the Discussion's softer word 'popular' is appropriate, but the abstract and contributions must be made consistent with what the data actually show.","section":"Abstract and Section I-B; Section III RQ2; Section IV-A"}],"minor_comments":[{"comment":"The reported percentages 54.55%, 38.18%, 5.45%, and 1.82% sum to 100% but correspond to a denominator of 55 respondents, not the 57 valid participants stated in Section II-B. Please report the exact counts (30/55, 21/55, 3/55, 1/55) and explain the two missing responses, so readers can verify the percentages.","section":"Section III, RQ2 (Likert frequency question)"},{"comment":"Several passages mix participant counts with percentage bases. For example, '11 participants (12.50%) working in Economics' uses a denominator of total domain responses (88), not the 57 participants, and similar issues appear in the programming-language and environment statistics. Please consistently state whether a percentage is of participants or of total selections, and add this clarification to the table captions for Tables I and II.","section":"Section III, RQ1 and RQ2 (percentage reporting)"},{"comment":"The percentage columns in both tables are proportions of total challenge selections, not of participants. Since the text does not state this explicitly, readers may misread the percentages as prevalence among the 57 respondents. Add a note such as 'Percentages are computed over the total number of challenge selections across all participants.'","section":"Tables I and II"},{"comment":"The survey instrument is said to be omitted for space and is available in the Zenodo artifact. Given that the artifact is a central part of the reproducibility story, consider including the full questionnaire as an appendix if the venue permits, or at least quoting the exact wording of the Likert and multiple-choice questions in the paper.","section":"Section II-A"},{"comment":"The phrase 'a fair number never use automated code quality tools' is vague; the preceding sentence gives 49.12%, so please use the percentage directly to avoid ambiguity.","section":"Section III, RQ2 (tool-use wording)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a sound preliminary descriptive study, and the artifact is a strong point in its favor. The main blocker is the unsupported temporal 'trend'/'shift' language about LLM adoption in the abstract, contribution list, and discussion; this must be corrected before publication. I recommend that the editor require the authors to either remove the temporal claim or rephrase it as current usage, and to add the missing denominators/clarifications described in the minor comments. With those changes, the paper would be acceptable as a preliminary empirical report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a genuinely useful exploratory survey with an honest limitations section, but the 'trend toward LLMs' in the abstract and contributions is not supported by the cross-sectional data.\n\nWhat's new: this is the first survey I know that asks research scientists directly about code comprehension practices, not just scientific software development broadly (Hannay et al. 2009, Prabhu et al. 2011). The documentation paradox—scientists rely on comments and docs while simultaneously reporting missing comments and docs as their top comprehension barriers—is a nice framing that gives tool-builders and educators a concrete wedge. The Zenodo artifact with the raw dataset and thematic coding is a real plus; it makes the descriptive numbers checkable. The authors are also appropriately cautious about generalizability in the threats section.\n\nWhere it's soft: the 'trend toward utilizing large language models' is the weakest claim. The data show that among the 51% who never use code quality tools, the remaining half who do use them mention LLMs in 41.38% of responses. That's a cross-sectional snapshot of current tool choices; no question asked about past behavior or change. The Discussion softens to 'popular,' but the abstract and contribution list still say 'trend' and 'shift toward.' That language should be deleted or reframed as 'use of LLMs is common among adopters.'\n\nThe sample is small and single-institution, and the authors acknowledge it. There is also a minor numeric hiccup: the reproducibility Likert percentages (45.16%, 38.60%) don't sum to 100 for 57 respondents; the counts do, so the prose numbers are off. Fixable. I don't see anything else manufactured; the descriptive statistics are consistent with the tables, and the citation pattern looks appropriate.\n\nWho this is for: anyone working on software engineering education for scientists, reproducibility infrastructure, or tooling for Jupyter/Python/R workflows. It won't move the needle for theory, but it provides a useful baseline and a few concrete research leads.\n\nRecommendation: worth sending to peer review. It's a well-scoped exploratory study with transparent method and a checkable artifact. The temporal overclaim needs to be cut or rephrased, and the numbers cleaned up, but that's minor relative to the value of the dataset and the documentation paradox. I'd accept it with those revisions.","headline":"Useful exploratory survey with honest limitations; the LLM 'trend' claim overreaches the cross-sectional data but the artifact and documentation paradox make it worth refereeing.","tokens_in":9631,"tokens_out":2560,"would_cite":true,"duration_ms":24056,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that research scientists mostly learn to program on their own, lack formal training in readable code, and rely on comments and AI assistants while struggling with documentation and identifier names.","keywords":["scientific software","code comprehension","code readability","survey study","identifier naming","documentation","research scientists","large language models"],"falsifier":"A probability-sampled survey of several hundred scientists across multiple institutions that found most had formal instruction in maintainable code, or that documentation was not among the top comprehension barriers, would undercut the paper's generalizations.","tokens_in":8858,"feed_emoji":"🧪","tokens_out":4294,"duration_ms":38411,"temperature":0.7,"pith_summary":"Surveying 57 research scientists, this paper tries to establish that scientific programmers are largely self-taught, with 57.9% reporting no formal instruction in writing readable code, and that their biggest comprehension problems are inadequate comments and documentation, poor identifier names, and weak project organization. It also claims that adoption of traditional code-quality tools is low, while large language models such as ChatGPT are becoming the de facto tool for improving code readability. A sympathetic reader would care because scientific software underpins reproducibility and collaboration, and these findings point to concrete gaps in training and tooling for scientists who code.","feed_headline":"Most scientists never trained in readable code","feed_subtitle":"Survey of 57 researchers shows documentation and naming are their top barriers, while AI assistants are filling the tool gap.","key_machinery":"The instrument is a 20-question survey administered through the Qualtrics platform, combining single-choice, multiple-choice, Likert-scale, and free-text items. Recruitment used convenience and snowball sampling from one university's departments, yielding 57 fully completed responses; three authors independently coded the free-text answers and resolved disagreements through consensus. That survey plus its thematic coding is the entire load-bearing mechanism, since every percentage and ranking in the results comes directly from it.","core_discovery":"On the paper's own terms, the central discovery is a profile: research scientists who write code rely on self-study and on-the-job learning, not formal software-engineering education; 33 of 57 respondents (57.9%) never received education or training in writing readable, maintainable code. When they try to understand others' code, the top reported obstacles are lack of comments (44 reports), missing project documentation (33), poor method/variable names (31), poor project structure (31), and unexplained hardcoded values (24). Nearly half (49.12%) say they never use automated code-quality tools, and among those who do, 41.38% use AI/LLM tools, mainly ChatGPT and Claude. All respondents rate code readability as at least slightly important for reproducibility, yet 54.55% say they only sometimes understand others' scientific code, and cryptic or short identifier names are the most common naming problem (40 reports).","pith_inferences":["If the self-selection bias runs in the opposite direction, with scientists who care about readability overrepresented among volunteers, the true level of formal training among all research scientists could be even lower than the reported 57.9%.","The heavy reliance on LLMs among scientists who do use quality tools suggests a natural experiment: comparing code quality and reproducibility of LLM-assisted versus traditional scientific code could quantify the risk of unchecked AI-generated code.","The prominence of cryptic and generic names in scientific code may reflect domain-specific shorthand, meaning generic linting tools may not catch names that are meaningful only within a lab and context-aware naming tools may be needed.","The documentation paradox could be an artifact of self-report: scientists may believe they document more than they actually do, so observational studies of real repositories would test whether the perceived gap is real."],"forward_implications":["Universities and research institutions should extend programming education for scientists beyond syntax to cover maintainability, naming, and documentation practices.","Code comprehension research needs to shift its focus from Java-centric studies to Python and R, the languages scientists actually use.","New tools for scientific code should target documentation generation and identifier-quality checks, the two most-reported pain points.","Because many scientists adopt LLMs for code quality without formal software-engineering training, guidance for critically evaluating AI-generated code is needed.","The gap between documentation's perceived importance and its inadequate delivery, the paper's 'documentation paradox,' warrants direct study of actual documentation behavior."],"supporting_citations":[{"why":"Supplies baseline evidence that scientists spend roughly a third of their time developing software and learn programming through practice.","marker":"[3]"},{"why":"Documents that scientific programs use the same languages as other software, grounding the language findings for Python and R.","marker":"[4]"},{"why":"Describes notebook-specific comprehension challenges such as non-linear execution and hidden state, supporting the paper's framing of scientific code difficulty.","marker":"[15]"},{"why":"Supports the claim that scientists prioritize domain knowledge over established software engineering practices.","marker":"[16]"},{"why":"Identifies the Qualtrics platform used to administer the 20-question survey.","marker":"[24]"},{"why":"Provides the survey-design guidelines the authors followed, including the pilot run with three scientists.","marker":"[25]"},{"why":"Justifies the convenience and snowball sampling approach and its methodological limits.","marker":"[27]"},{"why":"Underpins the premise that descriptive identifier names improve source code comprehension, motivating RQ3.","marker":"[29]"}],"fun_headline_variants":["Most scientists never trained to write readable code","57 scientists, 58% never taught readable code","Poor docs and names block scientific code understanding","AI fills code-quality gap for untrained scientists","Over half of scientist coders lack readability training"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 57 volunteers from one university, recruited by convenience and snowball sampling, represent research scientists broadly enough that the reported percentages reflect the wider community's practices.","fun_headline_variants_meta":{"raw":{"variants":["Most scientists never trained to write readable code","57 scientists, 58% never taught readable code","Poor docs and names block scientific code understanding","AI fills code-quality gap for untrained scientists","Over half of scientist coders lack readability training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000414,"raw_usage":{"total_tokens":2118,"prompt_tokens":905,"completion_tokens":1213,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":1143}},"tokens_in":521,"tokens_out":1213,"duration_ms":9701,"temperature":1.0,"reasoning_tokens":1143,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:23:30.295734+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A probability-sampled survey of several hundred scientists across multiple institutions that found most had formal instruction in maintainable code, or that documentation was not among the top comprehension barriers, would undercut the paper's generalizations.","supporting_citations":[{"cited_title":"How do scientists develop and use scientiﬁc soft ware?,","cited_arxiv_id":null,"evidence_quote":"Supplies baseline evidence that scientists spend roughly a third of their time developing software and learn programming through practice."},{"cited_title":"A survey of the practice of computational s cience,","cited_arxiv_id":null,"evidence_quote":"Documents that scientific programs use the same languages as other software, grounding the language findings for Python and R."},{"cited_title":"Exploration and Ex planation in Computational Notebooks,","cited_arxiv_id":null,"evidence_quote":"Describes notebook-specific comprehension challenges such as non-linear execution and hidden state, supporting the paper's framing of scientific code difficulty."},{"cited_title":"Scientists and software engineers: A tale of two cultures,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that scientists prioritize domain knowledge over established software engineering practices."},{"cited_title":"Qualtrics XM: The Leading Experience Management Soft ware","cited_arxiv_id":null,"evidence_quote":"Identifies the Qualtrics platform used to administer the 20-question survey."},{"cited_title":"Guidelines for conducting surveys in software engineering,","cited_arxiv_id":null,"evidence_quote":"Provides the survey-design guidelines the authors followed, including the pilot run with three scientists."},{"cited_title":"Sampling in software engineeri ng research: a critical review and guidelines,","cited_arxiv_id":null,"evidence_quote":"Justifies the convenience and snowball sampling approach and its methodological limits."},{"cited_title":"Descriptive compound identiﬁer names improve so urce code comprehension,","cited_arxiv_id":null,"evidence_quote":"Underpins the premise that descriptive identifier names improve source code comprehension, motivating RQ3."}],"review_version":1}