{"id":"86a6f1e2-ea04-44c4-a242-1c56c7a4999e","arxiv_id":"2411.19304","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SE researchers recommend best practices like hyperparameter tuning and human-in-the-loop evaluation, but these appear in under 20% of the 110 analyzed SE/ML papers.","lead":"This paper asks software engineering researchers how they use machine learning in research, teaching, and peer review, and compares their answers with what appears in published papers. It finds that practices researchers call best, like hyperparameter tuning and human review, show up in few papers, suggesting a gap between advice and practice.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gap estimate depends on treating unstated practices as absent; a replication-package audit can settle whether the 85%-vs-20% hyperparameter-tuning gap is real or a reporting artifact.","rationale":"The reader's weakest_assumption identifies the same load-bearing premise: the paper's gap conclusion is only as strong as the assumption that article text is a complete record of practices used. This is the condition most needed for the central claim, and it is the least secure. The authors themselves flag it in Section 6.2.2 ('it could also mean that they are implicitly assumed and not explicitly reported') and in the construct validity discussion, so the concern is not manufactured. The survey of article authors, which found hyperparameter tuning mentioned in 20-30% of responses, provides some supporting evidence that the low article prevalence is not purely an artifact, but the survey was intentionally brief and had a low response rate, so it does not settle the issue. The proposed replication-package audit is a concrete, low-cost way to test whether unstated practices were in fact performed. If the audit shows substantial unreported tuning, the gap is partly a reporting artifact and the paper's stronger wording in Section 7.1 ('gap between what is said to be done in research and how research is actually done') would need to be weakened to 'gap between acknowledged best practices and practices reported in articles.' Because the paper is exploratory, transparent about limitations, and provides a replication kit, the reader's CONDITIONAL verdict is appropriate; no change is needed, but the condition should be resolved by the audit before the quantitative gap is cited as established.","tokens_in":40961,"tokens_out":9944,"duration_ms":94714,"concrete_test":"Use the replication kits and online appendices for the 110 coded articles. For every article coded as not mentioning hyperparameter tuning, inspect its replication package for evidence of tuning: hyperparameter configuration files, grid/random search scripts, trial logs, or result tables across hyperparameter settings. Compute the fraction of 'absent' articles with such artifacts, and repeat the same check for human-evaluation protocols and exploratory-data-analysis scripts. If that fraction is substantial (e.g., >25%), the Section 6.2.2 gap estimates are inflated and the central claim requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central gap claim (e.g., Section 6.2.2: hyperparameter tuning in ~85% of interviews vs ~20% of articles) depends on treating the absence of an explicit textual mention in an article as evidence that the practice was not performed. The coding 'only considered explicit statements' (Section 7.1), and the authors concede in Section 6.2.2 that low article frequencies 'could also mean that [practices] are implicitly assumed and not explicitly reported.' SE methods sections are selective: hyperparameter tuning, exploratory data analysis, manual label validation, and human-in-the-loop checks are exactly the kind of routine steps that are often done informally and omitted from a paper's limited text. Thus the measured 'gap' may overstate the true gap between what researchers do and what they report. This is the load-bearing premise of the paper's headline conclusion; if a large share of the 110 articles performed these practices without saying so, the quantitative comparison collapses toward a reporting-style difference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods, pre-registered study of machine learning (ML) practices in software engineering (SE) research. The authors analyze 110 SE research articles from ASE, FSE, and ICSE published between 2011 and 2022, a survey of 769 first and last authors of those articles (58 responses, 47 usable), and 14 semi-structured interviews with purposively sampled ML4SE researchers. Using grounded-theory open and axial coding, they investigate four research questions concerning ML practices, challenges, review criteria, and education. The headline finding is that certain practices regarded as best practices, such as hyperparameter tuning, human-in-the-loop evaluation, exploratory data analysis, and manual label validation, are far less prevalent in the articles than in the interviews; for example, hyperparameter tuning appears in about 85% of interviews but only about 20% of the articles. The paper also reports that data-related challenges dominate, that existing review guidelines miss SE-specific aspects such as non-functional quality attributes, and that educators combine hands-on learning with traditional teaching methods.","tokens_in":41138,"tokens_out":5761,"duration_ms":51687,"significance":"The study addresses a genuine gap in the ML4SE literature by examining researchers, reviewers, and educators rather than industry practitioners. Its strengths include a pre-registered protocol, explicit inter-rater reliability reporting, dual coding for articles and interviews, a replication kit, and triangulation across three data sources. If the central gap claim is supportable, the paper provides actionable evidence for improving SE research practice, review checklists, and ML education. However, the significance is conditional: the headline gap between interviews and articles rests on treating the absence of an explicit textual mention as evidence that a practice was not performed. The authors themselves acknowledge in Section 6.2.2 that low article frequencies 'could also mean that they are implicitly assumed and not explicitly reported,' which undermines the strong interpretation in Section 7.1 that the gap is between 'what is said to be done in research and how research is actually done.' With a reframing toward reported practices, or an additional validation step such as a replication-package audit, the contribution would be solid and useful.","major_comments":[{"comment":"The central claim that best practices are 'often mentioned, but less often followed' is not sufficiently supported because the coding of articles 'only considered explicit statements' (Section 7.1). The paper itself concedes in Section 6.2.2 that low article frequencies 'could also mean that they are implicitly assumed and not explicitly reported.' Since SE methods sections often omit routine steps such as hyperparameter tuning, exploratory data analysis, manual label validation, and human-in-the-loop checks, the measured gap between interviews and articles may largely reflect reporting style rather than actual research practice. I recommend either weakening the conclusion to a 'reported practice gap' or validating a sample of the 110 articles against their replication packages and scripts to test whether the practices are truly absent.","section":"Section 7.1, Section 6.2.2"},{"comment":"The interview-article comparison is not matched on population, time period, or elicitation method. The 14 interviewees are influential ML4SE researchers selected via purposive sampling and academic contacts, while the 110 articles are ICSE/FSE/ASE papers from 2011 to 2022, most likely authored by a different set of researchers. The interviews ask participants about their general practices across research, reviewing, and teaching, whereas articles are written to report specific studies. Differences between the two data sources could therefore stem from population composition, temporal trends, or genre conventions, not from a gap between what researchers say and what they do. The paper should explicitly acknowledge this limitation or restructure the comparison, for example by interviewing a sample of authors of the coded articles about their specific papers.","section":"Section 5.1.1, Section 6.2.2"},{"comment":"The survey that informs RQ1.2 (best practices declared by authors) has a response rate of only 58 out of 769 contacted authors (7.5%), with 47 usable responses after filtering. This low and self-selected response rate severely limits the generalizability of the survey findings. The paper acknowledges the survey's limited insights in Section 6.2.7, yet it still uses survey percentages in Figures 2-12 and describes the survey as 'supporting their representativeness' of the interview findings. I recommend presenting the survey results as purely exploratory and refraining from using them to corroborate the central gap claim.","section":"Section 5.3.2, Section 6.2.7"},{"comment":"The comparison of ML pipeline stages between articles and interviews is structurally biased. The authors state that 'the interviews were designed to mention each stage described by Amershi et al. explicitly,' which means interviewees were prompted to discuss stages such as Model Monitoring and Model Deployment. Articles, in contrast, were coded from unguided text. The paper then reports that Model Monitoring is 'not considered at all' in articles and Model Deployment appears in only 2.7% of articles, treating this as evidence of a practice gap. This comparison conflates free reporting with prompted elicitation and should either be removed or substantially qualified.","section":"Section 6.2.4"}],"minor_comments":[{"comment":"The percentages in this section use inconsistent decimal notation, e.g., '2,7%' and '13,6%' alongside '95%' and '20%'. Please unify the decimal separator throughout the manuscript.","section":"Section 6.2.4"},{"comment":"The 'Avg. Length/duration' column mixes units: '48.6 min.' for articles, '11.95 Pp.' for surveys, and '259.06 chars' for interviews. Clarify what each value represents and consider using separate columns for each subject type.","section":"Table 3"},{"comment":"The authors report a deviation from the pre-registered protocol for interview coding, replacing the planned Krippendorff's alpha threshold with full double-checking by a second coder. No inter-rater agreement measure is reported for the interview coding. Please report the agreement obtained during the initial 20% double-coding or another reliability metric.","section":"Section 5.5"},{"comment":"The manuscript contains several typographical errors, including 'contect' (Section 6.1), 'the the' and 'for for' (Section 6.2.1), 'spiting' (Section 6.2.2), and 'an be sure' (Section 7.2.2). A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The ACM reference format on the title page still contains placeholder data ('Conference acronym XX', year 2018) and the abstract includes an awkward comma in 'to contribute to the knowledge, about the synergy.' These should be corrected before resubmission.","section":"Title page and Abstract"}],"recommendation":"major_revision","confidential_remarks":"This is a well-intentioned and methodologically detailed qualitative study with a pre-registered protocol and a replication kit, but its headline claim is currently overinterpreted. The gap between interviews and articles is vulnerable to the reporting-artifact concern, and the unmatched comparison populations make the 'saying vs. doing' framing difficult to defend. I do not see a reason for rejection: the underlying data are valuable, and the issue can be addressed by reframing the claim as a reported-practice gap or by adding a replication-package audit for a sample of the 110 articles. The authors should also be encouraged to make their analysis scripts and coding dictionaries available in the replication kit, which would strengthen reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a serious, transparent empirical study, but its headline gap number is softer than it looks because the paper only counts explicit mentions in articles. I'd send it to review, but the authors need to temper the 'actually done' language.\n\nWhat's actually new: the design—coding 110 SE articles, plus a survey of 58 authors (47 usable) and 14 expert interviews—is a genuinely new combination for ML4SE. It gives a multi-source picture of practices, challenges, review concerns, and teaching methods in one community. The protocol is pre-registered, the coding is double-coded with IRR reported, the replication kit is available, and the limitations section is unusually honest. The comparison with existing best-practice literature is useful; the point that hyperparameter tuning, human-in-the-loop evaluation, and EDA are discussed far more than they appear in papers is a real signal.\n\nThe soft spots are real but not fatal. The 85%-vs-20% hyperparameter tuning gap assumes that an article's text is a complete record of what the authors did. The authors themselves concede (Section 6.2.2) that low article frequencies could mean the practices are implicit and not reported. Since methods sections are selective and replication packages weren't coded, the measured gap overstates the practice gap; at minimum it's a reporting gap. The paper's Discussion wording—'less often followed' and 'how research is actually done'—goes beyond what the data can show. That's a load-bearing caveat, and it should be front and center in the conclusions. Also, the response rate (7.5%) and interview sample (14) are small, though the paper treats them as exploratory and doesn't overclaim. I'd also like uncertainty intervals around the key percentages; right now it's point estimates without error bars.\n\nThe survey part is weak by the authors' own admission—short free-text answers—and they handle it correctly by not leaning on it.\n\nWho this is for: anyone working on ML4SE guidelines, reviewers wanting empirical grounding on review criteria, and educators designing ML-for-SE courses. It's a useful reference even with the caveat.\n\nRecommendation: accept for peer review. A good referee should push for (1) revised wording that distinguishes reported practices from actual practices, and (2) some measure of uncertainty or sensitivity analysis on the gap. If those are addressed, this becomes a solid contribution.","headline":"Careful multi-source study of ML practices in SE; the headline gap is plausible but overstated by the explicit-mention assumption.","tokens_in":41657,"tokens_out":2105,"would_cite":true,"duration_ms":18907,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Software engineering researchers praise ML best practices they rarely follow in their own papers.","keywords":["machine learning practices","software engineering research","ML4SE","best practices gap","hyperparameter tuning","grounded theory","qualitative analysis","review guidelines"],"falsifier":"Inspect the replication packages and code repositories linked from the 110 analyzed articles and check whether hyperparameter tuning, human-in-the-loop evaluation, and exploratory data analysis occur there. If most packages show such steps, the article-versus-interview gap is a reporting artifact rather than a practice gap.","tokens_in":40789,"feed_emoji":"📉","tokens_out":2803,"duration_ms":26593,"temperature":0.7,"pith_summary":"This study asks how software engineering (SE) researchers actually use machine learning (ML), and finds a consistent split between what they say and what they publish. The authors analyze 110 SE research articles from top venues, plus 14 in-depth interviews with experienced ML4SE researchers and a survey of article authors. Their central finding: several practices widely regarded as best practice, including hyperparameter tuning, human-in-the-loop evaluation, exploratory data analysis, and manual label validation, appear in expert interviews at far higher rates than in the research articles themselves. The paper interprets this as a gap between what the SE research community believes is good ML practice and what it actually does in published work.","feed_headline":"Researchers talk up ML best practices more than they use them","feed_subtitle":"In SE papers, hyperparameter tuning appears in ~20% of articles but 85% of expert interviews.","key_machinery":"The central mechanism is a comparative coding instrument built on the nine-stage ML pipeline (model requirements, data collection, data cleaning, data labeling, feature engineering, model training, model evaluation, model deployment, model monitoring). The authors code research articles, short survey responses, and interview transcripts with grounded-theory open and axial coding, using variables for input, technique, purpose, SE task, quality attributes, challenges, reviewer perspective, and educator perspective. This lets them hold what researchers say they do next to what they actually write in papers.","core_discovery":"The paper establishes that much of the ML practice that SE researchers describe as important is not visible in the research articles they publish. Hyperparameter tuning is the clear quantitative example: 85% of interviewees discussed it, but explicit traces of it were found in about 20% of the 110 articles. Similar gaps appear for involving human experts in evaluation, performing exploratory data analysis, manual label validation, and evaluating models in specific scenarios. The authors argue that some reported differences reflect abstraction level, with articles giving more detail about training and evaluation steps while interviews give broader process-level views, but they still conclude that several widely endorsed best practices are 'often mentioned, but less often followed.'","pith_inferences":["The measured gap may overstate the real behavioral gap: if authors routinely perform steps like hyperparameter tuning but do not report them, the comparison conflates under-reporting with not doing. The authors acknowledge this because their coding only considered explicit statements.","A natural extension would be to examine replication packages and code repositories accompanying the same articles to see whether tuning, EDA, and human validation are present in the artifacts even when absent from the text.","The interview-versus-article comparison is also confounded by abstraction level: interviews encourage process-level answers while articles encourage method-level detail, so some of the 'gap' may reflect how each source communicates rather than how researchers behave.","The paper's results suggest a testable hypothesis for the SE community: if review checklists explicitly ask for hyperparameter tuning and human evaluation, the reported prevalence of these practices will rise in subsequent years, independent of any real change in research practice."],"forward_implications":["If the gap is real, SE research reporting standards should push for explicit documentation of hyperparameter tuning, human evaluation, and exploratory data analysis, since current articles often omit them.","Review guidelines for ML4SE should ask reviewers to check non-functional properties, qualitative analysis, and human involvement, areas the paper finds are not well covered by existing guidelines.","Model deployment and monitoring are nearly absent from SE research articles, suggesting that the research community is not producing much evidence about what happens after a model is trained.","Education inside SE relies heavily on hands-on experimental learning, but traditional formats like courses and text resources remain common, so the two approaches should be treated as complementary rather than competing.","The gap between stated best practices and observed practice gives concrete targets for future interventions: tune hyperparameters more systematically, involve human subjects in evaluations, and build practices for non-functional quality attributes into the standard ML workflow."],"supporting_citations":[{"why":"Provides the nine-stage ML pipeline that structures the interview script and serves as the coding frame for practices and stages.","marker":"[12]"},{"why":"Supplies a survey-based baseline of SE best practices for ML and their adoption in developer teams, which this study extends to researchers.","marker":"[86]"},{"why":"Systematic literature review of deep learning in SE that supplies practice categories and data-type classifications compared with the article coding here.","marker":"[97]"},{"why":"Consensus recommendations for machine-learning-based science that the paper compares against reviewers' stated concerns and finds partly incomplete.","marker":"[55]"},{"why":"A widely cited list of ML best practices that the paper uses as a reference point for practices mentioned in interviews and articles.","marker":"[101]"}],"fun_headline_variants":["Hyperparameter tuning: 85% of researchers recommend, only 20% of papers show it","SE researchers often recommend ML practices they rarely use in their own studies","ML best practices talked up more than followed in SE research","Researchers' ML advice in SE doesn't match their published work","Touted ML practices in SE papers: often mentioned, less often followed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central comparison treats each article's explicit text as a complete record of the practices the authors actually followed, so anything done but not written down is counted as absent.","fun_headline_variants_meta":{"raw":{"variants":["Hyperparameter tuning: 85% of researchers recommend, only 20% of papers show it","SE researchers often recommend ML practices they rarely use in their own studies","ML best practices talked up more than followed in SE research","Researchers' ML advice in SE doesn't match their published work","Touted ML practices in SE papers: often mentioned, less often followed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1625,"prompt_tokens":907,"completion_tokens":718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":623}},"tokens_in":523,"tokens_out":718,"duration_ms":6710,"temperature":1.0,"reasoning_tokens":623,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:18:41.429755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the replication packages and code repositories linked from the 110 analyzed articles and check whether hyperparameter tuning, human-in-the-loop evaluation, and exploratory data analysis occur there. If most packages show such steps, the article-versus-interview gap is a reporting artifact rather than a practice gap.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic literature review of deep learning in SE that supplies practice categories and data-type classifications compared with the article coding here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Consensus recommendations for machine-learning-based science that the paper compares against reviewers' stated concerns and finds partly incomplete."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A widely cited list of ML best practices that the paper uses as a reference point for practices mentioned in interviews and articles."}],"review_version":1}