Pith. sign in

REVIEW 4 major objections 5 minor 103 references

Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Software engineering researchers praise ML best practices they rarely follow in their own papers.

desk verdict Careful multi-source study of ML practices in SE; the headline gap is plausible but overstated by the explicit-mention assumption. read the letter →

arxiv 2411.19304 v1 pith:GACSDNX6 submitted 2024-11-28 cs.SE cs.LG

classification cs.SEcs.LG
keywords machinelearningpracticessoftwareengineeringresearchML4SEbestgaphyperparametertuninggroundedtheoryqualitativeanalysisreviewguidelines
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This study asks how software engineering (SE) researchers actually use machine learning (ML), and finds a consistent split between what they say and what they publish. The authors analyze 110 SE research articles from top venues, plus 14 in-depth interviews with experienced ML4SE researchers and a survey of article authors. Their central finding: several practices widely regarded as best practice, including hyperparameter tuning, human-in-the-loop evaluation, exploratory data analysis, and manual label validation, appear in expert interviews at far higher rates than in the research articles themselves. The paper interprets this as a gap between what the SE research community believes is good ML practice and what it actually does in published work.

What carries the argument

The central mechanism is a comparative coding instrument built on the nine-stage ML pipeline (model requirements, data collection, data cleaning, data labeling, feature engineering, model training, model evaluation, model deployment, model monitoring). The authors code research articles, short survey responses, and interview transcripts with grounded-theory open and axial coding, using variables for input, technique, purpose, SE task, quality attributes, challenges, reviewer perspective, and educator perspective. This lets them hold what researchers say they do next to what they actually write in papers.

What would settle it

Inspect the replication packages and code repositories linked from the 110 analyzed articles and check whether hyperparameter tuning, human-in-the-loop evaluation, and exploratory data analysis occur there. If most packages show such steps, the article-versus-interview gap is a reporting artifact rather than a practice gap.

Watch

Extended reading notes

Core claim

The paper establishes that much of the ML practice that SE researchers describe as important is not visible in the research articles they publish. Hyperparameter tuning is the clear quantitative example: 85% of interviewees discussed it, but explicit traces of it were found in about 20% of the 110 articles. Similar gaps appear for involving human experts in evaluation, performing exploratory data analysis, manual label validation, and evaluating models in specific scenarios. The authors argue that some reported differences reflect abstraction level, with articles giving more detail about training and evaluation steps while interviews give broader process-level views, but they still conclude that several widely endorsed best practices are 'often mentioned, but less often followed.'

Load-bearing premise

The paper's central comparison treats each article's explicit text as a complete record of the practices the authors actually followed, so anything done but not written down is counted as absent.

Editorial extensions

If this is right

  • If the gap is real, SE research reporting standards should push for explicit documentation of hyperparameter tuning, human evaluation, and exploratory data analysis, since current articles often omit them.
  • Review guidelines for ML4SE should ask reviewers to check non-functional properties, qualitative analysis, and human involvement, areas the paper finds are not well covered by existing guidelines.
  • Model deployment and monitoring are nearly absent from SE research articles, suggesting that the research community is not producing much evidence about what happens after a model is trained.
  • Education inside SE relies heavily on hands-on experimental learning, but traditional formats like courses and text resources remain common, so the two approaches should be treated as complementary rather than competing.
  • The gap between stated best practices and observed practice gives concrete targets for future interventions: tune hyperparameters more systematically, involve human subjects in evaluations, and build practices for non-functional quality attributes into the standard ML workflow.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The measured gap may overstate the real behavioral gap: if authors routinely perform steps like hyperparameter tuning but do not report them, the comparison conflates under-reporting with not doing. The authors acknowledge this because their coding only considered explicit statements.
  • A natural extension would be to examine replication packages and code repositories accompanying the same articles to see whether tuning, EDA, and human validation are present in the artifacts even when absent from the text.
  • The interview-versus-article comparison is also confounded by abstraction level: interviews encourage process-level answers while articles encourage method-level detail, so some of the 'gap' may reflect how each source communicates rather than how researchers behave.
  • The paper's results suggest a testable hypothesis for the SE community: if review checklists explicitly ask for hyperparameter tuning and human evaluation, the reported prevalence of these practices will rise in subsequent years, independent of any real change in research practice.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a mixed-methods, pre-registered study of machine learning (ML) practices in software engineering (SE) research. The authors analyze 110 SE research articles from ASE, FSE, and ICSE published between 2011 and 2022, a survey of 769 first and last authors of those articles (58 responses, 47 usable), and 14 semi-structured interviews with purposively sampled ML4SE researchers. Using grounded-theory open and axial coding, they investigate four research questions concerning ML practices, challenges, review criteria, and education. The headline finding is that certain practices regarded as best practices, such as hyperparameter tuning, human-in-the-loop evaluation, exploratory data analysis, and manual label validation, are far less prevalent in the articles than in the interviews; for example, hyperparameter tuning appears in about 85% of interviews but only about 20% of the articles. The paper also reports that data-related challenges dominate, that existing review guidelines miss SE-specific aspects such as non-functional quality attributes, and that educators combine hands-on learning with traditional teaching methods.

Significance. The study addresses a genuine gap in the ML4SE literature by examining researchers, reviewers, and educators rather than industry practitioners. Its strengths include a pre-registered protocol, explicit inter-rater reliability reporting, dual coding for articles and interviews, a replication kit, and triangulation across three data sources. If the central gap claim is supportable, the paper provides actionable evidence for improving SE research practice, review checklists, and ML education. However, the significance is conditional: the headline gap between interviews and articles rests on treating the absence of an explicit textual mention as evidence that a practice was not performed. The authors themselves acknowledge in Section 6.2.2 that low article frequencies 'could also mean that they are implicitly assumed and not explicitly reported,' which undermines the strong interpretation in Section 7.1 that the gap is between 'what is said to be done in research and how research is actually done.' With a reframing toward reported practices, or an additional validation step such as a replication-package audit, the contribution would be solid and useful.

major comments (4)
  1. [Section 7.1, Section 6.2.2] The central claim that best practices are 'often mentioned, but less often followed' is not sufficiently supported because the coding of articles 'only considered explicit statements' (Section 7.1). The paper itself concedes in Section 6.2.2 that low article frequencies 'could also mean that they are implicitly assumed and not explicitly reported.' Since SE methods sections often omit routine steps such as hyperparameter tuning, exploratory data analysis, manual label validation, and human-in-the-loop checks, the measured gap between interviews and articles may largely reflect reporting style rather than actual research practice. I recommend either weakening the conclusion to a 'reported practice gap' or validating a sample of the 110 articles against their replication packages and scripts to test whether the practices are truly absent.
  2. [Section 5.1.1, Section 6.2.2] The interview-article comparison is not matched on population, time period, or elicitation method. The 14 interviewees are influential ML4SE researchers selected via purposive sampling and academic contacts, while the 110 articles are ICSE/FSE/ASE papers from 2011 to 2022, most likely authored by a different set of researchers. The interviews ask participants about their general practices across research, reviewing, and teaching, whereas articles are written to report specific studies. Differences between the two data sources could therefore stem from population composition, temporal trends, or genre conventions, not from a gap between what researchers say and what they do. The paper should explicitly acknowledge this limitation or restructure the comparison, for example by interviewing a sample of authors of the coded articles about their specific papers.
  3. [Section 5.3.2, Section 6.2.7] The survey that informs RQ1.2 (best practices declared by authors) has a response rate of only 58 out of 769 contacted authors (7.5%), with 47 usable responses after filtering. This low and self-selected response rate severely limits the generalizability of the survey findings. The paper acknowledges the survey's limited insights in Section 6.2.7, yet it still uses survey percentages in Figures 2-12 and describes the survey as 'supporting their representativeness' of the interview findings. I recommend presenting the survey results as purely exploratory and refraining from using them to corroborate the central gap claim.
  4. [Section 6.2.4] The comparison of ML pipeline stages between articles and interviews is structurally biased. The authors state that 'the interviews were designed to mention each stage described by Amershi et al. explicitly,' which means interviewees were prompted to discuss stages such as Model Monitoring and Model Deployment. Articles, in contrast, were coded from unguided text. The paper then reports that Model Monitoring is 'not considered at all' in articles and Model Deployment appears in only 2.7% of articles, treating this as evidence of a practice gap. This comparison conflates free reporting with prompted elicitation and should either be removed or substantially qualified.
minor comments (5)
  1. [Section 6.2.4] The percentages in this section use inconsistent decimal notation, e.g., '2,7%' and '13,6%' alongside '95%' and '20%'. Please unify the decimal separator throughout the manuscript.
  2. [Table 3] The 'Avg. Length/duration' column mixes units: '48.6 min.' for articles, '11.95 Pp.' for surveys, and '259.06 chars' for interviews. Clarify what each value represents and consider using separate columns for each subject type.
  3. [Section 5.5] The authors report a deviation from the pre-registered protocol for interview coding, replacing the planned Krippendorff's alpha threshold with full double-checking by a second coder. No inter-rater agreement measure is reported for the interview coding. Please report the agreement obtained during the initial 20% double-coding or another reliability metric.
  4. [Throughout] The manuscript contains several typographical errors, including 'contect' (Section 6.1), 'the the' and 'for for' (Section 6.2.1), 'spiting' (Section 6.2.2), and 'an be sure' (Section 7.2.2). A careful proofreading pass is needed.
  5. [Title page and Abstract] The ACM reference format on the title page still contains placeholder data ('Conference acronym XX', year 2018) and the abstract includes an awkward comma in 'to contribute to the knowledge, about the synergy.' These should be corrected before resubmission.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the study's gap finding is an empirical comparison of two independently collected data sources, not a derivation that reduces to its inputs.

full rationale

The manuscript reports a qualitative empirical study, not a mathematical derivation. The central claim — that some practices (e.g., hyperparameter tuning) are common in expert interviews but rare in SE articles — rests on independently collected data: 110 articles coded with an externally published ML pipeline stage framework (Amershi et al. [12]), 47 survey responses, and 14 interviews. The coding rule that 'only considered explicit statements' is a construct-validity limitation that the authors explicitly acknowledge in Section 7.1 and Section 6.2.2; it means the article-level frequencies measure reported practices, which is precisely the construct the RQ1.1 variables target. That a practice may be performed but omitted from a paper's text is a possible threat to the interpretation, not a circular step: the measured percentages are not defined in terms of the conclusion, no fitted parameter is renamed as a prediction, and no load-bearing premise is justified solely by a self-citation. Self-citations (e.g., Mojica-Hanke et al. [69]) appear only as related work and are not used to establish the results. The comparison with external benchmarks such as Serban et al. and Amershi et al. further supports the independence of the analysis. Therefore the paper's analysis is self-contained and not circular.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are introduced because the study is qualitative and descriptive. The listed axioms are domain assumptions about the coding frame, the interpretation of explicit text, and sample representativeness. The paper explicitly discusses each of these as threats to validity.

assumptions (3)
  • domain assumption The ML pipeline stages defined by Amershi et al. [12] provide a valid taxonomy for coding ML practices in SE research and interviews.
    Used throughout the coding scheme (Section 5.2-5.4) to classify practices into stages; if this taxonomy is incomplete or biased, the frequencies of stages are affected.
  • domain assumption Researchers' explicit statements in articles and interviews are a valid indicator of the practices they used.
    The gap analysis in Section 7.1 treats absence of a practice in an article as non-use; the paper itself notes that implicit use is possible.
  • domain assumption The 14 interviewees and 47 survey respondents are sufficiently representative of the SE research community for the study's conclusions.
    Purposive sampling of influential researchers and low survey response rate (Section 5.1.1, Section 5.3.2); acknowledged as external validity threat.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education." pith.science (2026). https://pith.science/paper/GACSDNX6

@misc{pith2026241119304,
  author       = {Pith},
  title        = {Pith review of: Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GACSDNX6}},
  note         = {Machine review of arXiv:2411.19304}
}
read the original abstract

Context: Machine Learning (ML) significantly impacts Software Engineering (SE), but studies mainly focus on practitioners, neglecting researchers. This overlooks practices and challenges in teaching, researching, or reviewing ML applications in SE. Objective: This study aims to contribute to the knowledge, about the synergy between ML and SE from the perspective of SE researchers, by providing insights into the practices followed when researching, teaching, and reviewing SE studies that apply ML. Method: We analyzed SE researchers familiar with ML or who authored SE articles using ML, along with the articles themselves. We examined practices, SE tasks addressed with ML, challenges faced, and reviewers' and educators' perspectives using grounded theory coding and qualitative analysis. Results: We found diverse practices focusing on data collection, model training, and evaluation. Some recommended practices (e.g., hyperparameter tuning) appeared in less than 20\% of literature. Common challenges involve data handling, model evaluation (incl. non-functional properties), and involving human expertise in evaluation. Hands-on activities are common in education, though traditional methods persist. Conclusion: Despite accepted practices in applying ML to SE, significant gaps remain. By enhancing guidelines, adopting diverse teaching methods, and emphasizing underrepresented practices, the SE community can bridge these gaps and advance the field.

Figures

Figures reproduced from arXiv: 2411.19304 by the authors.

Figure 1
Figure 1. Research Protocol steps and their relation with the Variables (Var) and the Research Questions ( [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p038_2.png] view at source ↗
Figure 3
Figure 3. Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p039_3.png] view at source ↗
Figures from the paper (16 more)
Figure 4
Figure 4. Figure 4: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p040_4.png]
Figure 5
Figure 5. Figure 5: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p041_5.png]
Figure 6
Figure 6. Figure 6: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p042_6.png]
Figure 7
Figure 7. Figure 7: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p043_7.png]
Figure 8
Figure 8. Figure 8: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p044_8.png]
Figure 9
Figure 9. Figure 9: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p045_9.png]
Figure 10
Figure 10. Figure 10: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p046_10.png]
Figure 11
Figure 11. Figure 11: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p047_11.png]
Figure 12
Figure 12. Figure 12: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p048_12.png]
Figure 13
Figure 13. Figure 13: Percentage of subjects (i.e., papers, interviewees, and surveys) in which each [PITH_FULL_IMAGE:figures/full_fig_p049_13.png]
Figure 14
Figure 14. Figure 14: Percentage of interviews in which each SE task category appears. 0% 5% 10% 15% 20% 25% 30% 35% 40% testing implementation managment testing ML software representation requirements and specifications implementation ML maintenance team/developer aspects Other 40.0% (44)…
Figure 15
Figure 15. Figure 15: Percentage of papers in which each SE task category appears. Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p050_15.png]
Figure 16
Figure 16. Figure 16: Percentage of interviews in which each quality attribute category appears. C.5 Codes associated to the ML Challenges 0% 10% 20% 30% 40% 50% 60% data related challenges ground truth and evaluation resources needed/used conduct human studies model interpretability and t…
Figure 17
Figure 17. Figure 17: Percentage of interviews in which each challenge category appears. Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p051_17.png]
Figure 18
Figure 18. Figure 18: Percentage of interviews in which each reviewers’ perspective category appears. Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p052_18.png]
Figure 19
Figure 19. Figure 19: Percentage of interviews in which each educators’ perspective category appears. Received XXX; revised XXX; accepted XXX Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p053_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

103 extracted references · 66 canonical work pages

  1. [1]

    The ICSE 23 Workshop on Cloud Intelligence /AIOps

    2022. The ICSE 23 Workshop on Cloud Intelligence /AIOps. https://cloudintelligenceworkshop.org/CFP.html

  2. [2]

    Empirical Software Engineering Journal

    2023. Empirical Software Engineering Journal. https://www.springer.com/journal/10664/updates/19597712

  3. [3]

    Instituto de Ciências Matemáticas e de Computação, Bacharelado em Ciências de Computação

    2023. Instituto de Ciências Matemáticas e de Computação, Bacharelado em Ciências de Computação. https://uspdigital.usp.br/jupiterweb/ listarGradeCurricular?codcg=55&codcur=55041&codhab=0&tipo=N

  4. [4]

    Stack overflow

    2023. Stack overflow. https://stackoverflow.com/ Manuscript submitted to ACM Perspective of SE Researchers on ML Practices 31

  5. [5]

    ESEC/FSE 2023. 2022. ESEC/FSE 2023 - research papers - ESEC/FSE 2023. https://2023.esec-fse.org/track/fse-2023-research-papers

  6. [6]

    A-test2023. 2022. The workshop will be held in Kirchberg, Luxembourg. https://a-test.org/

  7. [7]

    Yaser S Abu-Mostafa, Malik Magdon-Ismail, and Hsuan-Tien Lin. [n. d.]. Learning from data - A short course. https://amlbook.com/

  8. [8]

    Viviana Acquaviva. 2022. Teaching Machine Learning for the Physical Sciences: A summary of lessons learned and challenges. In Proceedings of the Second Teaching Machine Learning and Artificial Intelligence Workshop . PMLR, 35–39

Show all 103 references
  1. [9]

    Ruben Acuña. 2023. Developing a Data Science Course to Support Software Engineering Students. In 2023 IEEE 35th International Conference on Software Engineering Education and Training (CSEE&T) . IEEE, 127–131

  2. [10]

    Alshangiti, H

    M. Alshangiti, H. Sapkota, P. K. Murukannaiah, X. Liu, and Q. Yu. 2019. Why is Developing Machine Learning Applications Challenging? A Study on Stack Overflow Posts. In 2019 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) . 1–11. https...

  3. [11]

    Amazon. 2023. codeguru. https://aws.amazon.com/codeguru/

  4. [12]

    Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software engineering for machine learning: A case study. In 2019 IEEE/ACM 41st International Conference on Software Engineerin...

  5. [13]

    Applitools. 2023. Applitools (V.1). https://applitools.com/

  6. [14]

    Daniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke, Fabio Pierazzi, Christian Wressnegger, Lorenzo Cavallaro, and Konrad Rieck

  7. [15]

    Anders Arpteg, Björn Brinne, Luka Crnkovic-Friis, and Jan Bosch. 2018. Software Engineering Challenges of Deep Learning. In 2018 44th Euromicro Conference on Software Engineering and Advanced Applications (SEAA) . 50–59. https://doi.org/10.1109/SEAA.2018.00018

  8. [16]

    Arham Arshad, Taher Ghaleb, and Paul Ralph. 2021. Towards a More Structured Peer Review Process with Empirical Standards. In Evaluation and Assessment in Software Engineering . 353–358

  9. [17]

    Abdul Ali Bangash, Hareem Sahar, Shaiful Chowdhury, Alexander William Wong, Abram Hindle, and Karim Ali. 2019. What do developers know about machine learning: a study of ml discussions on stackoverflow. In2019 IEEE/ACM 16th International Conference on Mining Software Repositor...

  10. [18]

    Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, and Stefano Russo. 2020. Learning-to-rank vs ranking-to-learn: Strategies for regression testing in continuous integration. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engin...

  11. [19]

    Stella Biderman and Walter J Scheirer. 2020. Pitfalls in machine learning research: Reexamining the development cycle. (2020)

  12. [20]

    Charles C Bonwell and James A Eison. 1991. Active learning: Creating excitement in the classroom. 1991 ASHE-ERIC higher education reports. ERIC

  13. [21]

    Pierre Bourque. 2014. https://ieeecs-media.computer.org/media/education/swebok/swebok-v3.pdf

  14. [22]

    Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D Sculley. 2017. The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (Big Data) . IEEE, 1123–1132

  15. [23]

    Rodney A Brooks. 1991. Intelligence without representation. Artificial intelligence 47, 1-3 (1991), 139–159

  16. [24]

    CAIN2023. 2022. Cain 2023 - Papers - Cain 2023. https://conf.researchr.org/track/cain-2023/cain-2023-call-for-papers

  17. [25]

    naturalizing

    Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu, and Baishakhi Ray. 2022. NatGen: Generative Pre-Training by “naturalizing” Source Code. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of S...

  18. [26]

    Preetha Chatterjee, Tushar Sharma, and Paul Ralph. 2022. Empirical standards for repository mining. In Proceedings of the 19th International Conference on Mining Software Repositories . 142–143

  19. [27]

    Elin Clemmedsson. 2018. Identifying Pitfalls in Machine Learning Implementation Projects - A Case Study of Four Technology-Intensive Organizations . MA Thesis. Royal Institue of Technology (KTH), SE-100 44 STOCKHOLM

  20. [28]

    Thomas D Cook, Donald Thomas Campbell, and Arles Day. 1979. Quasi-experimentation: Design & analysis issues for field settings . Vol. 351. Houghton Mifflin Boston

  21. [29]

    Juliet Corbin et al. 1990. Basics of qualitative research grounded theory procedures and techniques. (1990)

  22. [30]

    Marian Daun and Jennifer Brings. 2023. How ChatGPT will change software engineering education. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1 . 110–116

  23. [31]

    Davide Dell’Anna, Fatma Başak Aydemir, and Fabiano Dalpiaz. 2023. Evaluating classifiers in SE research: the ECSER pipeline and two replication studies. Empirical Software Engineering 28, 1 (2023), 3

  24. [32]

    Oxford English Dictionary. 2023. best practice, n. https://oed.com/dictionary/best-practce_n

  25. [33]

    Oxford English Dictionary. 2023. practice, n. https://oed.com/view/Entry/149226?rskey=GK6Iqx&result=1&isAdvanced=false

  26. [34]

    Oxford English Dictionary. 2023. tecnique. https://oed.com/view/Entry/198458?redirectedFrom=technique

  27. [35]

    Dynatrace LLC. 2023. dynatrace - Davis (V.1). https://www.dynatrace.com/platform/aiops/

  28. [36]

    Glen H Elder, Eliza K Pavalko, and Elizabeth Colerick Clipp. 1993. Working with archival data: Studying lives . Vol. 88. Sage

  29. [37]

    Hanya Elhashemy, Harold Abelson, and Tilman Michaeli. 2024. Adapting Computational Skills for AI Integration. In 2024 36th International Conference on Software Engineering Education and Training (CSEE&T) . IEEE, 1–5. Manuscript submitted to ACM 32 Mojica-Hanke et al

  30. [38]

    Association for Computing Machinery. 2023. ACM Transactions on Software Engineering and Methodology Reviewers: ACM Digital Library. https://dl.acm.org/journal/tosem/reviewers#iagfr

  31. [39]

    github. 2020. GitHub. https://github.com/

  32. [40]

    Github and OpenAI. 2023. Github copilot (V.1.100.0.0). https://github.com/features/copilot/

  33. [41]

    Alaleh Hamidi, Giuliano Antoniol, Foutse Khomh, Massimiliano Di Penta, and Mohammad Hamidi. 2021. Towards understanding developers’ machine-learning challenges: A multi-language study on stack overflow. In 2021 IEEE 21st International Working Conference on Source Code Analysis...

  34. [42]

    Junxiao Han, Emad Shihab, Zhiyuan Wan, Shuiguang Deng, and Xin Xia. 2020. What do Programmers Discuss about Deep Learning Frameworks. Empirical Software Engineering 25 (2020). https://doi.org/10.1007/s10664-020-09819-6

  35. [43]

    Hashemi, M

    Y. Hashemi, M. Nayebi, and G. Antoniol. 2020. Documentation of Machine Learning Software. In 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . 666–667. https://doi.org/10.1109/SANER48275.2020.9054844

  36. [44]

    Petra Heck and Gerard Schouten. 2021. Lessons learned from educating AI engineers. In 2021 IEEE/ACM 1st Workshop on AI Engineering-Software Engineering for AI (W AIN). IEEE, 1–4

  37. [45]

    Jan Herrington. 2005. Authentic learning environments in higher education . IGI Global

  38. [46]

    Angela Horneman, Andrew Mellinger, and Ipek Ozkaya. 2020. AI Engineering: 11 Foundational Practices . Technical Report. Carnegie Mellon University Pittsburgh United States

  39. [47]

    Lei Huang and Kuo-Sheng Ma. 2018. Introducing Machine Learning to First-year Undergraduate Engineering Students Through an Authentic and Active Learning Labware. In 2018 IEEE Frontiers in Education Conference (FIE) . 1–4. https://doi.org/10.1109/FIE.2018.8659308

  40. [48]

    ICSE2024. 2023. https://conf.researchr.org/track/icse-2024/icse-2024-research-track

  41. [49]

    IEEE202. 202AD. CSDL: IEEE Computer Society. https://www.computer.org/csdl/journal/ts

  42. [50]

    CORE INC. 2023. Core Rankings Portal. https://www.core.edu.au/conference-portal

  43. [51]

    M. J. Islam, H. Nguyen, Rangeet Pan, and H. Rajan. 2019. What Do Developers Ask About ML Libraries? A Large-scale Study Using Stack Overflow. ArXiv abs/1906.11940 (2019)

  44. [52]

    Devon Journal of Machine Learning Research. 2023. Guidelines for JMLR reviewers. https://www.jmlr.org/reviewer-guide.html

  45. [53]

    Christian Kaestner and Eunsuk Kang. 2020. Teaching software engineering for AI-enabled systems. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering: Software Engineering Education and Training . 45–48

  46. [54]

    Misha Kakkar, Sarika Jain, Abhay Bansal, and PS Grover. 2021. Combining data preprocessing methods with imputation techniques for software defect prediction. In Research Anthology on Recent Trends, Tools, and Implications of Computer Programming . IGI Global, 1792–1811

  47. [55]

    Sayash Kapoor, Emily M Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A Bail, Odd Erik Gundersen, Jake M Hofman, Jessica Hullman, Michael A Lones, Momin M Malik, et al . 2024. REFORMS: Consensus-based Recommendations for Machine-learning-based Science. Science Advances 10,...

  48. [56]

    Seohyun Kim, Jinman Zhao, Yuchi Tian, and Satish Chandra. 2021. Code prediction by feeding trees to transformers. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 150–162

  49. [57]

    Zoe Kotti, Rafaila Galanopoulou, and Diomidis Spinellis. 2023. Machine learning for software engineering: A tertiary study. Comput. Surveys 55, 12 (2023), 1–39

  50. [58]

    Dominik Kreuzberger, Niklas Kühl, and Sebastian Hirschl. 2023. Machine learning operations (mlops): Overview, definition, and architecture. In IEEE access, Vol. 11. IEEE, 31866–31879

  51. [59]

    Klaus Krippendorff. 2004. Content analysis: An introduction to its methodology . Sage publications

  52. [60]

    Filippo Lanubile, Silverio Martínez-Fernández, and Luigi Quaranta. 2023. Teaching MLOps in higher education through project-based learning. In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering Education and Training (ICSE-SEET) . IEEE, 95–100

  53. [61]

    Filippo Lanubile, Silverio Martínez-Fernández, and Luigi Quaranta. 2023. Training future ML engineers: a project-based course on MLOps. IEEE software (2023)

  54. [62]

    Nicole Leach Sankofa. 2022. Transformativist measurement development methodology: A mixed methods approach to scale construction. Journal of Mixed Methods Research 16, 3 (2022), 307–327

  55. [63]

    Lewis, Stephany Bellomo, and Ipek Ozkaya

    Grace A. Lewis, Stephany Bellomo, and Ipek Ozkaya. 2021. Characterizing and Detecting Mismatch in Machine-Learning-Enabled Systems. In 2021 IEEE/ACM 1st Workshop on AI Engineering - Software Engineering for AI (W AIN) . 133–140. https://doi.org/10.1109/WAIN52551.2021.00028

  56. [64]

    Michael A Lones. 2024. Avoiding common machine learning pitfalls. Patterns 5, 10 (2024)

  57. [65]

    Lívia S Marques, Christiane Gresse von Wangenheim, and Jean CR Hauck. 2020. Teaching machine learning in school: A systematic mapping of the state of the art. Informatics in Education (2020), 283–321

  58. [66]

    Silverio Martínez-Fernández, Justus Bogner, Xavier Franch, Marc Oriol, Julien Siebert, Adam Trendowicz, Anna Maria Vollmer, and Stefan Wagner

  59. [67]

    Atif Mashkoor, Wesley Klewerton Guez Assuncao, and Alexander Egyed. 2023. Teaching Engineering of AI-Intensive Systems. IEEE Software (2023)

  60. [68]

    ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 2 (2022), 1–59

    Software engineering for AI-based systems: a survey. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 2 (2022), 1–59

  61. [69]

    González

    Anamaria Mojica-Hanke, Andrea Bayona, Mario Linares-Vásquez, Steffen Herbold, and Fabio A. González. 2023. What are the Machine Learning best practices reported by practitioners on Stack Exchange? arXiv:2301.10516 [cs.SE]

  62. [70]

    Now I’m a bit{angry:}

    Peter Mayer, Yixin Zou, Florian Schaub, and Adam J Aviv. 2021. " Now I’m a bit{angry:}" Individuals’ Awareness, Perception, and Responses to Data Breaches that Affected Them. In 30th USENIX Security Symposium (USENIX Security 21) . 393–410. Manuscript submitted to ACM Perspect...

  63. [71]

    University of Colorado Boulder and Coursera. 2023. Academics - master of science in computer science boulder. https://www.coursera.org/ degrees/ms-computer-science-boulder/academics

  64. [72]

    University of Cape Town. 2023. Computer science courses: University of Cape Town. https://sit.uct.ac.za/our-degrees-undergraduates-bsc- degrees/computer-science-courses#CSC1010H

  65. [73]

    Graduate School of Information Science and The University of Tokyo Technology. 2019. Information on course registration, school credits and other procedures: Education: Graduate School of Information Science and Technology, the University of Tokyo. https://www.i.u-tokyo.ac.jp/...

  66. [74]

    The Institute of Electrical and Electronics Engineers. 1990. IEEE Standard Glossary of Software Engineering Terminology: IEEE Std 610.12-1990 . Institute of Electrical and Electronics Engineers

  67. [75]

    The University of Melbourne. 2023. Master of software engineering - The University of Melbourne. https://study.unimelb.edu.au/find/courses/ graduate/master-of-software-engineering/what-will-i-study/#sample-plans

  68. [76]

    University of London and Coursera. 2023. BSC Computer Science Curriculum and programme length. https://www.coursera.org/degrees/bachelor- of-science-computer-science-london/academics

  69. [77]

    Massachussetts Institute of Technology. 2022. Computer Science and Engineering (course 6-3). http://catalog.mit.edu/degree-charts/computer- science-engineering-course-6-3/

  70. [78]

    University of Oxford. [n. d.]. Computer science. https://www.ox.ac.uk/admissions/undergraduate/courses/course-listing/computer-science

  71. [79]

    Michael Quinn Patton. 2014. Qualitative research & evaluation methods: Integrating theory and practice . Sage publications

  72. [80]

    Fortieth International Conference on Machine Learning. 2023. ICML 2023 reviewer tutorial. https://icml.cc/Conferences/2023/ReviewerTutorial

  73. [81]

    Lori Pollock and Massimiliano Di Penta. 2023. ICSE 2023 Review Process and Guidelines-Short

  74. [82]

    Rolf Pfeifer and Josh Bongard. 2006. How the body shapes the way we think: a new view of intelligence . MIT press

  75. [83]

    Paul Ralph, Nauman bin Ali, Sebastian Baltes, Domenico Bianculli, Jessica Diaz, Yvonne Dittrich, Neil Ernst, Michael Felderer, Robert Feldt, Antonio Filieri, Breno Bernard Nicolau de França, Carlo Alberto Furia, Greg Gay, Nicolas Gold, Daniel Graziotin, Pinjia He, Rashina Hoda...

  76. [84]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning . PMLR, 28492–28518

  77. [85]

    Chinmay Sahu, Blaine Ayotte, and Mahesh K. Banavar. 2021. Integrating machine learning concepts into undergraduate classes. In 2021 IEEE Frontiers in Education Conference (FIE) . 1–5. https://doi.org/10.1109/FIE49875.2021.9637283

  78. [86]

    Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical Software Engineering 14, 2 (2009), 131–164

  79. [87]

    Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2021. Practices for engineering trustworthy machine learning applications. In 2021 IEEE/ACM 1st Workshop on AI engineering-software engineering for AI (W AIN). IEEE, 97–100

  80. [88]

    Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2020. Adoption and effects of software engineering best practices in machine learning. In Proceedings of the 14th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) . 1–12

  81. [89]

    A Strauss and J Corbin. 1998. Basics of qualitative research: grounded theory procedures and techniques., 2nd edn.(Sage: Thousand Oaks, CA). (1998)

  82. [90]

    Omar Shouman, Simon Fuchs, and Holger Wittges. 2022. Experiences from teaching practical machine learning courses to master’s students with mixed backgrounds. In Proceedings of the Second Teaching Machine Learning and Artificial Intelligence Workshop . PMLR, 62–67

  83. [91]

    Chakkrit Tantithamthavorn and Ahmed E Hassan. 2018. An experience report on defect modelling in practice: Pitfalls and challenges. InProceedings of the 40th international conference on software engineering: Software engineering in practice . 286–295

  84. [92]

    Tabnine. 2023. tabnine (V. 3.6.08). https://www.tabnine.com/

  85. [93]

    Lyu, Lei Ma, Jonathan I

    Sebastian Uchitel, Marsha Chechik, Massimiliano Di Penta, Bram Adams, Nazareno Aguirre, Gabriele Bavota, Domenico Bianculli, Kelly Blincoe, Ana Cavalcanti, Yvonne Dittrich, Filomena Ferrucci, Rashina Hoda, LiGuo Huang, David Lo, Michael R. Lyu, Lei Ma, Jonathan I. Maletic, Leo...

  86. [94]

    Matti Tedre, Tapani Toivonen, Juho Kahila, Henriikka Vartiainen, Teemu Valtonen, Ilkka Jormanainen, and Arnold Pears. 2021. Teaching machine learning in K–12 classroom: Pedagogical and technological trajectories for artificial intelligence education. IEEE access 9 (2021), 1105...

  87. [95]

    Bram Van Der Vlist, Rick Van De Westelaken, Christoph Bartneck, Jun Hu, Rene Ahn, Emilia Barakova, Frank Delbressine, and Loe Feijs. 2008. Teaching machine learning to design students. InTechnologies for E-Learning and Digital Entertainment: Third International Conference, Edu...

  88. [96]

    Arizona State University and Coursera. 2023. Curriculum and course information: ASU MCS. https://www.coursera.org/degrees/master-of- computer-science-asu/academics Manuscript submitted to ACM 34 Mojica-Hanke et al

  89. [97]

    Cody Watson, Nathan Cooper, David Nader Palacio, Kevin Moran, and Denys Poshyvanyk. 2022. A Systematic Literature Review on the Use of Deep Learning in Software Engineering Research. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 2 (2022), 1–58

  90. [98]

    Michael Vierhauser, Iris Groher, Tobias Antensteiner, and Clemens Sauerwein. 2024. Towards Integrating Emerging AI Applications in SE Education. arXiv preprint arXiv:2405.18062 (2024)

  91. [99]

    Mounika Yabaku and Sofia Ouhbi. 2024. University Students’ Perception and Expectations of Generative AI Tools for Software Engineering. In 2024 36th International Conference on Software Engineering Education and Training (CSEE&T) . IEEE, 1–5

  92. [100]

    Ohlsson, Björn Regnell, and Anders Wesslen

    Claes Wohlin, Per Runeson, Martin Höst, Magnus C. Ohlsson, Björn Regnell, and Anders Wesslen. 2012. Experimentation in Software Engineering . Springer Publishing Company, Incorporated

  93. [101]

    Martin Zinkevich. 2021. Rules of machine learning: Best Practices for ML Engineering. https://developers.google.com/machine-learning/guides/ rules-of-ml A Supplementary material for the Execution Plan with the SE Authors In this section of the Appendix, we provide the relevant...

  94. [102]

    S Ziglar. 2022. Teaching machine learning in k-12 computing education: Potential and pitfalls. arXiv preprint arXiv:2106.11034 (2022)

  95. [2022]

    Dos and Don’ts of Machine Learning in Computer Security. In Proc. of USENIX Security Symposium

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.