REVIEW 4 major objections 5 minor 43 references
An Empirical Study on Decision-Making Aspects in Responsible Software Engineering for AI
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This empirical study claims that the binding constraint on responsible AI software is not a lack of ethical principles but their operationalization in day-to-day engineering decisions.
desk verdict A decent exploratory study on responsible AI decision-making with honest limitations, but the headline gap claim is an interpretive synthesis rather than a directly measured result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study's central object is the decision-making process in responsible AI engineering, examined through mixed methods: semi-structured interviews analyzed by thematic and narrative analysis, a Likert-scale and open-ended survey, and static validation of findings with industry experts. The mechanism that carries the argument is the comparison between responsible AI development and conventional software development across organizational, team, and individual levels, with data as a new decision-driving element and probabilistic outcomes requiring continuous monitoring. Named constructs include H-shaped (also called π-shaped) competency, defined as professional depth in two distinct domains, one technical and one ethical, and the "principles-to-practice gap," the observed disconnect between stated ethical guidelines and operational decisions.
What would settle it
A concrete check is to audit a random sample of AI project artifacts—requirements documents, design records, test plans, and issue trackers—for documented ethical trade-off decisions, such as a fairness-accuracy choice or a privacy-utility trade-off. If such decisions are routinely recorded in specific, scenario-level terms across many organizations, the paper's central claim that guidelines are not operationalized would be contradicted; if such records are rare or vague, the claim is supported.
Extended reading notes
Core claim
The central discovery is a gap between the state of the art and the state of practice in responsible AI engineering: ethical issues in AI development largely mirror those in conventional software development, but they are more pronounced and are not operationalized. Practitioners' decisions are shaped by personal values, emerging specialized roles such as data engineers, machine learning engineers, and AI ethicists, and by organizational culture, but current ethical guidelines are too abstract to guide concrete choices about data, transparency, accountability, and post-deployment monitoring. The paper concludes that H-shaped competencies (deep skill in one technical domain plus deep skill in ethics), interdisciplinary collaboration, and an organizational culture of ethics are critical enablers, with transparency and accountability as the values practitioners weight most heavily.
Load-bearing premise
The study assumes that seven interviews, 51 self-selected survey responses, and four expert validation interviews, recruited through convenience and snowball sampling, fairly represent AI engineering practice, and that participants' descriptions of their decisions match what they actually do.
Editorial extensions
If this is right
- If the gap claim is right, adding scenario-driven ethical checklists and review boards to the software life cycle would do more than issuing additional principle statements.
- If H-shaped competency matters, hiring and training should reward dual technical-ethical depth rather than treating ethics as a separate staff function.
- If organizational culture is decisive, smaller companies may embed ethics more easily, while larger companies need explicit frameworks to keep ethics from becoming bureaucratized.
- If data is the key decision determinant in AI, the requirements and design phases deserve ethics review earlier than in conventional software development.
- If post-deployment monitoring depends on resources, responsible-AI guidance must include cost and capacity conditions, not just obligations.
Reading between the lines
- A direct extension is a testable check: count whether project artifacts such as requirements documents, design records, and test plans record explicit ethical trade-off decisions; the paper's evidence is self-reported, so artifact-level data would independently verify the gap.
- The H-shaped competency recommendation implies an educational pipeline change: software engineering curricula would need ethics modules deep enough to count as a second domain, not merely an elective.
- The study's emphasis on transparency and accountability suggests that future regulation may target decision documentation duties rather than model behavior alone.
- Because the sample skews toward professionals already engaged with AI, the gap may be even larger in organizations without self-selected interest in responsible AI; a random industry sample could test that.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an exploratory, interpretivist mixed-methods study of decision-making in responsible software engineering for AI (RSE for AI). The empirical basis is seven semi-structured interviews with practitioners, a survey of 51 respondents (Likert and open-ended questions), and four expert interviews used for static validation of findings. The authors use thematic and narrative analysis to derive themes under three research questions concerning the differences between responsible AI development and conventional software development, the influence of emerging roles, and the role of individual motivations and values. On this basis they claim a gap between the state of the art and industrial practice in RSE for AI, particularly a failure to operationalize ethical decision-making in the software engineering lifecycle, and they recommend interdisciplinary collaboration, H-shaped ethical-technical competences, and a stronger organizational ethics culture.
Significance. The paper addresses a timely and relevant topic and provides a useful descriptive account of how practitioners perceive ethical decision-making in AI development. Its main strengths are the mixed-method design, the explicit interpretivist stance, the use of both thematic and narrative analysis, and the public supplementary material. The reported quotes and tables give the reader a concrete view of participant responses. However, the contribution is primarily exploratory and hypothesis-generating; the sample is small and self-selected, the Likert items are not validated and are analyzed only descriptively, and the static validation is a member check rather than independent evidence. If the claims are reframed as participants' perceived experiences and recommendations, the paper can be a useful contribution to the empirical SE-for-AI literature; in its current form the abstract and conclusions state the central gap claim more strongly than the data support.
major comments (4)
- [Section IV-C and abstract/conclusion] The headline claim that 'current ethical guidelines are insufficiently implemented at the operational level' is not directly operationalized by the survey instrument. Q3 (organizational preparedness) and Q7 (effectiveness of methodologies) ask about perceived readiness and effectiveness, not about whether respondents actually apply ethical guidelines in concrete engineering tasks; the mildly positive Q3 and neutral Q7 distributions therefore do not measure the presence or absence of operationalization. The narrative themes in Sections IV-B2 and IV-B3 provide evidence that participants call for more structured support and describe implementation challenges, but these are perceived gaps, not a measured baseline of practice against a defined state of the art. I recommend rewording the abstract and Section VII to say that practitioners in this sample perceive a lack of operational frameworks and describe ethics as difficult to implement, rather than asserting as an empirical finding that guidelines are insufficiently implemented.
- [Section III-B and Section V] The research questions use causal language ('influence', 'impact') and the discussion repeatedly states that personal values 'can critically influence' decisions and that roles 'impact' decision-making, but the study design is cross-sectional self-report (interviews and a one-time survey). Such a design cannot establish causal influence or impact; it can only document participants' beliefs, attributions, and self-reported experiences. The empirical statements should be cast in terms of perceived influence or reported attribution, and the causal claims should be reserved for future longitudinal or intervention studies.
- [Section IV-D and Section VI] The static validation with four experts recruited from the authors' professional network is presented as evidence that the findings are 'confirmed' and applicable in industry, but it is a member check over the same findings, not independent corroboration. The paper itself concedes in Section VI that the validation interviews 'have the potential to introduce bias due to the possibility of leading questions'; given that the findings were presented to the experts before discussing them, this is a significant limitation. The conclusions should either present the validation as an expert credibility check with its own verbatim evidence, or drop the suggestion that four experts independently confirmed the findings.
- [Section III-C and Table I] The claim that data saturation was attained by the seventh interview is asserted without supporting detail (e.g., how saturation was assessed, which codes/themes became redundant), and the seven participants are heterogeneous in roles, company sizes, and experience; with convenience sampling from the authors' network, this cannot support generalizations about industry-wide practice. The 51 survey responses are also self-selected and no response rate is reported. While Section VI acknowledges limited generalizability, this limitation is in tension with the broad wording of the key finding in the abstract, so the conclusion should consistently restrict its claims to the sample and context.
minor comments (5)
- [Section I, first paragraph] The phrase 'in software engineering (software engineering) practices' appears to contain a duplicated parenthetical; the intended meaning should be clarified.
- [Table II] The column header 'Years in current role' appears twice and the current-company-size values are mixed with role tenure; the table should be cleaned so that each column has a unique header and consistent units.
- [Figures 2-4] The diverging stacked bar charts are described in the text but the axis labels and exact item wording are only in captions; for a journal version, ensure the charts carry numeric labels or a table of frequencies so that the distributions are independently inspectable.
- [Throughout] The terms 'responsible software engineering for AI', 'responsible AI engineering', and 'RAI' are used interchangeably; define the acronym once and use it consistently throughout the manuscript.
- [Section IV-C, Figures 2 and 4] Q2 under RQ1 and Q3 under RQ3 both ask about conflicting personal values, but they receive different wordings; clarify whether these are intended to be the same construct repeated for different units of analysis or two distinct items.
Circularity Check
No significant circularity: the paper's conclusions are interpretive syntheses of primary interview and survey data, not derived quantities, and its self-citations are contextual rather than load-bearing.
full rationale
This paper makes no formal derivation or prediction; it reports an interpretivist mixed-method empirical study. The central finding of a principles-to-practice gap is an inductive thematic and narrative synthesis from seven interviews, 51 survey responses, and four expert validation interviews. No quantity is fitted to data and then renamed as a prediction, and no equation equates an output with an input by construction. The survey does not directly measure operationalization, and the paper's self-acknowledged limitations in Section VI (social desirability bias, possible leading questions, small convenience sample, reliance on personal interpretation) are validity or generalizability concerns, not circularity. Self-citations such as the Behavioral Software Engineering framing in [9] and [22] and the supplementary material in [30] provide context and data access, but the empirical conclusions stand on primary qualitative and survey data rather than on those citations. The static validation with four industry experts is an external credibility check, not a derivation from the study's inputs. Therefore no circular step is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Seven interviews reached data saturation and therefore suffice to characterize the phenomenon.
- domain assumption Participants' self-reported accounts correspond to real decision-making processes in their organizations.
- domain assumption Interpretivist thematic and narrative analyses produce themes that are reliable across researchers and contexts.
- domain assumption Convenience and snowball samples from the authors' network and online forums can support generalizable practice recommendations.
Cite this review
Pith. "Pith review of An Empirical Study on Decision-Making Aspects in Responsible Software Engineering for AI." pith.science (2026). https://pith.science/paper/3WTHLHYI
@misc{pith2026250115691,
author = {Pith},
title = {Pith review of: An Empirical Study on Decision-Making Aspects in Responsible Software Engineering for AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/3WTHLHYI}},
note = {Machine review of arXiv:2501.15691}
}
read the original abstract
Incorporating responsible practices into software engineering (SE) for AI is essential to ensure ethical principles, societal impact, and accountability remain at the forefront of AI system design and deployment. This study investigates the ethical challenges and complexities inherent in responsible software engineering (RSE) for AI, underscoring the need for practical,scenario-driven operational guidelines. Given the complexity of AI and the relative inexperience of professionals in this rapidly evolving field, continuous learning and market adaptation are crucial. Through qualitative interviews with seven practitioners(conducted until saturation), quantitative surveys of 51 practitioners, and static validation of results with four industry experts in AI, this study explores how personal values, emerging roles, and awareness of AIs societal impact influence responsible decision-making in RSE for AI. A key finding is the gap between the current state of the art and actual practice in RSE for AI, particularly in the failure to operationalize ethical and responsible decision-making within the software engineering life cycle for AI. While ethical issues in RSE for AI largely mirror those found in broader SE process, the study highlights a distinct lack of operational frameworks and resources to guide RSE practices for AI effectively. The results reveal that current ethical guidelines are insufficiently implemented at the operational level, reinforcing the complexity of embedding ethics throughout the software engineering life cycle. The study concludes that interdisciplinary collaboration, H-shaped competencies(Ethical-Technical dual competence), and a strong organizational culture of ethics are critical for fostering RSE practices for AI, with a particular focus on transparency and accountability.
Reference graph
Works this paper leans on
-
[1]
Beyond code and algorithms: Navigating ethical complexities in artificial intelligence,
I. Dirgova´ Lupta´kova´, J. Posp´ıchal, and L. Huraj, “Beyond code and algorithms: Navigating ethical complexities in artificial intelligence,” in Proceedings of the Computational Methods in Systems and Software . Springer, 2023, pp. 316–332
work page 2023
-
[2]
Ai hazard management: A framework for the systematic management of root causes for ai risks,
R. Schnitzer, A. Hapfelmeier, S. Gaube, and S. Zillner, “Ai hazard management: A framework for the systematic management of root causes for ai risks,” in International Conference on Frontiers of Artificial Intelligence, Ethics, and Multidisciplinary Applications. Springer, 2023, pp. 359–375
work page 2023
-
[3]
E. Mendes, P. Rodriguez, V. Freitas, S. Baker, and M. A. Atoui, “Towards improving decision making and estimating the value of decisions in value- based software engineering: the value framework,” Software Quality Journal, vol. 26, pp. 607–656, 2018
work page 2018
-
[4]
Towards a roadmap on software engineering for responsible ai,
Q. Lu, L. Zhu, X. Xu, J. Whittle, and Z. Xing, “Towards a roadmap on software engineering for responsible ai,” in Proceedings of the 1st International Conference on AI Engineering: software Engineering for AI. New York, NY, USA: Association for Computing Machinery, 2022, p. 101–112. [Online]. Available: https://doi.org/10.1145/3522664.3528607
-
[5]
Software engineering as the linchpin of responsible ai,
L. Zhu, “Software engineering as the linchpin of responsible ai,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2023, pp. 3–4
work page 2023
-
[6]
Explaining the principles to practices gap in ai,
D. Schiff, B. Rakova, A. Ayesh, A. Fanti, and M. Lennon, “Explaining the principles to practices gap in ai,” IEEE Technology and Society Magazine, vol. 40, pp. 81–94, 06 2021
work page 2021
-
[7]
Tools and practices for responsible ai engineering,
R. Soklaski, J. Goodwin, O. Brown, M. Yee, and J. Matterer, “Tools and practices for responsible ai engineering,” arXiv preprint arXiv:2201.05647, 2022
arXiv 2022
-
[8]
Decision-making in software project management: A systematic literature review,
J. A. O. Cunha, H. P. Moura, and F. J. Vasconcellos, “Decision-making in software project management: A systematic literature review,” Procedia Computer Science, vol. 100, pp. 947–954, 2016
work page 2016
Show all 43 references
-
[9]
Behavioral software engineer- ing: A definition and systematic literature review,
P. Lenberg, R. Feldt, and L. G. Wallgren, “Behavioral software engineer- ing: A definition and systematic literature review,” Journal of Systems and software, vol. 107, pp. 15–37, 2015
2015
-
[10]
Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices,
B. Rakova, J. Yang, H. Cramer, and R. Chowdhury, “Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices,” Proceedings of the ACM on Human-Computer Interaction, vol. 5, no. CSCW1, pp. 1–23, 2021
2021
-
[11]
Ways of applying artificial intelligence in software engineering,
R. Feldt, F. G. de Oliveira Neto, and R. Torkar, “Ways of applying artificial intelligence in software engineering,” in Proceedings of the 6th International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering, 2018, pp. 35–41
2018
-
[12]
Applications of ai in classical software engineering,
M. Barenkamp, J. Rebstadt, and O. Thomas, “Applications of ai in classical software engineering,” AI Perspectives, vol. 2, no. 1, p. 1, 2020
2020
-
[13]
Responsible ai: requirements and challenges,
M. Ghallab, “Responsible ai: requirements and challenges,” AI Perspec- tives, vol. 1, no. 1, pp. 1–7, 2019
2019
-
[14]
Responsible ai: Bridging from ethics to practice,
B. Shneiderman, “Responsible ai: Bridging from ethics to practice,” Communications of the ACM, vol. 64, no. 8, pp. 32–35, 2021
2021
-
[15]
Ai and ethics—operationalizing responsible ai,
L. Zhu, X. Xu, Q. Lu, G. Governatori, and J. Whittle, “Ai and ethics—operationalizing responsible ai,” Humanity driven AI: Productiv- ity, well-being, sustainability and partnership, pp. 15–33, 2022
2022
-
[16]
The global landscape of ai ethics guidelines,
A. Jobin, M. Ienca, and E. Vayena, “The global landscape of ai ethics guidelines,” Nature machine intelligence , vol. 1, no. 9, pp. 389 –399, 2019
2019
-
[17]
Unveiling the black box: Bringing algorithmic trans - parency to ai,
G. Chaudhary, “Unveiling the black box: Bringing algorithmic trans - parency to ai,” Masaryk University Journal of Law and Technology , vol. 18, no. 1, pp. 93–122, 2024
2024
-
[18]
Towards ethical data -driven software: filling the gaps in ethics research and practice,
B. Johnson and J. Smith, “Towards ethical data -driven software: filling the gaps in ethics research and practice,” in 2021 IEEE/ACM 2nd International Workshop on Ethics in Software Engineering Research and Practice (SEthics). IEEE, 2021, pp. 18–25
2021
-
[19]
Responsible software engineering,
I. Schieferdecker, “Responsible software engineering,” The future of software quality assurance, pp. 137–146, 2020
2020
-
[20]
Ethics in the software development process: from codes of conduct to ethical deliberation,
J. Gogoll, N. Zuber, S. Kacianka, T. Greger, A. Pretschner, and J. Nida -Ru¨ melin, “Ethics in the software development process: from codes of conduct to ethical deliberation,” Philosophy and Technology , vol. 34, no. 4, p. 1085 –1108, Apr. 2021. [Online]. Available: http://dx...
2021 doi
-
[21]
Does acm’s code of ethics change ethical decision making in software development?
A. McNamara, J. Smith, and E. Murphy -Hill, “Does acm’s code of ethics change ethical decision making in software development?” in Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineerin...
2018
-
[22]
Human factors related challenges in software engineering –an industrial perspective,
P. Lenberg, R. Feldt, and L. G. Wallgren, “Human factors related challenges in software engineering –an industrial perspective,” in 2015 ieee/acm 8th international workshop on cooperative and human aspects of software engineering. IEEE, 2015, pp. 43–49
2015
-
[23]
The shifting sands of motivation: Revisiting what drives contributors in open source,
M. Gerosa, I. Wiese, B. Trinkenreich, G. Link, G. Robles, C. Treude, I. Steinmacher, and A. Sarma, “The shifting sands of motivation: Revisiting what drives contributors in open source,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 2021...
2021
-
[24]
Advancing the study of human values in software engineering,
E. Winter, S. Forshaw, L. Hunt, and M. A. Ferrario, “Advancing the study of human values in software engineering,” in 2019 IEEE/ACM 12th International Workshop on Cooperative and Human Aspects of Software Engineering (CHASE). IEEE, 2019, pp. 19–26
2019
-
[25]
Secondary studies on human aspects in software engineering: A tertiary study,
E. Zolduoarrati, S. A. Licorish, and N. Stanger, “Secondary studies on human aspects in software engineering: A tertiary study,” Journal of Systems and Software, vol. 200, p. 111654, 2023
2023
-
[26]
Motivation and what really drives human behavior,
B. Souders, “Motivation and what really drives human behavior,” Positive Psychology, 2020
2020
-
[27]
Towards a decision -making structure for selecting a research design in empirical software engineering,
C. Wohlin and A. Aurum, “Towards a decision -making structure for selecting a research design in empirical software engineering,” Empirical Software Engineering, vol. 20, pp. 1427–1455, 2015
2015
-
[28]
Development of qualitative semi - structured interview guide for case study research,
N. Naz, F. Gulab, and M. Aslam, “Development of qualitative semi - structured interview guide for case study research,” Competitive Social Science Research Journal, vol. 3, no. 2, pp. 42–52, 2022
2022
-
[29]
Selecting empirical methods for software engineering research,
S. Easterbrook, J. Singer, M. -A. Storey, and D. Damian, “Selecting empirical methods for software engineering research,” Guide to advanced empirical software engineering, pp. 285–311, 2008
2008
-
[30]
Supplementary material to
L. Murali Rani and F. Mohammadi, “Supplementary material to ”an empirical study on the decision making aspects in responsible ai development”,” 2024. [Online]. Available: https://doi.org/10.5281/zenodo.13771396
2024 doi
-
[31]
How many interviews are enough? an experiment with data saturation and variability,
G. Guest, A. Bunce, and L. Johnson, “How many interviews are enough? an experiment with data saturation and variability,” Field methods, vol. 18, no. 1, pp. 59–82, 2006
2006
-
[32]
Convenience and purposive sampling techniques: Are they the same,
E. I. Obilor, “Convenience and purposive sampling techniques: Are they the same,” International Journal of Innovative Social & Science Education Research, vol. 11, no. 1, pp. 1–7, 2023
2023
-
[33]
Using thematic analysis in psychology,
V. Braun and V. Clarke, “Using thematic analysis in psychology,” Qualitative research in psychology, vol. 3, no. 2, pp. 77–101, 2006
2006
-
[34]
Making the case for narrative methods in cross-cultural organizational research,
K. Soin and T. Scheytt, “Making the case for narrative methods in cross-cultural organizational research,” Organizational Research Methods, vol. 9, no. 1, pp. 55–77, 2006
2006
-
[35]
A comparative tale of two methods: How thematic and narrative analyses author the data story differently,
K. McAllum, S. Fox, M. Simpson, and C. Unson, “A comparative tale of two methods: How thematic and narrative analyses author the data story differently,” Communication Research and Practice, vol. 5, no. 4, pp. 358–375, 2019
2019
-
[36]
Qualitative software engineering research: Reflections and guidelines,
P. Lenberg, R. Feldt, L. Gren, L. G. Wallgren Tengberg, I. Tidefors, and D. Graziotin, “Qualitative software engineering research: Reflections and guidelines,” Journal of Software: Evolution and Process, p. e2607, 2023
2023
-
[37]
Narrative psychology,
M. Murray et al. , “Narrative psychology,” Qualitative psychology: A practical guide to research methods, pp. 85–107, 2015
2015
-
[38]
Design of diverging stacked bar charts for likert scales and other applications,
R. Heiberger and N. Robbins, “Design of diverging stacked bar charts for likert scales and other applications,” Journal of Statistical Software , vol. 57, pp. 1–32, 2014
2014
-
[39]
Plotting likert and other rating scales,
N. B. Robbins, R. M. Heiberger et al., “Plotting likert and other rating scales,” in Proceedings of the 2011 joint statistical meeting , vol. 1. American Statistical Association,, 2011
2011
-
[40]
Sustainability competencies and skills in software engineering: An industry perspective,
R. Heldal, N.-T. Nguyen, A. Moreira, P. Lago, L. Duboc, S. Betz, V. C. Coroama, B. Penzenstadler, J. Porras, R. Capilla et al. , “Sustainability competencies and skills in software engineering: An industry perspective,” arXiv preprint arXiv:2305.00436, 2023
2023 arXiv
-
[41]
Impact of business analytics and π-shaped skills on innovative performance: Findings from pls -sem and fsqca,
J. A. M. Hayajneh, M. B. H. Elayan, M. A. M. Abdellatif, and A. M. Abubakar, “Impact of business analytics and π-shaped skills on innovative performance: Findings from pls -sem and fsqca,” Technology in Society , vol. 68, p. 101914, 2022
2022
-
[42]
How does ai improve human decision-making? evidence from the ai -powered go program,
S. Choi, N. Kim, J. Kim, and H. Kang, “How does ai improve human decision-making? evidence from the ai -powered go program,” Evidence from the AI-Powered Go Program (April 2022). USC Marshall School of Business Research Paper Sponsored by iORB, No. Forthcoming, 2022
2022
-
[43]
Threats to validity in software engineering research: A critical reflection,
R. Verdecchia, E. Engstro¨m, P. Lago, P. Runeson, and Q. Song, “Threats to validity in software engineering research: A critical reflection,” Information and Software Technology, vol. 164, p. 107329, 2023
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.