{"id":"fcc6120e-3948-4804-ad04-7f1101d95f0b","arxiv_id":"2501.15691","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Practitioners in AI software development report that ethical guidelines are rarely operationalized, and decision-making is driven more by personal values and organizational culture than by formal frameworks.","lead":"This paper studies how software engineers make decisions when building AI systems responsibly, using interviews and surveys with practitioners. It finds a gap between written ethical principles and what happens in practice, and argues for better operational guidelines and organizational ethics culture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The principles-to-practice gap is asserted but never directly measured: the survey asks about preparation and effectiveness, not whether ethics guidelines are operationalized in the SDLC, so the central claim rests on interpretive narrative analysis rather than on the quantitative data.","rationale":"The reader's weakest assumption identifies sampling and self-report accuracy. Those are real limitations, but the more fundamental problem is that even granting a representative sample and truthful participants, the study's instruments do not measure the central construct. The abstract claims a gap between state of the art and practice, yet no survey item assesses whether ethics principles are operationalized in the engineering life cycle, and no state-of-the-art baseline is specified. The qualitative themes about needing resources and frameworks are reasonable and well-illustrated with quotes, but they are perceptions and preferences, not observations of a practice gap. This does not invalidate the study as an exploratory qualitative contribution; it means the strongest claim should be stated as a hypothesis or an interpretive theme, not as a measured finding. The existing CONDITIONAL verdict remains appropriate, so no verdict change is needed. The paper is honest about limitations and the qualitative analysis appears methodical, so I do not see a basis for rejection.","tokens_in":15678,"tokens_out":5134,"duration_ms":52965,"concrete_test":"Have two independent coders re-code the raw interview transcripts and open-ended survey responses using a pre-registered scheme that separates (a) explicit first-hand reports of missing or absent operational ethics frameworks in respondents' own projects from (b) general preferences for more guidance or resources; report counts and inter-rater agreement. If category (a) is not spontaneously and consistently present, the \"principles-to-practice gap\" is an analyst-driven theme rather than an empirical finding. Additionally, distribute a short follow-up survey with a direct item (\"In my team, AI ethics principles are translated into concrete engineering tasks\"; 1-5) to a similar population; if the distribution does not show a marked deficit, the headline claim lacks direct quantitative support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a gap exists between the state of the art and actual practice in RSE for AI because ethical guidelines are insufficiently implemented at the operational level. The study never operationalizes this gap. No baseline for \"state of the art\" is defined, and the survey instrument (Section IV-C) contains no item asking whether respondents actually apply ethical guidelines in concrete engineering tasks. Q3 asks how prepared the organization is, Q7 asks how effective existing methodologies are; neither captures the presence or absence of operationalization, and Q3's distribution is mildly positive while Q7 is neutral. The gap is therefore an interpretive synthesis of narrative themes about missing resources and support (Section IV-B2), not a directly observed or quantified result. Within an interpretivist design this can be a valid contribution, but it cannot support the abstract's causal-sounding statement that \"current ethical guidelines are insufficiently implemented at the operational level\" as an empirical finding. The static validation with four experts (Section IV-D) is a member check over the same findings, not independent evidence, and the paper concedes possible leading questions in Section VI. The load-bearing weakness is thus construct validity: even if the sample were perfectly representative and self-reports fully accurate, the data would not directly test the headline gap claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory, interpretivist mixed-methods study of decision-making in responsible software engineering for AI (RSE for AI). The empirical basis is seven semi-structured interviews with practitioners, a survey of 51 respondents (Likert and open-ended questions), and four expert interviews used for static validation of findings. The authors use thematic and narrative analysis to derive themes under three research questions concerning the differences between responsible AI development and conventional software development, the influence of emerging roles, and the role of individual motivations and values. On this basis they claim a gap between the state of the art and industrial practice in RSE for AI, particularly a failure to operationalize ethical decision-making in the software engineering lifecycle, and they recommend interdisciplinary collaboration, H-shaped ethical-technical competences, and a stronger organizational ethics culture.","tokens_in":16010,"tokens_out":5649,"duration_ms":51876,"significance":"The paper addresses a timely and relevant topic and provides a useful descriptive account of how practitioners perceive ethical decision-making in AI development. Its main strengths are the mixed-method design, the explicit interpretivist stance, the use of both thematic and narrative analysis, and the public supplementary material. The reported quotes and tables give the reader a concrete view of participant responses. However, the contribution is primarily exploratory and hypothesis-generating; the sample is small and self-selected, the Likert items are not validated and are analyzed only descriptively, and the static validation is a member check rather than independent evidence. If the claims are reframed as participants' perceived experiences and recommendations, the paper can be a useful contribution to the empirical SE-for-AI literature; in its current form the abstract and conclusions state the central gap claim more strongly than the data support.","major_comments":[{"comment":"The headline claim that 'current ethical guidelines are insufficiently implemented at the operational level' is not directly operationalized by the survey instrument. Q3 (organizational preparedness) and Q7 (effectiveness of methodologies) ask about perceived readiness and effectiveness, not about whether respondents actually apply ethical guidelines in concrete engineering tasks; the mildly positive Q3 and neutral Q7 distributions therefore do not measure the presence or absence of operationalization. The narrative themes in Sections IV-B2 and IV-B3 provide evidence that participants call for more structured support and describe implementation challenges, but these are perceived gaps, not a measured baseline of practice against a defined state of the art. I recommend rewording the abstract and Section VII to say that practitioners in this sample perceive a lack of operational frameworks and describe ethics as difficult to implement, rather than asserting as an empirical finding that guidelines are insufficiently implemented.","section":"Section IV-C and abstract/conclusion"},{"comment":"The research questions use causal language ('influence', 'impact') and the discussion repeatedly states that personal values 'can critically influence' decisions and that roles 'impact' decision-making, but the study design is cross-sectional self-report (interviews and a one-time survey). Such a design cannot establish causal influence or impact; it can only document participants' beliefs, attributions, and self-reported experiences. The empirical statements should be cast in terms of perceived influence or reported attribution, and the causal claims should be reserved for future longitudinal or intervention studies.","section":"Section III-B and Section V"},{"comment":"The static validation with four experts recruited from the authors' professional network is presented as evidence that the findings are 'confirmed' and applicable in industry, but it is a member check over the same findings, not independent corroboration. The paper itself concedes in Section VI that the validation interviews 'have the potential to introduce bias due to the possibility of leading questions'; given that the findings were presented to the experts before discussing them, this is a significant limitation. The conclusions should either present the validation as an expert credibility check with its own verbatim evidence, or drop the suggestion that four experts independently confirmed the findings.","section":"Section IV-D and Section VI"},{"comment":"The claim that data saturation was attained by the seventh interview is asserted without supporting detail (e.g., how saturation was assessed, which codes/themes became redundant), and the seven participants are heterogeneous in roles, company sizes, and experience; with convenience sampling from the authors' network, this cannot support generalizations about industry-wide practice. The 51 survey responses are also self-selected and no response rate is reported. While Section VI acknowledges limited generalizability, this limitation is in tension with the broad wording of the key finding in the abstract, so the conclusion should consistently restrict its claims to the sample and context.","section":"Section III-C and Table I"}],"minor_comments":[{"comment":"The phrase 'in software engineering (software engineering) practices' appears to contain a duplicated parenthetical; the intended meaning should be clarified.","section":"Section I, first paragraph"},{"comment":"The column header 'Years in current role' appears twice and the current-company-size values are mixed with role tenure; the table should be cleaned so that each column has a unique header and consistent units.","section":"Table II"},{"comment":"The diverging stacked bar charts are described in the text but the axis labels and exact item wording are only in captions; for a journal version, ensure the charts carry numeric labels or a table of frequencies so that the distributions are independently inspectable.","section":"Figures 2-4"},{"comment":"The terms 'responsible software engineering for AI', 'responsible AI engineering', and 'RAI' are used interchangeably; define the acronym once and use it consistently throughout the manuscript.","section":"Throughout"},{"comment":"Q2 under RQ1 and Q3 under RQ3 both ask about conflicting personal values, but they receive different wordings; clarify whether these are intended to be the same construct repeated for different units of analysis or two distinct items.","section":"Section IV-C, Figures 2 and 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about many limitations, and I found no evidence of fabrication or questionable research practices. My main concern is that the abstract and conclusion overstate what a small, self-selected, cross-sectional study can show. If the authors reframe the central claim as a perceived gap and reduce the causal language, I would support publication as an exploratory empirical study. The journal should also ensure that the supplementary material and figures are available for review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading this one. First, it is a legitimate, honestly reported exploratory study that gives a useful snapshot of practitioner perceptions about responsible AI decision-making. Second, the main claim — that ethical guidelines are insufficiently implemented at the operational level — is not directly measured by the survey; it is an interpretive synthesis from seven interviews and narrative analysis. The authors are transparent about this in their threats section, which is more than many qualitative papers do.\n\nWhat is actually new: a mixed-method dataset combining seven in-depth interviews, 51 survey responses, and four expert validation sessions, focused specifically on how personal values, emerging roles, and organizational context shape RAI decisions. The supplementary materials are open, the analysis follows established thematic and narrative methods, and the authors draw on relevant prior work without overclaiming novelty. The finding that interdisciplinary collaboration and H-shaped competencies are seen as valuable is consistent with previous RAI literature but adds practitioner voices from engineering roles.\n\nThe soft spots, in proportion: the sample is small and self-selected, the Likert items are unvalidated, and the static validation with four experts is a member check from the authors' own network, not an independent test. More importantly for the central claim, the stress-test note is right: Q3 asks how prepared the organization is, Q7 asks how effective methodologies are, but neither asks whether respondents actually apply ethical guidelines in concrete engineering tasks. The gap between principles and practice is therefore an interpretation of narrative themes about missing resources and support, not a directly quantified observation. That is acceptable in an interpretivist design, but the abstract's phrasing that \"results reveal current ethical guidelines are insufficiently implemented\" goes beyond what the data can support. The authors should soften that to reflect that this is a perceived gap reported by a small, non-representative sample.\n\nWho this is for: researchers in software engineering and responsible AI who want a compact, recent account of practitioner self-perceptions, and people designing larger studies on RAI decision-making. It is not a definitive empirical result, and the recommendations should be treated as hypothesis-generating.\n\nMy recommendation: send it to peer review. It is a coherent, well-scoped empirical study that deserves referee time, even though the framing needs adjustment. I would not desk-reject it, but the authors should be pushed to align the abstract and conclusion with the actual evidence.\n\nReading group: maybe, if you want a discussion piece on qualitative methods in SE. I would cite it if working on RAI practice gaps, though not as a cornerstone.","headline":"A decent exploratory study on responsible AI decision-making with honest limitations, but the headline gap claim is an interpretive synthesis rather than a directly measured result.","tokens_in":16397,"tokens_out":2013,"would_cite":true,"duration_ms":21168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This empirical study claims that the binding constraint on responsible AI software is not a lack of ethical principles but their operationalization in day-to-day engineering decisions.","keywords":["Responsible AI","Responsible software engineering","Ethical decision-making","Behavioral software engineering","Transparency","Accountability","Operational guidelines","Human factors"],"falsifier":"A concrete check is to audit a random sample of AI project artifacts—requirements documents, design records, test plans, and issue trackers—for documented ethical trade-off decisions, such as a fairness-accuracy choice or a privacy-utility trade-off. If such decisions are routinely recorded in specific, scenario-level terms across many organizations, the paper's central claim that guidelines are not operationalized would be contradicted; if such records are rare or vague, the claim is supported.","tokens_in":15485,"feed_emoji":"⚖️","tokens_out":5284,"duration_ms":44865,"temperature":0.7,"pith_summary":"This paper tries to establish that the main obstacle to responsible software engineering for AI is a failure to turn ethical principles into operational decisions inside the software life cycle. Drawing on seven interviews, 51 survey responses, and validation interviews with four industry experts, it argues that practitioners feel personal responsibility and report awareness of AI's societal impact, yet organizations lack concrete frameworks, checklists, and resources to apply those values at the project level. The authors propose that interdisciplinary collaboration, dual ethical-technical competence (\"H-shaped\" professionals), and a management-supported culture of ethics are what would close the gap. A sympathetic reader would care because the claim redirects attention from writing ethics codes to designing decision-support mechanisms for engineers.","feed_headline":"AI ethics guidance fails at the operational level, survey finds","feed_subtitle":"Interviews and surveys of 60+ practitioners show responsible AI principles stall between policy and practice.","key_machinery":"The study's central object is the decision-making process in responsible AI engineering, examined through mixed methods: semi-structured interviews analyzed by thematic and narrative analysis, a Likert-scale and open-ended survey, and static validation of findings with industry experts. The mechanism that carries the argument is the comparison between responsible AI development and conventional software development across organizational, team, and individual levels, with data as a new decision-driving element and probabilistic outcomes requiring continuous monitoring. Named constructs include H-shaped (also called π-shaped) competency, defined as professional depth in two distinct domains, one technical and one ethical, and the \"principles-to-practice gap,\" the observed disconnect between stated ethical guidelines and operational decisions.","core_discovery":"The central discovery is a gap between the state of the art and the state of practice in responsible AI engineering: ethical issues in AI development largely mirror those in conventional software development, but they are more pronounced and are not operationalized. Practitioners' decisions are shaped by personal values, emerging specialized roles such as data engineers, machine learning engineers, and AI ethicists, and by organizational culture, but current ethical guidelines are too abstract to guide concrete choices about data, transparency, accountability, and post-deployment monitoring. The paper concludes that H-shaped competencies (deep skill in one technical domain plus deep skill in ethics), interdisciplinary collaboration, and an organizational culture of ethics are critical enablers, with transparency and accountability as the values practitioners weight most heavily.","pith_inferences":["A direct extension is a testable check: count whether project artifacts such as requirements documents, design records, and test plans record explicit ethical trade-off decisions; the paper's evidence is self-reported, so artifact-level data would independently verify the gap.","The H-shaped competency recommendation implies an educational pipeline change: software engineering curricula would need ethics modules deep enough to count as a second domain, not merely an elective.","The study's emphasis on transparency and accountability suggests that future regulation may target decision documentation duties rather than model behavior alone.","Because the sample skews toward professionals already engaged with AI, the gap may be even larger in organizations without self-selected interest in responsible AI; a random industry sample could test that."],"forward_implications":["If the gap claim is right, adding scenario-driven ethical checklists and review boards to the software life cycle would do more than issuing additional principle statements.","If H-shaped competency matters, hiring and training should reward dual technical-ethical depth rather than treating ethics as a separate staff function.","If organizational culture is decisive, smaller companies may embed ethics more easily, while larger companies need explicit frameworks to keep ethics from becoming bureaucratized.","If data is the key decision determinant in AI, the requirements and design phases deserve ethics review earlier than in conventional software development.","If post-deployment monitoring depends on resources, responsible-AI guidance must include cost and capacity conditions, not just obligations."],"supporting_citations":[{"why":"Supplies the global convergence of ethical AI principles, the baseline the study measures operationalization against.","marker":"[16]"},{"why":"Argues codes of conduct underdetermine complex ethical situations, the theoretical basis for scenario-driven guidelines.","marker":"[20]"},{"why":"Empirically shows abstract codes are hard to apply, supporting the principles-to-practice gap finding.","marker":"[21]"},{"why":"Documents the principles-to-practices gap in AI that this study extends to software engineering decisions.","marker":"[6]"},{"why":"Calls for operationalizing responsible AI, the target the study says is not reached in practice.","marker":"[15]"},{"why":"Provides the value-based decision-making framing used to distinguish RAI decisions from conventional software decisions.","marker":"[3]"},{"why":"Supplies the behavioral perspective on software engineers' motivations and values used in the research questions.","marker":"[9]"},{"why":"Provides a roadmap for software engineering for responsible AI that the study contrasts with actual practice.","marker":"[4]"},{"why":"Frames decision-making in responsible AI as balancing trade-offs, the paper's central object of study.","marker":"[7]"}],"fun_headline_variants":["AI ethics advice too abstract for developers, study finds","Responsible AI: principles fail to reach engineering practice","Survey: ethical AI guidelines not operationalized in practice","AI ethics gap: policies don't translate to engineering decisions","Study: AI ethics guidance stalls between policy and practice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that seven interviews, 51 self-selected survey responses, and four expert validation interviews, recruited through convenience and snowball sampling, fairly represent AI engineering practice, and that participants' descriptions of their decisions match what they actually do.","fun_headline_variants_meta":{"raw":{"variants":["AI ethics advice too abstract for developers, study finds","Responsible AI: principles fail to reach engineering practice","Survey: ethical AI guidelines not operationalized in practice","AI ethics gap: policies don't translate to engineering decisions","Study: AI ethics guidance stalls between policy and practice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1444,"prompt_tokens":963,"completion_tokens":481,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":404}},"tokens_in":579,"tokens_out":481,"duration_ms":4729,"temperature":1.0,"reasoning_tokens":404,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:01:57.697721+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check is to audit a random sample of AI project artifacts—requirements documents, design records, test plans, and issue trackers—for documented ethical trade-off decisions, such as a fairness-accuracy choice or a privacy-utility trade-off. If such decisions are routinely recorded in specific, scenario-level terms across many organizations, the paper's central claim that guidelines are not operationalized would be contradicted; if such records are rare or vague, the claim is supported.","supporting_citations":[{"cited_title":"The global landscape of ai ethics guidelines,","cited_arxiv_id":null,"evidence_quote":"Supplies the global convergence of ethical AI principles, the baseline the study measures operationalization against."},{"cited_title":"Does acm’s code of ethics change ethical decision making in software development?","cited_arxiv_id":null,"evidence_quote":"Empirically shows abstract codes are hard to apply, supporting the principles-to-practice gap finding."},{"cited_title":"Explaining the principles to practices gap in ai,","cited_arxiv_id":null,"evidence_quote":"Documents the principles-to-practices gap in AI that this study extends to software engineering decisions."},{"cited_title":"Ai and ethics—operationalizing responsible ai,","cited_arxiv_id":null,"evidence_quote":"Calls for operationalizing responsible AI, the target the study says is not reached in practice."},{"cited_title":"Towards improving decision making and estimating the value of decisions in value- based software engineering: the value framework,","cited_arxiv_id":null,"evidence_quote":"Provides the value-based decision-making framing used to distinguish RAI decisions from conventional software decisions."},{"cited_title":"Behavioral software engineer- ing: A definition and systematic literature review,","cited_arxiv_id":null,"evidence_quote":"Supplies the behavioral perspective on software engineers' motivations and values used in the research questions."},{"cited_title":"Towards a roadmap on software engineering for responsible ai,","cited_arxiv_id":null,"evidence_quote":"Provides a roadmap for software engineering for responsible AI that the study contrasts with actual practice."}],"review_version":1}