REVIEW 3 major objections 5 minor 36 references
Explainability as a Compliance Requirement: What Regulated Industries Need from AI Tools for Design Artifact Generation
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Explainability is a precondition, not an option, for AI tools that generate design artifacts in regulated requirements engineering.
desk verdict Transparent qualitative study with a real new corpus, but the abstract outruns the n=10 evidence base; worth refereeing with revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on a formal, addressee-relative definition of explainability: a system S is explainable with respect to aspect X for addressee A in context C when an explainer E supplies information I that lets A understand X in C. The paper applies this to AI design-tool outputs, treating the requirements engineer as addressee and the regulatory context as part of the explanation's success condition. Empirically, the machinery is a four-part semi-structured interview protocol analysed with inductive coding by two independent coders; 59 codes across 1,163 excerpts were organised into a thematic map, and inter-coder agreement was about 0.76.
What would settle it
A study that measured validation effort in regulated requirements-engineering projects and found that opaque AI-generated design artifacts did not increase validation time or error rates relative to manual creation would contradict the core claim, as would a certification-body audit accepting AI-generated diagrams without source tracing or justification.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that non-explainable AI outputs in preliminary design integration impose hidden costs that outweigh their speed benefits: validation tasks took roughly three times longer than planned, teams treated tool-generated diagrams as rough drafts rather than deliverables, and in some cases managers displayed outputs that engineers could not defend to external stakeholders or auditors. Practitioners reported that tools "rearrange words into structured formats" without reasoning about system interactions, miss fail-safes and timing constraints, and misapply domain terms such as "load balancer" across industries. The paper identifies source tracing as the most critical missing feature, followed by contextual justifications, compliance-validation support, and collaboration features, and concludes that explainability is mandatory in regulated settings.
Load-bearing premise
The conclusions rest on ten practitioners recruited by convenience sampling, seven of them current AI-tool users, being representative of requirements engineering across the aerospace, automotive, medical-device, energy, telecom, and industrial-automation sectors.
Editorial extensions
If this is right
- Tool vendors that add clickable source tracing from each diagram element back to the originating requirement can remove the largest reported validation burden.
- Contextual justifications that use domain-specific terminology and standards would let teams show AI outputs to external stakeholders and auditors.
- Built-in compliance checks linked to standards such as ISO 26262, ISO 13485, and DO-178C would reduce the need for manual cross-checking and support certification.
- In safety-critical domains, the realistic deployment model is hybrid human-AI collaboration, not unattended full automation.
- Without explainability features, adopting AI artifact generators in regulated industries can increase net cost and project time rather than reduce them.
Reading between the lines
- If the reported three-times-longer validation estimate holds, explainability features have a concrete break-even point: vendors could measure generation time versus validation time and prioritise features accordingly.
- The addressee-relative definition implies that a single explanation will not satisfy all roles; engineers, compliance officers, and managers may need role-specific explanation views, a design direction the paper gestures at but does not develop.
- The inclusion of three discontinued users hints that adoption studies sampling only current users understate explainability's role; churn and abandonment data could quantify how many teams regress to manual methods.
- Automated source tracing may matter more than any other feature because it directly reconstructs the traceability chains that certification and audits already require, aligning AI tools with existing compliance infrastructure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative interview study of ten requirements engineering (RE) practitioners from six regulated industries, investigating why AI-based design-artifact generation tools face limited adoption. The authors identify challenges attributed to non-explainable outputs—extensive manual validation, reduced stakeholder trust, poor domain-specific understanding, workflow disruption, time delays, cost increases, and regulatory risks—along with mitigation strategies currently used in practice and desired tool features such as source tracing, contextual justifications, compliance validation, and collaboration support. The paper concludes that explainability is a compliance requirement and offers a practitioner-derived roadmap for tool improvement. The study is presented as the first cross-industry investigation of explainability in real-world design artifact generation, guided by five research questions and analyzed using inductive coding with two independent coders.
Significance. If the findings are taken as an exploratory, qualitative account, the paper is a useful contribution: it addresses a genuine gap in the RE literature by capturing practitioner experiences rather than only technical capabilities, and it does so with transparent methods—public interview protocol and codebook, dual coding with Cohen's Kappa of 0.76, participant-anchored themes, and an explicit threats-to-validity section. The proposed feature set (automated source tracing, contextual justification, domain-specific adaptation, compliance validation, collaborative features) is concrete and actionable for both tool developers and adopting organizations. However, the significance is limited by the small, convenience-sampled evidence base (n=10), and the title, abstract, and conclusion generalize beyond what the study design can support. As a hypothesis-generating study, it is valuable; as a population-level claim about 'what regulated industries need,' it currently overreaches.
major comments (3)
- [Section III-A, Table I, and Abstract] The title-level claim 'What Regulated Industries Need' and the abstract's statement that non-explainable outputs 'often negating the anticipated efficiency benefits' are population-level generalizations, but they rest on ten convenience-sampled LinkedIn recruits. Table I shows that aerospace, medical devices, and energy each have exactly one participant, and the sample spans six domains with varying regulatory regimes. The paper's own transferability caveat in Section VI-B concedes that the sample may not cover all regulatory or geographic variations, but this concession does not constrain the title or abstract. This is load-bearing because the headline finding is a prevalence claim ('often negates,' 'necessitate'), not merely a thematic existence claim. I ask the authors to either reframe the contribution as exploratory and hypothesis-generating, or provide robustness evidence (e.g., a leave-one-out analysis of the reported code frequencies) to show that the main patterns are not driven by single individuals.
- [Section III-C and Figure 1] The statement in Section III-C that data saturation was reached after six interviews supports theme discovery but does not support stable prevalence or frequency claims. Yet Figure 1 reports code frequencies (e.g., Manual Validation = 89, Trust Issues = 76, Source Tracing (Manual) = 45), and Section V-A states that 'validation tasks taking three times longer than planned' as a general finding. With n=10 and single-participant domains, these frequencies are highly sensitive to individual responses, and the 'three times longer' figure is presented without a direct participant quote or a range. The paper should either downgrade such numerical claims to qualitative prevalence language (e.g., 'most participants reported'), or add a stability analysis to demonstrate that the frequencies are not an artifact of one or two influential participants.
- [Section IV-C1 and Figure 1] There is an internal inconsistency in the quantitative reporting: Section IV-C1 states that the Source Tracing (Manual) code has '30 coded references across participants,' while Figure 1 reports 45 codes for the same category. Similarly, Section III-C reports 59 unique codes across 1163 coded excerpts, but I could not verify this total from Figure 1's displayed values. This inconsistency undermines the reliability of the frequency-based claims, which are used to identify which challenges and desired features are most prominent. The authors should correct the discrepancy and provide a consistent, auditable mapping between the codebook, the coded excerpts, and the reported counts.
minor comments (5)
- [Abstract] The phrase 'Our findings reveal that non-explainable AI outputs necessitate extensive manual validation' should be softened to 'participants reported that non-explainable AI outputs necessitated extensive manual validation' to accurately reflect the self-reported, interview-based nature of the evidence.
- [Section IV-B1] There is a formatting error in Participant E's quote: 'The requirement clearly stated that the system should pause, but the tool didn’t explain why it left it out. ' appears to have a mismatched quotation mark. Please correct the punctuation so the quote is clearly delimited.
- [Section V-A] The claim that 'validation tasks taking three times longer than planned' is presented as a numeric finding, but no participant quote or aggregation method is provided in that section. Please tie this to a specific quote or explicitly label it as a synthesis of multiple participant accounts.
- [Section II and References] Several references are incomplete: reference [5] lacks a year and page range, reference [31] lacks volume and page numbers, and reference [36] gives only a surname and initial without the full author list. Please verify all references conform to the venue's style.
- [Section III-B] The availability statement says the interview script and codebook are available at a figshare URL. Please provide a DOI or stable archival link rather than a plain URL, and check that the link is currently accessible.
Circularity Check
No circularity: this is a transparent interview study whose roadmap is a declared synthesis of participant input, with external validation explicitly deferred.
full rationale
The paper makes no formal derivation or predictive claim; its findings are qualitative interpretations of semi-structured interviews. RQ5 explicitly asks participants what features they want (Section I), and Section IV-E reports those stated preferences; Section V-G restates them as recommendations. This is the declared method of an inductive grounded-theory study, not a disguised prediction. The conclusion explicitly defers empirical evaluation of the proposed improvements ('Future research should empirically evaluate these proposed improvements,' Section VII), so the roadmap is not presented as an independently validated output. The external definition of explainability is adopted from Chazette et al. [14] and does not reduce to the paper's conclusions. No load-bearing self-citations appear in the reference list, and no uniqueness theorem or ansatz is imported from prior work by the authors. The main epistemic limitation is the generalizability of a ten-participant convenience sample, acknowledged in Sections VI-B and VI-C; that is an external-validity concern, not circularity. A minor numeric inconsistency between Section IV-C1 ('30 coded references') and Figure 1 (45 for Source Tracing (Manual)) is a reporting-quality issue but does not constitute a circular step.
Assumptions & free parameters
assumptions (4)
- domain assumption The Chazette et al. definition of explainability (a system is explainable to an addressee in a context if an explainer provides information enabling understanding of an aspect) is adopted as the study's operative definition.
- domain assumption Participant self-reports during semi-structured interviews are reliable evidence of real workflow costs and tool behavior.
- domain assumption Data saturation reached after six interviews implies thematic completeness of the ten-interview corpus.
- domain assumption Convenience sampling through LinkedIn yields a sample adequate for claims about regulated-industry RE practice.
Cite this review
Pith. "Pith review of Explainability as a Compliance Requirement: What Regulated Industries Need from AI Tools for Design Artifact Generation." pith.science (2026). https://pith.science/paper/7LSY5MWN
@misc{pith2026250709220,
author = {Pith},
title = {Pith review of: Explainability as a Compliance Requirement: What Regulated Industries Need from AI Tools for Design Artifact Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/7LSY5MWN}},
note = {Machine review of arXiv:2507.09220}
}
read the original abstract
Artificial Intelligence (AI) tools for automating design artifact generation are increasingly used in Requirements Engineering (RE) to transform textual requirements into structured diagrams and models. While these AI tools, particularly those based on Natural Language Processing (NLP), promise to improve efficiency, their adoption remains limited in regulated industries where transparency and traceability are essential. In this paper, we investigate the explainability gap in AI-driven design artifact generation through semi-structured interviews with ten practitioners from safety-critical industries. We examine how current AI-based tools are integrated into workflows and the challenges arising from their lack of explainability. We also explore mitigation strategies, their impact on project outcomes, and features needed to improve usability. Our findings reveal that non-explainable AI outputs necessitate extensive manual validation, reduce stakeholder trust, struggle to handle domain-specific terminology, disrupt team collaboration, and introduce regulatory compliance risks, often negating the anticipated efficiency benefits. To address these issues, we identify key improvements, including source tracing, providing clear justifications for tool-generated decisions, supporting domain-specific adaptation, and enabling compliance validation. This study outlines a practical roadmap for improving the transparency, reliability, and applicability of AI tools in requirements engineering workflows, particularly in regulated and safety-critical environments where explainability is crucial for adoption and certification.
Figures
Reference graph
Works this paper leans on
-
[1]
M. Landh ¨außer, S. J. K ¨orner, and W. F. Tichy, “From requirements to UML models and back: How automatic processing of text can support requirements engineering,” Software Quality J. , pp. 121–149, 2014
work page 2014
-
[2]
Requirements Engineering: A roadmap,
B. Nuseibeh and S. Easterbrook, “Requirements Engineering: A roadmap,” in Proc. Conf. on the Future of Software Engineering , 2000
work page 2000
-
[3]
B. Wang, R. Peng, Y . Li, H. Lai, and Z. Wang, “Requirements traceabil- ity technologies and technology transfer decision support: A systematic review,” Journal of Systems and Software , vol. 146, pp. 59–79, 2018
work page 2018
-
[4]
Requirements engineering in automotive development-experiences and challenges,
M. Weber and J. Weisbrod, “Requirements engineering in automotive development-experiences and challenges,” in Int. Conference. on Re- quirements Eng. IEEE, 2002, pp. 331–340
work page 2002
-
[5]
Challenges of working with artifacts in require- ments engineering and software engineering,
P. Ghazi and M. Glinz, “Challenges of working with artifacts in require- ments engineering and software engineering,”Requirements engineering
-
[6]
From requirements to code: A full model-driven development perspective,
´O. Pastor, M. Ruiz, and S. Espa ˜na, “From requirements to code: A full model-driven development perspective,” in International Conference on Software and Data Technologies . Springer, 2011, pp. 56–70
work page 2011
-
[7]
A. Aurum and C. Wohlin, Engineering and managing software require- ments. Springer, 2005, vol. 1
work page 2005
-
[8]
An NLP approach for cross-domain ambiguity detection in requirements engineering,
A. Ferrari and A. Esuli, “An NLP approach for cross-domain ambiguity detection in requirements engineering,” Automated Software Eng., 2019
work page 2019
Show all 36 references
-
[9]
Software testing with large language models: Survey, landscape, and vision,
J. Wang, Y . Huang, C. Chen, Z. Liu, S. Wang, and Q. Wang, “Software testing with large language models: Survey, landscape, and vision,”IEEE Transactions on Software Engineering , 2024
2024
-
[10]
Natural language processing for requirements engineering: A systematic mapping study,
L. Zhao, W. Alhoshan, A. Ferrari, K. J. Letsholo, M. A. Ajagbe, E.-V . Chioasca, and R. T. Batista-Navarro, “Natural language processing for requirements engineering: A systematic mapping study,” ACM Comput- ing Surveys (CSUR) , vol. 54, no. 3, pp. 1–41, 2021
2021
-
[11]
AI-based ques- tion answering assistance for analyzing natural-language requirements,
S. Ezzini, S. Abualhaija, C. Arora, and M. Sabetzadeh, “AI-based ques- tion answering assistance for analyzing natural-language requirements,” in Proc. International Conference. on Software Eng. (ICSE) , 2023
2023
-
[12]
Machine learning in requirements engineering: A mapping study,
K. Zamani, D. Zowghi, and C. Arora, “Machine learning in requirements engineering: A mapping study,” in International. Requirements Eng. Conference. Workshops (REW). IEEE, 2021, pp. 116–125
2021
-
[13]
Large language models for software engi- neering: A systematic literature review,
X. Hou, Y . Zhao, Y . Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large language models for software engi- neering: A systematic literature review,” ACM Transactions on Software Engineering and Methodology , no. 8, pp. 1–79, 2024
2024
-
[14]
Exploring explainability: A definition, a model, and a knowledge catalogue,
L. Chazette, W. Brunotte, and T. Speith, “Exploring explainability: A definition, a model, and a knowledge catalogue,” in Proc. Int. Requirements Eng. Conf. (RE) . IEEE, 2021, pp. 197–208
2021
-
[15]
Explainability as a non-functional requirement: challenges and recommendations,
L. Chazette and K. Schneider, “Explainability as a non-functional requirement: challenges and recommendations,” Requirements Engineer- ing, no. 4, pp. 493–514, 2020
2020
-
[16]
On the assessment of generative ai in modeling tasks: an experience report with chatgpt and uml,
J. C ´amara, J. Troya, L. Burgue ˜no, and A. Vallecillo, “On the assessment of generative ai in modeling tasks: an experience report with chatgpt and uml,” Software and Systems Modeling , no. 3, pp. 781–793, 2023
2023
-
[17]
Automated, interactive, and traceable domain modelling empowered by artificial intelligence,
R. Saini, G. Mussbacher, J. L. Guo, and J. Kienzle, “Automated, interactive, and traceable domain modelling empowered by artificial intelligence,” Software and Systems Modeling , 2022
2022
-
[18]
Generating sequence diagram from natural language requirements,
M. Jahan, Z. S. H. Abad, and B. Far, “Generating sequence diagram from natural language requirements,” inInternational Requirements Eng. Conference Workshops (REW), 2021, pp. 39–48
2021
-
[19]
An active learning approach for improving the accuracy of automated domain model extraction,
C. Arora, M. Sabetzadeh, S. Nejati, and L. Briand, “An active learning approach for improving the accuracy of automated domain model extraction,” ACM TOSEM, 2019
2019
-
[20]
Challenges in applying large language models to requirements engineering tasks,
J. J. Norheim, E. Rebentisch, D. Xiao, L. Draeger, A. Kerbrat, and O. L. de Weck, “Challenges in applying large language models to requirements engineering tasks,” Design Science, p. e16, 2024
2024
-
[21]
Detecting requirements defects with NLP patterns: An industrial experience in the railway domain,
A. Ferrari, G. Gori, B. Rosadini, I. Trotta, S. Bacherini, A. Fantechi, and S. Gnesi, “Detecting requirements defects with NLP patterns: An industrial experience in the railway domain,” Empirical Software Engineering, no. 6, pp. 3684–3733, 2018
2018
-
[22]
A systematic liter- ature review on using natural language processing in software require- ments engineering,
S.-C. Necula, F. Dumitriu, and V . Greavu-S ¸erban, “A systematic liter- ature review on using natural language processing in software require- ments engineering,” Electronics, vol. 13, no. 11, p. 2055, 2024
2024
-
[23]
Practitioners’ perceptions of the goals and visual explanations of defect prediction models,
J. Jiarpakdee, C. K. Tantithamthavorn, and J. Grundy, “Practitioners’ perceptions of the goals and visual explanations of defect prediction models,” in IEEE/ACM International Conference on Mining Software Repositories (MSR), 2021, pp. 432–443
2021
-
[24]
Transparency and explainability of AI systems: From ethical guidelines to requirements,
N. Balasubramaniam, M. Kauppinen, A. Rannisto, K. Hiekkanen, and S. Kujala, “Transparency and explainability of AI systems: From ethical guidelines to requirements,” Information and Software. Technology. , 2023
2023
-
[25]
Can requirements engineering support explainable artificial intelligence? towards a user-centric ap- proach for explainability requirements,
U.-E. Habiba, J. Bogner, and S. Wagner, “Can requirements engineering support explainable artificial intelligence? towards a user-centric ap- proach for explainability requirements,” in 2022 IEEE 30th international requirements engineering conference workshops (REW)
2022
-
[26]
Explainable AI for software engineering,
C. K. Tantithamthavorn and J. Jiarpakdee, “Explainable AI for software engineering,” in IEEE/ACM International Conference on Automated Software Engineering (ASE) , 2021, pp. 1–2
2021
-
[27]
Ai tool use and adoption in software development by individuals and organizations: a grounded theory study,
Z. S. Li, N. N. Arony, A. M. Awon, D. Damian, and B. Xu, “Ai tool use and adoption in software development by individuals and organizations: a grounded theory study,” arXiv preprint arXiv:2406.17325 , 2024
2024 arXiv
-
[28]
Navigating the complexity of generative ai adoption in software engineering,
D. Russo, “Navigating the complexity of generative ai adoption in software engineering,” ACM Transactions on Software Engineering and Methodology, pp. 1–50, 2024
2024
-
[29]
Classifying ambiguous requirements: An explainable approach in the railway indus- try,
L. Beqiri, C. S. Montero, A. Cicchetti, and A. Kruglyak, “Classifying ambiguous requirements: An explainable approach in the railway indus- try,” in Requirements Eng. Conference Workshops (REW) , 2024
2024
-
[30]
Designing NLP-based solutions for requirements variability management: Experiences from a design science study at visma,
P. Elahidoost, M. Unterkalmsteiner, D. Fucci, P. Liljenberg, and J. Fis- chbach, “Designing NLP-based solutions for requirements variability management: Experiences from a design science study at visma,” in Proc. International Working Conference on Requirements Eng.: Foun- dat...
2024
-
[31]
User stories and natural language processing: A systematic literature review,
I. K. Raharjana, D. Siahaan, and C. Fatichah, “User stories and natural language processing: A systematic literature review,” IEEE access
-
[32]
Applications of natural language process- ing in software traceability: A systematic mapping study,
Z. Pauzi and A. Capiluppi, “Applications of natural language process- ing in software traceability: A systematic mapping study,” Journal of Systems and Software , vol. 198, p. 111616, 2023
2023
-
[33]
The use of nlp-based text representation techniques to support requirement engineering tasks: A systematic mapping review,
R. Sonbol, G. Rebdawi, and N. Ghneim, “The use of nlp-based text representation techniques to support requirement engineering tasks: A systematic mapping review,” IEEE Access, pp. 62 811–62 830, 2022
2022
-
[34]
Convenience sampling, random sampling, and snowball sampling: How does sampling affect the validity of research?
R. W. Emerson, “Convenience sampling, random sampling, and snowball sampling: How does sampling affect the validity of research?” Journal of visual impairment & blindness , no. 2, pp. 164–168, 2015
2015
-
[35]
Charmaz, Constructing grounded theory: A practical guide through qualitative analysis
K. Charmaz, Constructing grounded theory: A practical guide through qualitative analysis. Sage, 2006
2006
-
[36]
Y . S. Lincoln, Naturalistic inquiry. Sage, 1985
1985
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.