{"id":"1160da47-fcbf-4ed4-b380-0da6c4227726","arxiv_id":"2505.07058","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature-based survey maps LIME, SHAP, counterfactual, rule extraction, and concept-based explanations onto SDLC phases, with no new empirical results.","lead":"This paper surveys common explainable AI methods, such as LIME, SHAP, and counterfactual explanations, and matches them to each stage of the software development lifecycle. A generalist might read it to see where AI-driven software work lacks transparency and which explanation tools are being proposed for each phase.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Phase-specific 'most effective' claims conflict with the paper's own statement that design and testing have not been researched, leaving the first-comprehensive claim unsupported.","rationale":"The survey's descriptions of XAI methods are generally accurate, and a phase-aware map could be a useful organizing device. The problem is evidentiary: the central novelty is 'first comprehensive', but the paper's own citation [10] has a title suggesting an SDLC-aligned XAI meta-review, and the paper does not compare itself to it. The internal mismatch between Section II.C and Sections III.B/III.D is the sharpest indicator that the phase recommendations are not all derived from surveyed applications. This does not make the paper worthless, but it means the conditional verdict is right: the contributions should be reframed as hypotheses and the SLR made auditable. The proposed check would settle whether the design and testing sections are grounded or extrapolated, and whether the novelty claim survives comparison with prior phase-aligned work.","tokens_in":10428,"tokens_out":4706,"duration_ms":47927,"concrete_test":"Reconstruct the SLR corpus using the reported keywords, databases, and 6-year inclusion window; tabulate primary studies per SDLC phase, especially for Design (Section III.B) and Testing (Section III.D). If zero studies exist for those phases, the 'most effective' claims are not survey results and must be reframed as hypotheses. In parallel, read reference [10] (XAIR) to determine whether it already provides an SDLC-aligned XAI mapping; if so, the 'first' claim requires a direct comparison showing what this survey adds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a phase-by-phase mapping of XAI techniques, and it claims to be the first comprehensive SDLC survey. This requires the mapping to be grounded in an auditable literature review. That condition is not met: Section II.C, summarizing [7], states 'Software design and testing have not been researched.' Yet Section III.B (Design) and Section III.D (Testing) assert that 'during the literature review, the most effective XAI techniques in addressing XAI challenges were found to be' LIME/SHAP, counterfactuals, rule extraction, and so on. If there are no primary studies for these phases, no literature review could establish which techniques are most effective for them; these are extrapolations or hypotheses presented as findings. The paper also does not provide its SLR protocol, screening counts, study list, or corpus table, so the 68%/16%/8%/8% distribution and every 'most effective' ranking are unverifiable. The claim of first comprehensiveness is therefore not supported by the evidence presented, and the phase-specific recommendations inherit this evidentiary gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys Explainable AI (XAI) techniques organized by Software Development Life Cycle (SDLC) phase. It claims to be the first comprehensive survey of XAI techniques covering every SDLC phase, and it maps techniques such as LIME, SHAP, counterfactual explanations, rule extraction, attention mechanisms, concept-based explanations, and example-based explanations to requirements elicitation, design, development, testing, deployment/monitoring, and maintenance/evolution. The paper also reports, from a prior systematic literature review [7], that 68% of XAI-in-SE research focuses on maintenance, 16% on development, and 8% each on management and requirements. It draws conclusions about which XAI techniques are 'most effective' in each phase and recommends phase-specific XAI adoption.","tokens_in":10615,"tokens_out":2742,"duration_ms":27278,"significance":"If the phase-specific mapping were backed by a transparent and auditable literature review, the paper would be a useful practitioner-oriented reference: it gives generally accurate descriptions of LIME, SHAP, counterfactuals, rule extraction, and attention mechanisms, and it organizes these methods by SDLC phase in a way that could help practitioners choose candidates for explainability. The expository descriptions of individual XAI methods (Sections II.B.1–II.B.4) are broadly correct, and the paper makes a reasonable high-level case that different SDLC phases pose different explainability challenges. However, the central claims—first comprehensiveness, the 68/16/8/8 distribution, and the per-phase 'most effective' rankings—rest on an unverifiable evidentiary base, since the paper does not provide a corpus table, screening counts, or a reproducible search protocol. The significance of the paper as a survey is therefore currently limited by its failure to substantiate its empirical claims.","major_comments":[{"comment":"There is a direct internal contradiction between the paper's stated evidence and its phase-specific effectiveness claims. Section II.C, summarizing reference [7], states that 'Software design and testing have not been researched.' Yet Section III.B (Design, item 3) and Section III.D (Testing, item 3) both assert that 'during the literature review, the most effective XAI techniques in addressing XAI challenges were found to be' LIME/SHAP, counterfactuals, rule extraction, etc. If no primary studies exist for these phases, no literature review could have 'found' which techniques are most effective; these are proposals or expert judgments presented as empirical findings. The authors must either reclassify these as recommended/plausible techniques rather than literature-derived findings, or provide the primary sources that support effectiveness for design and testing.","section":"Section II.C vs. Section III.B/III.D"},{"comment":"The 68/16/8/8 distribution of XAI-in-SE research across SDLC phases is presented three times (abstract, Introduction, and Section II.C) and motivates the entire paper, but the manuscript provides no way to verify it. The Introduction describes search venues and inclusion criteria at a high level, but the paper does not report the number of studies screened, the number included, exclusion criteria, or a study list/corpus table. The statistics are imported from [7], a single arXiv preprint, and the paper does not assess the quality or representativeness of that source. Without these details, the motivating statistics and the implied research gap are not auditable.","section":"Abstract, Section I, Section II.C"},{"comment":"The claim 'this is the first comprehensive survey of XAI techniques for every phase of the SDLC' is not supported by the evidence presented. The survey's own literature summary (Section II.C) states that design and testing have not been researched, so the content for those phases consists of proposed mappings rather than surveyed existing work. Additionally, the paper does not compare against the closely related prior work [10] (XAIR, a systematic metareview aligned to the software development process) beyond a citation, so the novelty claim is not established. Either provide a detailed comparison that demonstrates no prior phase-by-phase mapping exists, or soften the claim to 'first phase-specific proposal'.","section":"Abstract, Section I, Section V"},{"comment":"The phrase 'the most effective XAI techniques' is used repeatedly (e.g., Sections III.A.3, III.B.3, III.C.3, III.D.3, III.E.3, III.F.3) without any stated criterion for effectiveness. No evaluation framework, user study, fidelity measure, or comparative analysis is provided. If the paper intends to make a descriptive claim about the literature, it needs to show that the cited works actually evaluated these techniques and compared alternatives; if the claim is prescriptive, it should be worded as a recommendation. As written, the label 'most effective' is a comparative empirical claim with no supporting evidence.","section":"Section III (all phase subsections)"}],"minor_comments":[{"comment":"The SDLC phase naming is inconsistent: the abstract says 'requirements elicitation, design and development, testing and deployment, and evolution' while Section III separates design, development, testing, and deployment/monitoring into distinct subsections. Align the phase list across the abstract, introduction, and body.","section":"Abstract and Section III"},{"comment":"The bullet 'Other tasks: Software design and testing have not been researched' uses the word 'tasks,' but design and testing are SDLC phases in this paper's own taxonomy; rephrase to avoid confusion.","section":"Section II.C"},{"comment":"Figure 1 is referenced in Section II.B as a 'summarization and classification of standard XAI techniques,' but the figure content is not visible in the manuscript text and no figure caption is provided. Please include the figure or remove the reference.","section":"Figure 1"},{"comment":"The paper cites [10] (XAIR) as related work but does not explicitly state how its phase-level contribution differs from [10]'s alignment of XAI techniques to the software development process. A brief differentiation would strengthen the novelty discussion.","section":"Reference [10]"},{"comment":"Several references have incomplete author lists (e.g., [2] 'Cheligeer C, et. al' and [5] 'L. Chazette, et. al') and inconsistent rendering of 'et al.'; please correct to the journal's reference style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern in the reader's report is genuine and lands: the paper's central empirical claims—the distribution statistics and the per-phase 'most effective' rankings—are not verifiable from the manuscript, and the internal contradiction between Section II.C (design and testing not researched) and Sections III.B/III.D ('most effective' techniques for those phases) is a load-bearing problem that must be fixed. The issue is fixable by restructuring the claims as proposals/recommendations and adding an auditable review appendix (search protocol, screening counts, and a study corpus table), so major_revision rather than reject seems proportionate. Also note that the paper reads like an industry technical report; strengthen the academic framing if it is intended for a journal venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a survey paper, not a research result, and it should be read as one. What it does well: the descriptions of LIME, SHAP, counterfactuals, rule extraction, attention, and concept-based methods are accurate and accessible. The mapping of these techniques to SDLC phases is a useful organizing scaffold, especially for practitioners who want a quick overview of where XAI might apply. The writing is clean and the structure is sensible.\n\nWhat is not new: the claim of being the first comprehensive phase-aware survey does not hold up against the paper's own cited references. Mohammadkhani et al. [7] already aligns XAI with software engineering phases and provides the 68/16/8/8 distribution. Clement et al. [10] explicitly aligns XAI with the software development process. This paper mostly rebuilds that map with slightly different phase boundaries and adds example use cases. The technique summaries are textbook material. That is fine for a tutorial, but it is not a novel contribution.\n\nThe bigger problem is the evidence gap behind the central claims. The paper states in Section II.C, citing [7], that design and testing have not been researched. Then Sections III.B and III.D assert that \"the most effective XAI techniques\" for those phases are LIME/SHAP, counterfactuals, and rule extraction. If there are no primary studies in those phases, no literature review could establish effectiveness rankings. Those claims are extrapolations or hypotheses, but they are presented as findings. The stress-test note is right about this tension. Also, the paper says it conducted an SLR but provides no protocol, screening counts, study list, or corpus table. The distribution numbers and effectiveness rankings are therefore unverifiable, and the \"first comprehensive\" claim rests on that unverifiable base.\n\nIs it a bad paper? No. It is a useful secondary source for someone new to the area, and the phase-specific examples are often sensible. But the authors overreach. A revision that (a) removes or softens the novelty claim, (b) labels per-phase recommendations as hypotheses grounded in reasoning rather than systematic evidence, and (c) either reports the SLR details or explicitly depends on [7]'s corpus would make this an honest survey. As it stands, the load-bearing claims about effectiveness are not auditable.\n\nMy take: send it to peer review, but expect major revision. The core descriptions are sound and the phase map has practical value, but the evidence for the rankings and the first-comprehensive claim need fixing before publication. For my own work, I would cite the primary surveys, not this one.\n\nRecommendation: engage with it as a candidate survey after major revision; do not desk reject, but do not let the current claims through as-is.\n\nBest,\n[You]","headline":"A clearly written but over-claimed survey of XAI-for-SDLC phases; treat its phase-specific 'most effective' rankings as hypotheses, not findings.","tokens_in":11160,"tokens_out":1005,"would_cite":false,"duration_ms":12033,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first comprehensive, phase-specific mapping of explainable AI techniques onto every stage of the software development lifecycle, from requirements elicitation through maintenance.","keywords":["Explainable AI","Software Development Lifecycle","SDLC phases","LIME","SHAP","Counterfactual explanations","Rule extraction","Software engineering survey"],"falsifier":"A reader could re-run the search protocol the paper describes (peer-reviewed articles from the last six years across the major computing and scientific databases, using the stated keywords) and count how many XAI-in-software-engineering papers address each SDLC phase. If the resulting distribution differs materially from 68/16/8/8, or if design and testing turn out to contain substantial XAI research, the paper's gap-driven phase recommendations lose their evidential basis.","tokens_in":10258,"feed_emoji":"🤖","tokens_out":6416,"duration_ms":53506,"temperature":0.7,"pith_summary":"The paper sets out to show that explainability for AI-assisted software development is not a one-size-fits-all concern: each phase of the software development lifecycle (requirements, design, development, testing, deployment, maintenance) faces distinct explainability challenges, and different XAI techniques address them differently. It offers a phase-by-phase map that pairs techniques such as LIME, SHAP, rule extraction, counterfactuals, attention mechanisms, concept-based explanations, and example-based explanations with each phase's AI applications and explainability challenges. The paper also claims, citing prior systematic reviews, that XAI research in software engineering is heavily skewed toward maintenance (68%) and sparse in requirements and management (8% each), which motivates its emphasis on earlier phases. On top of the mapping, the paper claims to be the first comprehensive survey of XAI techniques for every SDLC phase. A sympathetic reader would take this as a practical guide and gap map intended to help practitioners choose XAI methods by phase and to steer research toward under-covered phases.","feed_headline":"First survey maps XAI onto every software phase","feed_subtitle":"It pairs LIME, SHAP, counterfactuals and rule extraction with requirements, design, testing, deployment and maintenance.","key_machinery":"The organizing device is a taxonomy-to-phase matrix. XAI techniques are classified by stage (ante-hoc vs post-hoc), scope (local vs global), and input/output format, then each SDLC phase is broken into AI applications, explainability challenges, and nominally most effective XAI techniques. The pairing does the argument's work: for example, counterfactual explanations answer 'what would need to differ in the input to change the outcome?' in design, testing, deployment, and maintenance, while LIME/SHAP attribute feature importance in requirements, development, testing, deployment, and maintenance. The mechanism is this repeatable technique-to-challenge pairing, underpinned by the distribution claim that most XAI-in-software-engineering research sits in maintenance.","core_discovery":"The central claim is that XAI techniques should be selected per SDLC phase rather than applied uniformly, and that every phase has workable XAI options. For requirements elicitation, LIME/SHAP can surface latent requirements and detect bias, while counterfactuals clarify trade-offs; for design, counterfactuals, rule extraction, and concept-based explanations justify architecture choices; for development, LIME/SHAP, example-based explanations, and counterfactuals explain code generation and debug suggestions; for testing, LIME/SHAP and counterfactuals diagnose test failures; for deployment and monitoring, the same attribution and counterfactual tools explain anomaly flags and scaling decisions; and for maintenance, LIME/SHAP, counterfactuals, and attention mechanisms explain bug prediction, fixes, and summaries. The paper asserts that this phase-specific coverage makes it the first comprehensive survey of XAI across the whole software development lifecycle.","pith_inferences":["The paper's per-phase 'most effective technique' rankings are narrative assignments rather than measured comparisons; a natural extension is a benchmark that runs the same attribution and counterfactual methods on requirements-elicitation and maintenance tasks to test whether the phase-specific recommendations survive.","If the 68/16/8/8 distribution is accurate, the highest-leverage move for the field is not a new XAI algorithm but shifting evaluation effort toward requirements and design, where the paper's own evidence says research is thinnest.","The taxonomy-to-phase matrix could be operationalized as a decision procedure: given an AI-assisted task, pick the technique by explanation scope (local vs global) and stage (ante-hoc vs post-hoc) rather than by technique familiarity.","A competing systematic map that counts papers per phase with a different protocol would directly test the survey's novelty and gap claims, since the distribution numbers are the load-bearing evidence for where XAI is missing."],"forward_implications":["If the mapping is correct, practitioners can choose XAI techniques per phase instead of relying on a single universal method.","Requirements and design phases become concrete targets for XAI adoption and research, countering the maintenance-heavy distribution.","Per-phase XAI can expose latent requirements and bias early, through SHAP-style feature attribution on user interactions and existing systems.","Counterfactual explanations provide a uniform way to expose trade-offs across design, testing, deployment, and maintenance, because the same 'what would change the outcome?' question applies everywhere.","Standardized evaluation metrics and benchmarking structures are needed to compare XAI methods across phases, which the paper identifies as necessary future work."],"supporting_citations":[{"why":"Supplies the phase-distribution statistics (68% maintenance, 16% development, 8% each management and requirements) that motivate the gap analysis and the paper's phase-specific recommendations.","marker":"[7]"},{"why":"Supplies the XAI classification taxonomy (stage, scope, input/output) used to organize and describe all techniques in the survey.","marker":"[30]"},{"why":"Source for the descriptions of LIME, concept-based explanations, and SHAP's role in surfacing latent requirements and measuring feature contributions.","marker":"[12]"},{"why":"Provides the AI applications in requirements elicitation (document analysis, chatbot elicitation, data mining) that the phase-specific XAI mapping builds on.","marker":"[2]"},{"why":"Supplies the survey of software engineering for AI-based systems, including design-phase AI applications and the black-box challenges used throughout the paper.","marker":"[3]"},{"why":"Source for LLM applications in requirements, testing, and maintenance that the paper pairs with XAI techniques in those phases.","marker":"[15]"},{"why":"Prior metareview of XAI aligned to the software development process, which the paper extends by claiming full per-phase coverage.","marker":"[10]"},{"why":"Grounds the explainable-software-systems framing from requirements analysis to system evaluation, used to argue for early-phase XAI.","marker":"[5]"}],"fun_headline_variants":["XAI for every SDLC phase: First comprehensive survey","From requirements to evolution: Phase-specific XAI guide","Survey reveals XAI lopsided: 68% on maintenance, 8% on early phases","LIME, SHAP, counterfactuals: XAI mapped to each software phase","XAI phase-by-phase: Closing the software engineering gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's recommendations rest on the assumption that the cited phase-distribution statistics—68% of XAI-in-software-engineering research in maintenance, 16% in development, and 8% each in management and requirements—are accurate and representative, so the claimed gaps in requirements and design are real gaps rather than artifacts of how the prior review searched.","fun_headline_variants_meta":{"raw":{"variants":["XAI for every SDLC phase: First comprehensive survey","From requirements to evolution: Phase-specific XAI guide","Survey reveals XAI lopsided: 68% on maintenance, 8% on early phases","LIME, SHAP, counterfactuals: XAI mapped to each software phase","XAI phase-by-phase: Closing the software engineering gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001535,"raw_usage":{"total_tokens":6174,"prompt_tokens":1005,"completion_tokens":5169,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":5074}},"tokens_in":621,"tokens_out":5169,"duration_ms":38544,"temperature":1.0,"reasoning_tokens":5074,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:25:23.308652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could re-run the search protocol the paper describes (peer-reviewed articles from the last six years across the major computing and scientific databases, using the stated keywords) and count how many XAI-in-software-engineering papers address each SDLC phase. If the resulting distribution differs materially from 68/16/8/8, or if design and testing turn out to contain substantial XAI research, the paper's gap-driven phase recommendations lose their evidential basis.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the phase-distribution statistics (68% maintenance, 16% development, 8% each management and requirements) that motivate the gap analysis and the paper's phase-specific recommendations."},{"cited_title":"Explainable software systems: from requirements analysis to system evaluation,","cited_arxiv_id":null,"evidence_quote":"Source for the descriptions of LIME, concept-based explanations, and SHAP's role in surfacing latent requirements and measuring feature contributions."},{"cited_title":"by analogy","cited_arxiv_id":null,"evidence_quote":"Provides the AI applications in requirements elicitation (document analysis, chatbot elicitation, data mining) that the phase-specific XAI mapping builds on."},{"cited_title":"They tell which features matter most","cited_arxiv_id":null,"evidence_quote":"Supplies the survey of software engineering for AI-based systems, including design-phase AI applications and the black-box challenges used throughout the paper."},{"cited_title":"striped” in an image dataset considerably influences the classification of images into “zebra","cited_arxiv_id":null,"evidence_quote":"Grounds the explainable-software-systems framing from requirements analysis to system evaluation, used to argue for early-phase XAI."}],"review_version":1}