{"id":"3eb35e73-3006-4977-ad61-8ed292f88af4","arxiv_id":"2506.16878","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 77 studies shows quantum and quantum-inspired optimization for classical software engineering is concentrated in software operations and testing, with most solutions reformulated as single-objective problems.","lead":"This paper surveys 77 research papers on using quantum and quantum-inspired optimization methods to solve software engineering problems such as test minimization and task scheduling. It maps where these techniques are being applied, which algorithms dominate, and which software engineering areas remain untouched.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 14 search queries differ semantically across databases, so the 77-study corpus and all reported distributions may be biased; the claimed map needs a normalized re-run.","rationale":"The reader's weakest assumption is that the conclusions depend on corpus retrieval completeness and unbiasedness, specifically the non-equivalent per-database queries in Table 14. My reading agrees: every quantitative result in the paper is downstream of the 77-study corpus, and half of those studies came from snowballing, which cannot compensate for a systematic miss in the initial six-query search. I did not find a more load-bearing flaw. The RQ11 findings bullet that lists six uncovered activities as covered is a real internal inconsistency, but it is contradicted by Table 11 and the surrounding prose, so it is fixable and less central than the retrieval bias. The duplicate-reference issues are also peripheral. Because the concern is concrete, reproducible, and fixable by re-running the search, the conditional verdict is appropriate; no change from the reader's verdict is needed.","tokens_in":37986,"tokens_out":6678,"duration_ms":69031,"concrete_test":"Re-run the search on all six databases on a fixed date with a single normalized query: Set 1–4 terms restricted to Title/Abstract/Keywords and the 'software' term also restricted to Title/Abstract/Keywords, using the intended phrase 'artificial intelligence OR ai'. Record the retrieved set and check whether all 77 primary studies are present. Then re-run the Section 3.3 inclusion/exclusion screen on the new pool and recompute Figures 5 and 6. If any primary study is missing from the new retrieval, or if the share for software engineering operations (currently 55.84%) shifts by more than a few percentage points, the corpus is not robust and the headline map must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The SLR's central empirical claims (e.g., 55.84% of primary studies target software engineering operations; the field has not covered six SWEBOK activities) are computed from a corpus of 77 primary studies assembled from the 2083 records retrieved by the six database queries in Table 14. Those queries are not equivalent instantiations of Listing 1. In Set 3, Springer uses 'artificial intelligence OR ai', while Wiley and Web of Science use 'artificial AND intelligence OR ai', which parses as '(artificial AND intelligence) OR ai' and can match different documents. More importantly, the field scopes differ: IEEE and ACM search 'software' in Full Text, Scopus and Web of Science in ALL fields, Wiley in Abstract, Springer with no field restriction at all for any term set. Consequently, the retrieved pools are not generated by one reproducible Boolean query (e.g., IEEE 557, Scopus 68, Springer 779). The inclusion/exclusion screen in Section 3.3 and snowballing in Section 3.4 cannot fully repair a systematic gap in the seed set: backward/forward snowballing only reaches papers connected to the 39 search-selected studies, so a relevant study missed by all six queries would likely remain absent. If the seed set is biased, the reported activity/problem distributions and the claimed gaps are not a reliable baseline map. This is a correctness risk in the corpus construction, not a stylistic issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic literature review (SLR) of research applying quantum and quantum-inspired optimization to classical software engineering (SE) problems. The authors searched six digital databases on 13 January 2025, retrieved 2083 records (2186 raw hits), and after screening and snowballing selected 77 primary studies. They answer 11 research questions covering bibliometric trends, targeted SE activities and problems, optimization reformulations, solution types (quantum, hybrid, quantum-inspired), quantum algorithms used, and open challenges. The key reported findings are that 55.84% of primary studies target 'software engineering operations', that testing/quality/security make up most of the rest, that 91.67% of reformulated problems are single-objective, and that six SWEBOK activities remain uncovered. The authors claim this is the first SLR specifically on quantum and quantum-inspired optimization for classical SE.","tokens_in":38191,"tokens_out":8487,"duration_ms":73873,"significance":"The review addresses a genuinely useful gap: the SBSE community currently lacks a systematic map of how quantum and quantum-inspired optimization have been applied to classical SE problems. If the corpus is representative, the paper's maps (venues, activities, problems, solution types, algorithms) provide a baseline for future research and for the QSE roadmap. Strengths include the public artifact (GitHub) with extracted data, the detailed search and snowballing protocol, the use of SWEBOK v4.0 as an organizing framework, and the explicit RQ-by-RQ findings with reproducible figures and tables. The main value is the empirical mapping rather than a new technique. My reservations concern internal inconsistencies in the reported findings and the reproducibility of the corpus construction; these are correctable but must be addressed before the paper's claims can be relied upon.","major_comments":[{"comment":"The finding states that 'six out of 15 SE activities have been covered by the primary studies, i.e., Software Architecture, Software Construction, SE Process, SE Models and Methods, SE Professional Practice, and SE Economics.' Table 11 marks exactly these six activities as ✗ (not covered), and the body text (e.g., 'there is no primary study that targets software architecture optimization') confirms they are absent. The RQ11 finding is thus the opposite of the evidence it summarizes. Please correct the finding to 'not covered' and ensure the count (6 of 15 uncovered, 9 covered) is stated consistently.","section":"§7.2.1 (Findings of RQ11)"},{"comment":"The text says: 'For software failure prediction (see ▶ in Table 5), all seven studies are categorized under software security.' In Table 5, software failure prediction appears under 'software quality' (7 papers), not under 'software security'; the overlap marker ▶ indicates it is shared between software quality and software engineering operations (paper [91]). The sentence should say 'software quality'. As written, the paragraph misreports the paper's own classification and confuses the subsequent discussion of the overlap.","section":"§5.2 (paragraph after Table 5)"},{"comment":"The six database queries are not semantically equivalent instantiations of Listing 1. For the technique set (Set 3), Springer uses 'artificial intelligence OR ai', whereas Wiley and Web of Science use 'artificial AND intelligence OR ai', which parse differently under standard precedence. The field scope for 'software' also varies: full text (IEEE, ACM), ALL fields (Scopus, WoS), Abstract (Wiley), and no field restriction (Springer). These differences mean the 2083-retrieved pool is not the output of a single reproducible Boolean query, and the subsequent screening and snowballing cannot fully repair a systematic gap in the seed set. Since the paper's central empirical claims (e.g., the 55.84% share for SE operations, the six uncovered activities) are computed from this corpus, the authors should either show that the query variants are equivalent after operator and field semantics are accounted for, or re-run the search with normalized queries (and report the new corpus) to demonstrate the stability of the reported distributions.","section":"§3.2 / Appendix Table 14"},{"comment":"The text states 'among the 17 hybrid quantum optimization solutions and 15 purely quantum optimization solutions', but Table 7 reports Hybrid 16 and Quantum 15. The discrepancy affects the interpretation of RQ9 counts (e.g., Figure 12 shows 32 quantum-approach instances). Please reconcile the count and state clearly how papers with multiple approaches (e.g., [20], [105], [130]) are counted.","section":"§6.3 vs Table 7"}],"minor_comments":[{"comment":"References [1] and [2] are identical (Abbas et al., 2024), as are [102] and [103] (Murillo et al., 2025). Please remove duplicates or renumber.","section":"References"},{"comment":"The sentence 'surpassing 27.27% in conferences, 10.39% in workshops, and 3 in open-access archive entries' mixes a count with percentages; state '3 papers (3.90%)' for consistency.","section":"§4.3 (Findings of RQ3)"},{"comment":"The word 'reproducability' should be 'reproducibility'.","section":"§1 (Introduction)"},{"comment":"The phrase 'were each quantum solution is associated' should be 'where each quantum solution is associated'.","section":"§7.2.2"},{"comment":"The claim 'we present the first SLR' should be qualified relative to Mandal et al. [96] and Murillo et al. [103], which are also 2025 SLRs with partial overlap; the paper already discusses the differences in scope in Section 9.1, but the abstract and intro should be worded to avoid an overclaim.","section":"§1 & Abstract"},{"comment":"Table 14 sums to 2186 raw hits, while the text reports 2083 unique papers after deduplication. Please clarify explicitly whether the 2083 number already excludes duplicates from the six-source union, as the current wording in Section 3.3 ('yielding 2083 papers') is ambiguous.","section":"§3.3 / Table 14"}],"recommendation":"major_revision","confidential_remarks":"The search-query normalization issue is the most serious methodological point; if the authors cannot justify the query variants or demonstrate corpus stability, the empirical claims would need to be revisited. The RQ11 finding is a simple but damaging error that must be fixed. Otherwise, the paper is a solid SLR with a useful artifact and detailed reporting; it fits the journal's scope and, after the corrections, could become a reliable baseline reference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. This is a useful SLR and, as far as I can tell, the first one aimed specifically at quantum and quantum-inspired optimization applied to classical SE. The authors assembled 77 studies, mapped them to SWEBOK activities, separated quantum, hybrid, and quantum-inspired approaches, and posted their data. That's a real service to the SBSE community; I'd hand this to a PhD student starting in the area.\n\nThe good stuff: the corpus is new, the classification is sensible enough, and they made a credible effort to differentiate themselves from Mandal et al.'s broader QSE review. The artifact helps reproducibility.\n\nThe soft spots, in order of severity.\n\nFirst, the search queries in Table 14 are not equivalent instantiations of Listing 1. Springer uses 'artificial intelligence OR ai' while Wiley and Web of Science use 'artificial AND intelligence OR ai'; 'software' gets searched in full text on some engines, in abstract on Wiley, and in no field at all on Springer. That means the 2083 initial records are not the output of a single reproducible query. Since snowballing only walks the neighborhood of the 39 search-selected papers, a relevant paper missed by all six queries would almost certainly be absent from the final 77. The authors do not mention this in their threats-to-validity section. I don't think this sinks the map—for an emerging field, an approximate corpus is still informative—but the reported distributions (55.84% for SE operations, for example) need more caution, and the authors should either normalize the queries, run a sensitivity analysis, or explicitly label this as a threat.\n\nSecond, there's a factual slip in Section 7.2.1. The findings for RQ11 say six SE activities 'have been covered' and then list Software Architecture, Software Construction, SE Process, etc., which Table 11 marks as not covered. The list is the right one for 'not covered'; the sentence is inverted. Simple fix.\n\nThird, references [1] and [2] are identical, and so are [102] and [103]. That's just sloppy bibliography work, but in an SLR it matters more than elsewhere because readers will chase those citations.\n\nI don't see a deeper methodological problem: the inclusion criteria are reasonable, the extraction rules are documented, and the conclusion that the field is pre-paradigmatic and concentrated in SE operations and testing follows from the data. The circularity concern about the authors' own papers doesn't bother me.\n\nSo: who should read this? SBSE researchers, newcomers to quantum-inspired optimization, and anyone writing a research agenda for quantum x SE. It deserves a serious referee. Send it to peer review with a request for minor-to-major revision on the query issue and the copyediting.","headline":"A useful first map of quantum(-inspired) optimization for classical SE, but the inconsistent search queries and a few sloppy errors should be fixed before the numbers are relied on.","tokens_in":38772,"tokens_out":3222,"would_cite":true,"duration_ms":33137,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents the first systematic literature review of quantum and quantum-inspired optimization applied to classical software engineering problems, finding 77 primary studies and showing where the field is concentrated and where…","keywords":["quantum optimization","quantum-inspired algorithms","software engineering","systematic literature review","search-based software engineering","single-objective reformulation","QUBO","test suite minimization"],"falsifier":"An independent reviewer could rerun the exact queries from the paper's appendix against the same six databases and apply the inclusion criteria; if the deduplicated total does not reproduce near 2083, or if the added studies shift the share of software engineering operations well below 55.84% and fill one of the reportedly empty activity areas, the review's quantitative map would be weakened.","tokens_in":37727,"feed_emoji":"📊","tokens_out":6483,"duration_ms":61421,"temperature":0.7,"pith_summary":"This paper sets out to establish a comprehensive, reproducible map of research that applies quantum or quantum-inspired optimization to classical software engineering problems. It claims to be the first systematic literature review of that specific intersection, screening 2083 candidate publications down to 77 primary studies. The review reports that the literature is concentrated on software engineering operations and software testing, that most optimization problems are reformulated into single-objective form even when the original SE problem had multiple objectives, and that quantum-inspired algorithms running on classical hardware currently outnumber pure quantum solutions. If the map is accurate, it gives the software engineering community a baseline for where quantum optimization already contributes and where it does not.","feed_headline":"First map of quantum optimization in software engineering","feed_subtitle":"From 2,083 candidates, 77 primary studies show testing and operations on top; six SE activities remain untouched.","key_machinery":"The carrying object is the systematic literature review protocol itself: searches across six digital databases using a combined five-set Boolean query, explicit inclusion and exclusion criteria, and two rounds of backward and forward snowballing. This protocol turns 2083 raw hits into 77 primary studies and produces the tables that support every gap claim. A second distinctive device is the classification scheme that distinguishes the original SE problem's objectives from the objectives in the reformulated optimization problem, which lets the paper quantify a simplification tendency: multi-objective and many-objective SE problems are frequently collapsed into single-objective quantum formulations. The same protocol uses a standardized software engineering body of knowledge to label activities, which is what makes the claim about uncovered areas a clean result.","core_discovery":"The paper's central discovery is the distributional shape of a young interdisciplinary field. It reports that, among 77 primary studies published between 2014 and 2025, 55.84% target software engineering operations, 20.78% target software testing, 12.99% target software quality, and 10.39% target software security; scheduling and test suite minimization are the two most frequently studied problems. It also reports that 65.75% of the SE problems are single-objective as originally posed, but after reformulation 91.67% of the optimization problems are single-objective, and that 48 of 80 solution instances are quantum-inspired approaches executed on classical hardware, with 17 hybrid and 15 purely quantum. Among the 32 approaches that use quantum algorithms, quantum annealing paired with a QUBO formulation dominates, and six of the fifteen activity chapters in the adopted software engineering body of knowledge have no primary study at all. The paper presents this as evidence that the field is real, growing, and still highly uneven.","pith_inferences":["If the reported distribution is representative, the gap list doubles as a research prioritization menu: testing and operations are crowded, while requirements selection, design, construction, and process improvement are nearly empty.","The observation that most primary studies appear outside SE venues implies an indexing problem that the review itself only partly solves; keeping the map current would require a cross-disciplinary watch rather than a single-community search.","A testable extension would repeat the protocol at a later date and track whether the share of hybrid and purely quantum solutions grows as hardware becomes more accessible, and whether the six uncovered activities gain studies.","The original-versus-reformulated objective count could become a standard reporting metric for future papers in this area, making the community's simplifying choices visible and comparable."],"forward_implications":["The field now has a reproducible baseline map; later reviewers can rerun the published queries and update the counts without starting from zero.","Six SE activity areas — architecture, construction, process, models and methods, professional practice, and economics — are essentially open terrain for quantum and quantum-inspired optimization research.","Because 91.67% of reformulated problems are single-objective, the trade-off structure of real multi-objective SE problems is largely unexplored; working on multi-objective quantum formulations is a direct research opening.","The dominance of quantum annealing with QUBO means that advances in gate-based algorithms or in hybrid classical-quantum decomposition could substantially change the field's center of gravity.","Most current solutions are quantum-inspired and run on classical hardware, so near-term practical use does not depend on waiting for fault-tolerant quantum computers."],"supporting_citations":[{"why":"Supplies the software engineering body of knowledge used to label SE activities and to identify uncovered areas.","marker":"[153]"},{"why":"Recent SLR whose SDLC and activity terminology informs the search term set for SE activities.","marker":"[89]"},{"why":"Defines single-, multi-, and many-objective optimization categories used to classify SE problems and their reformulations.","marker":"[124]"},{"why":"Supplies the many-objective definitions used in the same classification.","marker":"[83]"},{"why":"Provides the search-based software engineering context and the observation that many SE problems are multi-objective.","marker":"[58]"},{"why":"The most closely related prior SLR on quantum software engineering; the paper distinguishes its scope from it.","marker":"[96]"},{"why":"One of the high-quality recent SE SLRs cited to justify the choice of six digital databases.","marker":"[168]"},{"why":"Second recent SE SLR cited for the same database-selection justification.","marker":"[85]"}],"fun_headline_variants":["Quantum SE map finds six untouched activities","SE quantum survey: testing and ops on top","Most quantum SE work is quantum-inspired, not quantum","77 studies, one clear gap in quantum SE research","First quantum SE survey reveals uneven landscape"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusions stand on the assumption that its literature retrieval was complete and unbiased; if the database queries or the human inclusion and exclusion decisions missed a substantial share of relevant work, the reported percentages and the claimed gaps would no longer represent the field.","fun_headline_variants_meta":{"raw":{"variants":["Quantum SE map finds six untouched activities","SE quantum survey: testing and ops on top","Most quantum SE work is quantum-inspired, not quantum","77 studies, one clear gap in quantum SE research","First quantum SE survey reveals uneven landscape"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1711,"prompt_tokens":933,"completion_tokens":778,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":709}},"tokens_in":549,"tokens_out":778,"duration_ms":8566,"temperature":1.0,"reasoning_tokens":709,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:16:30.336341+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent reviewer could rerun the exact queries from the paper's appendix against the same six databases and apply the inclusion criteria; if the deduplicated total does not reproduce near 2083, or if the added studies shift the share of software engineering operations well below 55.84% and fill one of the reportedly empty activity areas, the review's quantitative map would be weakened.","supporting_citations":[{"cited_title":"2024.Guide to the Software Engineering Body of Knowledge (SWEBOK Guide), Version 4.0","cited_arxiv_id":null,"evidence_quote":"Supplies the software engineering body of knowledge used to label SE activities and to identify uncovered areas."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines single-, multi-, and many-objective optimization categories used to classify SE problems and their reformulations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The most closely related prior SLR on quantum software engineering; the paper distinguishes its scope from it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the high-quality recent SE SLRs cited to justify the choice of six digital databases."}],"review_version":2}