{"id":"64d9167c-6eb1-4aa8-a43b-f432c2d6dd74","arxiv_id":"2506.16101","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that categorizes 122 papers on regression test optimization and argues ROS-based autonomous systems need new semantic and neurosymbolic approaches.","lead":"This paper reviews 122 studies on regression testing optimization and maps them to test prioritization, minimization, and selection for ROS-based autonomous systems. It argues that existing methods fall short for autonomous systems and proposes new research directions such as semantic-aware coverage metrics and neurosymbolic reasoning.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first comprehensive ROSAS-tailored survey' claim is unverifiable and internally strained: no reproducible search protocol or included-paper list is provided, and most of the 122 studies are traditional or DNN-testing work by the authors' own classification.","rationale":"The reader's weakest assumption—completeness of the literature search—is the central issue, and I agree that the missing search protocol and missing list of included studies make the 'first comprehensive' claim unverifiable. I would add an internal-consistency concern: the manuscript's own classifications (Table 4: 88/122 traditional software testing) and its admission in §4.4 that most emerging techniques target deep learning models rather than ROSAS directly mean that even a complete corpus might not be a 'tailored' corpus. This does not change the conditional verdict: the survey is useful and largely coherent, but the authors should either provide an appendix with full study-level metadata and direct-ROSAS labels, or explicitly reframe the contribution as a transferability-oriented survey rather than the first ROSAS-tailored comprehensive review. The concrete test above would discriminate between these two framings. No fatal mathematical or internal logical error was found; the concern is about claim substantiation and framing, which is addressable in revision.","tokens_in":40657,"tokens_out":3377,"duration_ms":44762,"concrete_test":"Compile a complete appendix listing all 122 included papers with: (a) full citation, (b) source database, (c) optimization category, and (d) direct-ROSAS applicability—does the paper implement or evaluate on ROS/ROS2, a simulated robot or drone, or an autonomous vehicle, or is it only transferable? Then independently rerun the reported search in IEEE Xplore, ACM DL, SpringerLink, ScienceDirect, Scopus, and Google Scholar using the exact keyword combinations from §2.2. If any ROSAS-specific regression-testing optimization study is absent, or if fewer than half of the 122 studies are direct-ROSAS, the 'first comprehensive ROSAS-tailored survey' claim should be weakened to 'a survey of transferable regression-testing optimization techniques with a ROSAS research agenda.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is 'the first comprehensive survey systematically reviewing regression testing optimization techniques tailored for ROSAS' (Abstract, §1.3, Conclusion). That claim depends on two conditions: (1) the 122-study corpus is complete for regression-testing optimization in ROSAS, and (2) the corpus actually concerns ROSAS rather than merely being transferable to it. Neither condition is verifiable from the manuscript. §2.2 reports 2,100 candidates reduced to 122 but gives no full boolean search strings, no screening flow, no exclusion log, and no list of the 122 included papers; Table 5 only reports database provenance. More seriously, internal evidence strains the 'tailored' part. §4.4 states that most emerging techniques 'focus primarily on ... deep learning models rather than directly targeting the characteristics of ROSAS,' and Table 4 classifies 88 of 122 studies as 'Traditional Software Testing.' Inclusion criterion §2.3 explicitly admits studies 'with potential transferability' to ROSAS. If a large share of the corpus is general regression-testing optimization or DNN-focused testing that never touches ROS/ROS2, then the survey is a cross-domain roadmap with a ROSAS framing, not a comprehensive review of techniques tailored to ROSAS. That would weaken the foundational-reference claim even if no prior survey was missed. The missing evidence that would settle this is a per-paper mapping between each included study and the ROSAS-specific challenges listed in §3.2, plus a complete enumeration of the corpus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic literature review of regression testing optimization techniques for ROS-based autonomous systems (ROSAS). It reports a corpus of 122 studies, categorizes them into test case prioritization, test suite minimization, and test case selection, and introduces a taxonomy of techniques applicable to ROSAS. The survey also identifies challenges (coverage metrics, multi-modal data, non-determinism, and oracle problems) and proposes future directions built around frame-to-vector coverage metrics, multi-source foundation models, and neurosymbolic reasoning. The paper claims to be the first comprehensive survey of regression testing optimization techniques tailored specifically to ROSAS.","tokens_in":40876,"tokens_out":5526,"duration_ms":63111,"significance":"If the central claim were fully supported, the survey would be a useful foundational reference and roadmap for a growing research area. The paper's strengths are its broad corpus, an internally consistent set of categorization tables (the overlap in Table 2 is correctly explained, and Tables 3, 4, and 5 sum to 122), and a clear RQ-driven structure. The taxonomy in Figure 2 is a reasonable organizing device. However, the two load-bearing claims—'first comprehensive survey' and 'techniques tailored for ROSAS'—are not currently verifiable from the manuscript because the search protocol is not reproducible and the corpus is not demonstrably ROSAS-specific. The survey is potentially valuable as a cross-domain roadmap, but the stated contribution needs to be either substantiated with additional evidence or reframed.","major_comments":[{"comment":"The literature search is not reproducible. Section 2.2 reports 2,100 candidate papers reduced to 122 after applying the criteria in Section 2.3, but the manuscript does not provide the full boolean search strings used for each database, the date range of the search, a screening flow diagram, an inclusion/exclusion log, or a list of the 122 included studies. Table 5 reports only the database provenance of the final set, not which papers were included or excluded at each stage. Because the paper's 'first comprehensive survey' claim depends on the completeness of this search, the missing protocol is a load-bearing issue. The authors should provide a complete reproducibility package, including the query strings, screening decisions, and a numbered list of included studies.","section":"§2.2–2.3, Table 5"},{"comment":"The claim that the reviewed techniques are 'tailored for ROSAS' is not supported by the evidence presented. Table 4 classifies 88 of 122 studies as 'Traditional Software Testing,' and the inclusion criterion in Section 2.3 explicitly admits studies with 'potential transferability' to ROSAS. Section 4.4 then acknowledges that most emerging techniques 'focus primarily on the testing optimization of deep learning models rather than directly targeting the characteristics of ROSAS.' Without a per-paper mapping between each included study and the specific ROSAS challenges it addresses, the abstract's claim of a survey 'tailored for ROSAS' overstates what the corpus demonstrates. The authors should either add such a mapping (e.g., as an appendix table showing which studies address multi-modal data, asynchronous communication, real-time constraints, or safety requirements) or reframe the contribution as a cross-domain review with transferability analysis.","section":"§2.3, Table 4, §4.4"}],"minor_comments":[{"comment":"The text says 'In Chapter 3, we provide...' when it should say 'In Section 3'; the chapter/section terminology should be made consistent throughout.","section":"§4 opening paragraph"},{"comment":"The future research directions section relies heavily on the authors' own prior publications (STRaP [13], Zheng et al. [16], Neurostrata [65], GARL [14], Recover [194]) to define the open problems. While self-citation is not inherently improper, the survey should either include independent validation of these directions or explicitly disclose the relationship, so the proposed roadmap is not perceived as circular.","section":"§5.3"},{"comment":"The exclusion criteria mention 'non-peer-reviewed publications' but the final corpus appears to include a replication package entry (reference [125], 'LTM') and possibly other non-archival sources; this inconsistency should be clarified.","section":"§2.3"},{"comment":"The paper does not include a threats-to-validity or limitations subsection discussing the selection bias inherent in the chosen databases and keyword set, which would help readers calibrate the comprehensiveness claim.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a potentially useful survey, but the 'first comprehensive and ROSAS-tailored' claim needs stronger support before publication. I would ask for a reproducibility appendix (search strings, screening flow, included-study list) and either a per-paper ROSAS relevance mapping or a modified contribution statement. I do not see evidence of misconduct, but the concentration of self-citations in the future-directions section is worth monitoring in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a serviceable survey of regression testing optimization, but the 'first comprehensive survey tailored for ROSAS' claim is the weakest part of it. The taxonomy (TCP/TSM/TCS) is clear, the section on ROSAS-specific challenges is sensible, and the selection of 122 papers covers the standard literature you'd expect. It's genuinely useful as a starting point for someone entering this subfield.\n\nWhat's new is mostly the framing: previous surveys covered traditional software, CPS, or industrial practice, and none positioned the work explicitly around ROS. The authors also do a decent job of mapping traditional techniques onto ROSAS challenges, even if the mapping is often suggestive rather than deep. The future directions (frame-to-vector coverage, foundation models, neurosymbolic reasoning) are reasonable research bets, though they lean heavily on the authors' own prior papers.\n\nThe soft spots are real. The methodology in §2.2 is not reproducible: no actual search strings, no screening log, no list of the 122 included studies. That's a serious omission for a survey that claims comprehensiveness. More troubling, the internal evidence contradicts the 'tailored for ROSAS' framing. Table 4 shows 88 of 122 papers are traditional software testing, and §4.4 admits most emerging techniques focus on DL models rather than ROSAS characteristics. The inclusion criterion explicitly allows studies 'with potential transferability.' So the corpus is mostly general regression testing work, not ROSAS-specific research. The claim that this is the first comprehensive survey of techniques tailored to ROSAS is therefore overstated; it's better described as a survey of regression testing optimization with a ROSAS-oriented lens.\n\nNone of this kills the paper. The taxonomy and challenge analysis are still valuable. But the authors need to either soften the claim or provide the per-paper mapping that shows how each included study actually addresses ROSAS specifics. Without that, the foundational-reference status is not earned.\n\nMy take: worth engaging with, but needs revision before I'd trust it as the definitive survey. Send it to peer review, but expect major comments on methodology and scope.","headline":"Useful but overclaims its 'first comprehensive ROSAS-tailored' status; the taxonomy and challenge analysis are solid, but the methodology and corpus mapping need major strengthening.","tokens_in":41449,"tokens_out":1517,"would_cite":false,"duration_ms":17831,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first systematic survey devoted to regression testing optimization for ROS-based autonomous systems, sorting 122 studies into prioritization, minimization, and selection and identifying the gaps that remain.","keywords":["Robot Operating System","regression testing","test case prioritization","test suite minimization","test case selection","autonomous systems","safety-critical software","systematic survey"],"falsifier":"Re-running the reported search across the eight listed databases with the stated keyword combinations and locating a peer-reviewed survey or technique paper published before mid-2025 that specifically addresses regression testing optimization for ROS-based autonomous systems and is not among the 122 included studies would falsify the first-and-complete claim.","tokens_in":40401,"feed_emoji":"🤖","tokens_out":6300,"duration_ms":67884,"temperature":0.7,"pith_summary":"Regression testing re-runs tests after changes to catch regressions; this paper argues that doing it efficiently for ROS-based autonomous systems (ROSAS) is a distinct problem that existing optimization research has not systematically addressed. It claims to be the first survey dedicated to this topic, reviewing 122 studies and sorting them into three families: test case prioritization (ordering tests to reveal faults sooner), test suite minimization (cutting redundant tests), and test case selection (running only tests affected by a change). The survey maps each family onto the specific stresses of ROSAS—multi-modal sensor input, asynchronous distributed nodes, non-deterministic behavior, and real-time safety constraints—and concludes that most existing techniques assume deterministic, single-modal, non-real-time software. A sympathetic reader would take the contribution as a structured map of the field, a diagnosis of why current methods fall short, and a set of proposed research directions.","feed_headline":"First survey maps regression testing for ROS robots","feed_subtitle":"122 studies sorted into prioritization, minimization, and selection—and why they fall short for autonomous systems.","key_machinery":"Two devices carry the argument. The first is a three-part taxonomy: test case prioritization, test suite minimization, and test case selection, with each category subdivided by technique family, such as coverage-, history-, model-, search-, learning-, and confidence-based prioritization. The second is a ROSAS-specific challenge lens—multi-modal data, asynchronous and non-deterministic behavior, real-time and safety-critical constraints, and missing oracles—through which every surveyed method is judged. The taxonomy makes the gap visible by showing a large body of optimization work that assumes static, deterministic software, and it exposes where ROSAS needs semantics-aware, environment-aware, and adaptive methods. The paper's proposed future machinery, frame-to-vector coverage, which vectorizes sensor frames into semantic representations to measure what was actually exercised, is offered as the natural replacement for code-level coverage.","core_discovery":"The paper's central claim is that regression testing optimization for ROS-based autonomous systems is a genuinely different research territory, and that this review is the first to chart it. The organizing result is a taxonomy of 122 representative studies into test case prioritization, test suite minimization, and test case selection, with hybrid and emerging methods noted; the survey further identifies why each category transfers poorly to ROSAS. The headline diagnosis is that traditional coverage and dependency assumptions break down when tests must account for multi-modal sensor fusion, asynchronous publish-subscribe communication, stochastic learned components, and hard real-time and safety constraints. To close the gap, the paper proposes frame-to-vector semantic coverage metrics, integration of multi-source foundation models, and neurosymbolic reasoning, which combines learned perception with explicit logical rules, for detecting small but safety-critical violations.","pith_inferences":["Beyond the paper: the frame-to-vector coverage idea could be tested immediately on public ROS2 driving datasets by measuring whether semantic-vector similarity predicts fault detection better than statement coverage does.","Beyond the paper: because ROSAS generate rich logs, sensor streams, and build metadata, data-driven selection methods from continuous-integration testing could be adapted to ROSAS with relatively little new machinery.","Beyond the paper: if the survey's search is truly complete, the near absence of ROSAS-specific optimization papers before 2021 suggests the field is young enough that early methodological standards, such as common benchmarks and evaluation metrics, are still up for grabs."],"forward_implications":["A researcher entering the area gets a shared map: new techniques can be positioned as test case prioritization, minimization, selection, or hybrid, and compared against the same ROSAS challenge dimensions.","The gaps the survey identifies become concrete problem statements: semantic coverage metrics, multi-modal test data handling, and oracle construction are the places where progress is most needed.","The proposed future directions—frame-to-vector coverage, multi-source foundation models, and neurosymbolic reasoning—provide starting hypotheses for new methods rather than fully validated solutions.","For practitioners, the taxonomy doubles as a screening tool: it shows which traditional optimization methods are likely to transfer poorly to ROSAS and why.","If the 'first survey' claim holds, later work will likely cite this paper as the baseline against which new ROSAS regression-testing surveys are measured."],"supporting_citations":[{"why":"Documents challenges of testing robotic systems, motivating the need for a ROSAS-specific regression testing optimization review.","marker":"[18]"},{"why":"Classic survey of regression testing minimization, selection, and prioritization that the paper uses as the field's baseline and gap reference.","marker":"[19]"},{"why":"Systematic review of search-based test suite reduction, used to show prior work targets traditional software.","marker":"[20]"},{"why":"Systematic literature review of prioritization and regression test selection, serving as a comparison for the traditional-software focus.","marker":"[21]"},{"why":"Surveys machine-learning-based test selection and prioritization, used to show prior surveys omit ROSAS.","marker":"[22]"},{"why":"Surveys industrial testing optimization, used to contrast with autonomous-system needs.","marker":"[23]"},{"why":"Surveys test generation, selection, and prioritization for cyber-physical systems, the closest prior work but not ROSAS-specific.","marker":"[24]"},{"why":"Provides the systematic literature review guidelines that define the search and selection methodology.","marker":"[25]"},{"why":"Provides the snowballing guidelines that justify the forward and backward citation expansion.","marker":"[26]"}],"fun_headline_variants":["First survey of ROS regression testing: 122 studies sorted","ROS regression testing: why traditional methods break","New taxonomy exposes weaknesses in ROS test optimization","How to prioritize, minimize, select tests for ROS systems"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the database search, keywords, and snowballing captured every relevant prior study; if even one earlier survey or substantial body of work on regression testing optimization for ROSAS was missed, the claim to be the first survey of this area collapses.","fun_headline_variants_meta":{"raw":{"variants":["First survey of ROS regression testing: 122 studies sorted","ROS regression testing: why traditional methods break","New taxonomy exposes weaknesses in ROS test optimization","How to prioritize, minimize, select tests for ROS systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1315,"prompt_tokens":932,"completion_tokens":383,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":322}},"tokens_in":548,"tokens_out":383,"duration_ms":4592,"temperature":1.0,"reasoning_tokens":322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:43:44.546340+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the reported search across the eight listed databases with the stated keyword combinations and locating a peer-reviewed survey or technique paper published before mid-2025 that specifically addresses regression testing optimization for ROS-based autonomous systems and is not among the 122 included studies would falsify the first-and-complete claim.","supporting_citations":[],"review_version":1}