{"id":"6ad16bc7-b6d1-48a3-8851-bd7fea5db112","arxiv_id":"2501.08909","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic mapping study of 34 primary studies on XR software testing, providing a classification of topics, test facets, tools, datasets, and open research gaps.","lead":"This paper reviews 34 prior studies on software testing for extended reality (XR) applications and organizes them by topic, test activity, technique, and evaluation approach. It is useful as a structured map for researchers and practitioners who want to know what XR testing work already exists and where the gaps are.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The automated search string requires both a full XR term and its acronym in the title, so studies using only one form are systematically missed; several primary studies fall in this category.","rationale":"The reader identified search completeness as the weakest assumption, focusing on the use of OpenAlex as the sole database and on temporal bias. That concern is reasonable, but the more load-bearing and more concrete issue is the composition of the search string itself. The paper's own test set for validating the search string consists of eight studies that all happen to include both a full term and its acronym in their titles, for example 'VRTest: An Extensible Framework for Automatic Testing of Virtual Reality Scenes', so the test passes even though the restrictive AND condition excludes relevant titles containing only one form. Several final primary studies, such as PS9, PS10, PS12, and PS19, have titles without acronyms, implying they were found only through snowballing. Because the 34 primary studies are the empirical basis for every research question, a systematically incomplete automated search undermines the generalizability of the 'first systematic mapping study' claim unless snowballing demonstrably recovers the missing stratum. The proposed test would settle whether the search string is a true bottleneck: if a relaxed re-run yields no new relevant studies beyond the existing 34, the original verdict can stand; if it yields additional relevant studies, the paper should be revised to re-run the search or to temper its completeness claims. Given the importance of this check to the central claim, the verdict should be conditional rather than unchanged.","tokens_in":30633,"tokens_out":4376,"duration_ms":43855,"concrete_test":"Run the original search string in OpenAlex over the same date range, then run a relaxed variant in which the acronym requirement is relaxed to an OR, for example ($XR OR $XRacr) AND ($T OR $B), where $XR includes full terms and $XRacr includes acronyms in titles. Apply the paper's inclusion and exclusion criteria to both result sets and record how many of the 34 included primary studies are retrieved by the original string versus only by the relaxed string. If a substantial fraction, such as more than 20%, of the final primary studies are not retrieved by the original automated search and depend on snowballing, then the search string itself is a systematic source of incompleteness, and the mapping's coverage claim requires re-verification or explicit reporting of per-study sources.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim of presenting the first systematic mapping study of XR software testing depends on a comprehensive literature search. The search string in §4.1.2 is (($XR) AND ($XRacr)) AND ($T OR $B), where $XR contains full phrases such as \"virtual reality\" and $XRacr contains only acronyms such as VR, AR, XR, and MR, both restricted to titles. This structure requires a study's title to contain both a full term and the corresponding acronym, even though the full term alone or the acronym alone should suffice for relevance. Consequently, any relevant study whose title uses only \"Virtual Reality\" or only \"VR\" is excluded from the automated search. This is not hypothetical: several included primary studies have titles without acronyms, including PS9 \"Automated testing of virtual reality application interfaces\", PS10 \"Automated Usability Evaluation of Virtual Reality Applications\", PS12 \"Blackbox Testing on Virtual Reality Gamelan Saron Using Equivalence Partition Method\", and PS19 \"Less Cybersickness, Please: Demystifying and Detecting Stereoscopic Visual Inconsistencies in Virtual Reality Apps\". These studies must have been found by backward snowballing rather than by the automated search. The paper reports that snowballing identified 53 additional studies but does not disclose which of the final 34 primary studies came from the automated search versus snowballing, nor whether snowballing fully compensates for the systematic exclusion. If the automated search misses an entire class of relevant literature, the final corpus of 34 studies may be incomplete, and the reported trends, counts, and research-gap analysis derived from that corpus are at risk. This is a concrete internal methodological flaw that goes beyond the acknowledged OpenAlex coverage limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic mapping study (SMS) of software testing for Extended Reality (XR) applications, following the Petersen et al. (2015) guidelines. The authors define three research questions (with sub-questions) covering research status, test facets, and evaluation methodologies. They search OpenAlex with a structured query, screen candidates with three independent authors, perform three iterations of backward snowballing, and arrive at 34 primary studies. The selected studies are classified by venue, topic, research type, immersive technology, test activity, test objective, test target, test level, test type, test technique, evaluation metrics, and evaluation environment. The paper also catalogs datasets and tools referenced in the primary studies and proposes future research directions such as interaction formalisation, oracle automation, XR-specific testing, software-centric usability testing, and AI for XR testing.","tokens_in":30857,"tokens_out":5214,"duration_ms":55500,"significance":"If the search and selection process is complete, this is a useful contribution: it provides the first structured overview of XR software testing, an auditable classification scheme, a public repository of extracted data and tools, and a set of research directions grounded in the primary studies. The methodology has notable strengths: explicit inclusion/exclusion criteria, piloting of the selection criteria, independent screening by three authors, a data extraction form, and a dedicated threats-to-validity section. The paper also ships a public repository, which supports reproducibility. However, the significance is conditional on the completeness of the literature search; the search-string construction described in §4.1.2 has a structural flaw that directly affects which studies could be retrieved, and the paper does not report the provenance of the final 34 studies between the automated search and snowballing. This weakens confidence in the 'first comprehensive mapping study' claim and in the gap analysis built on the corpus.","major_comments":[{"comment":"The reported search string ($XR AND $XRacr) AND ($T OR $B), with $XRacr restricted to titles and $XR allowed in title or full text, requires every retrieved study to have both a full XR phrase and its acronym, and the acronym must appear in the title. This systematically excludes any relevant study whose title uses only the full phrase or only the acronym. This contradicts the search evaluation in §4.1.3: the test set includes Bierbaum et al. (PS9) with title 'Automated testing of virtual reality application interfaces' and Rafi et al. (PS23) with title 'PredART: Towards Automatic Oracle Prediction of Object Placements in Augmented Reality Testing', neither of which contains 'VR' or 'AR' in the title, so they cannot be matched by the stated query. Several final primary studies also have acronym-free titles (PS9, PS10, PS12, PS19), meaning they could only have been recovered by snowballing. The paper must either correct the reported search string, re-run the search with a disjunctive structure (full-term OR acronym), or provide a per-study provenance table and a recall analysis that justifies the current corpus.","section":"§4.1.2 and §4.1.3"},{"comment":"The flow from 1167 initial studies to 135 after selection and then to 34 final primary studies is not fully auditable. The text states that backward snowballing identified 53 additional studies, but it does not report how many of the final 34 primary studies came from the automated search versus snowballing, whether the 53 were added before or after full-text screening, or how many of the 53 were ultimately included. Because snowballing (Stage 3) starts from the 135 studies retained after the automated search, it cannot recover relevant studies that are not cited by those retained studies. A PRISMA-style flow diagram with per-path counts and a per-study provenance table is needed to support the completeness claim and to allow replication.","section":"§5 and Figure 5"},{"comment":"The threats-to-validity section acknowledges the single-database risk and the temporal bias of OpenAlex, but it does not acknowledge the structural limitation of the ANDed acronym requirement in the search string, which is a stronger and more specific threat to completeness. The statement that three iterations of snowballing 'significantly reduced the likelihood of missing important contributions' is only as strong as the starting set; if the automated search under-covers a sub-community whose papers are not co-cited with the retained studies, snowballing cannot compensate. The threats discussion should be revised to include this search-string limitation, and the 'first comprehensive' claim should be tempered to 'to the best of our knowledge' with an explicit statement of the residual completeness risk.","section":"§4.4"}],"minor_comments":[{"comment":"Please specify the exact OpenAlex query fields used (e.g., title.search, fulltext.search) and the full query string, since OpenAlex's coverage of full text varies by source and this affects reproducibility.","section":"§4.1.2"},{"comment":"'We want to know that multiple studies...' should read 'We note that multiple studies...'.","section":"§6.2.2"},{"comment":"Footnote 18 contains a typo: 'in their tile' should be 'in their title'.","section":"§5.1.1"},{"comment":"The word cloud in Figure 9 is difficult to audit and is not reproducible; a frequency table of test activities would be preferable for a mapping study.","section":"§5.2.1"},{"comment":"The statement about the Unity List dataset being no longer accessible is useful, but it would be clearer to state the access date and the URL checked, as done for other resources.","section":"§6.2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has clear strengths in transparency and methodology, but the search-string inconsistency is a load-bearing issue for a systematic mapping study. The authors should be given the opportunity to fix the search string, re-run the search, and report per-study provenance. This is a major revision rather than a rejection because the flaw is identifiable and fixable within the manuscript's scope, and the existing corpus and classification are likely to remain mostly valid after the correction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this is a carefully assembled mapping study of software testing for extended reality applications, the first dedicated one I know of. The authors follow the Petersen guidelines, use transparent selection criteria, three-author screening, and make the data extraction form and a tools/datasets repository publicly available. That repository is the real contribution: the tables of industrial and research tools (Airtest, AltUnity, VRGuide, TESTAR, etc.) and the dataset summaries give a newcomer to XR testing a solid entry point. The findings for RQ2 (functionality dominates the test objectives; machine learning is the most common technique; system-level testing is the rule) are plausible and grounded in the 34 primary studies. I would genuinely point someone at this paper for an orientation in the area.\n\nThe soft spot is the automated search string in Section 4.1.2. As written, the query requires an XR acronym (VR, AR, XR, MR) to appear in the title, together with a test/bug term in the title, and the full XR phrase somewhere in title or full text. That excludes any relevant paper whose title uses \"virtual reality\" or \"augmented reality\" without the acronym. Several of the final 34 primary studies have exactly that issue: PS8, PS9, PS10, PS12, PS19, and PS28. The paper does not disclose which of the 34 studies were found by the automated search versus snowballing, so it is impossible to tell whether the three snowballing iterations fully compensated. Worse, the search string as described cannot reproduce the authors' own test set: at least Bierbaum et al. (2003) and Corrêa Souza et al. (2018) have no acronym in their titles, yet the paper claims all eight test papers were retrieved. That is an internal inconsistency in the methodology, not just an acknowledged coverage gap.\n\nHow bad is that? It is not fatal. Backward snowballing, run three times, likely recovered a good chunk of the missed literature, and a mapping study is a synthesis rather than an experiment. But the \"first systematic mapping\" claim and the quantitative trends (59% VR, 26% AR, etc.) are only as solid as the search that produced them. The fix is straightforward: rerun the search with the acronym and the full term both optional in the title (or remove the title restriction on the acronym), then report any change in the final corpus. Until that is done, the completeness estimate should not be treated as definitive.\n\nMy recommendation: send it to peer review. The topic is timely, the presentation is clear, and the tools/datasets repository is a real asset. The reviewer guidelines should require the authors to correct the search string, disclose the automated-versus-snowballing split, and recheck their conclusions if the corpus changes. If you want a quick read of the current XR testing landscape, this is the best single document I know—just treat the counts as provisional.","headline":"Useful first map of XR testing with a real tools/datasets contribution, but the automated search string is internally inconsistent and undermines the completeness claim.","tokens_in":31448,"tokens_out":5935,"would_cite":true,"duration_ms":55716,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"XR software testing mapped: 34 studies, clear gaps","keywords":["software testing","extended reality","systematic mapping study","virtual reality","augmented reality","test automation","test oracle","XR applications"],"falsifier":"Run the paper's search string in IEEE Xplore, ACM Digital Library, and Scopus, apply the same inclusion and exclusion criteria, and compare the resulting set with the 34 primary studies. If the multi-database search yields additional primary studies that change the reported distributions (for example, more AR or integration-testing studies), the paper's 'first mapping' claim and its gap analysis would be shown to be incomplete.","tokens_in":30402,"feed_emoji":"🥽","tokens_out":4246,"duration_ms":38577,"temperature":0.7,"pith_summary":"This paper claims to be the first systematic mapping study of software testing for extended reality (XR) applications. It selects 34 primary studies from more than a thousand candidate works, classifies them by research topic, test activity, test concern, test technique, and evaluation method, and publishes the extracted data plus a repository of tools and datasets. The point is to give researchers and practitioners a structured, auditable picture of what has been tried in XR testing and where the open problems are. If the mapping is accepted, the field gains a shared baseline against which future work can be positioned.","feed_headline":"XR software testing mapped: 34 studies, clear gaps","feed_subtitle":"A first systematic review classifies the field and pinpoints missing pieces such as automated oracles and integration testing.","key_machinery":"The machinery is the systematic mapping methodology itself: a search string built from PICOC facets and synonym sets, selection criteria piloted on a test set, classification via systematic keywording, a data extraction form, and three rounds of backward snowballing. The load-bearing mechanism is the structured classification of each primary study along dimensions such as test activity, objective, target, level, type, technique, evaluation metric, and environment, which is what converts a list of papers into the claimed landscape and gap analysis.","core_discovery":"The central claim is that XR software testing is an emerging but underdeveloped field, and that the evidence supports a specific characterization: research grew from two studies in 2017 to ten in 2023; VR dominates (59% of studies) over AR (26%); automated testing is the most common topic; functionality is the leading test objective; system testing is the dominant level; machine learning is the most used technique; and most proposed solutions receive only preliminary validation. The paper further claims that test generation with automated oracles is the least explored activity, integration testing is nearly absent, and software-centric usability testing is underdeveloped.","pith_inferences":["A natural extension would be to replicate the search across IEEE Xplore, ACM Digital Library, and Scopus; if those databases surface many additional primary studies, the reported counts and gaps would need revision.","The gap analysis implies that oracle automation, especially visual oracles for six-degree-of-freedom interactions, is the bottleneck most likely to reward investment.","As XR development shifts toward new headsets and platforms, the paper's VR-centric evidence base may under-represent testing problems specific to mobile and web AR.","The mapping's emphasis on software-centric testing suggests a research program of converting user-study findings, such as cybersickness factors, into automated oracles."],"forward_implications":["A newcomer can use the 34-study map and the public repository as a starting point instead of re-searching the literature from scratch.","The dominance of machine-learning-based testing and image datasets suggests that future XR test tools will likely be oracle predictors trained on screenshots.","The near absence of integration testing (one study) and the rarity of unit testing point to concrete opportunities for new methods.","Because 60% of proposed solutions are validated only in controlled settings, industrial deployment remains largely unproven.","The growth since 2017 and the arrival of real-world evaluations in 2023 indicate a transition toward practice."],"supporting_citations":[{"why":"Provides the systematic mapping guidelines followed in planning, searching, classification, and reporting.","marker":"Petersen et al (2015)"},{"why":"Supplies the PICOC model and data extraction guidance used to design the search and forms.","marker":"Kitchenham and Charters (2007)"},{"why":"Defines the backward snowballing procedure used to supplement the automated search.","marker":"Wohlin (2014)"},{"why":"Describes OpenAlex, the bibliographic index that serves as the sole digital library for the search.","marker":"Priem et al (2022)"},{"why":"Provides the systematic keywording method used to build the classification scheme.","marker":"Petersen et al (2008)"},{"why":"An informal review of AR software engineering that motivates the need for a testing-specific map.","marker":"Börsting et al (2022)"},{"why":"A mapping study on VR quality metrics that shows existing secondary literature focuses on quality rather than testing.","marker":"Kuri et al (2021)"}],"fun_headline_variants":["First XR testing map: 34 studies, huge gaps remain","XR testing still immature: mapping reveals thin coverage","Systematic map: XR testing lacks automated oracles and integration","XR testing first map: VR dominates, validation sparse","34 XR testing studies mapped, key areas untouched"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole map depends on the assumption that searching OpenAlex plus three rounds of backward snowballing captured enough of the relevant literature, so the reported counts and gaps reflect the real field rather than the coverage of one database.","fun_headline_variants_meta":{"raw":{"variants":["First XR testing map: 34 studies, huge gaps remain","XR testing still immature: mapping reveals thin coverage","Systematic map: XR testing lacks automated oracles and integration","XR testing first map: VR dominates, validation sparse","34 XR testing studies mapped, key areas untouched"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000821,"raw_usage":{"total_tokens":3528,"prompt_tokens":813,"completion_tokens":2715,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":429,"completion_tokens_details":{"reasoning_tokens":2634}},"tokens_in":429,"tokens_out":2715,"duration_ms":21823,"temperature":1.0,"reasoning_tokens":2634,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:14:30.766957+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's search string in IEEE Xplore, ACM Digital Library, and Scopus, apply the same inclusion and exclusion criteria, and compare the resulting set with the 34 primary studies. If the multi-database search yields additional primary studies that change the reported distributions (for example, more AR or integration-testing studies), the paper's 'first mapping' claim and its gap analysis would be shown to be incomplete.","supporting_citations":[{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"A mapping study on VR quality metrics that shows existing secondary literature focuses on quality rather than testing."}],"review_version":1}