{"id":"b54c2752-1040-4e56-8e54-9d14dfaa83a6","arxiv_id":"2412.02853","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 177 papers identifies the main devices, applications, technologies, and grand challenges in immersive cultural heritage presentation.","lead":"This paper reviews 177 studies on how virtual, augmented, and mixed reality are used to present cultural heritage. It maps the devices, applications, and technologies in the field and names the main risks and research gaps.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5.3's three 'grand challenges' are not grounded in the reviewed corpus: almost all supporting citations ([221]–[240]) fall outside the 177 analyzed papers, so the headline gap-identification claim is not a systematic finding.","rationale":"Good-faith reading: the paper is a systematic scoping review with a large search (5,368 records), a PRISMA-style flow diagram, a data-availability link, a 20-expert screening panel, and detailed classification tables. These are real strengths, and the paper may be useful as an overview of the 177-paper corpus. The central claim, however, is that the review identifies under-explored grand challenges. For that to be true, the three challenges must be derived from the reviewed evidence. Section 5.3 is where the argument is least secure: its supporting citations are external to the corpus, one is a film-studies paper, and two are surveys that EC1 would exclude. This means the main contribution is a narrative synthesis presented as a systematic finding. The proposed check would settle it directly by coding the 177 papers for those themes and checking the provenance of the supporting citations. I agree with the reader that corpus representativeness is a concern, but I do not think it is the same concern; even a flawless selection would not fix the Section 5.3 grounding issue. The verdict should remain CONDITIONAL (not accepted as-is), with the condition made explicit: the authors must either provide traceable corpus evidence for the gaps or relabel Section 5.3 as expert-informed discussion, and reconcile EC1 with the presence of surveys in the argument.","tokens_in":32591,"tokens_out":11382,"duration_ms":116985,"concrete_test":"Create a coding sheet from Section 5.3's three gap themes (misinterpretation/over-interpretation; media purity; physical damage/light pollution) and apply it to the full text of each of the 177 included papers identified in the screening logs. Then cross-tabulate: how many of the 177 papers mention each theme, and how many of the Section 5.3 supporting citations ([221]–[240]) are among the 177. If, say, fewer than five papers mention any given theme and the supporting citations are external, the paper must be revised to label Section 5.3 as non-systematic expert commentary, and the abstract should no longer claim the gaps were found by the analysis. If the themes are genuinely discussed in the corpus with traceable citations, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing problem is not only corpus representativeness; it is that the three 'grand challenges' that constitute the paper's main contribution are not shown to emerge from the 177-paper corpus. The corpus is operationalized in Tables 2–4, but Section 5.3 ('Gaps and Risks') supports its claims almost entirely with references [221]–[240], none of which appear in those coding tables or in the included-paper list. The 'purity in media experience' subsection is anchored by [227], a film-studies article on '50 Shades of Grey' and 'Caroline', not an immersive-cultural-heritage study. The 'physical damage' subsection relies on external lighting case studies ([228]–[234]) rather than on any extracted finding from the 177 papers. Likewise, the reconstruction-risk argument cites [225] and [226], which are surveys, although EC1 excludes reviews. Consequently, the abstract's claim that 'Based on our analysis, we found...' these gaps is not traceable to the systematic data. The gap identification is better described as an expert narrative or traditional narrative review, presented as a result of the systematic review. This directly undermines the central claim that the paper identifies the grand challenges future research should address. The selection-bias concern from the reader is real, but this is more fundamental: even a perfectly representative corpus would not entail the Section 5.3 claims unless the corpus was coded for these themes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a scoping review of immersive technologies for cultural heritage. From 5,368 records retrieved from ACM Digital Library, IEEE Xplore, and Scopus, the authors screen down to 177 journal papers and code them along three dimensions—device, application, and technology—presenting frequency tables, a Sankey diagram, and publication-trend analysis. The review then summarizes strengths and challenges of immersive cultural heritage practice and identifies three 'grand challenges': misinterpretation of cultural heritage, the impact on 'purity in media experience,' and potential physical damage from immersive practices. The paper concludes with recommendations for future research and a data availability link.","tokens_in":32732,"tokens_out":9336,"duration_ms":85921,"significance":"If the methods and gap-identification claims were made reproducible and traceable to the analyzed corpus, this paper would be a useful reference for researchers and practitioners: it provides a reasonably large coded corpus, releases its data, uses root-based keyword search, employs a 20-expert screening panel, and presents a descriptive taxonomy of devices, applications, and technologies. The three proposed grand challenges are timely and could shape future work. However, the central contribution—the identification of those gaps—is not currently supported by the systematic data, and the screening protocol has internal inconsistencies. These issues are load-bearing for the paper's main claim, so the manuscript needs substantial revision before it can serve as a reliable systematic reference.","major_comments":[{"comment":"The three 'grand challenges' are not demonstrably derived from the 177-paper corpus. The references used in Section 5.3 ([221]–[240]) do not appear in the coding tables (Tables 2–4), which enumerate the included papers from roughly [32] to [207]; [227], the anchor for 'purity in media experience,' is a film-studies article on '50 Shades of Grey' and 'Caroline'; [225] and [226] are surveys excluded by EC1; and [228]–[234] are external lighting case studies. The abstract's statement 'Based on our analysis, we found...' is therefore not traceable to the systematic data. The same pattern appears in Sections 5.1 and 5.2, where several cited references (e.g., [210]–[220]) are not in the enumerated corpus. Please either recode the 177 papers for these themes and report the coding frequencies, or explicitly reframe Section 5.3 as expert interpretation rather than a systematic-review finding.","section":"Section 5.3 and Abstract"},{"comment":"The screening methodology is internally inconsistent and not reproducible. Conference papers are excluded because, after checking 623 English conference papers, the authors conclude that 'research based on conference papers is not very reliable,' yet Section 3.5 states that papers were not excluded on methodological quality. The filters 'page count, impact factor, citations, word of mouth, retractions, or duplication factors' are not operationalized with any thresholds. The inconsistency shows up in the results: [208], a conference paper, is cited in Section 4.2.1 as evidence about device use, and EC1 excludes reviews while Section 4.1.1 lists 'literature review studies' in the corpus and Tables 2–4 include [35], a state-of-the-art review. These issues prevent an auditor from reconstructing the included set and undermine the systematic-review claim.","section":"Sections 3.3–3.5"},{"comment":"The PRISMA counts do not reconcile. The initial search totals are 339 + 520 + 4,509 = 5,368; after the stated per-database filtering the numbers are 205 + 513 + 3,822 = 4,540, but Figure 1 reports 3,907 records after removing 1,461 duplicates and non-relevant entries (5,368 − 1,461 = 3,907). The figure also states that '1,448 journal papers and conference proceedings were screened' even though conference papers had already been excluded in Section 3.3. Please correct the numbers and the flow labels so that the screening process can be audited.","section":"Section 3.2 and Figure 1"}],"minor_comments":[{"comment":"Duplicate reference numbers appear in the coding tables: the VR row of Table 2 contains '[75, 75, 75–94]' and the Education row of Table 3 contains '[126, 126]'; please deduplicate and re-verify the underlying coding data.","section":"Tables 2 and 3"},{"comment":"Table 4's header reads 'Application Paper' but the table categorizes technologies, and Section 4.2.1 says 'we mainly extracted a total of 311 frequency of articles mentioning equipment' where 'technology' is meant; please correct both.","section":"Table 4 and Section 4.2.1"},{"comment":"The sentence 'or discuss technologies that have a direct link to the display of cultural heritage, will also be excluded' appears to state the opposite of the intended criterion and should be corrected to avoid ambiguity.","section":"Section 3.4, EC4"},{"comment":"Reference [209] duplicates reference [158] (Liritzis et al.), and [208], a conference paper, is cited in the Results despite the journal-only inclusion rule; please reconcile these citations with the stated criteria.","section":"References"},{"comment":"The last application row is labeled 'Care for special peopled'; this appears to be a typo for 'Care for special people.'","section":"Table 3"},{"comment":"The phrase 'for reasons that we do not know' is an informal aside; either remove it or provide a citable explanation for the suspension of the light show.","section":"Section 5.3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper would be more accurately positioned as an expert scoping overview than as a systematic review unless the gap-identification section is reworked and the screening protocol is made reproducible. I also note that [12] and [208] are the authors' own conference papers appearing in the background and results despite the journal-only inclusion rule; I do not see evidence that they determine the conclusions, but the inconsistency is worth checking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a scoping review with a decent corpus and genuinely useful categorization tables, but its headline contribution—the three \"grand challenges\" in Section 5.3—is not actually grounded in the systematic analysis. Almost every supporting citation there ([221]–[240]) falls outside the 177 analyzed papers and outside the coding tables. \"Purity in media experience\" leans on a film-studies article about 50 Shades of Grey; \"physical damage\" leans on external lighting case studies. The abstract says \"Based on our analysis, we found...\" and that claim is doing work the data can't support.\n\nWhat's genuinely good: the corpus is current (through January 2024), the device/application/technology breakdown is clearly presented, and the Sankey diagram gives a quick visual read on where the field concentrates. The data availability link is a real plus. The paper also emphasizes risks and \"dark sides\" more than many earlier surveys, which is fair and useful.\n\nThe soft spots beyond the gap-grounding problem: the screening methodology is shaky in ways that matter. Conference papers are dismissed in one sentence as \"not very reliable,\" the quality filters (page count, impact factor, \"word of mouth\") are subjective, and EC1 excludes reviews while review papers appear in the final corpus (e.g., [35], [225], [226]). No inter-rater reliability is reported. The writing is also rough in places, with citation and consistency slips.\n\nMy read: treat this as a well-intentioned scoping overview with directional insight, not a definitive evidence base. The categorization tables are worth using; the grand-challenges section should either be re-derived from the corpus or explicitly reframed as expert-informed gaps. This deserves a serious referee round—it's not a desk reject—but it needs major revision: make the gap claims traceable, tighten the screening description, and clean up the references.","headline":"A useful but flawed scoping review whose headline 'grand challenges' are not derived from the 177-paper corpus; the tables are worth using, the gap claims need re-grounding or reframing.","tokens_in":33393,"tokens_out":2050,"would_cite":false,"duration_ms":22840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 177-paper systematic review argues that immersive heritage displays face three under-studied grand challenges: misinterpretation, loss of media purity, and physical damage to the sites themselves.","keywords":["immersive technologies","cultural heritage","systematic review","virtual reality","augmented reality","mixed reality","grand challenges","misinterpretation"],"falsifier":"Re-run the screening while including the 623 English conference papers, or a random sample of them, and code them with the same device/application/technology scheme; if misinterpretation, media purity, or physical damage appear frequently in that body of work, the claim that these areas are under-explored collapses.","tokens_in":32258,"feed_emoji":"🏛️","tokens_out":3990,"duration_ms":38199,"temperature":0.7,"pith_summary":"This paper argues that prior surveys of immersive technologies for cultural heritage overstated benefits and neglected the downsides. From 5,368 retrieved articles the authors selected 177 for full-text analysis, classifying them by device, application, and technology. They claim that three areas remain under-explored: misinterpretation of cultural heritage through digital reconstruction, erosion of the 'purity in media experience,' and physical damage to heritage sites from immersive practices such as light shows. If these claims hold, future research and museum practice would need to shift from pure technological optimism toward risk-aware design.","feed_headline":"Immersive heritage tech has three blind spots","feed_subtitle":"A 177-paper review finds over-interpretation, lost media purity, and physical damage are under-studied risks.","key_machinery":"The machinery is a systematized scoping review: a keyword-based search in ACM Digital Library, IEEE Xplore, and Scopus yielding 5,368 records, a three-stage manual screening that reduces the set to 177 journal articles, and a coding scheme that classifies each article along three axes—device, application, and technology—visualized in a Sankey diagram. The screening criteria (exclusion of reviews, marginal technology mentions, and non-display contexts; inclusion of innovative applications, technology-heritage integration, and user interaction) are what allow the authors to claim their gaps are genuinely absent from the literature rather than an artifact of the search.","core_discovery":"The central claim is that the field's main open problems are not purely technical but interpretive and physical: over-interpretation or misinterpretation of heritage in 3D reconstructions, the degradation of an unmediated 'purity in media experience' when digital layers dominate, and damage to heritage itself through light and noise pollution or excessive interaction. The review supports this with case evidence such as the suspension of night tours at the Longmen Grottoes after lights attracted insects, and with a Sankey analysis showing that devices and technologies concentrate on exhibition and education while community engagement and care for special groups lag. On the paper's own terms, these gaps define the agenda for future research.","pith_inferences":["Beyond the paper: the 'purity in media experience' concept could be made measurable, for instance by tracking user attention to heritage content versus the technology itself, enabling empirical tests of whether over-mediation actually harms understanding.","Beyond the paper: because the corpus excludes all conference papers, the three identified gaps might partly reflect venue distribution, since recent immersive-heritage innovations appear heavily in conference proceedings.","Beyond the paper: the case-based evidence points toward a regulatory direction—heritage impact assessments for temporary light and projection installations, analogous to environmental impact assessments.","Beyond the paper: the device/application/technology coding scheme could serve as a reusable template for other reviewers, making cross-review comparison of heritage-technology literature feasible."],"forward_implications":["Immersive heritage displays should be evaluated not only for engagement but for fidelity to the original meaning and for distraction caused by over-mediation.","Conservation policy for fragile sites needs light and noise budgets informed by cases like the Longmen Grottoes suspension.","Future research should prioritize misinterpretation, media purity, and physical damage, because the current literature concentrates on exhibition enhancement and education.","Lack of standardization and interoperability is a documented barrier, implying that open data formats and shared workflows are needed for sustainable installations."],"supporting_citations":[{"why":"The prior survey of AR/VR/MR for cultural heritage that the paper positions as missing recent developments and negative impacts.","marker":"[1]"},{"why":"Establishes the challenges of digital preservation in heritage institutions, the baseline the paper extends toward presentation.","marker":"[2]"},{"why":"A bibliometric review restricted to VR, AR, and MR, which the paper argues is too narrow in technological scope.","marker":"[18]"},{"why":"A small-sample overview cited as insufficient for drawing broad conclusions about the field.","marker":"[19]"},{"why":"Grounds the claim that immersive representation can lead to over-interpretation or misunderstanding of heritage.","marker":"[221]"},{"why":"Supplies the notion of 'purity' in media that the paper adapts to the heritage-display context.","marker":"[227]"},{"why":"Documents insect attraction to ornamental lighting on heritage buildings, supporting the physical-damage claim.","marker":"[228]"},{"why":"The Longmen Grottoes case study on light-induced insect damage and the suspended night tour.","marker":"[231]"},{"why":"Discusses lighting approaches for heritage sites, cited to argue that tailored protection measures are required.","marker":"[233]"}],"fun_headline_variants":["Heritage VR's real risks are interpretive and physical","Three under-studied perils in immersive heritage: misreading, purity, harm","Immersive heritage: over-interpretation, lost purity, damage","Beyond tech: heritage VR's biggest challenges are not technical"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's trends and gaps all rest on the assumption that filtering out all conference papers and applying subjective quality screens (page count, impact factor, citations, word of mouth) does not bias the picture of what the field studies.","fun_headline_variants_meta":{"raw":{"variants":["Heritage VR's real risks are interpretive and physical","Three under-studied perils in immersive heritage: misreading, purity, harm","Immersive heritage: over-interpretation, lost purity, damage","Beyond tech: heritage VR's biggest challenges are not technical"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000254,"raw_usage":{"total_tokens":1503,"prompt_tokens":812,"completion_tokens":691,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":617}},"tokens_in":428,"tokens_out":691,"duration_ms":7658,"temperature":1.0,"reasoning_tokens":617,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:01:27.536971+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the screening while including the 623 English conference papers, or a random sample of them, and code them with the same device/application/technology scheme; if misinterpretation, media purity, or physical damage appear frequently in that body of work, the claim that these areas are under-explored collapses.","supporting_citations":[{"cited_title":"Journal of Librarianship and Information Science 43(3), 157–165 https://doi.org/10.1177/0961000611410585","cited_arxiv_id":null,"evidence_quote":"Establishes the challenges of digital preservation in heritage institutions, the baseline the paper extends toward presentation."},{"cited_title":"Sustainability 16(15), 6446 (2024) https://doi.org/10.3390/su16156446","cited_arxiv_id":null,"evidence_quote":"A bibliometric review restricted to VR, AR, and MR, which the paper argues is too narrow in technological scope."},{"cited_title":"Personal and Ubiquitous Computing 21, 187–189 (2017) https:// doi.org/10.1007/s00779-016-0984-y 28","cited_arxiv_id":null,"evidence_quote":"A small-sample overview cited as insufficient for drawing broad conclusions about the field."},{"cited_title":"Advanced Materials Research594, 155–160 (2012) https: //doi.org/10.4028/www.scientific.net/AMR.594-597.155","cited_arxiv_id":null,"evidence_quote":"The Longmen Grottoes case study on light-induced insect damage and the suspended night tour."},{"cited_title":"Sustainability13(5), 2720 (2021) https://doi.org/10.3390/su13052720","cited_arxiv_id":null,"evidence_quote":"Discusses lighting approaches for heritage sites, cited to argue that tailored protection measures are required."}],"review_version":1}