{"id":"1449de61-c730-47e9-a066-0731cdf27138","arxiv_id":"2608.01969","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a VR tourism app, users rated a mesh-based digital twin much higher on usability than a 3D Gaussian splatting twin, but found the 3DGS version slightly more realistic.","lead":"A 20-person lab study compared how people experience a 3D Gaussian splatting digital twin and a mesh-based digital twin of two towns inside a VR tourism app. The 3DGS version looked more realistic but scored badly on usability, while the mesh version was rated far better on practical quality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central trade-off claim rests on a confounded comparison; realism advantage is small and untested.","rationale":"The reader identifies the same core issue: the versions vary in location, navigation structure, and features, so the rendering technique is not isolated. This confound is acknowledged in §5.1, and the paper refrains from inferential comparisons. However, the headline narrative still leans on the comparative interpretation, which is the most load-bearing weakness. I also note that the experienced realism difference is very small (Δ=0.12), which further undermines the premise that 3DGS 'scored higher' in a meaningful way. Because the paper is already transparently framed as exploratory and the conclusion is carefully hedged, the appropriate verdict remains CONDITIONAL; my stress test does not change that assessment. The proposed concrete test—a controlled same-location, same-features comparison—would directly resolve whether the trade-off is real or an artifact of confounds.","tokens_in":13159,"tokens_out":5030,"duration_ms":54965,"concrete_test":"Conduct a within-subjects study using the same location (e.g., Etteln) reconstructed with both 3DGS and mesh pipelines, with identical navigation (continuous world plus teleportation) and identical interactive features (including IoT panels), and compare IPQ experienced realism and UEQ-S pragmatic quality. If the realism advantage of 3DGS and the pragmatic-quality disadvantage do not replicate, the current central claim is an artifact of confounds. As a minimal analytical check on the existing data, compute a paired 95% confidence interval for the realism difference; if it spans zero, the premise that 3DGS 'increased realism perception' is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—3DGS realism does not translate to UX—depends on attributing observed differences to the rendering technique. However, the two versions differ in multiple variables: location (Etteln vs Vilanova i la Geltrú), navigation (continuous interconnected world vs menu-selected isolated POIs), and interactive features (IoT sensor panels only in the Mesh version). §5.1 explicitly acknowledges that the study was not designed as a controlled experiment and that no inferential comparisons were made, but the abstract and conclusion still present the realism-versus-pragmatic-quality trade-off as the takeaway. This is not justified: the differences in pragmatic quality (UEQ-S M=1.70 vs 0.41) and even the experienced realism difference (IPQ M=-0.83 vs -0.71, Δ=0.12 on a [-3,3] scale) could be driven by location content or interaction richness rather than rendering. In particular, the realism advantage is tiny and would likely not survive a significance test; yet the entire conclusion rests on it. Without a controlled single-variable manipulation, the claim 'increased realism perception alone is not necessarily associated with an increased UX' is an interpretation, not a demonstrated finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory within-subjects study of two VR tourism digital-twin implementations: a mesh-based digital twin of Etteln and a 3D Gaussian splatting (3DGS) digital twin of Vilanova i la Geltrú. Twenty participants explored both versions and completed standardized questionnaires (UEQ-S, IPQ, CSQ-VR, Affective Slider, ATI) plus an unvalidated custom DT questionnaire. The authors present descriptive statistics and explicitly refrain from inferential comparisons between versions. They find positive affect in both versions, good UX ratings for the mesh version, poorer pragmatic quality for the 3DGS version, low-to-moderate presence, and mild cybersickness; the 3DGS version scored slightly higher on IPQ experienced realism. The conclusion frames this as suggesting that increased realism perception alone is not necessarily associated with increased UX.","tokens_in":13348,"tokens_out":6889,"duration_ms":79286,"significance":"Reported as an exploratory study, the paper offers a useful early data point on 3DGS versus mesh-based VR tourism. Strengths include reliance on standardized questionnaires with external benchmarks, transparent acknowledgement that the custom DT questionnaire is unvalidated, AB/BA order balancing, and conservative Holm–Bonferroni correction in the correlation analyses. The contribution is hypothesis-generating rather than confirmatory: the observed differences are descriptive, the realism advantage is small, and the two versions differ in multiple respects beyond rendering technique. With appropriate reframing, the paper could serve as a motivation for controlled follow-up studies.","major_comments":[{"comment":"The takeaway 'This indicates that increased realism perception alone is not necessarily associated with an increased UX' is not licensed by the design. The 3DGS and Mesh versions differ not only in rendering technique but also in location (Vilanova vs Etteln), navigation structure (isolated menu-selected POIs vs continuous interconnected environment), and interaction features (IoT panels only in the Mesh version). Section 5.1 correctly states that no direct conclusions about relative performance should be drawn, but the abstract and conclusion do not carry this qualification. Please reframe the conclusion as a descriptive observation or as a hypothesis for a controlled follow-up, and soften the corresponding abstract sentence.","section":"5.1 and Conclusion"},{"comment":"The narrative treats the 3DGS version's IPQ Experienced Realism advantage as a substantive finding, but the difference is 0.12 points on a [-3, 3] scale (M=-0.71, SD=1.06 vs M=-0.83, SD=0.96; n=20). The distributions overlap almost completely, and no inferential test or effect size is reported. In addition, the IPQ Experienced Realism subscale measures a subjective sense of realness, not necessarily photographic fidelity; the discussion equates it with the rendering technique's 'photorealistic appeal'. Please add uncertainty estimates and explicitly acknowledge the small magnitude before using this difference as the basis of the paper's central tension.","section":"§4 Results, Tables 2–3, Fig. 2"},{"comment":"The research gap in §2.4 is framed as determining 'what implications 3DGS-based DTs can have on UX and perception'. Given the multi-factor design, the study cannot support implications of the rendering technique; it can only describe two different systems. Consider narrowing the stated contribution to 'a first descriptive exploration that motivates controlled comparisons', or adding a follow-up design (e.g., same location, same navigation, only representation varied) to the future-work section. This would bring the aim, design, and conclusions into alignment.","section":"§2.4 / §3.1"}],"minor_comments":[{"comment":"The correlation result says 'Experience Realism' but the IPQ subscale is 'Experienced Realism'; please correct the terminology.","section":"§4.2"},{"comment":"The discussion says 'the point cloud data was reduced to simple mesh geometry', but §3.2.2 describes mesh extraction from a NeRF density field. Align the terminology with the reconstruction description.","section":"§5 Discussion"},{"comment":"The abstract says the 3DGS version showed 'clear weaknesses in pragmatic quality'. Since no statistical comparison is made, consider qualifying this as a descriptive sample-level observation rather than an unequivocal characterization.","section":"Abstract / §5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-scale exploratory study with a clear methodological limitation. The main fix is straightforward: align the headline claims with the stated limitations. If the authors remove or carefully hedge the causal/conclusion framing and report uncertainty, I would see this as publishable in the workshop/adjunct context. The lack of a controlled single-variable design is transparently acknowledged, so it is not a reason to reject; but it is a reason to require the conclusions to match the design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a transparent, nicely hedged exploratory lab study that gives the field a first descriptive data point on how 3DGS and mesh DTs are perceived in VR tourism. The central trade-off claim in the abstract, however, rests on a comparison that varies more than the rendering technique, so I wouldn't treat it as established.\n\nWhat's actually new: a within-subjects n=20 comparison of a real 3DGS reconstruction versus a mesh DT in a VR tourism app, using standard instruments (UEQ-S, IPQ, CSQ-VR, Affective Slider) plus a custom DT perception questionnaire. The authors report descriptive statistics only, clearly say no inferential comparison was made, and state the study wasn't designed to isolate a single variable. That's honest. The observed pattern—3DGS rated higher on experienced realism (IPQ -0.71 vs -0.83) but much lower on pragmatic quality (UEQ-S 0.41 vs 1.70)—is a plausible and useful preliminary hypothesis.\n\nThe soft spot is in the inference, not the reporting. The two versions differ in location (Etteln vs Vilanova), navigation (continuous world vs menu-selected POIs), and available interactions (IoT sensor panels only in the mesh version). The realism advantage is tiny: 0.12 on a -3 to 3 scale. With no inferential test, the abstract's \"increased realism perception alone is not necessarily associated with increased UX\" overreaches. What they can legitimately say is that in this implementation, the two conditions differed on many dimensions and the 3DGS version lagged on usability. The paper itself acknowledges this, but then the conclusion still slides back into the causal phrasing. Also, data and code aren't provided, and the sample is small and tech-savvy; both are stated. The custom DT questionnaire is unvalidated, and they flag it.\n\nBottom line: this is a workshop-quality exploratory study. For readers working on 3DGS for VR tourism, it's a reasonable prompt for a better-controlled follow-up. I'd send it to peer review because the topic is timely and the authors are transparent about limits, but I'd push for the claims to be scaled back to match the design. I wouldn't cite it as evidence for the trade-off, but I might mention it as a starting point.","headline":"Small exploratory VR tourism study with honest hedging, but the headline realism-vs-UX trade-off is not identifiable from a comparison that varies location, navigation, and features.","tokens_in":13916,"tokens_out":1798,"would_cite":false,"duration_ms":18405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In VR tourism, photorealistic 3D Gaussian splatting scores lower on user experience than a mesh-based digital twin, despite higher perceived realism.","keywords":["virtual reality tourism","3D Gaussian splatting","digital twins","user experience","presence","cybersickness","mesh rendering","realism perception"],"falsifier":"A controlled crossover experiment presenting the same location in both mesh and 3DGS forms, with identical navigation and interaction features, would settle the claim: if 3DGS pragmatic quality scores no longer fall below mesh scores in such a setup, the realism-versus-usability trade-off collapses.","tokens_in":13013,"feed_emoji":"🥽","tokens_out":2722,"duration_ms":31054,"temperature":0.7,"pith_summary":"This exploratory study compares two versions of a VR tourism application: one built from 3D meshes and one from 3D Gaussian splatting (3DGS). The central finding is that the 3DGS version felt more realistic but was rated worse on pragmatic quality and overall UX, while the mesh version was more usable and better received. The authors argue this shows that increased realism perception does not automatically improve user experience. They also report low presence and notable cybersickness in both versions, and they stress that the results are descriptive because the versions differ in location, navigation, and features.","feed_headline":"Splatting looks real, but mesh wins VR tourism UX","feed_subtitle":"Study finds photorealistic 3DGS scores lower on pragmatic quality than connected mesh environments.","key_machinery":"The comparison rests on two rendering pipelines: a mesh pipeline (NeRF-based reconstruction followed by isosurface extraction and manual post-processing) and a 3D Gaussian splatting pipeline (drone imagery, Structure-from-Motion, and optimization into 3D Gaussian primitives). These feed two versions of a Unity VR application evaluated with standardized questionnaires (UEQ-S, IPQ, CSQ-VR, Affective Slider) and a custom digital-twin perception questionnaire. The central mechanism is the contrast between the photorealism of splatting and the pragmatic usability of a coherent, freely explorable mesh world.","core_discovery":"In a within-subjects lab study with 20 participants, the mesh-based environment outperformed the 3DGS environment on pragmatic quality (UEQ-S mean 1.70 vs 0.41) and overall UX, while the 3DGS version scored higher on IPQ experienced realism (-0.71 vs -0.83 on the -3 to 3 scale). Cybersickness was similar across versions (19.5 vs 17.8 on CSQ-VR). The authors conclude that increased realism perception alone does not necessarily translate to an increased user experience. They emphasize that comparative statements are descriptive only, since the two versions also differed in replicated place, environment connectivity, and interaction features such as IoT sensor panels.","pith_inferences":["A likely implication left implicit is that the lower pragmatic quality of the 3DGS version may stem from its menu-driven, isolated POI navigation rather than from splatting itself; integrating splatted scenes into a continuous world could narrow the UX gap.","If 3DGS artifacts diminish with future algorithmic improvements, the realism advantage could grow without the usability penalty, potentially flipping the overall UX balance in favor of splatting.","Because users freely explored, the custom questionnaire results depend on which content they encountered (e.g., sensor panels may have been ignored), so guided tasks could yield more discriminating DT perception metrics.","A hybrid approach—splatting for distant photorealistic views and meshes for interactive close-up elements—could be a testable extension that combines realism with pragmatic usability."],"forward_implications":["If realism does not drive UX, developers of VR tourism should prioritize coherent navigation and interactive features over pure visual fidelity.","3DGS in its current form may introduce visual artifacts that contribute to cybersickness and lower pragmatic quality.","Both rendering approaches yield low presence scores, indicating that immersion design needs improvement independent of rendering technique.","The custom digital-twin questionnaire suggests simulation and understanding are strengths, while reliability and information access need work.","A controlled 2x2 crossover study with identical content is required to isolate the effect of the rendering technique on user experience."],"fun_headline_variants":["Mesh beats 3DGS on practical quality in VR tourism","Realism alone doesn't win VR tourism UX","Splatting's realism can't beat mesh's usability","VR tourism study: mesh wins on pragmatic quality","For VR tourism, pragmatic mesh beats flashy 3DGS"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's comparative interpretation assumes that the two versions differ primarily in rendering technique, when they also differ in location, navigation structure, and features (e.g., IoT panels), as the authors explicitly acknowledge.","fun_headline_variants_meta":{"raw":{"variants":["Mesh beats 3DGS on practical quality in VR tourism","Realism alone doesn't win VR tourism UX","Splatting's realism can't beat mesh's usability","VR tourism study: mesh wins on pragmatic quality","For VR tourism, pragmatic mesh beats flashy 3DGS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000363,"raw_usage":{"total_tokens":1812,"prompt_tokens":782,"completion_tokens":1030,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":950}},"tokens_in":526,"tokens_out":1030,"duration_ms":10029,"temperature":1.0,"reasoning_tokens":950,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:34:57.693719+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled crossover experiment presenting the same location in both mesh and 3DGS forms, with identical navigation and interaction features, would settle the claim: if 3DGS pragmatic quality scores no longer fall below mesh scores in such a setup, the realism-versus-usability trade-off collapses.","supporting_citations":[],"review_version":1}