{"id":"555d0848-1dce-4d8b-b32e-1196e243f7f1","arxiv_id":"2509.07740","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An exploratory evaluation of AR and VR digital-twin tourism apps finds generally positive enjoyment and low workload, with VR usability and cybersickness issues and user characteristics correlating with ratings.","lead":"A field study of two digital-twin tourism apps, one smartphone AR and one VR headset, tested with 84 people in Germany and Spain. It found the experiences enjoyable and low-effort, but VR caused usability and cybersickness issues, especially for novices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'low task load and high enjoyment' for both apps is unsupported: NASA-TLX was only deployed for AR-VT, and the enjoyment custom item only for the Etteln group.","rationale":"The reader's verdict is CONDITIONAL and I agree with that overall. The reader identified sample representativeness as the weakest assumption; I partially agree, but I find the instrumentation gap more directly load-bearing for the central claim. The paper is explicitly exploratory and the authors admit the omission in §3.6, so the appropriate remedy is a scoped abstract/conclusion rather than wholesale rejection. No independent verification is possible without data release, which reinforces the conditional stance. Thus the reader's verdict remains unchanged: conditional acceptance pending revision of the overbroad claim.","tokens_in":11518,"tokens_out":3662,"duration_ms":43377,"concrete_test":"Construct a two-way table from the paper: rows = NASA-TLX and enjoyment item; columns = VR-VT, AR-VT/Etteln, AR-VT/Fair, AR-VT/Vilanova; fill with 'administered' only where §3.6 and §4 report data. If the cells for VR×NASA-TLX and for AR×Vilanova/Fair×enjoyment are empty, the abstract's 'both applications' wording must be revised to 'AR-VT had low task load; the Etteln group reported high enjoyment in both apps.' This single check settles whether the overclaim exists.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.6 states NASA-TLX was used 'only to validate the AR application' and that the custom questionnaires were 'unintentionally not used for the Vilanova and Fair Group.' Accordingly, §4.1 reports no task-load values for VR-VT, and §4.2 reports enjoyment only for the Etteln Group (n=16). Yet the abstract and conclusion assert that 'both applications provided a low task load and high enjoyment.' The low-task-load claim for VR-VT and the high-enjoyment claim for the majority of AR-VT participants have no measurement behind them. This is not a question of sample representativeness; it is a coverage gap in the headline result. The reader's representativeness concern applies, but the more immediate load-bearing problem is that the central claim's scope exceeds the questionnaires actually administered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory in-field user study of two digital-twin-based XR tourism applications: an AR Virtual Tour (AR-VT) for on-site exploration and a VR Virtual Tour (VR-VT) for remote exploration. A total of 84 participants were recruited in three groups (Etteln n=16, Fair n=50, Vilanova n=18). UX, task load, presence, cybersickness, emotion, and custom feature-related items were collected with standardized questionnaires. The results show mixed outcomes: the VR app had moderate presence, low perceived realism, usability and cybersickness problems, and negative overall UEQ-S scores; the AR app's UX varied across groups, with only the Fair Group reaching 'good' UEQ-S ratings. The authors also report several Spearman correlations with age, XR experience, and affinity for technology interaction. The abstract and conclusion generalize the findings to 'both applications' with 'low task load and high enjoyment,' but the data do not support this claim as stated.","tokens_in":11733,"tokens_out":4695,"duration_ms":53426,"significance":"If the reported results were supported, the paper would provide useful early evidence about user experience with digital-twin-based XR tourism applications in real-world settings, including insights for tailoring interfaces to XR novices. The study addresses an underexplored area, uses standardized instruments, and is transparent about several of its own limitations, including the unintended omission of the custom questionnaire for two groups and the field-setting constraints. However, the central claims in the abstract outrun the measurements: task load was collected only for AR-VT, enjoyment was collected only for the small Etteln group, and the 'high UX' assertion for AR-VT does not hold for the Etteln and Vilanova groups. These gaps do not invalidate the whole study, but they require the headline conclusions to be substantially qualified.","major_comments":[{"comment":"The abstract's lead finding that 'both applications provided a low task load and high enjoyment' is not supported by the measurements. §3.6 states that NASA-TLX was used 'only to validate the AR application' and that the custom questionnaires were 'unintentionally not used for the Vilanova and Fair Group.' Consequently §4.1 reports no task-load data for VR-VT, and §4.2 reports enjoyment only for the Etteln Group (n=16). The low-task-load claim for VR-VT and the high-enjoyment claim for AR-VT participants outside Etteln are therefore without measurement. Please revise the abstract, §1, and §6 to state precisely which metrics were collected for which groups, or collect the missing data before asserting the headline result.","section":"Abstract and §3.6/§4.1"},{"comment":"The abstract's statement that 'the AR-based tour achieved high UX ratings' is contradicted by the group-level UEQ-S results in §4.2: only the Fair Group's overall mean (1.41) falls in the 'good' category; the Etteln Group (0.70) and Vilanova Group (1.11) are both classified 'below average.' Thus high AR UX is a subgroup result, not an overall result. Please qualify the claim, for example by reporting that the Fair Group rated the AR-VT as 'good' while the other two groups rated it below average.","section":"Abstract and §4.2"},{"comment":"The VR-VT 'enjoyment' claim rests on a single custom item collected only from the Etteln Group (n=16), with mean 4.44 on a 7-point scale. The UEQ-S overall mean for the same application is -0.10, described as 'bad' in §5.1. Calling the experience 'high enjoyment' in the abstract is not warranted. Please describe the VR-VT result as mixed: moderate presence, low realism, usability problems, and moderate cybersickness, with an enjoyment item that was only administered to one small group.","section":"§5.1, §4.1, §3.6"},{"comment":"The correlation analysis is reported as the basis for 'significant relationships' with age, XR experience, and ATI, but it is an uncorrected exploratory search across many variables. In the Etteln group (n=16), many correlations are tested; by chance alone one would expect several p<0.05 findings. For the claim that 'correlation analysis revealed significant relationships ... for both applications,' please provide the full correlation matrix, state the number of comparisons made, and either apply a correction or explicitly present the findings as hypothesis-generating. Also note that the scatter-plot regression lines in Figs. 3-4 imply linear associations while Spearman rank correlations were reported.","section":"§3.8, §4.1, §4.2"}],"minor_comments":[{"comment":"The subgroup labels 'elderly (n=7), young (n=5), and middle-aged (n=4)' are used without age brackets; given the overall age range of 16-79 years, please define these categories.","section":"§3.2"},{"comment":"UEQ-S benchmark classifications ('below average', 'good', etc.) are used without a citation to the UEQ-S benchmark data; please provide the reference so readers can interpret these labels.","section":"§4.2"},{"comment":"NASA-TLX was altered from its original continuous response format to a 7-point scale. Please add a sentence in the limitations section noting that this modification may affect comparability with other NASA-TLX studies.","section":"§3.6"},{"comment":"The Fair Group was tested at a fairground rather than in front of the museum, and different smartphone models were used across groups. These are acknowledged in §6.2, but the conclusion could more explicitly avoid generalizing the AR-VT results to the intended tourist scenario.","section":"§3.3/§3.4"},{"comment":"The VR-VT procedure included a training phase, while the AR-VT procedure did not describe an equivalent familiarization step. This asymmetry should be acknowledged as a possible confound when comparing task demand across the two applications.","section":"§3.7"},{"comment":"The statement about expanding evaluation methodologies to include 'DT-specific metrics' would benefit from a concrete example, since the paper does not propose what such a metric would measure.","section":"§6.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its limitations and the unintended non-administration of the custom questionnaire to two groups. The main problem is that the abstract and conclusion overstate the empirical coverage. This is fixable by rewording the central claims to match the data: low task load only for AR-VT, enjoyment only for the Etteln subsample, and AR-VT UX high only in the Fair Group. The exploratory correlations also need tempering. I see no evidence of fabrication or circularity; the paper should be reconsidered after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper has real, useful data on two new digital-twin XR tourism applications, but its abstract overreaches. The claim that both apps produced low task load and high enjoyment is not actually supported by the instruments. NASA-TLX was only used for AR-VT, and the enjoyment item was only collected in the Etteln group (n=16). So the 'both' in the headline is doing a lot of work.\n\nWhat the paper does well: it reports a field evaluation of an AR mobile portal tour and a VR headset tour, built as part of the DIDYMOS-XR project. As far as I can tell, these specific apps have not been evaluated before, and the study gives a useful interim benchmark. The methods are standard questionnaires (UEQ-S, SAM, IPQ, CSQ-VR, NASA-TLX) applied in naturalistic settings. The authors are transparent about the limitations—they note the custom questionnaire was unintentionally dropped for two groups, and they explicitly avoid comparative statistics between groups because of contextual confounds. The correlation results (age, experience, tech affinity) largely replicate known XR findings, and the authors frame them as exploratory, which is appropriate.\n\nWhere it gets soft: the abstract and conclusion go beyond the measurements. The VR-VT UEQ-S mean is -0.10, which is negative, and no task load data were collected for VR. The enjoyment mean for VR comes from the 16 Etteln participants only. The samples are small and convenience-based; two groups fall below their planned 20. The correlations are uncorrected for multiple comparisons, and no data are released. These are specific, fixable problems rather than a broken central result.\n\nWho should read it: anyone building or evaluating XR tourism applications, especially within digital-twin projects. It is a pilot-scale study, not a definitive answer. I would send it to peer review—a serious referee could ask for a revised abstract that matches the actual coverage, and maybe supplemental data. With those changes it would be a reasonable contribution to an applied HCI venue.","headline":"Solid pilot data for two XR tourism apps, but the abstract's 'low task load and high enjoyment' overstates what was measured.","tokens_in":12192,"tokens_out":2833,"would_cite":true,"duration_ms":28309,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Digital-twin XR tourism apps are enjoyable and low-effort, but VR beginners hit usability and cybersickness walls.","keywords":["digital twins","extended reality","tourism","user experience","augmented reality","virtual reality","cybersickness","presence"],"falsifier":"Run the same AR-VT and VR-VT applications with a pre-registered, balanced sample of actual tourists at the intended museum sidewalk and village locations, with at least 20 per group and matched devices. The central claim weakens if novices no longer show lower pragmatic quality and higher cybersickness, or if the AR app's UX advantage over VR disappears.","tokens_in":11465,"feed_emoji":"🥽","tokens_out":4784,"duration_ms":46969,"temperature":0.7,"pith_summary":"This paper reports an in-field user study of two digital-twin-based tourism applications: an AR tour that lets people stand outside a building and look inside through a virtual portal, and a VR tour that lets them explore a reconstructed village remotely. The authors are trying to establish that both formats can deliver enjoyable, low-workload experiences, and that the main obstacles are specific and partly predictable from user traits. Across 84 participants in Germany and Spain, the AR app earned high UX ratings, while the VR app produced a stronger sense of presence but worse usability and some cybersickness. Age, prior XR experience, and affinity for technology were significantly correlated with how people rated the apps. If these relationships hold, designers of XR tourism should treat novice users as a distinct audience and prioritize visual fidelity and comfortable locomotion.","feed_headline":"Digital-twin XR tours: fun and low-effort, with VR caveats","feed_subtitle":"Across 84 users, AR tours scored high UX while VR novices faced controls and motion sickness.","key_machinery":"The carrying mechanism is the paired application design plus a standardized questionnaire battery. AR-VT places a user on the sidewalk with a smartphone and lets them tap a virtual portal into the building's digital twin; VR-VT places the user inside the village twin with locomotion, weather control, and snapshot features. The evaluation relies on UEQ-S for UX, NASA-TLX for task load, SAM for emotion, IPQ for presence, CSQ-VR for cybersickness, and ATI for technological affinity. These instruments convert subjective experience into comparable scores, and Spearman correlations link those scores to age, prior XR experience, and tech affinity.","core_discovery":"The central empirical claim is that digital-twin XR tourism applications are viable and generally well received, but the two formats differ sharply. The AR-VT, a smartphone app that opens a virtual portal into a building's digital twin, produced high UX scores and low task load across all three participant groups, with high enjoyment. The VR-VT, a headset experience of a reconstructed village, produced moderate presence but pragmatic-quality scores below the UEQ benchmark and cybersickness symptoms concentrated among less experienced users. The paper further claims that user factors explain a meaningful share of response: prior XR experience correlated strongly with higher pragmatic quality","pith_inferences":["The same data imply that cybersickness may be partly trainable: if prior experience is strongly protective, brief acclimation sessions could reduce symptoms before the tour, a testable design change the paper does not propose.","A controlled within-subject comparison of AR-VT and VR-VT on the same city would isolate format effects; the present field design cannot separate format from location and sample.","If visual fidelity was the main cause of low perceived realism, improving the 3D reconstruction or rendering should raise IPQ realism and possibly pragmatic scores, which could be tested by varying digital-twin quality while holding content constant.","The ATI correlations suggest the same app may need adaptive UI: simpler gestures for low-affinity users and richer features for high-affinity users."],"forward_implications":["If the claims hold, AR digital-twin tours can serve as low-effort on-site enhancements even outside opening hours, since task load was low and UX high across groups.","VR digital-twin tours should not be assumed ready for the general public: usability and cybersickness scores suggest novices need onboarding, comfort options, and simpler locomotion before broad deployment.","Prior XR experience is a usable predictor: designers can expect novices to rate pragmatics lower and report more cybersickness, so evaluation should oversample novices.","Age and technology affinity are relevant covariates in XR tourism studies; ignoring them can misattribute app quality to user traits or vice versa.","A follow-up iteration should focus on visual fidelity and richer interaction, since qualitative feedback flagged movement, zoom, and information content as weak points."],"supporting_citations":[{"why":"Supplies the user-centric evaluation methodology and the AR-VT/VR-VT use-case descriptions on which the study is built.","marker":"[18]"},{"why":"Supplies the UEQ-S instrument used to measure pragmatic and hedonic UX.","marker":"[17]"},{"why":"Supplies the NASA-TLX used to measure task load for the AR application.","marker":"[9]"},{"why":"Supplies the IPQ used to measure presence in the VR application.","marker":"[10]"},{"why":"Supplies the CSQ-VR used to measure cybersickness symptoms in VR.","marker":"[12]"},{"why":"Supplies the ATI scale used to measure technological affinity and its correlations with UX.","marker":"[7]"},{"why":"Supplies the SAM used to measure valence, arousal, and dominance.","marker":"[2]"},{"why":"Supports the interpretation that repeated VR exposure reduces cybersickness.","marker":"[5]"},{"why":"Supports the age-related presence findings for older adults in immersive VR.","marker":"[4]"}],"fun_headline_variants":["AR digital-twin tours beat VR in usability, but VR immerses","Digital-twin XR: AR shines, VR suffers from cybersickness","XR tourism: AR easy, VR immersive but sickening for novices","84 users test XR twin tours: AR high UX, VR rough on novices"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the convenience samples and field locations represent the broader tourist population well enough that the measured UX levels and correlations reflect the applications rather than recruitment, setting, or device differences.","fun_headline_variants_meta":{"raw":{"variants":["AR digital-twin tours beat VR in usability, but VR immerses","Digital-twin XR: AR shines, VR suffers from cybersickness","XR tourism: AR easy, VR immersive but sickening for novices","84 users test XR twin tours: AR high UX, VR rough on novices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000694,"raw_usage":{"total_tokens":2945,"prompt_tokens":681,"completion_tokens":2264,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":2181}},"tokens_in":425,"tokens_out":2264,"duration_ms":15958,"temperature":1.0,"reasoning_tokens":2181,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T21:40:46.027007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same AR-VT and VR-VT applications with a pre-registered, balanced sample of actual tourists at the intended museum sidewalk and village locations, with at least 20 per group and matched devices. The central claim weakens if novices no longer show lower pragmatic quality and higher cybersickness, or if the AR app's UX advantage over VR disappears.","supporting_citations":[{"cited_title":"In: 2025 IEEE In- ternational Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR)","cited_arxiv_id":null,"evidence_quote":"Supplies the user-centric evaluation methodology and the AR-VT/VR-VT use-case descriptions on which the study is built."},{"cited_title":"Online, www.igroup.org/pq/ipq/items.php","cited_arxiv_id":null,"evidence_quote":"Supplies the IPQ used to measure presence in the VR application."},{"cited_title":"https://doi.org/10.13140/RG.2.2.36571.03362","cited_arxiv_id":null,"evidence_quote":"Supplies the CSQ-VR used to measure cybersickness symptoms in VR."},{"cited_title":"International Journal of Human-Computer Interaction35(6), 456–467 (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the ATI scale used to measure technological affinity and its correlations with UX."},{"cited_title":"Journal of behavior therapy and experimental psychiatry 25(1), 49–59 (1994)","cited_arxiv_id":null,"evidence_quote":"Supplies the SAM used to measure valence, arousal, and dominance."},{"cited_title":"IEEE Transactions on Visualization and Computer Graphics pp","cited_arxiv_id":null,"evidence_quote":"Supports the interpretation that repeated VR exposure reduces cybersickness."},{"cited_title":"Frontiers in Virtual Reality2(Oct 2021)","cited_arxiv_id":null,"evidence_quote":"Supports the age-related presence findings for older adults in immersive VR."}],"review_version":1}