{"id":"2447cdde-75e1-47e8-bcf1-261ddb0c549f","arxiv_id":"2411.18438","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Moderate adaptive AI guidance in VR pizza-making increased gaze on the tutor and reduced head movement versus a non-adaptive baseline.","lead":"Researchers built a VR pizza-making experience with an AI tutor that adapts its guidance to user choices and demographics. In a 54-person study, moderate adaptivity boosted attention to the tutor and cut exploratory head movement, suggesting balanced AI support works best.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI response latency and longer tutor utterances in adaptive conditions may confound the gaze and head-movement metrics, so the central engagement claim is not yet secure.","rationale":"The reader's weakest-assumption analysis identified the same core threat: the behavioral metrics may be biased by AI response timing and longer tutor speech in the adaptive conditions. This is indeed the most load-bearing issue. The paper itself flags AI variability in Section 5.4 but does not control it, and the discussion in Section 5.1 only addresses the temporal metric, not dwell time or head movement. A re-analysis that excludes or adjusts for waiting/speaking intervals would either confirm the effects as genuine engagement signals or reveal them as artifacts. The additional gap—'optimal' versus pairwise non-significance between moderate and high—is secondary but reinforces the need for a conditional verdict rather than full acceptance. The authors' open-source release makes the proposed check feasible. My recommendation matches the reader's conditional stance because the central claim can likely be repaired with an appropriate re-analysis, but as reported it is not fully established.","tokens_in":16910,"tokens_out":2165,"duration_ms":23603,"concrete_test":"Using the session logs in the open-source repository, partition each participant's timeline into system states: active task interaction, participant speech capture (STT), GPT-4 inference wait, and tutor TTS playback. Recompute relative avatar dwell time and mean head-movement speed using only active-task intervals, and also compute per-condition total wait plus TTS duration. Run the same Kruskal-Wallis and Dunn tests on the filtered metrics. If the No-vs-Moderate differences lose significance or the effect sizes drop by more than half, the original claim is confounded by latency and speech length. As a secondary check, include per-participant total non-task time as a covariate in an ANCOVA or regression on the unfiltered metrics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—moderate adaptivity increases visual attention on the AI avatar and reduces exploratory head movement—rests on two behavioral metrics, relative avatar dwell time and average head movement speed. Both are vulnerable to an uncontrolled confound: the adaptive system introduces longer waiting and speaking intervals. Section 3.1 states that each STT interaction records 5 seconds of participant speech and GPT-4 responses arrive after 2–5 seconds; Section 3.4 notes the tutor always answers via TTS after optional speech. Appendix Table 2 shows that moderate and high conditions produce substantially longer tutor utterances than the baseline. Section 3.7 computes dwell time as a percentage of total session time and head speed as a mean over the entire session. If participants orient toward the avatar while waiting for a response or stand relatively still during idle periods, the moderate condition's higher dwell time and lower head speed could be artifacts of system latency and utterance length, not of engagement. The paper acknowledges this for the temporal metric in Section 5.1 ('variability in AI response times may influence temporal metrics') and for general AI variability in Section 5.4, but it does not establish that the gaze and head-movement metrics are exempt. The 'optimal' label is also stronger than the data: Sections 4.2 and 4.3 show no significant difference between moderate and high adaptivity on any key metric, so the inverted-U interpretation is not directly supported by pairwise tests. The load-bearing assumption is that the behavioral metrics isolate engagement from system-timing artifacts; this assumption is currently untested and is contradicted by the system description.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a between-subjects VR user study (N=54, three conditions: no, moderate, and high adaptive Gen-AI guidance in a Neapolitan pizza-making task) and reports multimodal behavioral measures of engagement: total session time, relative avatar dwell time from eye tracking, average head movement speed, and presence of verbal interaction. The authors conclude that moderate adaptivity optimally enhances engagement, increasing visual attention on the AI tutor and reducing exploratory head movement. The paper also contributes an open-source VR system and design recommendations for adaptive educational technologies.","tokens_in":17128,"tokens_out":4455,"duration_ms":44206,"significance":"If the claimed effects are valid, the paper would be useful for the HCI and adaptive-learning communities: it demonstrates a concrete adaptive Gen-AI tutor in a procedural ICH context, uses multimodal behavioral metrics rather than self-report alone, and offers practical design guidance. The study is competently analyzed with appropriate nonparametric tests and effect sizes, and the open-source contribution is a strength. However, the central conclusion rests on two behavioral metrics that are vulnerable to a timing confound, and the 'optimal'/'sweet-spot' characterization is stronger than the pairwise test results support. These issues are addressable with additional analyses, but they currently block acceptance.","major_comments":[{"comment":"The two central metrics—relative avatar dwell time and average head movement speed—are computed over the entire session, but adaptive conditions contain systematically more waiting and listening time. Section 3.1 states that each STT interaction records 5 seconds of participant speech and GPT-4 responses arrive after 2–5 s, and Appendix Table 2 shows longer tutor utterances in the moderate and high conditions. As a result, adaptive sessions last about 200 s longer (§4.1), so participants have more time to gaze at the avatar while waiting or listening and more time with reduced head movement. Since dwell time is a percentage of total session time and head speed is averaged over the whole session, the observed differences could arise mechanically from interface latency and utterance length rather than from engagement. The manuscript notes timing variability for temporal metrics (§5.1) and AI variability in general (§5.4), but it does not establish that the gaze and head-movement metrics are exempt. Please report event-locked analyses (e.g., dwell time during tutor speech only, head speed during active task segments) or include response latency / utterance length as covariates to disentangle the confound.","section":"§3.1, §3.7, §5.1"},{"comment":"The claim that moderate adaptivity is 'optimal' and reflects a non-monotonic sweet spot is not supported by the pairwise comparisons. For dwell time, only No vs Moderate is significant (p=.005); No vs High is not (p=.137) and Moderate vs High is not (p=.178). For head speed, only No vs Moderate is significant (p=.029); No vs High is marginal (p=.058) and Moderate vs High is not (p=.687). Total session time also shows no significant difference between Moderate and High (p=.927). The data therefore support at most a moderate-vs-baseline effect, not an inverted-U relationship. Please soften the optimality claim or provide explicit evidence of a nonlinear trend (e.g., a contrast test or a model with a quadratic term) before asserting that moderate adaptivity is superior to high adaptivity.","section":"§4.2, §4.3, §5.1"},{"comment":"Lower average head movement speed is interpreted as 'reduced unnecessary exploratory behaviour,' but the metric is an undifferentiated whole-session average. Reduced speed could also indicate passivity, reduced physical engagement, or simply more time spent listening to longer tutor utterances. Given that the high-adaptivity condition shows a similar but non-significant reduction, the interpretation depends on a finer-grained analysis. The authors should either validate the head-movement measure against task-relevant exploration (e.g., head movement during active ingredient selection and dough preparation) or temper the claim that exploratory behaviour specifically was reduced.","section":"§4.3, §5.1"}],"minor_comments":[{"comment":"The text reports p<.001 for the Welch ANOVA, but the figure annotation uses four asterisks, which the caption maps to p<.0001; please align the reported p-value and the significance notation.","section":"Figure 5 caption"},{"comment":"The caption states 'Dunn's post-hoc' but does not report which pairwise comparison is significant or its direction and p-value; please provide the complete post-hoc result.","section":"Appendix Figure 9 caption"},{"comment":"The sentence 'Each STT interaction recorded 5 seconds of participant speech' is ambiguous: it could mean a fixed five-second recording window, a maximum duration, or a post-hoc extraction. Please clarify the recording procedure and whether its length varied across conditions or interactions.","section":"§3.1"},{"comment":"Several reference entries are malformed, e.g., [5] lists 'Pamela Beach and Jen McConnel and. 2019' and [60] lists 'K. Renninger and Suzanne Hidi. 2016' with missing author names; please correct the author lists.","section":"References"},{"comment":"The experimental design would benefit from a manipulation check: participants were not told about the Gen-AI agent, but no measure reports whether they noticed differences in adaptivity across conditions. Adding a brief post-experiment awareness question would strengthen the construct validity of the adaptivity manipulation.","section":"§3.4"}],"recommendation":"major_revision","confidential_remarks":"The timing confound is the key issue. It is potentially fixable through reanalysis, and the paper's open-source release would allow such analyses to be checked. I would not reject outright, but if the authors cannot provide event-level or covariate-controlled analyses, the central engagement claim should be substantially weakened to reflect that the effects may be partly attributable to system latency and utterance length."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competently run between-subjects VR study comparing no, moderate, and high adaptive Gen-AI guidance in a pizza-making task. The new bit is the three-level adaptivity comparison in a culinary cultural-heritage context, and the pattern that moderate beats baseline on two multimodal engagement metrics. That empirical finding is worth taking seriously, but the central claim is not secure yet, for two reasons.\n\nFirst, the confound the paper partly admits: adaptive conditions have GPT-4 delays of 2–5 s and longer generated TTS utterances. Both avatar dwell time and head-movement speed are computed over the whole session. If participants look at a talking avatar and stand still while listening to longer speech, those metrics will move in the observed direction without meaningfully signaling 'engagement' in the intended sense. The authors acknowledge this for temporal metrics but assume the gaze and head-movement measures are exempt, and they do not test that assumption. Second, the 'moderate is optimal' framing is stronger than the data: pairwise tests show no significant difference between moderate and high on any key metric. The inverted-U story is plausible but not directly supported.\n\nCredit where it is due: the experiment is cleanly executed, the statistical tests and effect sizes are appropriate, the sample is decent for a between-subjects VR study, and the limitations section is candid. The open-source link is a plus. The literature coverage is reasonable, and the self-citations are background, not load-bearing.\n\nWho this is for: HCI/VR education researchers interested in adaptive AI evaluation. The confound is addressable in revision—re-analyze with speech duration as a covariate, compute metrics per interaction phase, or restrict gaze analysis to non-speaking intervals—and the 'optimal' conclusion should be softened. I would send it to review; it deserves a serious referee, but the current version should not pass as is.","headline":"Competent VR user study with a real confound: AI response latency and utterance length may drive the gaze and head-movement effects, and the 'moderate is optimal' claim outruns the statistics.","tokens_in":17677,"tokens_out":2226,"would_cite":false,"duration_ms":23208,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Moderate adaptive AI guidance in VR reliably raises engagement measures compared with no adaptivity, and high adaptivity adds no further gain.","keywords":["intangible cultural heritage","virtual reality","generative artificial intelligence","adaptive learning","multimodal engagement","eye tracking","head movement","procedural learning"],"falsifier":"A replication that fixes AI response latency and tutor speech duration across the three conditions would settle the issue: if the moderate-adaptivity advantage in avatar dwell time and the reduction in head-movement speed disappear, the claimed engagement effect is largely an artifact of pacing.","tokens_in":16715,"feed_emoji":"🍕","tokens_out":5143,"duration_ms":43191,"temperature":0.7,"pith_summary":"The paper asks whether the amount of adaptivity in a generative-AI tutor changes how engaged learners are in a virtual-reality procedural task. It builds a VR Neapolitan pizza-making environment with a GPT-4 tutor and compares three between-subject conditions: a fixed script, adaptation to ingredient choices, and adaptation to both ingredient choices and demographics. Using eye tracking, head movement, session time, and optional speech as behavioural proxies, it finds that moderate adaptivity produces the highest avatar dwell time and the lowest exploratory head movement relative to baseline, while high adaptivity does not improve on moderate. The authors read this as evidence for a sweet spot: real-time, action-based adaptation supports attention and focus, but adding demographic-based personalization gives no extra engagement.","feed_headline":"Moderate AI guidance beats no adaptivity in VR pizza training","feed_subtitle":"Adapting to ingredient choices lifts gaze and cuts wandering; demographic tailoring adds nothing.","key_machinery":"The central mechanism is a three-level adaptivity manipulation embedded in a GPT-4-based virtual tutor: moderate adaptivity responds to ingredient choices, high adaptivity adds demographic personalization, and the baseline uses a fixed script. Engagement is measured through multimodal behavioural proxies — relative avatar dwell time from eye-tracking, average head-movement speed, total session duration, and a binary verbal-interaction indicator — which together are meant to capture visual attention, physical exploration, and social engagement while avoiding the confound of subjective questionnaires.","core_discovery":"On the paper's own terms, the central discovery is that moderate adaptivity — generated responses tailored only to real-time ingredient choices — significantly increases the percentage of session time spent gazing at the AI tutor and decreases average head-movement speed compared with a non-adaptive baseline, whereas high adaptivity (adding demographic-based personalization) yields no further significant gains. The paper also reports significantly longer total session time in both adaptive conditions relative to baseline and no significant differences in optional verbal interaction. It interprets these patterns as showing that balanced adaptivity sustains engagement and focused attention without removing learner agency, and it generalizes this to procedural cultural-heritage learning.","pith_inferences":["Beyond the paper: because adaptive conditions also had longer AI response delays and longer tutor utterances, part of the higher avatar dwell time and slower head movement may reflect waiting behavior rather than engagement; a replication that equalizes response latency and speech duration would test this.","Beyond the paper: the sweet-spot result suggests a design heuristic for other procedural heritage tasks — adapt to the learner's current action rather than to static demographic profiles — but only if the engagement proxy is validated against learning outcomes.","Beyond the paper: the absence of verbal-interaction differences could imply that proactive avatar speech dominates the dialogue; letting users drive questions might reveal different engagement patterns."],"forward_implications":["Systems with moderate, action-based adaptivity can be expected to produce longer VR sessions than fixed-script tutors, a direct corollary of the reported session-time effect.","Designers should favor a bounded adaptivity level that responds to user choices rather than maximal personalization, since high adaptivity did not outperform moderate adaptivity on any engagement metric.","Multimodal behavioural metrics such as gaze and head motion can reveal engagement differences that a single temporal metric might confound with AI response latency.","Adaptive Gen-AI tutors that proactively provide guidance may not need to prompt users to speak, since verbal interaction was low and not condition-dependent."],"supporting_citations":[{"why":"Supplies head-movement speed as a measure of physical exploration and spatial engagement in VR.","marker":"[78]"},{"why":"Supports using multimodal eye-tracking and head-movement metrics to assess cognitive engagement in VR learning.","marker":"[20]"},{"why":"Grounds dwell time on task-relevant areas as an indicator of attention in immersive VR education.","marker":"[49]"},{"why":"Provides the guidance-fading effect used to explain why high adaptivity failed to improve engagement.","marker":"[67]"},{"why":"Defines engagement in human-agent interaction, the conceptual target of the study.","marker":"[26]"},{"why":"Supports the interpretation that proactive social cues reduce users' reliance on explicit verbal interaction.","marker":"[23]"},{"why":"Supplies the gaze-ray casting AOI method used to compute dwell time on the avatar.","marker":"[9]"}],"fun_headline_variants":["Moderate AI tutoring best for VR pizza engagement","Too much AI adaptivity doesn't help VR pizza learners","Moderate Gen-AI adaptivity boosts VR pizza focus","Balanced AI guidance wins in VR pizza training","Moderate adaptivity best for engagement in VR pizza"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gaze and head-movement measures are unbiased proxies for engagement, rather than artifacts of the longer wait times and longer AI speech that the adaptive conditions introduce.","fun_headline_variants_meta":{"raw":{"variants":["Moderate AI tutoring best for VR pizza engagement","Too much AI adaptivity doesn't help VR pizza learners","Moderate Gen-AI adaptivity boosts VR pizza focus","Balanced AI guidance wins in VR pizza training","Moderate adaptivity best for engagement in VR pizza"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000353,"raw_usage":{"total_tokens":1866,"prompt_tokens":832,"completion_tokens":1034,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":958}},"tokens_in":448,"tokens_out":1034,"duration_ms":6838,"temperature":1.0,"reasoning_tokens":958,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:12:05.048105+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that fixes AI response latency and tutor speech duration across the three conditions would settle the issue: if the moderate-adaptivity advantage in avatar dwell time and the reduction in head-movement speed disappear, the claimed engagement effect is largely an artifact of pacing.","supporting_citations":[{"cited_title":"Yaremych and Susan Persky","cited_arxiv_id":null,"evidence_quote":"Supplies head-movement speed as a measure of physical exploration and spatial engagement in VR."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds dwell time on task-relevant areas as an indicator of attention in immersive VR education."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the guidance-fading effect used to explain why high adaptivity failed to improve engagement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines engagement in human-agent interaction, the conceptual target of the study."}],"review_version":1}