{"id":"4da1745b-cde8-45ba-b453-8a7493253ad2","arxiv_id":"2505.07377","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM-driven peer questions in a VR classroom increase students' fixation on instructional content and cognitive load for complex topics, without significant learning gains.","lead":"Researchers built a virtual classroom where the teacher and student peers are powered by LLMs, and watched how real students looked and learned when AI peers asked questions. The study suggests peer questions focus attention on the lesson content, especially for hard topics, without adding distracting mental effort.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline eye-tracking and pupil p-values appear to treat events or samples, not participants, as independent; if corrected, the scanpath and cognitive-load claims may not survive.","rationale":"The reader's verdict of CONDITIONAL is reasonable, but the specific weakest assumption identified, the topic/condition confound, is only partially accurate. The same-topic comparisons are between-subjects: for the Double-Slit topic, participants in Peer-QnA and Peer-NoQnA are different people, so topic content is held constant. The confound is genuine for the cross-topic complexity comparison, but the authors explicitly label those differences as exploratory. A more load-bearing threat is the unit-of-analysis problem. The reported SDs for fixation duration, saccade amplitude, and pupil diameter are only plausible if the data were pooled across eye-movement events or time samples rather than aggregated to participants. The pupil p<.001 is mathematically inconsistent with a participant-level t-test given the reported means, SDs, and group sizes. This directly undermines the central claims that peer questions produce more targeted scanpaths and higher cognitive engagement. If the participant-level reanalysis fails, the remaining support is the normalized total fixation duration on instructional content and NASA-TLX in one topic, which is a much narrower result than the abstract and design recommendations suggest. The paper is transparent about its small sample and exploratory nature, and the authors disclose that topic complexity was not manipulated, so this is not a rejection-level flaw. It should remain CONDITIONAL, requiring the authors to provide participant-level analyses and to soften claims that depend on the current p-values.","tokens_in":20114,"tokens_out":10850,"duration_ms":101728,"concrete_test":"Reanalyze the raw eye-tracking and pupil data with participant as the unit of analysis: for each participant and condition, compute mean fixation duration, mean saccade amplitude, and mean pupil diameter; then recompute the Double-Slit and History comparisons with t-tests or mixed-effects models including participant random intercepts and order as a covariate. If the 5 ms fixation difference, the roughly 2 degree saccade difference, and the pupil p<.001 no longer reach significance, the scanpath and cognitive-load claims are a unit-of-analysis artifact. As a quick arithmetic check, compute the two-sample t-statistic from the reported Double-Slit pupil means, SDs, and group sizes; p<.001 is not consistent with n around 9-10, confirming that the current inference is not participant-level.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is not the topic/condition pairing per se, but the unit of analysis in the headline eye-tracking and pupil statistics. For the Double-Slit topic, the paper reports mean fixation duration 233 vs 228 ms with SDs around 103/101 ms and p<.001, and saccade amplitude about 136 vs 138 degrees with SDs around 69 degrees and p=.015. Those SDs describe raw fixation/saccade events, not participant-level means, so the tests appear to treat thousands of non-independent eye movements as independent observations. At the participant level, with roughly 9-10 participants per condition, a 5 ms difference would not be significant. The pupil-diameter result is even more telling: M=.59, SD=.11 vs M=.51, SD=.19, p<.001 is not reproducible from an independent t-test with these group sizes; the p-value implies either sample-level analysis or an error. Consequently, the paper's contribution (4), that peer questions led to longer mean fixation duration and shorter average saccade amplitude, and the pupil-based cognitive-load evidence (contribution 3) rest on a likely unit-of-analysis artifact. The normalized total fixation duration on instructional content (0.72 vs 0.60, p=.02) and NASA-TLX (60.50 vs 41.73, p=.026) appear to be participant-level and could remain valid, but they are the only two headline results that would survive a correct analysis. The topic/condition confound is real for the cross-topic complexity comparison, which the authors acknowledge is exploratory, but the same-topic comparisons are between-subjects and thus less confounded by topic content.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a virtual-reality classroom study in which all instructors and peers are driven by large language models. Nineteen participants experienced two conditions—Peer-QnA, where LLM-driven peer avatars ask questions, and Peer-NoQnA, where they do not—across two topics (Double-Slit Experiment and History of Video Games) in a counterbalanced within-subjects design. Using head-mounted eye tracking, NASA-TLX, pupil diameter, and knowledge questionnaires, the paper reports that in the Double-Slit topic the Peer-QnA condition increased cognitive load and directed attention to primary instructional content, with longer mean fixation durations and shorter saccade amplitudes, while such differences were absent in the History of Video Games topic. The authors interpret the results through cognitive load and signaling theory and derive design recommendations for LLM-driven VR learning environments. The manuscript acknowledges that topic complexity was not manipulated and that the study is exploratory on that dimension.","tokens_in":20324,"tokens_out":4868,"duration_ms":46708,"significance":"If the central results were valid, the paper would provide an initial empirical account of how LLM-driven peer interactions affect attention and cognitive engagement in VR classrooms, a topic of active interest in human-computer interaction and AI in education. The fully LLM-driven environment is a useful proof of concept, and the combination of eye tracking, subjective workload, and learning-outcome measures is appropriate for this research direction. The paper also gives explicit design implications that could guide future system builders. However, the evidential weight of the reported findings is presently limited by two structural problems: the statistical analyses appear to use non-independent events or samples as the unit of analysis for several headline results, and the experimental design confounds interaction condition with topic content and order. No data or analysis code are provided, so the reported statistics cannot be independently verified.","major_comments":[{"comment":"The p-values for mean fixation duration, saccade amplitude, and pupil diameter appear to be computed on individual eye-movement events or pupil samples rather than on participant-level aggregates. For example, in §4.1.2 the mean fixation durations are 233 ms vs. 228 ms with standard deviations of 103 ms and 101 ms; with roughly 9–10 participants per condition, a participant-level t-test on a 5 ms difference cannot produce p<.001. Similarly, in §4.1.1 the pupil-diameter means are 0.59 (SD=0.11) vs. 0.51 (SD=0.19) with p<.001, which is not consistent with an independent-samples t-test at the participant level (a rough calculation gives p≈.27). The paper must state the unit of analysis for each test. If eye-movement events or pupil samples were treated as independent observations, the tests are invalid because those observations are strongly non-independent within participants, and the reported significance levels are inflated. This directly affects contributions (3) and (4), which rest on pupil-based cognitive load and on the longer-fixation/shorter-saccade findings.","section":"§4.1.1, §4.1.2, §4.2.2"},{"comment":"The design confounds interaction condition with topic. Each participant experienced Peer-QnA on one topic and Peer-NoQnA on the other, and the topic–condition pairing was counterbalanced across only four cases. Consequently, the comparisons within a given topic (e.g., Double-Slit Peer-QnA vs. Peer-NoQnA) are between-subjects, and the participant groups differ not only in the interaction condition but also in the condition assigned to the other topic and in topic order. The observed differences in the Double-Slit Experiment therefore cannot be uniquely attributed to the question-asking manipulation. The paper itself states that topic complexity was not manipulated, so the cross-topic interpretation in §5.1 (that peer questions are more effective in complex subjects) is exploratory. The design recommendations in §5.2, which claim that in more complex subjects peer interactions more effectively guide attention, go beyond what this design can support.","section":"§3.3, §5.1"},{"comment":"The statistical reporting is incomplete and internally inconsistent. The analysis section says that independent t-tests or Wilcoxon signed-rank tests were used depending on normality, but for most results the paper does not state which test was used or whether the comparison was paired or unpaired. Given the within-subjects structure of the overall design, paired analyses would be expected for many comparisons, but the reported means and standard deviations appear to be presented as between-subjects group summaries. Additionally, no correction for multiple comparisons is applied across the many metrics tested (fixation duration, saccade amplitude, saccade velocity, pupil diameter, NASA-TLX, knowledge scores) for each topic, inflating the risk of false positives. The paper should provide a complete statistical table with test names, sample sizes, and effect sizes, and should re-analyze the data with participant as a random effect when aggregating over events.","section":"§3.7, §4.1–§4.2"}],"minor_comments":[{"comment":"The saccade mean velocity values are reported as M=1.28°/s and M=1.26°/s in the History of Video Games topic, whereas the Double-Slit values in §4.1.2 are around 127–128°/s; this appears to be a decimal-place error and should be corrected.","section":"§4.2.2"},{"comment":"The degrees of freedom for the regression and correlation analyses are inconsistent: the text reports r(17) and F(1,17) in one place and r(18) and F(1,18) in another for the same Double-Slit knowledge-questionnaire analysis; these should be reconciled.","section":"§4.1.3"},{"comment":"The I-VT velocity and duration thresholds (Table 1) and the 1-second baseline correction for pupil diameter are analysis parameters; a sensitivity analysis or a justification for the chosen values would strengthen the paper, since several reported conclusions depend on these preprocessing choices.","section":"§3.6"},{"comment":"The text states both that higher saccade velocity indicates cognitive load and that higher average saccade velocity is associated with increased stress and reduced concentration; these statements should be clarified so the reader understands the expected direction for the experimental conditions.","section":"§3.5.1"},{"comment":"Calling the design 'within-subjects' throughout is misleading because the key comparisons within each topic are between participants; the paper should explicitly describe which comparisons are within-subject and which are between-subject.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The unit-of-analysis issue in the eye-tracking and pupil statistics is the most serious problem. If a re-analysis with participant-level means does not preserve the significant effects, contributions (3) and (4) will collapse, and the paper will be left with only the normalized total fixation duration and NASA-TLX results as support for its central claims. Even with re-analysis, the topic/condition confound limits the strength of the complexity-related conclusions. I would advise the editor that the manuscript needs a substantively revised empirical section before it can be considered for publication, and that the authors should be asked to provide the de-identified data and analysis scripts to verify the corrected statistics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe real news here is that they built a fully LLM-driven VR classroom—teacher and peers all LLM—and measured eye tracking, cognitive load, and learning. That is a legitimate first-step contribution for HCI/educational technology, and I think the system itself deserves to be seen. They also disclose the main limitations honestly: small sample, topic complexity not manipulated, LLM interaction quality unexamined. Credit where it's due.\n\nThe soft spots are not minor. The unit of analysis in several headline eye-tracking statistics appears to be individual fixations/saccades, not participants. Mean fixation duration 233 vs 228 ms with SDs around 103/101 ms across roughly nine people per condition cannot give p<.001. The pupil diameter result (M=.59 vs .51, SDs .11/.19, p<.001) is even more obviously a non-reproducible test at participant level. So contributions (3) and (4)—that peer questions increased cognitive load (pupil) and produced longer fixations/shorter saccades—likely rest on an artifact. What survives is the normalized total fixation on instructional content (0.72 vs 0.60, p=.02) and NASA-TLX (60.5 vs 41.7, p=.026). Those are participant-level and could be real, but they are modest and come from a small, between-subjects comparison within the Double-Slit topic.\n\nThe topic/condition confound is real but secondary. They counterbalanced which topic went with which condition, so the same-topic comparison is not confounded by topic content; the cross-topic “complexity” story is exploratory, and they say so. I wouldn't reject the paper just for that.\n\nMy take: this paper deserves a serious referee, but the current manuscript is not acceptable as-is. The authors need to reanalyze all eye-tracking and pupil data at participant level, report per-participant effect sizes, and soften any causal claims accordingly. If they do that, there may be a useful empirical nugget here. Send it to review, but the referee should be told to check the statistics carefully.","headline":"The system is worth reporting, but the headline eye-tracking and pupil p-values probably treat fixations as independent observations; only the normalized fixation-duration and NASA-TLX results look like they could survive a participant-level reanalysis.","tokens_in":717,"tokens_out":2657,"would_cite":false,"duration_ms":39529,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a fully LLM-driven virtual classroom, questions from AI peer avatars steer students' gaze toward the instructional content and raise cognitive engagement without adding extraneous load.","keywords":["LLM-driven virtual classroom","peer question-asking","eye tracking","cognitive load","NASA-TLX","virtual reality education","visual attention","LLM agents"],"falsifier":"A replication that holds the topic fixed—for example, a between-subjects design in which one group gets peer questions and another does not on the same technical lesson—would falsify the attentional-signal claim if normalized fixation on the instructional content no longer differs between groups.","tokens_in":19834,"feed_emoji":"👁️","tokens_out":9777,"duration_ms":81874,"temperature":0.7,"pith_summary":"This paper reports on a virtual-reality classroom in which the teacher and the fellow students are all powered by a large language model, and asks whether questions asked by those AI peers change how a human student attends and learns. The authors compared a condition where AI peers asked the teacher questions after each slide with a condition where only the participant could ask, across two topics: the Double-Slit Experiment and the History of Video Games. Their central claim is that peer questions act as attentional signals: in the more technical topic, the question condition produced more fixation on the instructional content ($M=0.72$ vs $M=0.60$ of normalized fixation time, $p=.02$), higher NASA-TLX workload, and larger pupil diameter, and the extra cognitive load correlated with attention to the material rather than with distraction. The paper concludes that LLM-driven peer questions can focus attention and deepen engagement on complex subjects without introducing extraneous load, and that this should guide the design of VR learning spaces.","feed_headline":"AI peer questions keep VR students' eyes on the lesson","feed_subtitle":"In an LLM-driven virtual classroom, AI classmates' questions raised attention to the material, mainly on hard topics.","key_machinery":"The load-bearing object is the fully LLM-driven virtual classroom itself: an avatar teacher delivers slide content and AI peer avatars either ask questions or stay silent, with the speech generated by an LLM in real time. The argument runs through eye-tracking metrics—normalized total fixation duration on the mainboard and teacher, mean fixation duration, saccade amplitude, and pupil diameter—plus NASA-TLX workload scores. These measures are used, within cognitive load theory, to separate extraneous load from germane load; the key interpretive move is that fixation time on the instructional content tracks germane processing, so the load increase from peer questions counts as engagement rather than distraction.","core_discovery":"The discovery the paper pursues is that a fully LLM-driven virtual classroom is not merely a believable simulation; the behavior of AI peers measurably steers a human learner's visual attention. In the Peer-QnA condition, AI student avatars raised their hands and asked questions after each slide, while in Peer-NoQnA they stayed silent. For the Double-Slit Experiment, the paper reports significantly longer normalized total fixation duration on the mainboard and teacher, shorter saccade amplitudes, higher NASA-TLX workload ($M=60.50$ vs $M=41.73$, $p=.026$), and higher pupil diameter ($p<.001$), with a positive correlation between cognitive load and fixation on the instructional content ($r(18)=0.60$, $p=.0067$). The authors interpret this as evidence that the questions acted as signals directing attention to the material, so the added load was germane rather than extraneous. For the less technical History of Video Games topic, no attention or load differences appeared and quiz scores only trended higher ($p=.056$), which the paper reads as evidence that content complexity moderates the effect.","pith_inferences":["Because the Peer-QnA condition was always paired with one topic and Peer-NoQnA with the other, the Double-Slit results are entangled with topic content and order; a fully crossed or between-subjects replication is needed to confirm the effect is about peer questions rather than the topic.","If the attentional-signal account is correct, the effect should scale with question relevance: AI peer questions aimed at a specific slide element should produce an even sharper gaze shift than generic clarifying questions.","A testable extension would be to manipulate topic complexity as an independent variable and compare scripted versus generated questions, separating the signal value of a peer question from its linguistic content."],"forward_implications":["In fully LLM-driven VR classrooms, having AI peers ask questions after each slide can be used as a design tool to direct students' visual attention to the teacher and mainboard, at least in technical lessons.","The cognitive load increase that accompanies peer questions should be interpreted as engagement with the material, because it correlates with time spent fixating the instructional content.","For simpler topics, peer questions may leave attention and load unchanged, while the observed near-significant rise in quiz scores suggests they can still support learning through other routes.","Designers can follow the paper's recommendation to include peer question-asking for complex content while monitoring speech-to-text reliability and other technical issues that reduce the experience."],"supporting_citations":[{"why":"The prior VR-classroom study whose eye-tracking setup and areas-of-interest approach this experiment reuses.","marker":"[24]"},{"why":"The direct predecessor system with an LLM-driven active student in VR that this study builds on.","marker":"[58]"},{"why":"Cognitive load theory, the framework used to interpret load as extraneous versus germane.","marker":"[94]"},{"why":"Signaling in multimedia learning, the theory cited for treating peer questions as attention-directing cues.","marker":"[64]"},{"why":"Instructional cues in prose processing, used to support the claim that questions guide attention.","marker":"[29]"},{"why":"Eye-tracking review linking fixation duration to attention and cognitive processing.","marker":"[30]"},{"why":"The source for interpreting saccade amplitude as an index of mental effort and focused attention.","marker":"[17]"},{"why":"The model used to generate the teacher's and peers' speech in the virtual classroom.","marker":"[69]"},{"why":"The I-VT algorithm used to classify fixations and saccades from raw gaze data.","marker":"[84]"},{"why":"The baseline-correction method applied to pupil diameter before analysis.","marker":"[63]"}],"fun_headline_variants":["LLM peers' questions focus VR learners' eyes on content","AI classmates' queries sharpen attention in VR lessons","Virtual AI peers steer gaze to hard material in VR study","In VR classrooms, AI questions direct attention to slides"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that differences between the conditions are caused by the peer questions, even though each participant experienced the question condition on one topic and the no-question condition on the other, so topic and condition are entangled.","fun_headline_variants_meta":{"raw":{"variants":["LLM peers' questions focus VR learners' eyes on content","AI classmates' queries sharpen attention in VR lessons","Virtual AI peers steer gaze to hard material in VR study","In VR classrooms, AI questions direct attention to slides"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":1159,"prompt_tokens":945,"completion_tokens":214,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":149}},"tokens_in":561,"tokens_out":214,"duration_ms":2474,"temperature":1.0,"reasoning_tokens":149,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:16:38.130452+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that holds the topic fixed—for example, a between-subjects design in which one group gets peer questions and another does not on the same technical lesson—would falsify the attentional-signal claim if normalized fixation on the instructional content no longer differs between groups.","supporting_citations":[{"cited_title":"Lewandowska, I","cited_arxiv_id":null,"evidence_quote":"The direct predecessor system with an LLM-driven active student in VR that this study builds on."},{"cited_title":"Shemshack and J","cited_arxiv_id":null,"evidence_quote":"Cognitive load theory, the framework used to interpret load as extraneous versus germane."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Signaling in multimedia learning, the theory cited for treating peer questions as attention-directing cues."},{"cited_title":"Mathˆ ot, J","cited_arxiv_id":null,"evidence_quote":"The model used to generate the teacher's and peers' speech in the virtual classroom."},{"cited_title":"Radianti, T","cited_arxiv_id":null,"evidence_quote":"The I-VT algorithm used to classify fixations and saccades from raw gaze data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The baseline-correction method applied to pupil diameter before analysis."}],"review_version":1}