Pith. sign in

REVIEW 3 major objections 5 minor 112 references

Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read In a fully LLM-driven virtual classroom, questions from AI peer avatars steer students' gaze toward the instructional content and raise cognitive engagement without adding extraneous load.

desk verdict The system is worth reporting, but the headline eye-tracking and pupil p-values probably treat fixations as independent observations; only the normalized fixation-duration and NASA-TLX results look like they could survive a participant-level reanalysis. read the letter →

arxiv 2505.07377 v1 pith:HWMOXN5Y submitted 2025-05-12 cs.HC cs.AI

classification cs.HCcs.AI
keywords LLM-drivenvirtualclassroompeerquestion-askingeyetrackingcognitiveloadNASA-TLXrealityeducationvisualattentionLLMagents
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports on a virtual-reality classroom in which the teacher and the fellow students are all powered by a large language model, and asks whether questions asked by those AI peers change how a human student attends and learns. The authors compared a condition where AI peers asked the teacher questions after each slide with a condition where only the participant could ask, across two topics: the Double-Slit Experiment and the History of Video Games. Their central claim is that peer questions act as attentional signals: in the more technical topic, the question condition produced more fixation on the instructional content ($M=0.72$ vs $M=0.60$ of normalized fixation time, $p=.02$), higher NASA-TLX workload, and larger pupil diameter, and the extra cognitive load correlated with attention to the material rather than with distraction. The paper concludes that LLM-driven peer questions can focus attention and deepen engagement on complex subjects without introducing extraneous load, and that this should guide the design of VR learning spaces.

What carries the argument

The load-bearing object is the fully LLM-driven virtual classroom itself: an avatar teacher delivers slide content and AI peer avatars either ask questions or stay silent, with the speech generated by an LLM in real time. The argument runs through eye-tracking metrics—normalized total fixation duration on the mainboard and teacher, mean fixation duration, saccade amplitude, and pupil diameter—plus NASA-TLX workload scores. These measures are used, within cognitive load theory, to separate extraneous load from germane load; the key interpretive move is that fixation time on the instructional content tracks germane processing, so the load increase from peer questions counts as engagement rather than distraction.

What would settle it

A replication that holds the topic fixed—for example, a between-subjects design in which one group gets peer questions and another does not on the same technical lesson—would falsify the attentional-signal claim if normalized fixation on the instructional content no longer differs between groups.

Watch

Extended reading notes

Core claim

The discovery the paper pursues is that a fully LLM-driven virtual classroom is not merely a believable simulation; the behavior of AI peers measurably steers a human learner's visual attention. In the Peer-QnA condition, AI student avatars raised their hands and asked questions after each slide, while in Peer-NoQnA they stayed silent. For the Double-Slit Experiment, the paper reports significantly longer normalized total fixation duration on the mainboard and teacher, shorter saccade amplitudes, higher NASA-TLX workload ($M=60.50$ vs $M=41.73$, $p=.026$), and higher pupil diameter ($p<.001$), with a positive correlation between cognitive load and fixation on the instructional content ($r(18)=0.60$, $p=.0067$). The authors interpret this as evidence that the questions acted as signals directing attention to the material, so the added load was germane rather than extraneous. For the less technical History of Video Games topic, no attention or load differences appeared and quiz scores only trended higher ($p=.056$), which the paper reads as evidence that content complexity moderates the effect.

Load-bearing premise

The central comparison assumes that differences between the conditions are caused by the peer questions, even though each participant experienced the question condition on one topic and the no-question condition on the other, so topic and condition are entangled.

Editorial extensions

If this is right

  • In fully LLM-driven VR classrooms, having AI peers ask questions after each slide can be used as a design tool to direct students' visual attention to the teacher and mainboard, at least in technical lessons.
  • The cognitive load increase that accompanies peer questions should be interpreted as engagement with the material, because it correlates with time spent fixating the instructional content.
  • For simpler topics, peer questions may leave attention and load unchanged, while the observed near-significant rise in quiz scores suggests they can still support learning through other routes.
  • Designers can follow the paper's recommendation to include peer question-asking for complex content while monitoring speech-to-text reliability and other technical issues that reduce the experience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the Peer-QnA condition was always paired with one topic and Peer-NoQnA with the other, the Double-Slit results are entangled with topic content and order; a fully crossed or between-subjects replication is needed to confirm the effect is about peer questions rather than the topic.
  • If the attentional-signal account is correct, the effect should scale with question relevance: AI peer questions aimed at a specific slide element should produce an even sharper gaze shift than generic clarifying questions.
  • A testable extension would be to manipulate topic complexity as an independent variable and compare scripted versus generated questions, separating the signal value of a peer question from its linguistic content.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a virtual-reality classroom study in which all instructors and peers are driven by large language models. Nineteen participants experienced two conditions—Peer-QnA, where LLM-driven peer avatars ask questions, and Peer-NoQnA, where they do not—across two topics (Double-Slit Experiment and History of Video Games) in a counterbalanced within-subjects design. Using head-mounted eye tracking, NASA-TLX, pupil diameter, and knowledge questionnaires, the paper reports that in the Double-Slit topic the Peer-QnA condition increased cognitive load and directed attention to primary instructional content, with longer mean fixation durations and shorter saccade amplitudes, while such differences were absent in the History of Video Games topic. The authors interpret the results through cognitive load and signaling theory and derive design recommendations for LLM-driven VR learning environments. The manuscript acknowledges that topic complexity was not manipulated and that the study is exploratory on that dimension.

Significance. If the central results were valid, the paper would provide an initial empirical account of how LLM-driven peer interactions affect attention and cognitive engagement in VR classrooms, a topic of active interest in human-computer interaction and AI in education. The fully LLM-driven environment is a useful proof of concept, and the combination of eye tracking, subjective workload, and learning-outcome measures is appropriate for this research direction. The paper also gives explicit design implications that could guide future system builders. However, the evidential weight of the reported findings is presently limited by two structural problems: the statistical analyses appear to use non-independent events or samples as the unit of analysis for several headline results, and the experimental design confounds interaction condition with topic content and order. No data or analysis code are provided, so the reported statistics cannot be independently verified.

major comments (3)
  1. [§4.1.1, §4.1.2, §4.2.2] The p-values for mean fixation duration, saccade amplitude, and pupil diameter appear to be computed on individual eye-movement events or pupil samples rather than on participant-level aggregates. For example, in §4.1.2 the mean fixation durations are 233 ms vs. 228 ms with standard deviations of 103 ms and 101 ms; with roughly 9–10 participants per condition, a participant-level t-test on a 5 ms difference cannot produce p<.001. Similarly, in §4.1.1 the pupil-diameter means are 0.59 (SD=0.11) vs. 0.51 (SD=0.19) with p<.001, which is not consistent with an independent-samples t-test at the participant level (a rough calculation gives p≈.27). The paper must state the unit of analysis for each test. If eye-movement events or pupil samples were treated as independent observations, the tests are invalid because those observations are strongly non-independent within participants, and the reported significance levels are inflated. This directly affects contributions (3) and (4), which rest on pupil-based cognitive load and on the longer-fixation/shorter-saccade findings.
  2. [§3.3, §5.1] The design confounds interaction condition with topic. Each participant experienced Peer-QnA on one topic and Peer-NoQnA on the other, and the topic–condition pairing was counterbalanced across only four cases. Consequently, the comparisons within a given topic (e.g., Double-Slit Peer-QnA vs. Peer-NoQnA) are between-subjects, and the participant groups differ not only in the interaction condition but also in the condition assigned to the other topic and in topic order. The observed differences in the Double-Slit Experiment therefore cannot be uniquely attributed to the question-asking manipulation. The paper itself states that topic complexity was not manipulated, so the cross-topic interpretation in §5.1 (that peer questions are more effective in complex subjects) is exploratory. The design recommendations in §5.2, which claim that in more complex subjects peer interactions more effectively guide attention, go beyond what this design can support.
  3. [§3.7, §4.1–§4.2] The statistical reporting is incomplete and internally inconsistent. The analysis section says that independent t-tests or Wilcoxon signed-rank tests were used depending on normality, but for most results the paper does not state which test was used or whether the comparison was paired or unpaired. Given the within-subjects structure of the overall design, paired analyses would be expected for many comparisons, but the reported means and standard deviations appear to be presented as between-subjects group summaries. Additionally, no correction for multiple comparisons is applied across the many metrics tested (fixation duration, saccade amplitude, saccade velocity, pupil diameter, NASA-TLX, knowledge scores) for each topic, inflating the risk of false positives. The paper should provide a complete statistical table with test names, sample sizes, and effect sizes, and should re-analyze the data with participant as a random effect when aggregating over events.
minor comments (5)
  1. [§4.2.2] The saccade mean velocity values are reported as M=1.28°/s and M=1.26°/s in the History of Video Games topic, whereas the Double-Slit values in §4.1.2 are around 127–128°/s; this appears to be a decimal-place error and should be corrected.
  2. [§4.1.3] The degrees of freedom for the regression and correlation analyses are inconsistent: the text reports r(17) and F(1,17) in one place and r(18) and F(1,18) in another for the same Double-Slit knowledge-questionnaire analysis; these should be reconciled.
  3. [§3.6] The I-VT velocity and duration thresholds (Table 1) and the 1-second baseline correction for pupil diameter are analysis parameters; a sensitivity analysis or a justification for the chosen values would strengthen the paper, since several reported conclusions depend on these preprocessing choices.
  4. [§3.5.1] The text states both that higher saccade velocity indicates cognitive load and that higher average saccade velocity is associated with increased stress and reduced concentration; these statements should be clarified so the reader understands the expected direction for the experimental conditions.
  5. [§3.3] Calling the design 'within-subjects' throughout is misleading because the key comparisons within each topic are between participants; the paper should explicitly describe which comparisons are within-subject and which are between-subject.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical comparisons of measured outcomes, not derivations that reduce to their inputs.

full rationale

The paper reports a within-subjects VR eye-tracking experiment comparing Peer-QnA and Peer-NoQnA conditions. Its central claims—higher normalized fixation duration on primary content, higher NASA-TLX scores and pupil diameters, longer mean fixation durations, and shorter saccade amplitudes under Peer-QnA—are direct comparisons of recorded measurements, not quantities fitted from or definitionally equivalent to the same measurements. The regression of cognitive load on fixation duration is descriptive and explicitly fit on the same participants; although the word 'predict' is used, no out-of-sample or cross-validated prediction is claimed, so it does not constitute a fitted-input-called-prediction. The interpretation that increased load is germane rather than extraneous is post hoc reasoning, which may be questionable but is not circularity. The acknowledged topic-condition pairing confound and the likely unit-of-analysis problem in the per-event eye-tracking statistics are serious validity concerns, but they concern statistical independence and causal attribution, not whether the derivation reduces to its inputs. Self-citations are present for standard eye-tracking and VR-classroom methodology, but the main results are not justified by these citations; they rest on the collected data. No uniqueness theorem, ansatz, or renamed-known-result is invoked. Therefore no circular step can be identified.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No new physical entities are introduced. The free parameters are standard eye-tracking thresholds and preprocessing choices. The assumptions are about what counts as instructional content, what topic complexity means, pupil diameter as cognitive load, and LLM behavior stability. The main unverified premise is the topic-complexity interpretation, which the authors explicitly acknowledge as exploratory.

free parameters (2)
  • I-VT velocity and duration thresholds = head velocity 7 deg/s, gaze velocity 30 deg/s fixation, 40 deg/s saccade; fixation duration 100-500 ms, saccade…
    Chosen by the authors based on prior work (Table 1). These thresholds directly set what counts as a fixation and saccade, thus shaping all attention metrics and the main results. They are not fitted to the current data, but they are hand-selected parameters.
  • Baseline correction duration for pupil diameter = 1 second
    Divisive baseline correction with a baseline duration of 1 second (Section 3.6). This choice affects pupil diameter comparisons and the cognitive load inference.
assumptions (4)
  • domain assumption The mainboard and teacher are the only primary instructional content, and other avatars are peripheral.
    The paper defines the mainboard and teacher as the primary instructional content (Section 3.5.1) and uses this to compute the main attention measure. If peer avatars' questions were themselves informative, excluding them from 'primary content' partly determines the observed shift toward the mainboard and teacher.
  • domain assumption Topic complexity is embodied by the two topic choices, with the Double-Slit Experiment more complex than History of Video Games.
    The paper repeatedly interprets topic differences as complexity effects (Sections 4 and 5), even though it states that complexity was not manipulated or measured. The conclusion that peer questions matter more in complex subjects rests on this unverified assumption.
  • domain assumption Pupil diameter changes reflect cognitive load rather than lighting, arousal, or other confounds.
    The paper uses pupil diameter as an objective measure of cognitive load (Section 3.5.2) without controlling for luminance changes in the VR scene or other factors. This is a standard assumption in eye-tracking research, but it is load-bearing for the cognitive load claim.
  • domain assumption LLM-driven avatars behave consistently across participants and conditions.
    The paper uses ChatGPT-4o in real time. The number, timing, and phrasing of AI peer questions vary because they are generated by the LLM, not scripted. This variability is not controlled in the analysis and could introduce noise or systematic differences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms." pith.science (2026). https://pith.science/paper/HWMOXN5Y

@misc{pith2026250507377,
  author       = {Pith},
  title        = {Pith review of: Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HWMOXN5Y}},
  note         = {Machine review of arXiv:2505.07377}
}
read the original abstract

Transforming educational technologies through the integration of large language models (LLMs) and virtual reality (VR) offers the potential for immersive and interactive learning experiences. However, the effects of LLMs on user engagement and attention in educational environments remain open questions. In this study, we utilized a fully LLM-driven virtual learning environment, where peers and teachers were LLM-driven, to examine how students behaved in such settings. Specifically, we investigate how peer question-asking behaviors influenced student engagement, attention, cognitive load, and learning outcomes and found that, in conditions where LLM-driven peer learners asked questions, students exhibited more targeted visual scanpaths, with their attention directed toward the learning content, particularly in complex subjects. Our results suggest that peer questions did not introduce extraneous cognitive load directly, as the cognitive load is strongly correlated with increased attention to the learning material. Considering these findings, we provide design recommendations for optimizing VR learning spaces.

Figures

Figures reproduced from arXiv: 2505.07377 by the authors.

Figure 1
Figure 1. Views from the LLM-driven virtual classroom environment. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Results for the Double-Slit Experiment across Peer-QnA and Peer-NoQnA conditions. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Linear regression results for the Double-Slit Exper [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Knowledge questionnaire scores for Peer-QnA and [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

112 extracted references · 74 canonical work pages

  1. [1]

    Virtual reality (VR) technologies have become increas- ingly prevalent in this transformation

    INTRODUCTION Education is undergoing a significant digital transforma- tion, accelerated by technological advancements and further driven by the COVID-19 pandemic, which necessitated a shift from in-person to digital learning environments [103, 15]. Virtual reality (VR) technologies have become increas- ingly prevalent in this transformation. Advances in ...

  2. [2]

    In the follow- ing subsections, we review the literature in these two areas, highlighting how each contributes to advancements in edu- cational technology

    RELA TED WORK In the evolving educational technology landscape, immer- sive VR and LLMs have become transformative tools that significantly impact learning environments. In the follow- ing subsections, we review the literature in these two areas, highlighting how each contributes to advancements in edu- cational technology. 2.1 VR Environments in Educatio...

  3. [3]

    ChatGPT-4o

    METHODOLOGY As LLM-driven classrooms become more common, it is im- portant to understand how students interact and learn in these new environments. The main purpose of this study is to evaluate student behaviors in fully LLM-driven classroom settings and to analyze the impact of LLM-driven peer ques- tions on engagement, attention, cognitive load, and lea...

  4. [4]

    Following this, we provided the findings from a general questionnaire, where user feedback on their overall experience was collected

    RESULTS We presented the results for each topic’s cognitive load anal- ysis, visual-scanpath analysis, and learning outcomes sep- arately. Following this, we provided the findings from a general questionnaire, where user feedback on their overall experience was collected. 4.1 Topic 1: Double-Slit Experiment 4.1.1 Cognitive Load Analysis In the Double-Slit...

  5. [5]

    I expe- rienced technical issues during the session. (R)

    DISCUSSION This section discusses the impact of LLM-driven peer inter- actions on cognitive load, attention, and learning outcomes. We explore how peer-driven questions and content complex- ity influence these factors on attention, engagement, cog- nitive load, and learning outcomes. Then, user feedback on the LLM-driven VR classroom environment provides ...

  6. [6]

    CONCLUSION In this study, we designed an individual learning environ- ment with a fully LLM-driven virtual classroom, where stu- dents could interact with LLM-driven teachers and engage in a classroom setting with LLM-driven peers who also inter- acted with the instructor. We investigated student behav- ior using eye-tracking data and cognitive load asses...

  7. [7]

    Abd-Alrazaq, R

    A. Abd-Alrazaq, R. AlSaad, D. Alhuwail, A. Ahmed, P. M. Healy, S. Latifi, S. Aziz, R. Damseh, S. Alabed Alrazak, and J. Sheikh. Large language models in medical education: Opportunities, challenges, and future directions. JMIR Medical Education, 9:1–9, 2023

  8. [8]

    Agtzidis, M

    I. Agtzidis, M. Startsev, and M. Dorr. 360-degree video gaze behaviour: A ground-truth data set and a classification algorithm for eye movements. In Proceedings of the 27th ACM international conference on multimedia, pages 1007–1015, 2019

Show all 112 references
  1. [9]

    Akkoyunlu and M

    B. Akkoyunlu and M. Y. Soylu. A study of student’s perceptions in a blended learning environment based on different learning styles. Journal of Educational Technology & Society, 11(1):183–193, 2008

  2. [10]

    Albus, A

    P. Albus, A. Vogt, and T. Seufert. Signaling in virtual reality influences learning outcome and cognitive load. Computers & Education , 166:104154, 2021

  3. [11]

    S. M. Alharbi, A. I. Elfeky, and E. S. Ahmed. The effect of e-collaborative learning environment on development of critical thinking and higher order thinking skills. Journal of Positive School Psychology , 6(6):6848–6854, 2022

  4. [12]

    S. Z. A. Ansari, V. K. Shukla, K. Saxena, and B. Filomeno. Implementing virtual reality in entertainment industry. In Cyber Intelligence and Information Retrieval: Proceedings of CIIR 2021 , pages 561–570. Springer, 2022

  5. [13]

    J. Beatty. Task-evoked pupillary responses, processing load, and the structure of processing resources. Psychological bulletin, 91(2):276, 1982

  6. [14]

    Behroozi, A

    M. Behroozi, A. Lui, I. Moore, D. Ford, and C. Parnin. Dazed: measuring the cognitive load of solving technical interview problems at the whiteboard. In Proceedings of the 40th International Conference on Software Engineering: New Ideas and Emerging Results, pages 93–96, 2018

  7. [15]

    Bozkir, D

    E. Bozkir, D. Geisler, and E. Kasneci. Assessment of driver attention during a safety critical situation in vr to generate vr-based training. In ACM Symposium on Applied Perception 2019 , pages 1–5, 2019

  8. [16]

    Bozkir, D

    E. Bozkir, D. Geisler, and E. Kasneci. Person independent, privacy preserving, and real time assessment of cognitive load using eye tracking in a virtual reality setup. In 2019 IEEE conference on virtual reality and 3D user interfaces (VR) , pages 1834–1837. IEEE, 2019

  9. [17]

    Bozkir, S

    E. Bozkir, S. ¨Ozdel, K. H. C. Lau, M. Wang, H. Gao, and E. Kasneci. Embedding large language models into extended reality: Opportunities and challenges for inclusion, engagement, and privacy. In Proceedings of the 6th ACM Conference on Conversational User Interfaces , pages 1–7, 2024

  10. [18]

    Bozkir, S

    E. Bozkir, S. ¨Ozdel, M. Wang, B. David-John, H. Gao, K. Butler, E. Jain, and E. Kasneci. Eye-tracked virtual reality: a comprehensive survey on methods and privacy challenges. arXiv preprint arXiv:2305.14080, 2023

  11. [19]

    Bozkir, P

    E. Bozkir, P. Stark, H. Gao, L. Hasenbein, J.-U. Hahn, E. Kasneci, and R. G ¨ollner. Exploiting object-of-interest information to understand attention in vr classrooms. In 2021 IEEE Virtual Reality and 3D User Interfaces (VR) , pages 597–605. IEEE, 2021

  12. [20]

    J. A. Bueno-Vesga, X. Xu, and H. He. The effects of cognitive load on engagement in a virtual reality learning environment. In 2021 IEEE Virtual Reality and 3D User Interfaces (VR) , pages 645–652. IEEE, 2021

  13. [21]

    Bygstad, E

    B. Bygstad, E. Øvrelid, S. Ludvigsen, and M. Dæhlen. From dual digitalization to digital learning space: Exploring the digital transformation of higher education. Computers & Education , 182:104463, 2022

  14. [22]

    S. N. Chakrabartty. Best split-half and maximum reliability. IOSR Journal of Research & Method in Education, 3(1):1–8, 2013

  15. [23]

    S. Chen, J. Epps, N. Ruiz, and F. Chen. Eye activity as a measure of human mental effort in hci. In Proceedings of the 16th international conference on Intelligent user interfaces , pages 315–318, 2011

  16. [24]

    Chen, H.-C

    S.-C. Chen, H.-C. She, M.-H. Chuang, J.-Y. Wu, J.-L. Tsai, and T.-P. Jung. Eye movements predict students’ computer-based assessment performance of physics concepts in different presentation modalities. Computers & Education , 74:61–72, 2014

  17. [25]

    Chirico, F

    A. Chirico, F. Lucidi, M. De Laurentiis, C. Milanese, A. Napoli, and A. Giordano. Virtual reality in health system: beyond entertainment. a mini-review on the efficacy of vr during cancer treatment. Journal of cellular physiology, 231(2):275–287, 2016

  18. [26]

    Christopoulos, M

    A. Christopoulos, M. Conrad, and M. Shukla. Increasing student engagement through virtual interactions: How? Virtual Reality, 22(4):353–369, 2018

  19. [27]

    Criollo-C, J

    S. Criollo-C, J. Cerezo, A. Guerrero-Arias, A. D. Samala, S. Rawas, and S. Luj´ an-Mora. Analysis of the mental workload associated with the use of virtual reality technology as support in the higher educational model. IEEE Access, 2024

  20. [28]

    Ferdinand, H

    J. Ferdinand, H. Gao, P. Stark, E. Bozkir, J.-U. Hahn, E. Kasneci, and R. G ¨ollner. The impact of a usefulness intervention on students’ learning achievement in a virtual biology lesson: An eye-tracking-based approach. Learning and Instruction, 90:101867, 2024

  21. [29]

    Freina and M

    L. Freina and M. Ott. A literature review on immersive virtual reality in education: state of the art and perspectives. In The international scientific conference elearning and software for education , pages 10–1007, 2015

  22. [30]

    H. Gao, E. Bozkir, L. Hasenbein, J.-U. Hahn, R. G¨ollner, and E. Kasneci. Digital transformations of classrooms in virtual reality. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–10, 2021

  23. [31]

    H. Gao, L. Hasenbein, E. Bozkir, R. G ¨ollner, and E. Kasneci. Evaluating the effects of virtual human animation on students in an immersive vr classroom using eye movements. In Proceedings of the 28th ACM Symposium on Virtual Reality Software and Technology, VRST ’22, New Yor...

  24. [32]

    H. Gao, L. Hasenbein, E. Bozkir, R. G ¨ollner, and E. Kasneci. Exploring gender differences in computational thinking learning in a vr classroom: Developing machine learning models using eye-tracking data and explaining the models. International Journal of Artificial Intellige...

  25. [33]

    H. Gao, H. Huai, S. Yildiz-Degirmenci, M. Bannert, and E. Kasneci. Datalivr: Transformation of data literacy education through virtual reality with chatgpt-powered enhancements. In 2024 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pages 120–129, 2024

  26. [34]

    Gibaldi and S

    A. Gibaldi and S. P. Sabatini. The saccade main sequence revised: A fast and repeatable tool for oculomotor analysis. Behavior Research Methods, 53:167–187, 2021

  27. [35]

    S. M. Glynn and F. J. Di Vesta. Control of prose processing via instructional and typographical cues. Journal of Educational Psychology , 71(5):595, 1979

  28. [36]

    Hahn and P

    L. Hahn and P. Klein. Eye tracking in physics education research: A systematic literature review. Physical Review Physics Education Research, 18(1):013102, 2022

  29. [37]

    A. Han, X. Zhou, Z. Cai, S. Han, R. Ko, S. Corrigan, and K. A. Peppler. Teachers, parents, and students’ perspectives on integrating generative ai into elementary literacy education. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–17, 2024

  30. [38]

    Hasenbein, P

    L. Hasenbein, P. Stark, U. Trautwein, H. Gao, E. Kasneci, and R. G ¨ollner. Investigating social comparison behaviour in an immersive virtual reality classroom based on eye-movement data. Scientific Reports, 13(1):14672, 2023

  31. [39]

    Hasenbein, P

    L. Hasenbein, P. Stark, U. Trautwein, A. C. M. Queiroz, J. Bailenson, J.-U. Hahn, and R. G ¨ollner. Learning with simulated virtual classmates: Effects of social-related configurations on students’ visual attention and learning experiences in an immersive virtual reality class...

  32. [40]

    Huang, W

    L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232 , 2023

  33. [41]

    Huang, E

    Y. Huang, E. Richter, T. Kleickmann, and D. Richter. Class size affects preservice teachers’ physiological and psychological stress reactions: An experiment in a virtual reality classroom. Computers & education, 184:104503, 2022

  34. [42]

    Huang, E

    Y. Huang, E. Richter, T. Kleickmann, and D. Richter. Virtual reality in teacher education from 2010 to 2020. Bildung f¨ur eine digitale Zukunft , pages 399–441, 2023

  35. [43]

    Huang, E

    Y. Huang, E. Richter, T. Kleickmann, A. Wiepke, and D. Richter. Classroom complexity affects student teachers’ behavior in a vr classroom. Computers & Education, 163:104100, 2021

  36. [44]

    Izquierdo-Domenech, J

    J. Izquierdo-Domenech, J. Linares-Pellicer, and I. Ferri-Molla. Virtual reality and language models: A new frontier in learning. International Journal of Interactive Multimedia and Artificial Intelligence , 8(5), 2024

  37. [45]

    Javaid and A

    M. Javaid and A. Haleem. Virtual reality applications toward medical field. Clinical Epidemiology and Global Health , 8(2):600–605, 2020

  38. [46]

    Jeon and S

    J. Jeon and S. Lee. Large language models in education: A focus on the complementary relationship between human teachers and chatgpt. Education and Information Technologies, 28(12):15873–15892, 2023

  39. [47]

    Q. Jin, Y. Liu, S. Yarosh, B. Han, and F. Qian. How will vr enter university classrooms? multi-stakeholders investigation of vr in higher education. In Proceedings of the 2022 CHI conference on human factors in computing systems , pages 1–17, 2022

  40. [48]

    Joshi, S

    A. Joshi, S. Kale, S. Chandel, and D. K. Pal. Likert scale: Explored and explained. British journal of applied science & technology, 7(4):396–403, 2015

  41. [49]

    M. A. Just and P. A. Carpenter. Eye fixations and cognitive processes. Cognitive psychology, 8(4):441–480, 1976

  42. [50]

    Kapadia, S

    N. Kapadia, S. Gokhale, A. Nepomuceno, W. Cheng, S. Bothwell, M. Mathews, J. S. Shallat, C. Schultz, and A. Gupta. Evaluation of large language model generated dialogues for an ai based vr nurse training simulator. In International Conference on Human-Computer Interaction, pag...

  43. [51]

    Kasneci, H

    E. Kasneci, H. Gao, S. Ozdel, V. Maquiling, E. Thaqi, C. Lau, Y. Rong, G. Kasneci, and E. Bozkir. Introduction to eye tracking: A hands-on tutorial for students and practitioners. arXiv preprint arXiv:2404.15435, 2024

  44. [52]

    Kasneci, K

    E. Kasneci, K. Sessler, S. K ¨uchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. G¨unnemann, E. H¨ullermeier, S. Krusche, G. Kutyniok, T. Michaeli, C. Nerdel, J. Pfeffer, O. Poquet, M. Sailer, A. Schmidt, T. Seidel, M. Stadler, J. Weller, J. Kuhn, and G. K...

  45. [53]

    Kiefer, I

    P. Kiefer, I. Giannopoulos, A. Duchowski, and M. Raubal. Measuring cognitive load for map tasks through pupil diameter. In Geographic Information Science: 9th International Conference, GIScience 2016, Montreal, QC, Canada, September 27-30, 2016, Proceedings 9, pages 323–337. S...

  46. [54]

    A. King. Effects of self-questioning training on college students’ comprehension of lectures. Contemporary Educational Psychology, 14(4):366–381, 1989

  47. [55]

    P. A. Kirschner. Cognitive load theory: Implications of cognitive load theory on the design of learning, 2002

  48. [56]

    M. M. Kouijzer, H. Kip, Y. H. Bouman, and S. M. Kelders. Implementation of virtual reality in healthcare: a scoping review on the implementation process of virtual reality in various healthcare settings. Implementation science communications, 4(1):67, 2023

  49. [57]

    Krejtz, A

    K. Krejtz, A. T. Duchowski, A. Niedzielska, C. Biele, and I. Krejtz. Eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze. PloS one, 13(9):e0203629, 2018

  50. [58]

    Lewandowska, I

    A. Lewandowska, I. Rejer, K. Bortko, and J. Jankowski. Eye-tracker study of influence of affective disruptive content on user’s visual attention and emotional state. Sensors, 22(2):547, 2022

  51. [59]

    Lewis, E

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt¨aschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems , 33:9459–9474, 2020

  52. [60]

    Lieb and T

    A. Lieb and T. Goel. Student interaction with newtbot: An llm-as-tutor chatbot for secondary physics education. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024

  53. [61]

    H. W. Lilliefors. On the kolmogorov-smirnov test for normality with mean and variance unknown. Journal of the American statistical Association , 62(318):399–402, 1967

  54. [62]

    X. P. Lin, B. B. Li, Z. N. Yao, Z. Yang, and M. Zhang. The impact of virtual reality on student engagement in the classroom–a critical review of the literature. Frontiers in Psychology, 15:1360574, 2024

  55. [63]

    R. Liu, L. Wang, T. A. Koszalka, and K. Wan. Effects of immersive virtual reality classrooms on students’ academic achievement, motivation and cognitive load in science lessons. Journal of Computer Assisted Learning, 38(5):1422–1433, 2022

  56. [64]

    Z. Liu, Z. Zhu, L. Zhu, E. Jiang, X. Hu, K. Peppler, and K. Ramani. Classmeta: Designing interactive virtual classmate to promote vr classroom participation. In Conference on Human Factors in Computing Systems - Proceedings. Association for Computing Machinery, 2024

  57. [65]

    N. L. Loman and R. E. Mayer. Signaling techniques that increase the understandability of expository prose. Journal of Educational psychology , 75(3):402, 1983

  58. [66]

    Y. Lou, P. C. Abrami, and S. d’Apollonia. Small group and individual learning with technology: A meta-analysis. Review of educational research, 71(3):449–521, 2001

  59. [67]

    Lu and X

    X. Lu and X. Wang. Generative students: Using llm-simulated student profiles to support question item evaluation. In Proceedings of the Eleventh ACM Conference on Learning@ Scale, pages 16–27, 2024

  60. [68]

    C. P. Malkewitz, P. Schwall, C. Meesters, and J. Hardt. Estimating reliability: A comparison of cronbach’sα, mcdonald’s ωt and the greatest lower bound. Social Sciences & Humanities Open , 7(1):100368, 2023

  61. [69]

    Mathˆ ot, J

    S. Mathˆ ot, J. Fabius, E. Van Heusden, and S. Van der Stigchel. Safe and sensible preprocessing and baseline correction of pupil-size data. Behavior research methods, 50:94–106, 2018

  62. [70]

    P. D. Mautone and R. E. Mayer. Signaling as a cognitive guide in multimedia learning. Journal of educational Psychology, 93(2):377, 2001

  63. [71]

    Meyer, T

    J. Meyer, T. Jansen, R. Schiller, L. W. Liebenow, M. Steinbach, A. Horbach, and J. Fleckenstein. Using llms to bring evidence-based feedback into the classroom: Ai-generated feedback increases secondary students’ text revision, motivation, and positive emotions. Computers and ...

  64. [72]

    E. R. Mollick and L. Mollick. Using ai to implement effective teaching strategies in classrooms: Five strategies, including prompts. The Wharton School Research Paper, 2023

  65. [73]

    Mystakidis

    S. Mystakidis. Distance education gamification in social virtual reality: A case study on student engagement. In 2020 11th International Conference on Information, Intelligence, Systems and Applications (IISA, pages 1–6. IEEE, 2020

  66. [74]

    Naimi-Akbar, M

    I. Naimi-Akbar, M. Weurlander, and L. Barman. Teaching-learning in virtual learning environments: a matter of forced compromises away from student-centredness? Teaching in Higher Education, pages 1–17, 2023

  67. [75]

    Hello gpt-4o, 2024

    OpenAI. Hello gpt-4o, 2024. Accessed: 2024-08-30

  68. [76]

    Whisper api, 2024

    OpenAI. Whisper api, 2024. Accessed: 2024-09-12

  69. [77]

    J. Pallant. SPSS survival manual: A step by step guide to data analysis using IBM SPSS . Routledge, 2020

  70. [78]

    Papaioannou, M.-G

    G. Papaioannou, M.-G. Volakaki, S. Kokolakis, and D. Vouyioukas. Learning spaces in higher education: a state-of-the-art review. Trends in Higher Education, 2(3):526–545, 2023

  71. [79]

    Park and D

    H. Park and D. Ahn. The promise and peril of chatgpt in higher education: Opportunities, challenges, and design implications. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–21, 2024

  72. [80]

    Peterson-Ahmad

    M. Peterson-Ahmad. Enhancing pre-service special educator preparation through combined use of virtual simulation and instructional coaching. Education Sciences, 8(1):10, 2018

  73. [81]

    Poole and L

    A. Poole and L. J. Ball. Eye tracking in hci and usability research. In Encyclopedia of human computer interaction, pages 211–219. IGI global, 2006

  74. [82]

    J. Pottle. Virtual reality and the transformation of medical education. Future healthcare journal, 6(3):181–185, 2019

  75. [83]

    M. P. Pratama, R. Sampelolo, and H. Lura. Revolutionizing education: harnessing the power of artificial intelligence for personalized learning. Klasikal: Journal of education, language teaching and science, 5(2):350–357, 2023

  76. [84]

    Radianti, T

    J. Radianti, T. A. Majchrzak, J. Fromm, and I. Wohlgenannt. A systematic review of immersive virtual reality applications for higher education: Design elements, lessons learned, and research agenda. Computers & education, 147:103778, 2020

  77. [85]

    N. A. Rappa, S. Ledger, T. Teo, K. Wai Wong, B. Power, and B. Hilliard. The use of eye tracking technology to explore learning and performance within virtual reality and mixed reality settings: a scoping review. Interactive Learning Environments, 30(7):1338–1350, 2022

  78. [86]

    K. Rayner. Eye movements in reading and information processing: 20 years of research. Psychological bulletin, 124(3):372, 1998

  79. [87]

    N. M. Razali, Y. B. Wah, et al. Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests. Journal of statistical modeling and analytics, 2(1):21–33, 2011

  80. [88]

    M. A. Rojas-S´ anchez, P. R. Palos-S´ anchez, and J. A. Folgado-Fern´ andez. Systematic literature review and bibliometric analysis on virtual reality and education. Education and Information Technologies, 28(1):155–192, 2023

  81. [89]

    J. P. Royston. An extension of shapiro and wilk’s w test for normality to large samples. Journal of the Royal Statistical Society: Series C (Applied Statistics), 31(2):115–124, 1982

  82. [90]

    D. D. Salvucci and J. H. Goldberg. Identifying fixations and saccades in eye-tracking protocols. In Proceedings of the 2000 symposium on Eye tracking research & applications, pages 71–78, 2000

  83. [91]

    Savitzky and M

    A. Savitzky and M. J. Golay. Smoothing and differentiation of data by simplified least squares procedures. Analytical chemistry, 36(8):1627–1639, 1964

  84. [92]

    A. W. Services. Amazon polly: Text-to-speech service, 2024. Accessed: 2024-09-12

  85. [93]

    T. Seufert. Supporting coherence formation in learning from multiple representations. Learning and instruction, 13(2):227–237, 2003

  86. [94]

    Shemshack and J

    A. Shemshack and J. M. Spector. A systematic literature review of personalized learning terms. Smart Learning Environments, 7(1):33, 2020

  87. [95]

    J. M. Spector. The potential of smart technologies for learning and instruction. International Journal of Smart Technology and Learning, 1(1):21–32, 2016

  88. [96]

    New class among first taught entirely in virtual reality, 2021

    Stanford News. New class among first taught entirely in virtual reality, 2021. Accessed: 2024-08-20

  89. [97]

    P. Stark. Towards Effective Virtual Reality Learning Environments: Assessment of Information Processing and Learning through Eye Tracking. PhD thesis, Universit¨at T¨ubingen, 2024

  90. [98]

    Stark, A

    P. Stark, A. J. Jung, J.-U. Hahn, E. Kasneci, and R. G¨ollner. Using gaze transition entropy to detect classroom discourse in a virtual reality classroom. In Proceedings of the 2024 Symposium on Eye Tracking Research and Applications, pages 1–11, 2024

  91. [99]

    J. Sweller. Element interactivity and intrinsic, extraneous, and germane cognitive load. Educational psychology review, 22:123–138, 2010

  92. [100]

    J. Sweller. Cognitive load theory. In Psychology of learning and motivation , volume 55, pages 37–76. Elsevier, 2011

  93. [101]

    Tavakol and R

    M. Tavakol and R. Dennick. Making sense of cronbach’s alpha. International journal of medical education, 2:53, 2011

  94. [102]

    Technologies

    V. Technologies. Varjo xr-3: Mixed reality headset,

  95. [103]

    Zawacki-Richter

    O. Zawacki-Richter. The current state and impact of covid-19 on digital higher education in germany. Human Behavior and Emerging Technologies , 3(1):218–226, 2021

  96. [104]

    H. R. Tenenbaum, N. E. Winstone, P. J. Leman, and R. E. Avery. How effective is peer interaction in facilitating learning? a meta-analysis. Journal of Educational Psychology, 112(7):1303, 2020

  97. [105]

    A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting. Large language models in medicine. Nature medicine, 29(8):1930–1940, 2023

  98. [106]

    J. J. Van Merri ¨enboer and J. Sweller. Cognitive load theory in health professional education: design principles and strategies. Medical education, 44(1):85–93, 2010

  99. [107]

    Westphal, E

    A. Westphal, E. Richter, R. Lazarides, and Y. Huang. More i-talk in student teachers’ written reflections indicates higher stress during vr teaching. Computers & Education , 212:104987, 2024

  100. [108]

    L. Wu, Z. Zheng, Z. Qiu, H. Wang, H. Gu, T. Shen, C. Qin, C. Zhu, H. Zhu, Q. Liu, et al. A survey on large language models for recommendation. World Wide Web, 27(5):60, 2024

  101. [109]

    L. Yan, L. Sha, L. Zhao, Y. Li, R. Martinez-Maldonado, G. Chen, X. Li, Y. Jin, and D. Gaˇ sevi´ c. Practical and ethical challenges of large language models in education: A systematic scoping review, 1 2024

  102. [111]

    Zhang, L

    H. Zhang, L. Yu, M. Ji, Y. Cui, D. Liu, Y. Li, H. Liu, and Y. Wang. Investigating high school students’ perceptions and presences under vr learning environment. In Cross Reality (XR) and Immersive Learning Environments (ILEs) in Education , pages 97–117. Routledge, 2023

  103. [112]

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023

  104. [2024]

    Accessed: 2024-09-12

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.