Pith. sign in

REVIEW 3 major objections 5 minor 89 references

Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Explainable XR presents a single action-centric recording format, plus LLM-generated insights, to make XR user analytics consistent across AR, VR, and MR.

desk verdict Solid XR analytics framework with a genuinely useful action-centric schema; LLM insight claims outrun their evidence, but the paper is honest and worth reviewing. read the letter →

arxiv 2501.13778 v2 pith:WECHJGYN submitted 2025-01-23 cs.HC cs.CL

classification cs.HCcs.CL
keywords ExtendedRealityUserBehaviorAnalyticsLargeLanguageModelsVisualCross-VirtualityActionDescriptorMultimodalDataCollectionMulti-userXR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Explainable XR aims to make user-behavior analysis in extended reality as straightforward as reviewing a spreadsheet. The paper proposes a single recording format, the User Action Descriptor, that stores every action—who did what, when, where, how, why, and what was targeted—along with a snapshot of the surrounding scene, so the same pipeline works for AR, VR, and MR sessions, alone or in groups. A visual analytics interface then uses multiple specialized large-language-model agents to summarize sessions, estimate users' intentions, classify physical objects users touched, and highlight where each insight came from. The authors demonstrate the system on five applications spanning individual and collaborative, synchronous and asynchronous XR use, and report that 14 study participants found the interface easy to use and the LLM insights helpful, while also acknowledging that the LLM's arithmetic and some extrapolations can be wrong. If the framework works as claimed, researchers would no longer need custom per-study logging tools to collect, compare, and explain immersive user data.

What carries the argument

The central object is the User Action Descriptor (UAD), a 5W1H-inspired schema—When, Where, Who, What, Why, How, plus referent and context—that makes each user action the trigger for information capture. The machinery also includes a platform-agnostic session recorder with template and direct logging, a post-hoc processor that builds context point clouds and uses LLM agents for context description, intention estimation, and referent classification, and a linked visual analytics interface with spatial, temporal, data, plot, and insight viewers whose LLM insights are anchored to source actions by Analysis-of-Interest markers. The UAD is what carries the argument: because every datapoint is tied to an action and its context, the same pipeline can analyze a VR game, an MR selection task, and an AR collaborative session without per-application customization.

What would settle it

Take a recorded session with known ground truth—predefined actions, verified referent objects, and hand-computed task-completion times—run Explainable XR's LLM agents and insight generator on it, and count how many generated intentions, referent labels, and numeric claims match the ground truth; if a substantial share are wrong, or if the interface's conclusions change no more than chance when the LLM component is removed, the paper's central claim would be falsified.

Watch

Extended reading notes

Core claim

The central discovery is that XR user analytics can be organized around a single action-centric data structure rather than around raw device streams or task-specific logs. The User Action Descriptor binds each logged user action to its type, timestamp, 6DoF location, trigger device, target referent, and a reconstructed point-cloud context, so every piece of multimodal data is interpretable as part of a user's moment-by-moment behavior. The paper argues that this structure is virtuality-agnostic and task-agnostic, and that it scales to multi-user sessions because every action carries its own user identity and context. Large language models then turn the structured logs into analyst-facing insights: they describe action contexts, infer intentions for actions whose meaning depends on conversation or scene context, classify physical referents in AR and MR, and generate up to ten insight summaries tailored to an analyst's stated Analysis-of-Interest. Five prototype applications across VR, MR, and AR, including a collaborative AR analytics task, are used to show the framework's range, with a technical overhead evaluation and a 14-participant user study supporting the usability claims.

Load-bearing premise

The load-bearing premise is that LLM-generated insights, intention estimates, referent classifications, and numeric summaries are accurate enough for analysts to rely on them; the paper acknowledges LLM hallucination and one participant's report that 'some of the maths are wrong' in the temporal-pattern analysis.

Editorial extensions

If this is right

  • Researchers can collect and compare XR user behavior across AR, VR, and MR studies without designing a new logging format for each experiment.
  • Multi-user collaborative sessions can be analyzed at the level of individual actions, making it possible to trace who contributed what to a shared task.
  • LLM-generated insights with linked markers can reduce data overload by giving analysts a starting point and pointing back to the exact actions behind each claim.
  • The framework's recorder overhead is low enough for interactive use, with the main cost being referent export in a portable 3D format.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the UAD could become a common interchange format for XR user-behavior datasets, enabling cross-study comparison that task-specific logs do not support.
  • Editorial extension: the reported arithmetic errors point toward a concrete architecture change—letting LLM agents delegate numeric computation to deterministic code—that would raise reliability without altering the framework's design.
  • Editorial extension: as headset platforms restrict raw camera access, the context-point-cloud component will likely need to shift to OS-provided scene meshes, and a UAD variant that stores those meshes natively would keep the pipeline viable.
  • Editorial extension: the dependence on Analysis-of-Interest prompt phrasing suggests a measurable improvement path—adding a prompt-refinement step so casual researcher questions yield structured, useful insights.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Explainable XR (EXR), an end-to-end framework for recording, processing, and visualizing user behavior in XR sessions across virtualities (AR, VR, MR), including multi-user and cross-virtuality scenarios. The main contributions are: (1) the User Action Descriptor (UAD), an action-centric structured schema that captures the who/what/when/where/why/how of each user action along with the visual referent and scene context; (2) a Unity-based action recorder with template-based and direct logging; (3) a web-based visual analytics interface with spatial, temporal, plot, data, and insight views; and (4) LLM-assisted analytics that generate insights from the recorded UAD data via multiple specialized agents. The authors report a system evaluation of recording overhead across three devices, a comparison of multi-agent versus single-agent insight generation using self-evaluating LLM metrics, and a user study with 14 participants who performed four analysis tasks across five prototype XR applications.

Significance. If substantiated, EXR would be a valuable general-purpose analytics tool for XR research, filling a gap left by task-specific frameworks like MRAT, ARGUS, and ReLive. The paper's strengths include the principled action-centric UAD schema (with the 5W1H mapping), the demonstrated versatility across five prototype applications, the public release of the source code, and the measured overhead of the recording pipeline across different XR devices. The cross-virtuality and multi-user support, combined with the unified visual interface, are useful contributions independent of the LLM component. However, the central claim of 'highly usable ... delivering multifaceted, actionable insights' rests on the quality and reliability of the LLM-generated insights, and the current evidence for that component is not sufficient: the insight-quality evaluation is based on a self-evaluating agent from the same LLM family, the user study has no control condition and no inferential statistics, and the manuscript itself acknowledges hallucination and arithmetic inconsistencies in the LLM outputs.

major comments (3)
  1. [§4.4, Table 4] The quality of LLM-generated insights is evaluated using a self-evaluating agent with SEVQ-inspired metrics (C1–C5), where the scoring LLM is from the same family (or the same model) that produced the insights. This design cannot establish insight accuracy or actionability, because the producer and the judge share the same biases and failure modes, including the arithmetic inconsistencies acknowledged in §5. I recommend an independent evaluation—for example, human expert ratings of insight correctness against the ground-truth UAD logs, or at minimum a judge from a different LLM family—and a comparison against a non-LLM baseline derived directly from the recorded data.
  2. [§4.1–§4.3] The usability and usefulness claims are based on a 14-participant study with Likert-scale ratings reported as means (e.g., µ=4.5, µ=4.6) without confidence intervals, significance tests, or a control condition. In particular, there is no no-LLM condition in which the interface shows only the recorded data without the Analytics Insights, so the participants' positive ratings cannot isolate the added value of the LLM-assisted component. The phrase 'highly usable' and claims such as 'participants evaluated the usefulness of the Analytics insights highly' are therefore descriptive only; providing a comparison condition or precise statistical reporting would materially strengthen these claims.
  3. [§5 and §4.3] The manuscript acknowledges two load-bearing reliability problems: §5 states that LLM extrapolations can produce inaccurate insights 'due to the hallucination of the LLM,' and §4.3 reports participant P13's observation that 'some of the maths are wrong,' with the authors conceding 'occasional inconsistencies in the agent's math computations.' Since the central value proposition is that EXR 'delivers ... actionable insights into user behaviors,' the paper should report the frequency and severity of such failure modes, or otherwise bound the reliability of the generated insights, rather than only listing future plans (multi-agent debate, confidence scores). Without this, the claim that analysts can trust the LLM-generated insights is not supported.
minor comments (5)
  1. [§4.4] There is a typo in the first sentence: 'accuractely' should be 'accurately'.
  2. [§3.3.2 and §4.4] The text lists six analytical aspects of the extracted insights (space, time, action, intent, context, user) but the evaluation in §4.4 uses five criteria (C1–C5). The relationship between the six aspects and the five criteria is not explained; clarifying this mapping would help readers interpret the evaluation.
  3. [§4.4, Table 3] The overhead comparison with ReLive reports single average values over 100 calls without standard deviations or significance testing. For the Log+R case, EXR takes 101.44 ms versus ReLive's 1.01 ms, and the asynchronous fallback of 1.13 ms is described without reporting the measurement protocol; please specify the variance and the conditions under which the asynchronous number was measured.
  4. [§4.2] The statement 'All participants agreed (µ=4.1)' is imprecise; 'agreed' implies a consensus threshold that is not defined. Reporting the response distribution or a criterion (e.g., the percentage of participants scoring above 4) would be more informative.
  5. [§3.3.1 and Supplementary Materials] The paper repeatedly refers to Supplementary Materials for prompt details and prototype application specifics. For reproducibility, the main text should state which LLM models were used, the temperature or other sampling parameters, and the number of independent runs for the multi-agent insight generation.

Circularity Check

1 steps flagged · score 4.0 of 10

LLM insight quality is validated by a self-evaluating LLM agent, making the 'actionable insights' claim partially circular; the UAD/recorder/visualizer contributions are independently grounded.

  1. other [Section 4.4 ('System Evaluation'), LLM insight quality evaluation; abstract conclusion]
    "To assess the output quality, we apply the concept of self-evaluating agent [20, 27, 38, 47]. We develop the evaluation metrics inspired by the SEVQ metrics of LIDA [20]."

    The paper's central value claim—'delivering multifaceted, actionable insights into user behaviors'—is supported by Table 4 scores (C1–C5, μ=9.04/8.90) that are produced by a self-evaluating LLM agent. The generator and the judge are both LLM-based, so the scores measure self-consistency within the same model family rather than external correctness. No independent human-coded ground truth or no-LLM control is used for insight accuracy. The paper also concedes that the judge misses failures: Section 5 states 'these extrapolation can suggest insights that are inaccurately derived, due to the hallucination of the LLM', and participant P13 is quoted as 'I feel some of the maths are wrong'.

full rationale

The core framework contributions—the UAD schema, the Unity recorder, the point-cloud context generation, and the visual analytics interface—are self-contained and do not reduce to their inputs. The user study (Sections 4.2–4.3) provides independent evidence for the usability and usefulness of the visual interface, and the overhead measurements in Section 4.4 are external performance benchmarks. The circularity is confined to the LLM-assisted insight component: its quality is evaluated by a 'self-evaluating agent' rooted in LIDA's SEVQ metrics, meaning the producer and judge are the same kind of model. That is a genuinely self-referential validation loop for the 'actionable insights' claim, and the paper's own acknowledgements of hallucination and incorrect math computations show that this loop does not catch important errors. Because the framework's data-recording and visualization contributions have independent support, the overall circularity is partial rather than total; a score of 4 reflects one significant self-referential evaluation that does not invalidate the central non-LLM framework.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The framework rests primarily on design assumptions about the sufficiency of action-centric logging, the reliability of LLM synthesis, and the adequacy of low-resolution context capture. The hand-chosen parameters (snapshot resolution, number of insights, number of agents) affect the quality of the output and the reported performance, but they are not fitted in a statistical sense. No new physical or conceptual entities are postulated beyond the UAD data schema.

free parameters (3)
  • Context snapshot resolution = 480x270 (75% downsampled from 1920x1080)
    Hand-chosen to reduce logging overhead from 143.94ms to 28.20ms; the paper acknowledges this degrades point cloud quality (Section 5).
  • Number of Analytics Insights = up to 10
    Arbitrary cap on LLM-generated insights presented to the analyst, with no justification for the limit (Section 3.3.2).
  • Number of LLM agents = 6 (Spatial, Temporal, User, Action, Context, Intent)
    Design choice for multi-agent decomposition; no principled derivation of the number or the specific role split (Section 3.3.2).
assumptions (4)
  • domain assumption Unity's Input System provides a complete and consistent interface to all trackable XR device sensors across AR, VR, and MR platforms.
    The Action Recorder relies on Unity's Input System for TriggerSource; the paper notes that camera access is restricted on VisionOS and Meta Horizon OS, so the assumption does not fully hold (Sections 3.2 and 5).
  • domain assumption User action-centric logging captures all information needed to understand user behavior, and events not linked to user actions can be safely excluded.
    UAD filters out non-action events such as autonomously moving objects outside the user's frustum, which could be relevant for some analyses (Section 3.1).
  • domain assumption Large language models can correctly synthesize multimodal UAD data (gaze, gestures, point clouds, audio transcripts) into accurate insights.
    The Analytics Assistant depends on LLM performance for intention estimation, referent classification, and insight generation; the paper reports hallucination and math errors in Section 5, indicating this assumption is fragile.
  • domain assumption Semi-dense point clouds reconstructed from low-resolution (480x270) snapshots preserve sufficient spatial detail for context analysis.
    The Context field uses these point clouds, and the paper acknowledges that quality degradation is a known limitation of the downsampling (Sections 3.3.1 and 5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework." pith.science (2026). https://pith.science/paper/WECHJGYN

@misc{pith2026250113778,
  author       = {Pith},
  title        = {Pith review of: Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WECHJGYN}},
  note         = {Machine review of arXiv:2501.13778}
}
read the original abstract

We present Explainable XR, an end-to-end framework for analyzing user behavior in diverse eXtended Reality (XR) environments by leveraging Large Language Models (LLMs) for data interpretation assistance. Existing XR user analytics frameworks face challenges in handling cross-virtuality - AR, VR, MR - transitions, multi-user collaborative application scenarios, and the complexity of multimodal data. Explainable XR addresses these challenges by providing a virtuality-agnostic solution for the collection, analysis, and visualization of immersive sessions. We propose three main components in our framework: (1) A novel user data recording schema, called User Action Descriptor (UAD), that can capture the users' multimodal actions, along with their intents and the contexts; (2) a platform-agnostic XR session recorder, and (3) a visual analytics interface that offers LLM-assisted insights tailored to the analysts' perspectives, facilitating the exploration and analysis of the recorded XR session data. We demonstrate the versatility of Explainable XR by demonstrating five use-case scenarios, in both individual and collaborative XR applications across virtualities. Our technical evaluation and user studies show that Explainable XR provides a highly usable analytics solution for understanding user actions and delivering multifaceted, actionable insights into user behaviors in immersive environments.

Figures

Figures reproduced from arXiv: 2501.13778 by the authors.

Figure 1
Figure 1. Explainable XR provides a streamlined pipeline to record, visualize, and analyze users of an immersive session, facilitating [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Explainable XR Pipeline Overview: The blue arrows denote the internal calls and flows of Explainable XR, and red arrows denote the inputs of the researcher in our framework. The pipeline initiates by recording the multimodal interactions of the subjects in XR sessions, and importing it into our Action Visual Analyzer via User Prompt Interface. The researcher can perform analytical tasks in our Visual Analytics Inter… view at source ↗
Figure 5
Figure 5. Virtuality-agnostic Session Reconstruction: UAD binds each [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figures from the paper (5 more)
Figure 3
Figure 3. Figure 3: Action Template Logging Editor and its auto-generated code: Our [PITH_FULL_IMAGE:figures/full_fig_p005_3.png]
Figure 4
Figure 4. Figure 4: Structure of Logging Function: It conforms to the User Action [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: LLM-Analytics Assistance: Given the prompt for the direction [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Visual Analytics Interface: (A) The Data Manager includes Data Filter that filters the visualized action across the viewers, and Data Viewer [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: General-purpose Analytics Framework: EXR can be utilized for diverse analytics tasks such as pattern finding, contextual visualization, and [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 68 canonical work pages

  1. [1]

    Developer

    Apple. Developer. https://developer.apple.com/videos/play/ wwdc2024/10139/, 2024. Sep. 14. 2024. 9

  2. [2]

    Arnold, M

    M. Arnold, M. Goldschmitt, and T. Rigotti. Dealing with information overload: a comprehensive review. Frontiers in Psychology, 14:1122200,

  3. [3]

    Bohus, S

    D. Bohus, S. Andrist, N. Saw, A. Paradiso, I. Chakraborty, and M. Rad. Sigma: An open-source interactive system for mixed-reality task assistance research–extended abstract. In Prof. of VRW, pp. 889–890, 2024. 2

  4. [4]

    Boorboor, M

    S. Boorboor, M. S. Castellana, Y . Kim, C. Zhu-Tian, J. Beyer, H. Pfis- ter, and A. E. Kaufman. V oxAR: adaptive visualization of volume ren- dered objects in optical see-through augmented reality. IEEE TVCG, 30(10):6801–6812, 2023. 1

  5. [5]

    Boorboor, Y

    S. Boorboor, Y . Kim, P. Hu, J. M. Moses, B. A. Colle, and A. E. Kaufman. Submerse: Visualizing storm surge flooding simulations in immersive display ecologies. IEEE TVCG, 30(09):6365–6377, 2024. 1

  6. [6]

    Brandstätter and A

    K. Brandstätter and A. Steed. Dialogues for one: Single-user content creation using immersive record and replay. In Proc. of VRST, pp. 1–11,

  7. [7]

    Brookes, M

    J. Brookes, M. Warburton, M. Alghadier, M. Mon-Williams, and F. Mush- taq. Studying human behaviour with virtual reality: The unity experiment framework. bioRxiv, 2018. 2, 3

  8. [8]

    Brudy, S

    F. Brudy, S. Suwanwatcharachat, W. Zhang, S. Houben, and N. Marquardt. Eagleview: A video analysis tool for visualising and querying spatial interactions of people and devices. In Proc. of ISS, pp. 61–72, 2018. 2, 3

Show all 89 references
  1. [9]

    Büschel, A

    W. Büschel, A. Lehmann, and R. Dachselt. Miria: A mixed reality toolkit for the in-situ visualization and analysis of spatio-temporal interaction data. In Proc. of CHI, pp. 1–15, 2021. 2

  2. [10]

    Y . Cao, Y . Lan, F. Zhai, and P. Li. 5w1h extraction with large language models. arXiv preprint arXiv:2405.16150, 2024. 4

  3. [11]

    Castelo, J

    S. Castelo, J. Rulff, E. McGowan, B. Steers, G. Wu, S. Chen, I. Roman, R. Lopez, E. Brewer, C. Zhao, et al. ARGUS: Visualization of AI-assisted task guidance in AR. IEEE TVCG, 2023. 1, 2, 3

  4. [12]

    Castelo, J

    S. Castelo, J. Rulff, P. Solunke, E. McGowan, G. Wu, I. Roman, R. Lopez, B. Steers, Q. Sun, J. Bello, et al. Hubar: A visual analytics tool to explore human behavior based on fnirs in ar guidance systems. IEEE TVCG, 2024. 3

  5. [13]

    K. Choe, C. Lee, S. Lee, J. Song, A. Cho, N. W. Kim, and J. Seo. En- hancing data literacy on-demand: Llms as guides for novices in chart interpretation. IEEE TVCG, 2024. 2

  6. [14]

    Cognitive3d

    Cognitive3D. Cognitive3d. https://cognitive3d.com/product/ objectives/. Mar. 26. 2024. 2, 3

  7. [15]

    Cools, X

    R. Cools, X. Zhang, and A. L. Simeone. Crest: Design and evaluation of the cross-reality study tool. In Proc. of MUM, pp. 409–419, 2023. 2

  8. [16]

    W. Cui, X. Zhang, Y . Wang, H. Huang, B. Chen, L. Fang, H. Zhang, J. Lou, and D. Zhang. Text-to-viz: Automatic generation of infographics from proportion-related natural language statements. IEEE TVCG, 26(01):906– 916, 2020. 3

  9. [17]

    David-John, C

    B. David-John, C. Peacock, T. Zhang, T. S. Murdison, H. Benko, and T. R. Jonker. Towards gaze-based prediction of the intent to interact in virtual reality. In Proc. of ETRA, pp. 1–7, 2021. 2

  10. [18]

    De La Torre, C

    F. De La Torre, C. M. Fang, H. Huang, A. Banburski-Fahey, J. Amores Fer- nandez, and J. Lanier. Llmr: Real-time prompting of interactive worlds using large language models. In Proc. of CHI, pp. 1–22, 2024. 3

  11. [19]

    Deuchler, W

    J. Deuchler, W. Hettmann, D. Hepperle, and M. Wölfel. Streamlining physiological observations in immersive virtual reality studies with the virtual reality scientific toolkit. In Proc. of VRW, pp. 485–488, 2023. 2

  12. [20]

    V . Dibia. Lida: A tool for automatic generation of grammar-agnostic visu- alizations and infographics using large language models. arXiv preprint arXiv:2303.02927, 2023. 2, 8

  13. [21]

    M. D. Dogan, E. J. Gonzalez, K. Ahuja, R. Du, A. Colaço, J. Lee, M. Gonzalez-Franco, and D. Kim. Augmented object intelligence with xr-objects. In Proc. of UIST, pp. 1–15, 2024. 3

  14. [22]

    H. Duan, Y . Yang, and K. Y . Tam. Do llms know about hallucina- tion? an empirical investigation of llm’s hidden states. arXiv preprint arXiv:2402.09733, 2024. 9

  15. [23]

    Enriquez, W

    D. Enriquez, W. Tong, C. North, H. Qu, and Y . Yang. Evaluating layout dimensionalities in pc+ vr asymmetric collaborative decision making. In Proc. of ISS, pp. 112–132, 2024. 1

  16. [24]

    Feldman, J

    P. Feldman, J. R. Foulds, and S. Pan. Trapping llm hallucinations using tagged context prompts. arXiv preprint arXiv:2306.06085, 2023. 9

  17. [25]

    Gasques, J

    D. Gasques, J. G. Johnson, T. Sharkey, Y . Feng, R. Wang, Z. R. Xu, E. Zavala, Y . Zhang, W. Xie, X. Zhang, et al. Artemis: A collaborative mixed-reality system for immersive surgical telementoring. In Proc. of CHI, pp. 1–14, 2021. 1

  18. [26]

    Gorisse, O

    G. Gorisse, O. Christmann, and C. Dubosc. Rec: A unity tool to replay, ex- port and capture tracked movements for 3d and virtual reality applications. In Proc. of AVI, pp. 1–3, 2022. 2

  19. [27]

    Z. Guo, R. Jin, C. Liu, Y . Huang, D. Shi, L. Yu, Y . Liu, J. Li, B. Xiong, D. Xiong, et al. Evaluating large language models: A comprehensive survey. arXiv preprint arXiv:2310.19736, 2023. 8

  20. [28]

    P. Hu, S. Boorboor, S. Jadhav, J. Marino, S. Mirhosseini, and A. E. Kauf- man. Spatial perception in immersive visualization: A study and findings. In Proc. of ISMAR-Adjunct, pp. 369–372, 2022. 1

  21. [29]

    P. Hu, Q. Sun, P. Didyk, L.-Y . Wei, and A. E. Kaufman. Reducing simulator sickness with perceptual camera control. ACM ToG, 38(6):1–12,

  22. [30]

    Huang, Z

    H. Huang, Z. Lin, Z. Wang, X. Chen, K. Ding, and J. Zhao. Towards llm-powered verilog rtl assistant: Self-verification and self-correction. arXiv preprint arXiv:2406.00115, 2024. 9

  23. [31]

    Hubenschmid, J

    S. Hubenschmid, J. Wieland, D. I. Fink, A. Batch, J. Zagermann, N. Elmqvist, and H. Reiterer. Relive: Bridging in-situ and ex-situ vi- sual analytics for analyzing mixed reality user studies. In Proc. of CHI, pp. 1–20, 2022. 1, 2, 3

  24. [32]

    S. Jana, D. Molnar, A. Moshchuk, A. Dunn, B. Livshits, H. J. Wang, and E. Ofek. Enabling {Fine-Grained} permissions for augmented reality applications with recognizers. In Proc. of USENIX Security, pp. 415–430,

  25. [33]

    Jang, E.-J

    S. Jang, E.-J. Ko, and W. Woo. Unified user-centric context: Who, where, when, what, how and why. Proc. of ubiPCMM, 149, 01 2005. 4

  26. [34]

    Jansen, J

    P. Jansen, J. Britten, A. Häusele, T. Segschneider, M. Colley, and E. Rukzio. Autovis: Enabling mixed-immersive analysis of automotive user interface interaction studies. In Proc. of CHI, pp. 1–23, 2023. 2

  27. [35]

    Javerliat, S

    C. Javerliat, S. Villenave, P. Raimbaud, and G. Lavoué. Plume: Record, replay, analyze and share user behavior in 6dof xr experiences. IEEE TVCG, 2024. 1, 2, 3

  28. [36]

    Z. Ji, T. Yu, Y . Xu, N. Lee, E. Ishii, and P. Fung. Towards mitigating hallucination in large language models via self-reflection. arXiv preprint arXiv:2310.06271, 2023. 9

  29. [37]

    Q. Jin, Y . Liu, S. Yarosh, B. Han, and F. Qian. How will vr enter university classrooms? multi-stakeholders investigation of vr in higher education. In Proc. of CHI, pp. 1–17, 2022. 1

  30. [38]

    Kadavath, T

    S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Dodds, N. DasSarma, E. Tran-Johnson, S. Johnston, S. El- Showk, A. Jones, N. Elhage, T. Hume, A. Chen, Y . Bai, S. Bowman, S. Fort, and J. Kaplan. Language models (mostly) know what they know....

  31. [39]

    Kasahara, V

    S. Kasahara, V . Heun, A. S. Lee, and H. Ishii. Second surface: multi- user spatial collaboration system based on augmented reality. In Proc. of SIGGRAPH Asia, pp. 1–4, 2012. 1

  32. [40]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM ToG, 42(4):139–1,

  33. [41]

    Khurana, M

    A. Khurana, M. Glueck, and P. K. Chilana. Do I Just Tap My Head- set? How Novice Users Discover Gestural Interactions with Consumer Augmented Reality Applications. Proc. of IMWUT, 7(4):1–28, 2024. 1

  34. [42]

    Y . Kim, S. Boorboor, A. Rahmati, and A. Kaufman. Design of privacy preservation system in augmented reality. In Proc. of VizSec, 2021. 9

  35. [43]

    Y . Kim, S. Goutam, A. Rahmati, and A. Kaufman. Erebus: Access control for augmented reality systems. In Proc. of USENIX Security, pp. 929–946,

  36. [44]

    B. Lee, X. Hu, M. Cordeil, A. Prouzeau, B. Jenny, and T. Dwyer. Shared surfaces and spaces: Collaborative data visualisation in a co-located im- mersive environment. IEEE TVCG, 27(2):1171–1181, 2020. 1

  37. [45]

    Y . Lee, B. Yoo, and S.-H. Lee. Sharing ambient objects using real-time point cloud streaming in web-based xr remote collaboration. In Proc. of 10 © 2025 IEEE. This is the author’s version of the article that has been published in IEEE Transactions on Visualization and Compute...

  38. [46]

    Y . Li, Y . Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 9

  39. [47]

    S. Lin, J. Hilton, and O. Evans. Teaching models to express their uncer- tainty in words. arXiv preprint arXiv:2205.14334, 2022. 8

  40. [48]

    Lubos, T

    S. Lubos, T. N. T. Tran, A. Felfernig, S. Polat Erdeniz, and V .-M. Le. Llm-generated explanations for recommender systems. In Proc. of UMAP Adjunct, pp. 276–285, 2024. 3

  41. [49]

    W. Luo, Z. Yu, R. Rzayev, M. Satkowski, S. Gumhold, M. McGinity, and R. Dachselt. Pearl: Physical environment based augmented reality lenses for in-situ human movement analysis. In Proc. of CHI, pp. 1–15, 2023. 2

  42. [50]

    P. Ma, R. Ding, S. Wang, S. Han, and D. Zhang. Insightpilot: An llm- empowered automated data exploration system. In Proc. of EMNLP, pp. 346–352, 2023. 2

  43. [51]

    M. N. Mahdi, A. R. Ahmad, R. Ismail, M. A. Subhi, M. M. Abdulrazzaq, and Q. S. Qassim. Information overload: the effects of large amounts of information. In Proc. of IT-ELA, pp. 154–159, 2020. 5

  44. [52]

    E. S. Martinez, A. A. Malik, and R. P. McMahan. Clovr: Collecting and logging openvr data from steamvr applications. In Proc. of VRW, pp. 485–492, 2024. 2

  45. [53]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Proc. of ECCV, 65(1):99–106, 2021. 4

  46. [54]

    Mirhosseini, P

    S. Mirhosseini, P. Ghahremani, S. Ojal, J. Marino, and A. Kaufman. Exploration of large omnidirectional images in immersive environments. In Proc. of VR, pp. 413–422, 2019. 1

  47. [55]

    Mirhosseini, I

    S. Mirhosseini, I. Gutenko, S. Ojal, J. Marino, and A. Kaufman. Immersive virtual colonoscopy. IEEE TVCG, 25(5):2011–2021, 2019. 1

  48. [56]

    V . Nair, W. Guo, J. Mattern, R. Wang, J. F. O’Brien, L. Rosenberg, and D. Song. Unique Identification of 50,000+ Virtual Reality Users from Head & Hand Motion Data. In Proc. of USENIX Security, pp. 895–910,

  49. [57]

    D. Nam, A. Macvean, V . Hellendoorn, B. Vasilescu, and B. Myers. Using an llm to help with code understanding. In Proc. of ICSE, pp. 1–13, 2024. 2

  50. [58]

    Nebeling, M

    M. Nebeling, M. Speicher, X. Wang, S. Rajaram, B. D. Hall, Z. Xie, A. R. Raistrick, M. Aebersold, E. G. Happ, J. Wang, et al. Mrat: The mixed reality analytics toolkit. In Proc. of CHI, pp. 1–12, 2020. 1, 2, 3, 5

  51. [59]

    Numan and A

    N. Numan and A. Steed. Exploring user behaviour in asymmetric collabo- rative mixed reality. In Proc. of VRST, pp. 1–11, 2022. 2

  52. [60]

    Learning to reason with llms

    OpenAI. Learning to reason with llms. https://openai.com/index/ learning-to-reason-with-llms/ , 2024. Sep. 14. 2024. 9

  53. [61]

    Openai o1 system card

    OpenAI. Openai o1 system card. https://assets. ctfassets.net/kftzwdyauwt9/67qJD51Aur3eIc96iOfeOP/ 71551c3d223cd97e591aa89567306912/o1_system_card.pdf,

  54. [62]

    Piumsomboon, Y

    T. Piumsomboon, Y . Lee, G. Lee, and M. Billinghurst. Covar: a collabo- rative virtual and augmented reality system for remote collaboration. In Proc. of SIGGRAPH Asia, pp. 1–2, 2017. 1

  55. [63]

    H. Qu, Y . Cai, and J. Liu. Llms are good action recognizers. InProc. of CVPR, pp. 18395–18406, 2024. 2

  56. [64]

    Robert, H.-Y

    F. Robert, H.-Y . Wu, L. Sassatelli, S. Ramanoel, A. Gros, and M. Winck- ler. An integrated framework for understanding multimodal embodied experiences in interactive virtual reality. In Proc. of IMX, pp. 14–26, 2023. 2

  57. [65]

    Roesner, D

    F. Roesner, D. Molnar, A. Moshchuk, T. Kohno, and H. J. Wang. World- driven access control for continuous sensing. In Proc. of CCS, pp. 1169– 1181, 2014. 9

  58. [66]

    Romero, R

    D. Romero, R. J. Patel, A. Markopolou, and S. Elmalaki. GaitGuard: Towards Private Gait in Mixed Reality. arXiv preprint arXiv:2312.04470,

  59. [67]

    Saffo, S

    D. Saffo, S. Di Bartolomeo, C. Yildirim, and C. Dunne. Remote and collaborative virtual reality experiments via social vr platforms. In Proc. of CHI, pp. 1–15, 2021. 1

  60. [68]

    L. Shen, H. Li, Y . Wang, and H. Qu. From data to story: Towards automatic animated data video creation with llm-based multi-agent systems. IEEE TVCG, 2024. 3

  61. [69]

    L. Shen, E. Shen, Y . Luo, X. Yang, X. Hu, X. Zhang, Z. Tai, and J. Wang. Towards natural language interfaces for data visualization: A survey.IEEE TVCG, 29(6):3121–3144, 2022. 2

  62. [70]

    Slocum, Y

    C. Slocum, Y . Zhang, N. Abu-Ghazaleh, and J. Chen. Going through the motions:ar/vr keylogging from user head motions. In Proc. of USENIX Security, pp. 159–174, 2023. 2

  63. [71]

    Steed, L

    A. Steed, L. Izzouzi, K. Brandstätter, S. Friston, B. Congdon, O. Olkkonen, D. Giunchi, N. Numan, and D. Swapp. Ubiq-exp: A toolkit to build and run remote and distributed mixed reality experiments. Frontiers in VR, 3:912078, 2022. 2

  64. [72]

    Sultanum, M

    N. Sultanum, M. Brudno, D. Wigdor, and F. Chevalier. More text please! understanding and supporting the use of visualization for clinical text overview. In Proc. of CHI, pp. 1–13, 2018. 5

  65. [73]

    Q. Sun, A. Patney, L.-Y . Wei, O. Shapira, J. Lu, P. Asente, S. Zhu, M. McGuire, D. Luebke, and A. Kaufman. Towards virtual reality infinite walking: dynamic saccadic redirection. ACM ToG, 37(4):1–13, 2018. 1

  66. [74]

    H. Tian, G. A. Lee, H. Bai, and M. Billinghurst. Using virtual replicas to improve mixed reality remote collaboration. IEEE TVCG, 29(5):2785– 2795, 2023. 1, 4

  67. [75]

    Input system

    Unity. Input system. https://docs.unity3d.com/Packages/com. unity.inputsystem@1.10/manual/index.html, 2024. Aug. 27

  68. [76]

    Unity engine

    Unity. Unity engine. https://unity.com/products/unity-engine,

  69. [77]

    Villenave, J

    S. Villenave, J. Cabezas, P. Baert, F. Dupont, and G. Lavoué. Xrecho: A unity plug-in to record and visualize user behavior during xr sessions. In Proc. of MMSys, pp. 341–346, 2022. 2

  70. [78]

    C. Y . Wang, D. Saffo, B. Moriarty, and B. Maclntyre. Collabxr: Bridging realities in collaborative workspaces with dynamic plugin and collabora- tive tools integration. In IEEE VRW, pp. 454–457, 2024. 1

  71. [79]

    E. Wen, T. I. Kaluarachchi, S. Siriwardhana, V . Tang, M. Billinghurst, R. W. Lindeman, R. Yao, J. Lin, and S. Nanayakkara. Vrhook: A data collection tool for vr motion sickness research. In Proc. of UIST, pp. 1–9,

  72. [80]

    D. Wolf, J. J. Dudley, and P. O. Kristensson. Performance envelopes of in-air direct and smartwatch indirect control for head-mounted aug- mented reality. In 2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp. 347–354, 2018. 1

  73. [81]

    A. Wu, Y . Wang, X. Shu, D. Moritz, W. Cui, H. Zhang, D. Zhang, and H. Qu. Ai4vis: Survey on artificial intelligence approaches for data visualization. IEEE TVCG, 28(12):5049–5070, 2021. 2, 5

  74. [82]

    Z. Xu, S. Jain, and M. Kankanhalli. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817,

  75. [83]

    Yao, K.-P

    J.-Y . Yao, K.-P. Ning, Z.-H. Liu, M.-N. Ning, and L. Yuan. Llm lies: Hallucinations are not bugs, but features as adversarial examples. arXiv preprint arXiv:2310.01469, 2023. 9

  76. [84]

    K. Yu, U. Eck, F. Pankratz, M. Lazarovici, D. Wilhelm, and N. Navab. Duplicated reality for co-located augmented reality collaboration. IEEE TVCG, 28(5):2190–2200, 2022. 1

  77. [85]

    Yu and Y

    Y . Yu and Y . Bi. A study on “5w1h” user analysis on interaction design of interface. In Proc. of CAIDCD, vol. 1, pp. 329–332, 2010. 4

  78. [86]

    Z. Yu, D. Zeidler, V . Victor, and M. Mcginity. Dynascape: Immersive authoring of real-world dynamic scenes with spatially tracked rgb-d videos. In Proc. of VRST, pp. 1–12, 2023. 2

  79. [87]

    Zhang, C

    P. Zhang, C. Li, and C. Wang. Viscode: Embedding information in visual- ization images using encoder-decoder network. IEEE TVCG, 27(2):326– 336, 2020. 5

  80. [88]

    Zhang, C

    Y . Zhang, C. Slocum, J. Chen, and N. Abu-Ghazaleh. It’s all in your head (set): Side-channel attacks on ar/vr systems. In Proc. of USENIX Security, pp. 3979–3996, 2023. 2

  81. [89]

    Y . Zhao, Y . Zhang, Y . Zhang, X. Zhao, J. Wang, Z. Shao, C. Turkay, and S. Chen. Leva: Using large language models to enhance visual analytics. IEEE TVCG, 2024. 2 11

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.