{"id":"f173ae18-4b07-42d0-8e87-94fd26789d16","arxiv_id":"2501.13778","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Explainable XR provides a unified, action-centric recording and visualization framework with LLM-generated insights for analyzing user behavior across AR, VR, and MR.","lead":"This paper introduces Explainable XR, a framework that records what people do in augmented, virtual, and mixed reality, and uses AI language models to summarize and explain their actions. It offers researchers a standardized tool for studying user behavior across different immersive environments without building custom logging and analysis pipelines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM output accuracy is not independently validated; the central 'actionable insights' claim rests on self-evaluation and uncorroborated user perceptions.","rationale":"The reader's weakest_assumption identifies LLM insight accuracy as the load-bearing assumption, and I agree. The paper's self-evaluation metrics in Section 4.4 are explicitly scored by the same LLM family ('self-evaluating agent'), which does not validate factual correctness. The user study in Sections 4.2-4.3 measures perceived usefulness, not accuracy, and lacks a no-LLM baseline, so it cannot support the claim that LLM-generated insights are reliable or that they add value beyond the visual interface alone. The acknowledged math errors in Section 4.3 and hallucinations in Section 5 are concrete evidence that the accuracy concern is real, not hypothetical. I do not recommend changing the verdict because this is a systems paper with a plausible and useful framework; the insufficiency is in the evidence for the accuracy of the AI-generated components, not in the fundamental soundness of the approach. A CONDITIONAL verdict asking for independent validation matches the evidence level. The concrete test I propose would directly settle whether the accuracy concern is substantive.","tokens_in":21854,"tokens_out":1348,"duration_ms":11308,"concrete_test":"Have independent human annotators (not the authors, not the study participants, and not the LLM family used) score the LLM-generated insights and intention labels for factual accuracy against the raw UAD data, including all numeric claims like 'average task completion time.' If a substantial fraction of numeric or referential claims are wrong, the 'actionable insights' claim requires qualification. Additionally, run a within-subjects user study comparing task completion with the Insight Viewer enabled versus a version with placeholder or no LLM insights, measuring decision correctness and time. If users perform no better with real LLM insights than with placeholders or no insights, the claimed added value is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that EXR 'delivers actionable insights into user behaviors' depends on LLM-generated insights, intention estimations, and referent classifications being accurate enough to ground analysts' conclusions. The paper's own evidence does not establish this. Section 4.4 evaluates insight quality using 'self-evaluating agent' metrics (C1-C5) scored by the same LLM family that produced the insights, which is circular. The user study (Sections 4.2-4.3) has no no-LLM control condition, so participants' positive ratings of the Insight Viewer cannot isolate the added value of the LLM outputs. The paper explicitly acknowledges failure modes: Section 5 reports 'hallucination of the LLM' in extrapolations, and Section 4.3 quotes participant P13: 'I feel some of the maths are wrong,' with the authors conceding 'occasional inconsistencies in the agent's math computations.' These admissions undermine the reliability component of the 'actionable insights' claim. The framework's logging schema and visual interface are plausible contributions, but the value proposition that analysts can trust LLM-generated insights is not yet supported by independent evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Explainable XR (EXR), an end-to-end framework for recording, processing, and visualizing user behavior in XR sessions across virtualities (AR, VR, MR), including multi-user and cross-virtuality scenarios. The main contributions are: (1) the User Action Descriptor (UAD), an action-centric structured schema that captures the who/what/when/where/why/how of each user action along with the visual referent and scene context; (2) a Unity-based action recorder with template-based and direct logging; (3) a web-based visual analytics interface with spatial, temporal, plot, data, and insight views; and (4) LLM-assisted analytics that generate insights from the recorded UAD data via multiple specialized agents. The authors report a system evaluation of recording overhead across three devices, a comparison of multi-agent versus single-agent insight generation using self-evaluating LLM metrics, and a user study with 14 participants who performed four analysis tasks across five prototype XR applications.","tokens_in":22110,"tokens_out":4164,"duration_ms":38200,"significance":"If substantiated, EXR would be a valuable general-purpose analytics tool for XR research, filling a gap left by task-specific frameworks like MRAT, ARGUS, and ReLive. The paper's strengths include the principled action-centric UAD schema (with the 5W1H mapping), the demonstrated versatility across five prototype applications, the public release of the source code, and the measured overhead of the recording pipeline across different XR devices. The cross-virtuality and multi-user support, combined with the unified visual interface, are useful contributions independent of the LLM component. However, the central claim of 'highly usable ... delivering multifaceted, actionable insights' rests on the quality and reliability of the LLM-generated insights, and the current evidence for that component is not sufficient: the insight-quality evaluation is based on a self-evaluating agent from the same LLM family, the user study has no control condition and no inferential statistics, and the manuscript itself acknowledges hallucination and arithmetic inconsistencies in the LLM outputs.","major_comments":[{"comment":"The quality of LLM-generated insights is evaluated using a self-evaluating agent with SEVQ-inspired metrics (C1–C5), where the scoring LLM is from the same family (or the same model) that produced the insights. This design cannot establish insight accuracy or actionability, because the producer and the judge share the same biases and failure modes, including the arithmetic inconsistencies acknowledged in §5. I recommend an independent evaluation—for example, human expert ratings of insight correctness against the ground-truth UAD logs, or at minimum a judge from a different LLM family—and a comparison against a non-LLM baseline derived directly from the recorded data.","section":"§4.4, Table 4"},{"comment":"The usability and usefulness claims are based on a 14-participant study with Likert-scale ratings reported as means (e.g., µ=4.5, µ=4.6) without confidence intervals, significance tests, or a control condition. In particular, there is no no-LLM condition in which the interface shows only the recorded data without the Analytics Insights, so the participants' positive ratings cannot isolate the added value of the LLM-assisted component. The phrase 'highly usable' and claims such as 'participants evaluated the usefulness of the Analytics insights highly' are therefore descriptive only; providing a comparison condition or precise statistical reporting would materially strengthen these claims.","section":"§4.1–§4.3"},{"comment":"The manuscript acknowledges two load-bearing reliability problems: §5 states that LLM extrapolations can produce inaccurate insights 'due to the hallucination of the LLM,' and §4.3 reports participant P13's observation that 'some of the maths are wrong,' with the authors conceding 'occasional inconsistencies in the agent's math computations.' Since the central value proposition is that EXR 'delivers ... actionable insights into user behaviors,' the paper should report the frequency and severity of such failure modes, or otherwise bound the reliability of the generated insights, rather than only listing future plans (multi-agent debate, confidence scores). Without this, the claim that analysts can trust the LLM-generated insights is not supported.","section":"§5 and §4.3"}],"minor_comments":[{"comment":"There is a typo in the first sentence: 'accuractely' should be 'accurately'.","section":"§4.4"},{"comment":"The text lists six analytical aspects of the extracted insights (space, time, action, intent, context, user) but the evaluation in §4.4 uses five criteria (C1–C5). The relationship between the six aspects and the five criteria is not explained; clarifying this mapping would help readers interpret the evaluation.","section":"§3.3.2 and §4.4"},{"comment":"The overhead comparison with ReLive reports single average values over 100 calls without standard deviations or significance testing. For the Log+R case, EXR takes 101.44 ms versus ReLive's 1.01 ms, and the asynchronous fallback of 1.13 ms is described without reporting the measurement protocol; please specify the variance and the conditions under which the asynchronous number was measured.","section":"§4.4, Table 3"},{"comment":"The statement 'All participants agreed (µ=4.1)' is imprecise; 'agreed' implies a consensus threshold that is not defined. Reporting the response distribution or a criterion (e.g., the percentage of participants scoring above 4) would be more informative.","section":"§4.2"},{"comment":"The paper repeatedly refers to Supplementary Materials for prompt details and prototype application specifics. For reproducibility, the main text should state which LLM models were used, the temperature or other sampling parameters, and the number of independent runs for the multi-agent insight generation.","section":"§3.3.1 and Supplementary Materials"}],"recommendation":"major_revision","confidential_remarks":"The reader's conditional verdict is in line with my assessment. The core infrastructure (UAD schema, recorder, visual interface) is a solid contribution and the open-source availability strengthens the paper. The main risk is the circularity of the LLM evaluation and the lack of independent validation of insight reliability; these are fixable with additional experiments, so I would not recommend rejection. The paper fits TVCG's scope well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Know this: EXR is a solid systems paper that delivers a genuinely useful action-centric logging schema and a clean visual analytics interface for XR sessions. The LLM-assisted insight layer is the soft spot; the evidence for its correctness is circular, and the authors admit as much in their limitations.\n\nWhat's actually new: the User Action Descriptor schema binds every multimodal data stream (gaze, controller, audio, 6DoF, referent, context point cloud) to a single user action with intent and trigger source. That is a real step beyond ReLive, PLUME, ARGUS, and MRAT, which log event streams or task-based data. The Unity recorder and web-based visualizer are well designed, and the five use cases (VR game, MR selection, AR scene reconstruction, collaborative AR, AR inspection) demonstrate the cross-virtuality and multi-user claim convincingly. The overhead table is concrete and useful: 0.08–0.14 ms for base logging, with the GLB referent export as the one expensive path, and they mitigate with async calls. That's the kind of engineering detail reviewers want.\n\nWhere it's soft: the quality of LLM insights is evaluated by the same LLM family that produced them, using LIDA-inspired SEVQ criteria. That is circular. The user study (n=14) has no no-LLM control, so participants' positive ratings (4.2–4.6 on 5-point Likert) can't isolate the value added by the AI outputs. One participant said 'some of the maths are wrong,' and Section 5 owns that extrapolation can hallucinate. Also, ambiguous AoI prompts yield useless referent lists. These are real limits, but the paper states them openly, and they don't invalidate the framework as a system contribution. The abstract's 'actionable insights' claim overshoots what the evidence supports; that should be tempered.\n\nThis is a strong candidate for a serious referee: it's an honest, reproducible systems paper with a novel schema and useful implementation. I'd recommend peer review, with the expectation that reviewers ask for either independent evaluation of the LLM outputs or softened claims. If I were working on XR user analytics, I'd cite the UAD schema and reuse the logging design; it's a practical contribution.","headline":"Solid XR analytics framework with a genuinely useful action-centric schema; LLM insight claims outrun their evidence, but the paper is honest and worth reviewing.","tokens_in":22615,"tokens_out":2515,"would_cite":true,"duration_ms":22634,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Explainable XR presents a single action-centric recording format, plus LLM-generated insights, to make XR user analytics consistent across AR, VR, and MR.","keywords":["Extended Reality","User Behavior Analytics","Large Language Models","Visual Analytics","Cross-Virtuality","User Action Descriptor","Multimodal Data Collection","Multi-user XR"],"falsifier":"Take a recorded session with known ground truth—predefined actions, verified referent objects, and hand-computed task-completion times—run Explainable XR's LLM agents and insight generator on it, and count how many generated intentions, referent labels, and numeric claims match the ground truth; if a substantial share are wrong, or if the interface's conclusions change no more than chance when the LLM component is removed, the paper's central claim would be falsified.","tokens_in":21596,"feed_emoji":"🥽","tokens_out":8874,"duration_ms":72002,"temperature":0.7,"pith_summary":"Explainable XR aims to make user-behavior analysis in extended reality as straightforward as reviewing a spreadsheet. The paper proposes a single recording format, the User Action Descriptor, that stores every action—who did what, when, where, how, why, and what was targeted—along with a snapshot of the surrounding scene, so the same pipeline works for AR, VR, and MR sessions, alone or in groups. A visual analytics interface then uses multiple specialized large-language-model agents to summarize sessions, estimate users' intentions, classify physical objects users touched, and highlight where each insight came from. The authors demonstrate the system on five applications spanning individual and collaborative, synchronous and asynchronous XR use, and report that 14 study participants found the interface easy to use and the LLM insights helpful, while also acknowledging that the LLM's arithmetic and some extrapolations can be wrong. If the framework works as claimed, researchers would no longer need custom per-study logging tools to collect, compare, and explain immersive user data.","feed_headline":"One action log unifies XR user analytics across AR, VR, and MR","feed_subtitle":"A who-what-when-where-why-how action record plus LLM insights makes immersive sessions readable and analyzable.","key_machinery":"The central object is the User Action Descriptor (UAD), a 5W1H-inspired schema—When, Where, Who, What, Why, How, plus referent and context—that makes each user action the trigger for information capture. The machinery also includes a platform-agnostic session recorder with template and direct logging, a post-hoc processor that builds context point clouds and uses LLM agents for context description, intention estimation, and referent classification, and a linked visual analytics interface with spatial, temporal, data, plot, and insight viewers whose LLM insights are anchored to source actions by Analysis-of-Interest markers. The UAD is what carries the argument: because every datapoint is tied to an action and its context, the same pipeline can analyze a VR game, an MR selection task, and an AR collaborative session without per-application customization.","core_discovery":"The central discovery is that XR user analytics can be organized around a single action-centric data structure rather than around raw device streams or task-specific logs. The User Action Descriptor binds each logged user action to its type, timestamp, 6DoF location, trigger device, target referent, and a reconstructed point-cloud context, so every piece of multimodal data is interpretable as part of a user's moment-by-moment behavior. The paper argues that this structure is virtuality-agnostic and task-agnostic, and that it scales to multi-user sessions because every action carries its own user identity and context. Large language models then turn the structured logs into analyst-facing insights: they describe action contexts, infer intentions for actions whose meaning depends on conversation or scene context, classify physical referents in AR and MR, and generate up to ten insight summaries tailored to an analyst's stated Analysis-of-Interest. Five prototype applications across VR, MR, and AR, including a collaborative AR analytics task, are used to show the framework's range, with a technical overhead evaluation and a 14-participant user study supporting the usability claims.","pith_inferences":["Editorial extension: the UAD could become a common interchange format for XR user-behavior datasets, enabling cross-study comparison that task-specific logs do not support.","Editorial extension: the reported arithmetic errors point toward a concrete architecture change—letting LLM agents delegate numeric computation to deterministic code—that would raise reliability without altering the framework's design.","Editorial extension: as headset platforms restrict raw camera access, the context-point-cloud component will likely need to shift to OS-provided scene meshes, and a UAD variant that stores those meshes natively would keep the pipeline viable.","Editorial extension: the dependence on Analysis-of-Interest prompt phrasing suggests a measurable improvement path—adding a prompt-refinement step so casual researcher questions yield structured, useful insights."],"forward_implications":["Researchers can collect and compare XR user behavior across AR, VR, and MR studies without designing a new logging format for each experiment.","Multi-user collaborative sessions can be analyzed at the level of individual actions, making it possible to trace who contributed what to a shared task.","LLM-generated insights with linked markers can reduce data overload by giving analysts a starting point and pointing back to the exact actions behind each claim.","The framework's recorder overhead is low enough for interactive use, with the main cost being referent export in a portable 3D format."],"supporting_citations":[{"why":"Provides the open-ended 6DoF recording and replay baseline that Explainable XR extends toward multi-user and AR/MR physical scenes.","marker":"[35]"},{"why":"Supplies the in-situ/ex-situ visual analytics design and the performance comparison baseline used in the system evaluation.","marker":"[31]"},{"why":"Establishes task-based AR/MR logging and intention inference for verbal actions, which the UAD generalizes into an action-centric format.","marker":"[58]"},{"why":"Demonstrates an AR analytics interface with spatiotemporal action tracking, one of the main comparison points for the visual analytics design.","marker":"[11]"},{"why":"Provides the self-evaluation metrics and single-agent versus multi-agent comparison approach used to assess the LLM-generated insights.","marker":"[20]"},{"why":"Supplies the 5W1H user-context model that the UAD schema is explicitly based on.","marker":"[33]"}],"fun_headline_variants":["One action log makes every XR user move interpretable","LLM-assisted XR analytics unify AR, VR, and MR behavior","Explainable XR: LLM turns session logs into user insights","Cross-virtuality user analytics via a single action descriptor","A universal action record enables LLM-driven XR behavior analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that LLM-generated insights, intention estimates, referent classifications, and numeric summaries are accurate enough for analysts to rely on them; the paper acknowledges LLM hallucination and one participant's report that 'some of the maths are wrong' in the temporal-pattern analysis.","fun_headline_variants_meta":{"raw":{"variants":["One action log makes every XR user move interpretable","LLM-assisted XR analytics unify AR, VR, and MR behavior","Explainable XR: LLM turns session logs into user insights","Cross-virtuality user analytics via a single action descriptor","A universal action record enables LLM-driven XR behavior analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000917,"raw_usage":{"total_tokens":3969,"prompt_tokens":1014,"completion_tokens":2955,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":2867}},"tokens_in":630,"tokens_out":2955,"duration_ms":19487,"temperature":1.0,"reasoning_tokens":2867,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:35:22.323788+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a recorded session with known ground truth—predefined actions, verified referent objects, and hand-computed task-completion times—run Explainable XR's LLM agents and insight generator on it, and count how many generated intentions, referent labels, and numeric claims match the ground truth; if a substantial share are wrong, or if the interface's conclusions change no more than chance when the LLM component is removed, the paper's central claim would be falsified.","supporting_citations":[{"cited_title":"Javerliat, S","cited_arxiv_id":null,"evidence_quote":"Provides the open-ended 6DoF recording and replay baseline that Explainable XR extends toward multi-user and AR/MR physical scenes."},{"cited_title":"Hubenschmid, J","cited_arxiv_id":null,"evidence_quote":"Supplies the in-situ/ex-situ visual analytics design and the performance comparison baseline used in the system evaluation."},{"cited_title":"Nebeling, M","cited_arxiv_id":null,"evidence_quote":"Establishes task-based AR/MR logging and intention inference for verbal actions, which the UAD generalizes into an action-centric format."},{"cited_title":"Castelo, J","cited_arxiv_id":null,"evidence_quote":"Demonstrates an AR analytics interface with spatiotemporal action tracking, one of the main comparison points for the visual analytics design."},{"cited_title":"Jang, E.-J","cited_arxiv_id":null,"evidence_quote":"Supplies the 5W1H user-context model that the UAD schema is explicitly based on."}],"review_version":1}