{"id":"35fcc52c-a054-4904-9a52-bcb91ef3bacd","arxiv_id":"2411.08409","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A cross-modal transformer using heterogeneous scene graphs predicts VR user trajectories more accurately than gaze- and point-cloud-only baselines on the CREATTIVE3D dataset.","lead":"This paper introduces DiVR, a model that predicts where people will walk in virtual reality scenes by combining motion history, eye gaze, and a structured graph of the scene's objects and spaces. It reports lower prediction errors than earlier methods on the CREATTIVE3D VR dataset, especially in complex tasks and simulated low vision, suggesting VR data can support context-aware movement prediction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Same-condition DiVR-Het numbers differ across Tables 4 and 7 (and Table 3), so the headline accuracy claim is not tied to a single reproducible evaluation protocol.","rationale":"The reader's verdict was CONDITIONAL, and my read does not move that verdict; the concern is real but addressable through the public code and the requested clarifications. I identify the internal numerical inconsistency as the most load-bearing concern because it directly undermines the quantitative basis of the central claim, independent of any assumptions about baseline tuning. The reader's weakest_assumption focused on baseline adaptation, but the reader's rationale also mentioned inconsistent numbers across Tables 3 and 7; my emphasis is on that inconsistency, so agreement is partial. The concrete test is designed to resolve the inconsistency by reproduction and to check whether DiVR's advantage survives seed-level variation and fair baseline adaptation. Until those checks are reported, CONDITIONAL acceptance remains appropriate, so no verdict adjustment is needed.","tokens_in":10753,"tokens_out":6382,"duration_ms":64785,"concrete_test":"Clone https://gitlab.inria.fr/ffrancog/creattive3d-divr-model and run the repository's training/evaluation pipeline with the stated 70/15/15 split. First verify whether a single configuration can reproduce the DiVR-Het NV+ST entries in Table 4 (0.604/0.696) and Table 7 (0.535/0.705); if it cannot produce both, the inconsistency is confirmed. Then run MLP, TRACK, GIMO, DiVR-Hom, and DiVR-Het for at least five random seeds with identical data splits and preprocessing, reporting mean +/- std ADE/FDE for every condition in Tables 3-7. If DiVR-Het does not consistently beat GIMO and TRACK at the per-seed level in the conditions where it is claimed to win, the central claim should be weakened. Also inspect the baseline scripts to confirm GIMO's point-cloud and gaze pipeline was retrained on CREATTIVE3D with comparable epochs and hyperparameter search.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is quantitative: DiVR is said to achieve higher accuracy through heterogeneous graph context. In the manuscript, DiVR-Het under the named condition NV+ST is reported as ADE/FDE 0.604/0.696 in Table 4 and 0.535/0.705 in Table 7; Table 3 reports DiVR-Het as 0.588/0.842 without naming the condition, and it is these numbers that produce the claimed 31.2% and 44.3% reductions over MLP. If these entries correspond to the same 70/15/15 split and evaluation protocol, two of them cannot both be correct; if they correspond to different protocols, the paper does not state which protocol supports the headline. Either way, the numerical evidence for the central claim is not anchored. This concern is more load-bearing than baseline tuning alone because it involves the paper's own reported values: even a fairly tuned baseline comparison would remain unfalsifiable while the headline numbers are ambiguous. The absence of error bars or seed-level statistics further prevents distinguishing real improvements from noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DiVR, a cross-modal transformer for human trajectory prediction in VR scenes, built on the PerceiverIO architecture and integrating three modalities: past motion, gaze-interpolated point clouds, and temporal heterogeneous graph context extracted with a TemporalGCN. Using the CREATTIVE3D dataset, the authors report ADE/FDE comparisons against MLP, TRACK, and GIMO baselines, plus ablation studies and generalization tests across users, scenes, and tasks. The central claim is that integrating heterogeneous static and dynamic scene context improves prediction accuracy and adaptability over existing baselines and over homogeneous graphs. Source code is made publicly available.","tokens_in":10959,"tokens_out":6439,"duration_ms":60962,"significance":"If the quantitative claims are validated, the paper would make a useful contribution by demonstrating that high-level scene graphs, rather than raw point clouds, provide effective context for human trajectory prediction, and by extending this line of work to controlled VR environments with simulated low vision. The release of code and the use of a publicly available VR dataset are concrete assets. However, the empirical evidence as presented is not currently anchored: the same model is reported with materially different errors in different tables, baseline adaptation is underspecified, and no uncertainty quantification is provided. The significance of the work therefore depends on the authors reconciling these issues in revision.","major_comments":[{"comment":"The same DiVR-Het configuration is reported with inconsistent ADE/FDE values across tables. Under the named condition NV+ST, Table 4 reports DiVR-Het as 0.604/0.696, Table 7 reports 0.535/0.705, and Table 3 reports 0.588/0.842 without naming the condition; the Table 3 numbers are the basis for the 31.2% and 44.3% reductions stated in §4.2. Similarly, the complex-task row for DiVR-Het is 0.910/1.661 in Table 4 but 0.695/1.218 in Table 7. Baselines also differ across tables (e.g., MLP is 0.854/1.512 in Table 3 and 0.652/1.139 in Table 4). The paper must state which protocol, data split, and training condition each table refers to, and the headline claim must be tied to a single reproducible evaluation protocol.","section":"§4.2, Tables 3, 4, and 7"},{"comment":"The adaptation of baselines to the VR setting is not specified. The paper cites original papers for MLP, TRACK, and GIMO but does not state whether they were retrained on CREATTIVE3D, how their hyperparameters were chosen, or how GIMO's gaze and point-cloud pipeline was adapted from real-world motion capture to the 2D head-position VR data used here. If the baselines were not comparably tuned or preprocessed, DiVR's reported gains could be an artifact of the comparison rather than a real modeling improvement. Please report the exact adaptation protocol, hyperparameter selection, and tuning budget for each baseline.","section":"§4.1"},{"comment":"The abstract's blanket claim that DiVR achieves higher accuracy than other models is not supported by several reported rows. In Table 4, under NV+ST, GIMO and TRACK achieve lower ADE (0.509 and 0.515) than DiVR-Het (0.604). In Table 6, under scene generalization GIMO has a lower ADE (0.997 vs 1.073), and under task generalization GIMO has both a lower ADE (1.296 vs 1.384) and a lower FDE (2.100 vs 2.605) than DiVR-Het. The superiority claims should be qualified to the conditions and metrics where DiVR is actually better, or the numerical results should be corrected.","section":"§4.2, Tables 4 and 6"},{"comment":"No error bars, confidence intervals, or significance tests are reported. All tables present single-run point estimates, and the user-generalization experiment is based on one random selection of 10 test users. Many of the differences used to support the central claim are small (e.g., Table 4 low-vision ADE: DiVR-Het 0.615 vs GIMO 0.630 vs TRACK 0.636), so without repeated runs or resampled splits it is impossible to distinguish real gains from noise. Please report results over multiple seeds and/or user-split resamplings with standard deviations and a significance test.","section":"All tables and §4.2"},{"comment":"It is not stated whether the 70/15/15 train/validation/test split is user-disjoint. Section 4.2 defines user generalization separately by holding out 10 users, which suggests the standard context-evaluation split may mix trajectories from the same user across train and test. If so, the reported accuracy and generalization numbers would be inflated. The paper must clarify the split granularity and, if the goal is generalization, ensure that the standard split is also user-disjoint.","section":"§4.1 and §4.2"}],"minor_comments":[{"comment":"In the Graph-based techniques paragraph, 'trajectory precition' should be 'trajectory prediction'.","section":"§2"},{"comment":"The caption writes 'DiVT-het' where the model name should be 'DiVR-Het'.","section":"Figure 4 caption"},{"comment":"The loss function components Ltrans, Lrec, and Ldes_trans are named but not defined; please provide exact equations and specify the relative weights used in training.","section":"§3.2"},{"comment":"The rules for assigning edge types (approach, adjacent, interaction) and their temporal attributes from the CREATTIVE3D annotations should be stated precisely, as they determine both the graph representation and the input to the TemporalGCN.","section":"§3.1"},{"comment":"The sentence 'DiVR-Hom reduces its ADE from 0.592 to 0.470 and its FDE from 0.908 to 0.816' compares results under different test conditions (NV+ST vs Complex Task) rather than the same condition under different training regimes; please report the intended like-for-like comparison.","section":"§4.2, Table 5 discussion"},{"comment":"Since the Perceiver architecture is motivated by efficiency, reporting model parameter counts and training/inference time would strengthen the presentation.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's motivation and dataset/code release are valuable, and the qualitative examples are suggestive. The main obstacle is that the numerical evidence for the central claim is not reproducible as presented, due to cross-table inconsistencies and missing uncertainty quantification. I recommend major revision rather than rejection because these issues are fixable within the manuscript's scope by standardizing the evaluation protocol, rerunning baselines under controlled conditions, and reporting variance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this paper is worth reading for the application and the public code, but the central quantitative claim is not currently anchored. DiVR is the first to apply heterogeneous temporal graphs to real-human trajectory prediction in interactive VR with simulated low vision. That is a genuine, if incremental, new combination of known components, and the graph-context ablation shows the graph branch carries real weight. The code and dataset are public, which makes the work reproducible in principle, and the generalization experiments (user, scene, task) are a thoughtful addition.\n\nThe soft spots are proportional and addressable. The same configuration (DiVR-Het, NV+ST) appears with different numbers across tables: 0.588/0.842 in Table 3, 0.604/0.696 in Table 4, and 0.535/0.705 in Table 7. Those cannot all describe the same split and protocol, and the paper never explains which protocol supports the headline 31.2%/44.3% reductions. That is a load-bearing inconsistency. Beyond that, there are no error bars or significance tests, so even a single consistent number would leave the gains unquantified. The baseline description is also thin: we are not told whether MLP, TRACK, and GIMO were retrained on CREATTIVE3D or how their hyperparameters were chosen, which leaves room for an undertuned-comparison artifact.\n\nThe abstract overclaims. In Table 4, GIMO and TRACK both beat DiVR-Het on ADE under NV+ST (0.509 and 0.515 vs. 0.604), and in Table 6 DiVR-Het loses to GIMO on scene and task generalization. So “higher accuracy” is not universally true; the strength is in complex-task and low-vision conditions and in graph ablation. The authors do acknowledge reliance on high-quality scene-graph data as a limitation, and the internal reasoning is otherwise coherent.\n\nWho this is for: researchers working on context-aware trajectory prediction, especially in VR and assisted-navigation settings. The paper deserves a serious referee, but only with a clear request for revision: reconcile the table numbers, add repeated-seed statistics, document baseline adaptation, and rewrite the abstract to match the evidence. I would bring it to reading group because the methodological inconsistency is a useful cautionary case, and I would cite the dataset or the graph-context ablation if I build on that line. Recommend: send to peer review with major-revision expectations.","headline":"DiVR is a solid first application of heterogeneous graph context to VR trajectory prediction, but its headline numbers are internally inconsistent and the blanket superiority claim is not supported by all its own tables.","tokens_in":11557,"tokens_out":1612,"would_cite":true,"duration_ms":17431,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiVR, a cross-modal transformer that represents VR scenes as heterogeneous temporal graphs, predicts human trajectories more accurately than MLP, gaze-only, and point-cloud baselines on the CREATTIVE3D dataset.","keywords":["virtual reality","trajectory prediction","heterogeneous graphs","cross-modal transformer","gaze","point cloud","low vision","scene understanding"],"falsifier":"A direct test would be to retrain MLP, TRACK, and GIMO on CREATTIVE3D with the identical data splits, input sampling, and a hyperparameter search (including learning rate, epoch count, and context preprocessing), then measure whether DiVR-Het still beats the best-tuned baseline by the reported margins. A second check would be to replace the heterogeneous graph in DiVR-Het with a homogeneous graph or random edge features: if the ADE/FDE gap over a no-graph version disappears, the improvement is due to the graph structure; if it persists, the gain comes from the architecture or other context.","tokens_in":10530,"feed_emoji":"🥽","tokens_out":9956,"duration_ms":84074,"temperature":0.7,"pith_summary":"The paper introduces DiVR, a trajectory-prediction model for virtual-reality scenes that fuses three sources of information about a user: their past head motion, gaze-directed scene point clouds, and a heterogeneous graph describing the scene's static layout and dynamic interactions. The central claim is that the graph context, in particular, improves prediction accuracy over models that use only motion, gaze, or raw point clouds, and that it helps the model adapt to new users, tasks, scenes, and simulated low vision. On the CREATTIVE3D dataset, DiVR's heterogeneous-graph variant reduces average displacement error by 31.2% and final displacement error by 44.3% compared with an MLP baseline, and also beats gaze-only and point-cloud baselines in the main evaluation. The authors further show through ablations that removing the graph context hurts accuracy more than removing gaze. The work matters because VR environments offer a controlled way to collect rich behavioral data, and if the graph-context advantage transfers, it could improve pedestrian prediction in autonomous driving and accessible navigation.","feed_headline":"Heterogeneous graphs cut VR trajectory error by up to 44%","feed_subtitle":"DiVR combines motion, gaze, and scene graphs to beat MLP, TRACK, and GIMO in VR path prediction.","key_machinery":"The central mechanism is the heterogeneous temporal scene graph, where nodes and edges carry type labels that let the model distinguish a pedestrian waiting at a traffic light from one crossing, and a location from a moving vehicle. A Temporal Graph Convolutional Network with edge-conditioned filters converts the evolving graph into a latent context vector for each timestep; global mean pooling and concatenation produce a spatiotemporal embedding. This context vector is combined in a cross-modal transformer with a PointNet++-encoded gaze-interpolated point cloud and a PerceiverIO motion encoder, so that the final prediction is conditioned on both the raw sensory input and the structured relational description of the environment.","core_discovery":"The paper's core claim is that representing a virtual urban scene as a directed heterogeneous temporal graph—with typed nodes for locations, pedestrians, vehicles, traffic lights, and buttons, and typed edges for proximity, adjacency, and active interactions—and encoding that graph with a temporal graph convolutional network yields a context feature that, when fused with motion and gaze via a PerceiverIO cross-modal transformer, reduces trajectory prediction error relative to existing baselines. In the balanced CREATTIVE3D evaluation, DiVR-Het reaches an ADE of 0.588 m and an FDE of 0.842 m, versus 0.854/1.512 for the MLP, 0.759/1.247 for TRACK, and 0.757/1.240 for GIMO. The paper also reports that DiVR-Het improves over GIMO on complex-task and low-vision generalization tests when trained on simple tasks, and that graph ablation incurs a larger error increase than gaze ablation, indicating the graph is the more important context modality.","pith_inferences":["If heterogeneous graph context is indeed the main carrier of the improvement, then a practical route for real-world deployment would be to replace manually annotated VR scene graphs with automatically parsed scene graphs from HD maps and sensor data; a testable extension is to evaluate DiVR on real pedestrian datasets with predicted graphs.","The ablation gap between graph and gaze suggests that for trajectory prediction, semantic structure (what objects are and how they relate) may matter more than raw perceptual features; a follow-up could test this by feeding the same graph embedding into a simple concatenation model instead of the full cross-modal transformer.","The mixed results in task generalization hint that the heterogeneous graph's benefit depends on the specific relation types: the model improved in complex tasks involving traffic lights and buttons but not in the symmetric 'return to start' task; an analysis of which edge types drive the gain could guide graph design.","Because the CREATTIVE3D dataset provides oracle annotations, a natural stress test is to measure how DiVR degrades when the graph is corrupted (e.g., random edges or missing types), which would quantify how much of the gain relies on perfect scene understanding."],"forward_implications":["DiVR-Het outperforms all three baselines on the primary CREATTIVE3D evaluation, with a 31.2% lower ADE and 44.3% lower FDE than an MLP that sees only past positions.","Ablation results show that replacing the heterogeneous graph with zeros increases ADE by up to 35.8% under complex tasks, implying the graph contributes more than gaze to prediction accuracy.","Training on both simple and complex tasks (or on both normal and low vision) improves DiVR's complex-task and low-vision performance beyond the baselines, suggesting the model benefits from exposure to diverse conditions.","In user generalization, DiVR-Hom achieves the best ADE (0.513) and DiVR-Het the best FDE (0.828) among compared models; task generalization remains difficult, with all models' FDE around 2, but DiVR-Hom still leads in ADE.","The paper positions heterogeneous graph context as a complement to gaze and point-cloud modalities, and the qualitative examples show DiVR correctly pausing for a traffic light or walking toward a button where GIMO predicts an immediate crossing."],"supporting_citations":[{"why":"Supplies the GIMO baseline that DiVR must beat and the gaze-plus-point-cloud cross-modal design DiVR extends.","marker":"[30]"},{"why":"Provides the PerceiverIO architecture used for the motion, gaze, and graph branches.","marker":"[10]"},{"why":"Introduces the heterogeneous graph representation for urban scene motion prediction that DiVR adapts to VR.","marker":"[2]"},{"why":"Provides the TRACK gaze-based baseline and the notion of gaze-modulated context for motion prediction.","marker":"[16]"},{"why":"Provides the MLP baseline that gives the paper's primary accuracy comparison.","marker":"[9]"},{"why":"Supplies the CREATTIVE3D VR dataset with scene annotations, gaze, tasks, and low-vision conditions used in all experiments.","marker":"[23]"},{"why":"Establishes message-passing graph neural networks, cited as the basis for the TemporalGCN.","marker":"[8]"},{"why":"Defines dynamic edge-conditioned convolution, the specific mechanism used in the TemporalGCN to handle typed edges.","marker":"[18]"}],"fun_headline_variants":["DiVR's heterogeneous graph cuts VR trajectory error 44%","Graph beats gaze: DiVR cross-modal transformer lowers VR path error","DiVR uses graph context to predict VR paths better than gaze alone","Heterogeneous graph fusion reduces VR trajectory error 44%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported accuracy gains assume that the MLP, TRACK, and GIMO baselines were adapted to the CREATTIVE3D VR setting with equally careful hyperparameter tuning and preprocessing; the experimental section only points to the baselines' original papers and does not state whether they were retrained on the VR data or how their hyperparameters were selected, so DiVR's advantage could partly stem from undertuned comparisons rather than the model itself.","fun_headline_variants_meta":{"raw":{"variants":["DiVR's heterogeneous graph cuts VR trajectory error 44%","Graph beats gaze: DiVR cross-modal transformer lowers VR path error","DiVR uses graph context to predict VR paths better than gaze alone","Heterogeneous graph fusion reduces VR trajectory error 44%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3557,"prompt_tokens":978,"completion_tokens":2579,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":2518}},"tokens_in":594,"tokens_out":2579,"duration_ms":20286,"temperature":1.0,"reasoning_tokens":2518,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:37:41.790327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to retrain MLP, TRACK, and GIMO on CREATTIVE3D with the identical data splits, input sampling, and a hyperparameter search (including learning rate, epoch count, and context preprocessing), then measure whether DiVR-Het still beats the best-tuned baseline by the reported margins. A second check would be to replace the heterogeneous graph in DiVR-Het with a homogeneous graph or random edge features: if the ADE/FDE gap over a no-graph version disappears, the improvement is due to the graph structure; if it persists, the gain comes from the architecture or other context.","supporting_citations":[{"cited_title":"In: European Conference on Computer Vision (2022)","cited_arxiv_id":null,"evidence_quote":"Supplies the GIMO baseline that DiVR must beat and the gaze-plus-point-cloud cross-modal design DiVR extends."},{"cited_title":"Transportation Research Part C: Emerging Technologies157, 104405 (2023)","cited_arxiv_id":null,"evidence_quote":"Introduces the heterogeneous graph representation for urban scene motion prediction that DiVR adapts to VR."},{"cited_title":"IEEE Transactions on Pattern Analysis and Machine Intelligence 44(9), 5681–5699 (2021) 16 F.Franco et al","cited_arxiv_id":null,"evidence_quote":"Provides the TRACK gaze-based baseline and the notion of gaze-modulated context for motion prediction."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"Provides the MLP baseline that gives the paper's primary accuracy comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CREATTIVE3D VR dataset with scene annotations, gaze, tasks, and low-vision conditions used in all experiments."},{"cited_title":"In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","cited_arxiv_id":null,"evidence_quote":"Defines dynamic edge-conditioned convolution, the specific mechanism used in the TemporalGCN to handle typed edges."}],"review_version":1}