Pith. sign in

REVIEW 4 major objections 6 minor 152 references

Given only a current view, an agent can mentally simulate a panoramic trajectory to a target situation and answer spatial what-if questions without physical exploration.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 13:49 UTC pith:BMF7UTPU

load-bearing objection Useful new dataset for target-directed panoramic mental trajectories, but path-phase QA gains from generated videos are still near-zero, so the core emulative claim is only half-shown. the 4 major comments →

arxiv 2603.06445 v3 pith:BMF7UTPU submitted 2026-03-06 cs.CV

What if? Emulative Simulation with World Models for Situated Reasoning

classification cs.CV
keywords emulative simulationworld modelssituated reasoningpanoramic videomental explorationwhat-if questionssim-to-real transfer
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Situated reasoning usually depends on active exploration, which robots cannot always perform and people with visual impairments may not risk. This paper asks whether an agent that sees only its present surroundings can still answer spatial what-if questions by imagining the journey to a target situation. It introduces WanderDream: 15.8K panoramic videos of imagined paths across 1,088 real scenes, plus 158K questions about start states, paths, and end states. Experiments show that intermediate imagined frames improve end-state answers beyond start-and-end frames alone, that better trajectory generation correlates with stronger reasoning, and that training on these imagined paths transfers to real head-mounted panoramic captures. The work aims to let agents reason about inaccessible places through mental exploration rather than physical movement.

Core claim

Mental exploration is essential for situated reasoning. Even when a target location can be inferred from the current view, intermediate frames produced by world models strengthen end-state answers; higher-quality generated trajectories yield stronger reasoning; and models trained on WanderDream transfer to real panoramic recordings despite agent occlusions and non-shortest real paths.

What carries the argument

WanderDream, a dual dataset: WanderDream-Gen supplies 15.8K panoramic videos of trajectories from current viewpoints to robot navigation or human action situations, and WanderDream-QA supplies 158K questions spanning start-state, path, and end-state reasoning along those trajectories.

Load-bearing premise

The paper treats shortest-path trajectories planned in simulation as a faithful enough stand-in for how agents mentally simulate going somewhere.

What would settle it

If intermediate frames from a high-quality generated trajectory never raise end-state QA scores over giving only start and end frames on held-out real panoramic captures, the claim that mental exploration is essential would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Agents can answer situated questions about places they cannot physically reach by generating an imagined trajectory first.
  • World-model video quality on target-directed trajectories predicts downstream spatial reasoning accuracy.
  • Training on simulated shortest-path panoramic videos improves both generation and QA on real head-mounted captures.
  • A generate-then-reason pipeline can support emulative simulation until unified video-and-text models exist.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If intermediate frames drive the end-state gains, shorter or sparsely sampled imagined clips may still suffice for many what-if queries, cutting generation cost.
  • The same imagination engine could support human–robot collaboration when each party faces different embodiment limits.
  • Occlusion and anchor-confusion failures suggest that releasing depth and semantics with the trajectories will be needed before the method is reliable in cluttered homes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces WanderDream, a large-scale dataset and benchmark for emulative simulation: mentally generating panoramic trajectories from a current egocentric view to a target situation and answering spatial what-if questions along that path without physical exploration. WanderDream-Gen provides 15.8K panoramic videos over 1,088 real scenes (HM3D robotic navigation, ScanNet++ human situations, plus a small real-world set); WanderDream-QA supplies 158K QA pairs spanning start, path, and end phases (10 types). Sequential world-model + MLLM frameworks (prompt extension or fine-tuning of Wan/CogVideoX/HunyuanVideo; Qwen3-VL reasoning) and a closed-loop baseline (MindJourney) are evaluated on generation metrics (FVD, End-FID, S-SSIM, LPIPS) and LLM-as-judge QA correctness. The authors conclude that mental exploration is essential, world models perform competitively on Gen, imagination facilitates QA, and the data transfer to real panoramic captures.

Significance. If the claims hold, the work cleanly separates emulative (experience-oriented) from instrumental (task-oriented) simulation and supplies the first large panoramic trajectory + path-aligned QA resource for that setting, with clear relevance to robots under embodiment constraints and assistive systems for visually impaired users. Strengths include scale and multi-source construction, a four-axis experimental design (GT frame ablations, Gen metrics, QA with generated videos, sim-to-real), high human–GPT judge correlation (Spearman 0.972 on 800 pairs), and Likert quality scores above 4.7. The dataset and problem framing are a genuine contribution even if current world-model gains remain modest; release of code and data would further raise impact.

major comments (4)
  1. [Abstract; §5.2; Tables 4–5] Abstract claim (3) and §5.2 state that imagination “substantially facilitates reasoning on WanderDream-QA.” Tables 4–5 show path-phase averages (the defining emulative segment: LS/SE/OR/DC or RP) rising only from 36.4→37.6 (ScanNet++) and 53.2→53.4 (HM3D) under the best generated s∆5 inputs—gains of ~0–1.2 points—while end-state gains are modest (+2.0 / +2.2). The s0-only baseline already dominates start-state questions and remains competitive on path. The GT necessity result in Fig. 7 therefore does not transfer to the imperfect imaginations the framework actually produces. The abstract, introduction, and conclusion should restate claim (3) to match the measured effect sizes (e.g., modest end-state support correlated with End-FID; negligible path-phase facilitation with current generators).
  2. [§5.2; Tables 3–5] The paper asserts that “models with higher video generation quality also provide stronger support for reasoning along the trajectory” (§5.2). Path-phase scores barely move across large FVD/End-FID spreads (Table 3 vs. Tables 4–5), whereas end-state scores track End-FID more than trajectory coherence (FVD). This weakens the claimed link between emulative trajectory quality and path reasoning—the core of the proposed setting. A quantitative correlation (or ablation that isolates intermediate-frame fidelity) should be reported, or the claim narrowed to end-state prediction.
  3. [§3.1; §3.4; Tab. 6; Fig. 8] §3.1 grounds trajectories in shortest-path planners (Habitat; PRM+Dijkstra / linear interpolation) citing Epstein et al. on cognitive maps. The real-world set (§3.4, Tab. 6) already notes non-shortest, variable-velocity human motion and agent occlusions; FVD degrades and MindJourney underperforms s0-only. The weakest modeling assumption—that shortest-path panoramic videos are a faithful enough proxy for mental simulation—is load-bearing for the transfer claim yet only lightly stress-tested. Either expand real-world diversity (agents, non-shortest routes) or explicitly bound the claim to “shortest-path emulative proxies.”
  4. [§5.2; Tables 4–6; Fig. 7] QA improvements lack error bars, multiple seeds, or significance tests (Tables 4–6, Fig. 7). Given effect sizes of 0–2 points on a 0–100 LLM-judge scale and known judge variance, it is unclear whether reported gains exceed noise. At minimum, report standard deviations over questions/scenes or bootstrap intervals for the headline deltas used to support claims (1) and (3).
minor comments (6)
  1. [Table 1] Table 1 header and body contain spacing artifacts (“W anderDream”, “V enue”).
  2. [Table 4] Table 4 lists “HuyuanVideo+LoRA” (missing ‘n’); align naming with Table 3/5 (HunyuanVideo).
  3. [§2] Related Work: “rea- world settings” → “real-world settings.”
  4. [Fig. 7] Fig. 7 is central to claim (1) but is hard to read in grayscale; consider distinct markers/line styles and numerical annotations for the end-state s0,sT vs s∆5 comparison.
  5. [§6] Latency numbers in §6 (35–283 s per trajectory) are useful; stating whether they include MLLM reasoning and at what resolution would aid reproducibility.
  6. [§3.1] Clarify whether depth/semantic maps released with Gen are used by any reported model or only for future work (§3.1, §E).

Circularity Check

0 steps flagged

No circularity: empirical dataset + generation/QA benchmarks; no prediction reduces to a fitted free parameter or self-citation chain.

full rationale

WanderDream is an empirical systems paper. WanderDream-Gen supplies held-out panoramic trajectories (shortest-path Habitat planner on HM3D; PRM/Dijkstra or linear interpolation on ScanNet++) and evaluates world models with standard video metrics (FVD, End-FID, S-SSIM, LPIPS) against ground-truth video. WanderDream-QA supplies independently GPT-generated answers from SoM annotations and trajectory metadata; MLLM answers are scored by a separate LLM-as-judge protocol whose Spearman correlation with human ratings is reported (0.972). The necessity-of-imagination claim is an ablation over ground-truth frame subsets (s0 vs s0+sT vs s∆5 vs full video), not a derivation. The claim that better generation correlates with better QA is a post-hoc empirical correlation across models, not a fitted parameter renamed as a prediction. The shortest-path assumption is an external citation to Epstein et al. (cognitive maps), not a self-citation uniqueness theorem. Real-world transfer is measured on held-out panoramic captures. There is no equation, free parameter, or load-bearing uniqueness result that reduces a claimed prediction to its own inputs by construction. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 1 invented entities

The central empirical claims rest on standard simulation and cognitive-map assumptions plus a few modeling choices that are not independently validated outside this paper. No free parameters are fitted to force the main result; the free choices are design decisions for data generation.

free parameters (2)
  • max path length / sampling distances
    HM3D paths capped at 5 m; ScanNet++ start distances fixed to 1.5–3 m; video length fixed to 21 frames (4N+1). These are hand-chosen design constants that define the distribution of the dataset.
  • head-pitch range ±30°
    Random pitch at start and constrained pitch at end for human situations are chosen by the authors to mimic real egocentric distortion; not derived from measurement.
axioms (3)
  • domain assumption Imagined trajectories inside a cognitive map follow the shortest path (Epstein et al. 2017).
    Used in §3.1 to justify Habitat shortest-path planner and PRM/Dijkstra construction of ground-truth videos.
  • domain assumption GPT-5 with Set-of-Mark annotations produces sufficiently accurate long-form spatial QA for training and evaluation.
    QA generation pipeline (§3.2); human Likert check on 800 pairs is post-hoc validation, not an independent derivation.
  • domain assumption LLM-as-a-judge (GPT-4o-mini) scores of 1–5 faithfully measure factual correctness of free-form answers.
    Evaluation metric Eq. (1) and §4; supported by Spearman 0.97 with humans but still an external model assumption.
invented entities (1)
  • emulative simulation (as operationalized by WanderDream) no independent evidence
    purpose: Name the experience-oriented layer of mental imagination that produces a full visual trajectory to a target situation and supports what-if QA along the path.
    The term is taken from Moulton & Kosslyn 2011 but is given a concrete dataset and evaluation protocol here; no new physical entity is postulated.

pith-pipeline@v1.1.0-grok45 · 30406 in / 2585 out tokens · 24675 ms · 2026-07-15T13:49:31.943829+00:00 · methodology

0 comments
read the original abstract

Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints of robots or safety concerns of visually impaired users. Given only a limited observation, can an agent mentally simulate a future trajectory toward a target situation and answer spatial what-if questions? We introduce WanderDream, the first large-scale dataset designed for the emulative simulation of mental exploration, enabling models to reason without active exploration. WanderDream-Gen comprises 15.8K panoramic videos across 1,088 real scenes from HM3D, ScanNet++, and real-world captures, depicting imagined trajectories from current viewpoints to target situations. WanderDream-QA contains 158K question-answer pairs, covering starting states, paths, and end states along each trajectory to comprehensively evaluate exploration-based reasoning. Extensive experiments with world models and MLLMs demonstrate (1) that mental exploration is essential for situated reasoning, (2) that world models achieve compelling performance on WanderDream-Gen, (3) that imagination substantially facilitates reasoning on WanderDream-QA, and (4) that WanderDream data exhibit remarkable transferability to real-world scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

152 extracted references · 30 linked inside Pith

  1. [1]

    In: ECCV (2020)

    Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., Guibas, L.J.: ReferIt3D: Neural listeners for fine-grained 3D object identification in real-world scenes. In: ECCV (2020)

  2. [2]

    arXiv preprint arXiv:2509.23661 (2025)

    An, X., Xie, Y., Yang, K., Zhang, W., Zhao, X., Cheng, Z., Wang, Y., Xu, S., Chen, C., Wu, C., Tan, H., Li, C., Yang, J., Yu, J., Wang, X., Qin, B., Wang, Y., Yan, Z., Feng, Z., Liu, Z., Li, B., Deng, J.: LLaVA-OneVision-1.5: Fully open framework for democratized multimodal training. arXiv preprint arXiv:2509.23661 (2025)

  3. [3]

    Anthropic: Claude.https://www.anthropic.com(2024)

  4. [4]

    In: CVPR (2025) 16 Ruiping Liuet al

    Bar, A., Zhou, G., Tran, D., Darrell, T., LeCun, Y.: Navigation world models. In: CVPR (2025) 16 Ruiping Liuet al

  5. [5]

    arXiv preprint arXiv:2311.15127 (2023)

    Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)

  6. [6]

    In: ICCV (2025)

    Çelen, A., Pollefeys, M., Barath, D., Armeni, I.: HouseTour: A virtual real estate A(I)gent. In: ICCV (2025)

  7. [7]

    arXiv preprint arXiv:2506.21539 (2025)

    Cen, J., Yu, C., Yuan, H., Jiang, Y., Huang, S., Guo, J., Li, X., Song, Y., Luo, H., Wang, F., Zhao, D., Chen, H.: WorldVLA: Towards autoregressive action world model. arXiv preprint arXiv:2506.21539 (2025)

  8. [8]

    In: CVPR (2024)

    Chen, H., Hou, Y., Qu, C., Testini, I., Hong, X., Jiao, J.: 360+x: A panoptic multi-modal scene understanding dataset. In: CVPR (2024)

  9. [9]

    In: NeurIPS (2024)

    Chen, L., Wei, X., Li, J., Dong, X., Zhang, P., Zang, Y., Chen, Z., Duan, H., Bin, L., Tang, Z., Yuan, L., Qiao, Y., Lin, D., Zhao, F., Wang, J.: ShareGPT4Video: Improving video understanding and generation with better captions. In: NeurIPS (2024)

  10. [10]

    arXiv preprint arXiv:2509.08519 (2025)

    Chen, L., Ma, T., Liu, J., Li, B., Chen, Z., Liu, L., He, X., Li, G., He, Q., Wu, Z.: HuMo: Human-centric video generation via collaborative multi-modal conditioning. arXiv preprint arXiv:2509.08519 (2025)

  11. [11]

    In: NeurIPS (2025)

    Chen, L., Zhou, Z., Zhao, M., Wang, Y., Zhang, G., Huang, W., Sun, H., Wen, J.R., Li, C.: FlexWorld: Progressively expanding 3D scenes for flexible-view ex- ploration. In: NeurIPS (2025)

  12. [12]

    In: ICCV (2025)

    Chen, Y., Wang, Y., Zhang, Z.: DrivingGPT: Unifying driving world modeling and planning with multi-modal autoregressive transformers. In: ICCV (2025)

  13. [13]

    arXiv preprint arXiv:2509.13317 (2025)

    Cheng, A.C., Fu, Y., Chen, Y., Liu, Z., Li, X., Radhakrishnan, S., Han, S., Lu, Y., Kautz, J., Molchanov, P., Yin, H., Wang, X., Liu, S.: 3D aware region prompted vision language model. arXiv preprint arXiv:2509.13317 (2025)

  14. [14]

    Cambridge University Press, Cambridge, UK (1997)

    Clancey, W.J.: Situated Cognition: On Human Knowledge and Computer Repre- sentations. Cambridge University Press, Cambridge, UK (1997)

  15. [15]

    ACM Computing Surveys (2025)

    Ding, J., Zhang, Y., Shang, Y., Zhang, Y., Zong, Z., Feng, J., Yuan, Y., Su, H., Li, N., Sukiennik, N., Xu, F., Li, Y.: Understanding world or predicting future? A comprehensive survey of world models. ACM Computing Surveys (2025)

  16. [16]

    Nature Neuroscience (2017)

    Epstein, R.A., Patai, E.Z., Julian, J.B., Spiers, H.J.: The cognitive map in hu- mans: spatial navigation and beyond. Nature Neuroscience (2017)

  17. [17]

    In: CVPR (2025)

    Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., Zhang, M., Chen, P., Li, Y., Lin, S., Zhao, S., Li, K., Xu, T., Zheng, X., Chen, E., Shan, C., He, R., Sun, X.: Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis. In: CVPR (2025)

  18. [18]

    In: ICML (2025)

    Gao, S., Zhou, S., Du, Y., Zhang, J., Gan, C.: AdaWorld: Learning adaptable world models with latent actions. In: ICML (2025)

  19. [19]

    In: NeurIPS (2025)

    Gui, D., Guo, X., Zhou, W., Lu, Y.: Image as a world: Generating interactive world from single image via panoramic video generation. In: NeurIPS (2025)

  20. [20]

    In: NeurIPS (2018)

    Ha, D., Schmidhuber, J.: Recurrent world models facilitate policy evolution. In: NeurIPS (2018)

  21. [21]

    In: IROS (2024)

    Hao, Y., Yang, F., Fang, N., Liu, Y.S.: EMBOSR: Embodied spatial reasoning for enhanced situated question answering in 3D scenes. In: IROS (2024)

  22. [22]

    In: CVPR (2025) What if? Emulative Simulation with World Models for Situated Reasoning 17

    Hassan, M., Stapf, S., Rahimi, A., Rezende, P.M.B., Haghighi, Y., Brüggemann, D., Katircioglu, I., Zhang, L., Chen, X., Saha, S., Cannici, M., Aljalbout, E., Ye, B., Wang, X., Davtyan, A., Salzmann, M., Scaramuzza, D., Pollefeys, M., Favaro, P., Alahi, A.: GEM: A generalizable ego-vision multimodal world model for fine- grained ego-motion, object dynamics...

  23. [23]

    In: NeurIPS (2022)

    Ho, J., Salimans, T., Gritsenko, A.A., Chan, W., Norouzi, M., Fleet, D.J.: Video diffusion models. In: NeurIPS (2022)

  24. [24]

    In: NeurIPS (2023)

    Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., Gan, C.: 3d-llm: Injecting the 3d world into large language models. In: NeurIPS (2023)

  25. [25]

    arXiv preprint arXiv:2511.07222 (2025)

    Hu, J., Zhao, S., Chen, Q.G., Qiu, X., Liu, J., Xu, Z., Luo, W., Zhang, K., Lu, Y.: Omni-View: Unlocking how generation facilitates understanding in unified 3D model based on multiview images. arXiv preprint arXiv:2511.07222 (2025)

  26. [26]

    arXiv preprint arXiv:2510.14965 (2025)

    Hu, M., Huang, Z., Wang, T., Pang, J., Lin, D., Zheng, N., Xu, R.: ChangingGrounding: 3D visual grounding in changing scenes. arXiv preprint arXiv:2510.14965 (2025)

  27. [27]

    In: CVPR (2022)

    Hu, Y., Luo, C., Chen, Z.: Make it move: Controllable image-to-video generation with text descriptions. In: CVPR (2022)

  28. [28]

    arXiv preprint arXiv:2503.04641 (2025)

    Hu,Y.,Wang,L.,Liu,X.,Chen,L.H.,Guo,Y.,Shi,Y.,Liu,C.,Rao,A.,Wang,Z., Xiong, H.: Simulating the real world: A unified survey of multimodal generative models. arXiv preprint arXiv:2503.04641 (2025)

  29. [29]

    IEEE Robotics and Automation Letters (2025)

    Huang, C., Yan, S., Burgard, W.: BYE: Build your encoder with one sequence of exploration data for long-term dynamic scene understanding. IEEE Robotics and Automation Letters (2025)

  30. [30]

    In: ICML (2024)

    Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.C., Jia, B., Huang, S.: An embodied generalist agent in 3D world. In: ICML (2024)

  31. [31]

    arXiv preprint arXiv:2507.07781 (2025)

    Huang, J., Li, Z., Zhang, H., Chen, R., He, X., Guo, Y., Wang, W., Liu, T., Gong, M.: SURPRISE3D: A dataset for spatial understanding and reasoning in complex 3D scenes. arXiv preprint arXiv:2507.07781 (2025)

  32. [32]

    In: ECCV (2024)

    Jia,B.,Chen,Y.,Yu,H.,Wang,Y.,Niu,X.,Liu,T.,Li,Q.,Huang,S.:SceneVerse: Scaling 3D vision-language learning for grounded scene understanding. In: ECCV (2024)

  33. [33]

    arXiv preprint arXiv:2506.03135 (2025)

    Jia, M., Qi, Z., Zhang, S., Zhang, W., Yu, X., He, J., Wang, H., Yi, L.: OmniS- patial: Towards comprehensive spatial reasoning benchmark for vision language models. arXiv preprint arXiv:2506.03135 (2025)

  34. [34]

    In: AAAI (2026)

    Jin, Q., Wu, Y., Chen, C.: PanoNav: Mapless zero-shot object navigation with panoramic scene parsing and dynamic memory. In: AAAI (2026)

  35. [35]

    In: ICCV (2021)

    Koh, J.Y., Lee, H., Yang, Y., Baldridge, J., Anderson, P.: Pathdreamer: A world model for indoor navigation. In: ICCV (2021)

  36. [36]

    IEEE Transactions on Multimedia (2024)

    Köksal, A., Ak, K.E., Sun, Y., Rajan, D., Lim, J.H.: Controllable video generation with text-based instructions. IEEE Transactions on Multimedia (2024)

  37. [37]

    arXiv preprint arXiv:2412.03603 (2024)

    Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Yuan, J., Long, Y., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., H...

  38. [38]

    Labs, P.: Pika: Ideas to video.https://pika.art(2024), accessed: 2024

  39. [39]

    Labs, W.: Marble: A multimodal 3D world model.https://www.worldlabs.ai (2025)

  40. [40]

    OpenReview (2022),https://openreview.net/pdf?id=BZ5a1r-kVsf

    LeCun, Y.: A path towards autonomous machine intelligence. OpenReview (2022),https://openreview.net/pdf?id=BZ5a1r-kVsf

  41. [41]

    In: EMNLP (2025) 18 Ruiping Liuet al

    Li, D., Jiang, B., Huang, L., Beigi, A., Zhao, C., Tan, Z., Bhattacharjee, A., Jiang, Y., Chen, C., Wu, T., Shu, K., Cheng, L., Liu, H.: From generation to judgment: Opportunities and challenges of LLM-as-a-judge. In: EMNLP (2025) 18 Ruiping Liuet al

  42. [42]

    arXiv preprint arXiv:2510.08531 (2025)

    Li, H., Li, D., Wang, Z., Yan, Y., Wu, H., Zhang, W., Shen, Y., Lu, W., Xiao, J., Zhuang, Y.: SpatialLadder: Progressive training for spatial reasoning in vision- language models. arXiv preprint arXiv:2510.08531 (2025)

  43. [43]

    arXiv preprint arXiv:2510.21682 (2025)

    Li, S., Yang, C., Fang, J., Yi, T., Lu, J., Cen, J., Xie, L., Shen, W., Tian, Q.: WorldGrow: Generating infinite 3D world. arXiv preprint arXiv:2510.21682 (2025)

  44. [44]

    In: ICCV (2025)

    Li, Y., Torralba, A.: Multimodal action conditioned video simulation. In: ICCV (2025)

  45. [45]

    arXiv preprint arXiv:2509.22415 (2025)

    Liang, J., Chen, R., Jiao, X., Liang, S., Liu, S., Zhang, Q., Hu, Z., Cao, X.: Explaining multimodal LLMs via intra-modal token interactions. arXiv preprint arXiv:2509.22415 (2025)

  46. [46]

    In: NeurIPS (2024)

    Linghu, X., Huang, J., Niu, X., Ma, X.S., Jia, B., Huang, S.: Multi-modal situated reasoning in 3D scenes. In: NeurIPS (2024)

  47. [47]

    In: SSRR (2005)

    Liu, J., Wang, Y., Ma, S., Li, B.: Analysis of stairs-climbing ability for a tracked reconfigurable modular robot. In: SSRR (2005)

  48. [48]

    In: CVPR (2025)

    Liu, J., Lin, S., Li, Y., Yang, M.H.: DynamicScaler: Seamless and scalable video generation for panoramic scenes. In: CVPR (2025)

  49. [49]

    arXiv preprint arXiv:2509.10884 (2025)

    Liu, Q., Huang, T., Zhang, Z., Tang, H.: Nav-R1: Reasoning and navigation in embodied scenes. arXiv preprint arXiv:2509.10884 (2025)

  50. [50]

    arXiv preprint arXiv:2412.03118 (2024)

    Liu, R., Zhang, J., Schön, A., Müller, K., Zheng, J., Yang, K., Guo, A., Ger- ling, K., Stiefelhagen, R.: ObjectFinder: An open-vocabulary assistive system for interactive object search by blind people. arXiv preprint arXiv:2412.03118 (2024)

  51. [51]

    In: NeurIPS (2025)

    Liu, R., Zheng, J., Chen, Y., Wang, Z., Peng, K., Yang, K., Zhang, J., Pollefeys, M., Stiefelhagen, R.: Situat3DChange: Situated 3D change understanding dataset for multimodal large language model. In: NeurIPS (2025)

  52. [52]

    In: ICLR (2025)

    Lu, T., Shu, T., Yuille, A., Khashabi, D., Chen, J.: Generative world explorer. In: ICLR (2025)

  53. [53]

    In: ICLR (2023)

    Ma, X., Yong, S., Zheng, Z., Li, Q., Liang, Y., Zhu, S.C., Huang, S.: SQA3D: Situated question answering in 3D scenes. In: ICLR (2023)

  54. [54]

    In: CVPR (2024)

    Majumdar, A., Ajay, A., Zhang, X., Putta, P., Yenamandra, S., Henaff, M., Sil- wal, S., Mcvay, P., Maksymets, O., Arnaud, S., Yadav, K., Li, Q., Newman, B., Sharma, M., Berges, V., Zhang, S., Agrawal, P., Bisk, Y., Batra, D., Kalakrish- nan, M., Meier, F., Paxton, C., Sax, A., Rajeswaran, A.: OpenEQA: Embodied question answering in the era of foundation m...

  55. [55]

    In: CVPR (2024)

    Man, Y., Gui, L.Y., Wang, Y.X.: Situational awareness matters in 3D vision language reasoning. In: CVPR (2024)

  56. [56]

    In: Predictions in the Brain: Using Our Past to Generate a Future

    Moulton, S.T., Kosslyn, S.M.: Imagining predictions: Mental imagery as mental emulation. In: Predictions in the Brain: Using Our Past to Generate a Future. Oxford University Press (2011)

  57. [57]

    Mullins, C.D.: Cognitive Behavioral Group Therapy for Blind and Visually Im- paired Adults: Acceptance, Problem-Solving, and Cognitive Distortions. Ph.D. thesis, Philadelphia College of Osteopathic Medicine (2019)

  58. [58]

    BMC Health Services Research (2021)

    van Munster, E.P.J., van der Aa, H.P.A., Verstraten, P., van Rens, G.H.M.B., van Nispen, R.M.A.: Barriers and facilitators to recognize and discuss depression and anxiety experienced by adults with vision impairment or blindness: a qualitative study. BMC Health Services Research (2021)

  59. [59]

    In: IROS (2025)

    Nie, D., Guo, X., Duan, Y., Zhang, R., Chen, L.: WMNav: Integrating vision- language models into world models for object goal navigation. In: IROS (2025)

  60. [60]

    OpenAI: Sora: Creating video from text. Tech. rep., OpenAI (2024),https:// openai.com/sora What if? Emulative Simulation with World Models for Situated Reasoning 19

  61. [61]

    OpenAI: Function calling (Aug 2025),https://developers.openai.com/api/ docs/guides/function-calling/, openAI API documentation

  62. [62]

    OpenAI: Gpt-5 technical overview.https://openai.com/index/gpt-5(2025)

  63. [63]

    Parker-Holder, J., Fruchter, S.: Genie 3: A new frontier for world models.https: //deepmind.google/blog/genie-3-a-new-frontier-for-world-models(2025)

  64. [64]

    arXiv preprint arXiv:2410.18072 (2024)

    Qin, Y., Shi, Z., Yu, J., Wang, X., Zhou, E., Li, L., Yin, Z., Liu, X., Sheng, L., Shao, J., Bai, L., Ouyang, W., Zhang, R.: WorldSimBench: Towards video generation models as world simulators. arXiv preprint arXiv:2410.18072 (2024)

  65. [65]

    In: NeurIPS (2021)

    Ramakrishnan, S.K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A.X., Savva, M., Zhao, Y., Batra, D.: Habitat-matterport 3D dataset (HM3D): 1000 large-scale 3D environments for embodied AI. In: NeurIPS (2021)

  66. [66]

    In: COLM (2025)

    Ray, A., Duan, J., Brown, E., Tan, R., Bashkirova, D., Hendrix, R., Ehsani, K., Kembhavi, A., Plummer, B.A., Krishna, R., Zeng, K.H., Saenko, K.: SAT: Dynamic spatial aptitude training for multimodal language models. In: COLM (2025)

  67. [67]

    In: ICLR (2024)

    Richens, J., Everitt, T.: Robust agents learn causal world models. In: ICLR (2024)

  68. [68]

    com / blog / overcoming - challenges - robotics - warehouse (2016)

    inViaRobotics,I.:Overcomingthechallengesofroboticsinthewarehouse.https: / / inviarobotics . com / blog / overcoming - challenges - robotics - warehouse (2016)

  69. [69]

    In: ICCV (2019)

    Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D.: Habitat: A Platform for Embodied AI Research. In: ICCV (2019)

  70. [70]

    IEEE Access (2023)

    Seo, T., Ryu, S., Won, J.H., Kim, Y., Kim, H.S.: Stair-climbing robots: A review on mechanism, sensing, and performance evaluation. IEEE Access (2023)

  71. [71]

    In: WACV (2025)

    Song, I., Lee, S., Joo, M., Lee, J.: Anomaly detection for people with visual impairments using an egocentric 360-degree camera. In: WACV (2025)

  72. [72]

    In: WACV (2025)

    Sugandhika, C., Li, C., Rajan, D., Fernando, B.: Situational scene graph for struc- tured human-centric situation understanding. In: WACV (2025)

  73. [73]

    arXiv preprint arXiv:2507.06119 (2025)

    Tan, Z., Yang, H., Qin, L., Gong, J., Yang, M., Li, H.: Omni-Video: Democratiz- ing unified video understanding and generation. arXiv preprint arXiv:2507.06119 (2025)

  74. [74]

    arXiv preprint arXiv:2312.11805 (2023)

    Team, G., Anil, R., Borgeaud, S., Alayrac, J.B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., Millican, K., et al.: Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)

  75. [75]

    arXiv preprint arXiv:2507.21809 (2025)

    Team, H., Wang, Z., Liu, Y., Wu, J., Gu, Z., Wang, H., Zuo, X., Huang, T., Li, W., Zhang, S., Lian, Y., Tsai, Y., Wang, L., Liu, S., Jiang, P., Yang, X., Guo, D., Tang, Y., Mao, X., Yu, J., Yu, J., Zhang, J., Chen, M., Dong, L., Jia, Y., Zhang, C., Tan, Y., Zhang, H., Ye, Z., He, P., Wu, R., Chen, M., Li, Z., Qin, W., Wang, L., Sun, Y., Niu, L., Yuan, X.,...

  76. [76]

    arXiv preprint arXiv:2510.22443 (2025)

    Veerabadran, V., Xiao, F., Kamra, N., Matias, P., Chen, J., Drooff, C., Roads, B.D., Williams, R., Henderson, E., Zhao, X., Carlberg, K., Tighe, J., Ridgeway, K.: Benchmarking egocentric multimodal goal inference for assistive wearable agents. arXiv preprint arXiv:2510.22443 (2025)

  77. [77]

    In: CVPR (2024) 20 Ruiping Liuet al

    Wang, A., Wu, B., Chen, S., Chen, Z., Guan, H., Lee, W.N., Li, L.E., Gan, C.: SOK-Bench: A situated video reasoning benchmark with aligned open-world knowledge. In: CVPR (2024) 20 Ruiping Liuet al

  78. [78]

    arXiv preprint arXiv:2503.20314 (2025)

    Wang, A., Ai, B., Wen, B., Mao, C., Xie, C., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Meng, X., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W., Wang, W., ...

  79. [79]

    In: ICCV (2023)

    Wang, H., Liang, W., Van Gool, L., Wang, W.: Dreamwalker: Mental planning for continuous vision-language navigation. In: ICCV (2023)

  80. [80]

    arXiv preprint arXiv:2509.09676 (2025)

    Wang, J., Yuan, Y., Zheng, R., Lin, Y., Gao, J., Chen, L., Bao, Y., Zhang, Y., Zeng, C., Zhou, Y., Long, X., Zhu, H., Zhang, Z., Cao, X., Yao, Y.: Spa- tialVID: A large-scale video dataset with spatial annotations. arXiv preprint arXiv:2509.09676 (2025)

Showing first 80 references.