Pith. sign in

REVIEW 3 major objections 6 minor 58 references

Off-the-shelf vision-language models cannot yet act through a body: the best full success rate on a long-horizon humanoid bench is 16.8%, and the bottleneck is embodied self-awareness, not seeing the target.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 13:50 UTC pith:FTGAKXJR

load-bearing objection Clean middle-layer benchmark: frozen VLMs fail after recognition on body-relative tracking; the diagnosis is useful but interface-bound. the 3 major comments →

arxiv 2607.27180 v1 pith:FTGAKXJR submitted 2026-07-29 cs.CV cs.RO

HumanCLAW: Can Vision-Language Models Act Through a Body?

classification cs.CV cs.RO
keywords vision-language modelsembodied AIaction intelligencehumanoid motionegocentric navigationembodied self-awarenessbenchmarkhalf-physics simulation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether a frozen vision-language model can decide, moment by moment, what a physical human body should do next when each choice becomes real motion with gravity and collisions. It introduces HumanCLAW, a closed-loop setup that lets the model issue atomic whole-body skills while a motion generator and half-physics simulator execute them, so failures read as bad decisions rather than lost balance. On HumanCLAW-Bench—1,218 egocentric find-navigate-sit episodes in 41 indoor houses—nine frontier models all fail; the strongest reaches only 16.8% full success. Once a target is truly in view, recognition is largely intact. What collapses is tracking the body itself: where it is, whether it has arrived, and whether it has hit something. A sympathetic reader cares because this isolates action intelligence as a distinct gap between describing a scene and inhabiting one.

Core claim

No current off-the-shelf VLM solves closed-loop whole-body find-navigate-interact under physical consequence: the best model completes the full progression on only 16.8% of episodes, and after target recognition the dominant failures are egocentric body and self-localization errors—stopping while far, not noticing arrival or jamming, and sitting into empty space—rather than perception or motor execution. The missing capacity is embodied self-awareness.

What carries the argument

HumanCLAW: a harnessed VLM issues one atomic parameterized skill each half-second; a skill-conditioned motion generator turns it into continuous full-body motion; a half-physics simulator applies contact, gravity, and collisions while factoring out balance and motor-tracking failures, so each failure attributes at the decision level.

Load-bearing premise

The setup assumes that with only egocentric video and text history—and no felt body pose or contact signal—failed episodes still mainly show a missing internal sense of the body, not an interface that never tells the model what its limbs are doing.

What would settle it

Add proprioceptive body state and contact feedback to the same frozen VLMs on the same episodes: if full find-navigate-interact success then jumps from the mid-teens toward reliable completion while collision and false-arrival errors collapse, the diagnosis of missing embodied self-awareness as an internal faculty is weakened; if the gap remains, it is strengthened.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Benchmarking embodied VLMs must score closed-loop body decisions under physical outcome, not only open-loop plans or camera navigation.
  • Gains in image recognition alone will not close long-horizon humanoid task success if self-localization and termination stay broken.
  • Atomic skill interfaces plus reusable motion priors let new skills and tasks be added without retraining the decision maker.
  • Progress should target persistent body-state estimates, calibrated stop/arrive decisions, and anticipation of where the body will land after each skill.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If body awareness is the bottleneck, training or architectural work that forces models to predict their own limb occupancy and contact from egocentric history may transfer across embodiments more than collecting more action trajectories.
  • The same self-localization failures would likely appear in any first-person agent that must stop next to and then precisely place a body relative to furniture, not only humanoids.
  • A fair next experiment is a controlled ablation that restores proprioception and touch while holding the skill vocabulary fixed, to separate missing input from missing faculty.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces HumanCLAW, a closed-loop evaluation framework that lets frozen off-the-shelf VLMs issue atomic whole-body skill commands (walk, turn, climb, sit, etc.), which a skill-conditioned motion DiT realizes as 0.5s full-body chunks executed in a half-physics Habitat/Bullet simulator that preserves contact, gravity, and collisions while factoring out balance and motor-tracking failures. On HumanCLAW-Bench (1,218 egocentric find–navigate–interact episodes over 41 HSSD houses), nine frontier VLMs are evaluated with staged objective and acknowledged success metrics, action-quality and disturbance measures, harness ablations, skill-fidelity checks, and automated root-cause attribution. None solves the benchmark; the best full InteractSR is 16.8% (Gemini-3.1). The authors argue that target recognition is largely intact once the object is rendered, and that post-recognition failures concentrate in egocentric self-localization, arrival/termination, and body placement—termed missing embodied self-awareness.

Significance. If the empirical picture holds, the work supplies a useful middle layer between symbolic agent benchmarks and end-to-end VLAs: continuous full-body motion with physical scene consequences, decision-level failure attribution, and a frozen generalist decision maker. Strengths include multi-model staged metrics with objective vs. acknowledged splits (Table 4), skill functional fidelity vs. MoMask (Table 1), harness ablations (Table 3), collision-by-body-part breakdowns (Table 5), and fully automated, reproducible root-cause rules (Appendix B, Table 7). The plug-and-play skill ControlNet design and half-physics decoupling are concrete systems contributions. The headroom (best full success 16.8%, NavSR peaking at 42.4%) and the focus on closed-loop action intelligence rather than open-loop spatial QA make the benchmark a credible stress test for the next generation of embodied foundation models.

major comments (3)
  1. [Abstract; Findings 4–6; §2.4; §6] Abstract and Findings 4–6 attribute post-recognition failures primarily to a missing internal faculty of “embodied self-awareness,” yet §2.4, Finding 6, and §6 state that the agent receives only egocentric RGB and text history, with no proprioceptive pose or contact/tactile channel, and that collisions “are never felt.” This interface choice is load-bearing for the diagnosis: many reported errors (unaware jammed, stop while far, sit on air, leg/arm collisions in Table 5 and Fig. 9) are exactly what an information-starved controller would produce. The empirical claim that none of nine VLMs solves the bench remains intact, but the abstract’s punchline should be scoped to “under egocentric vision-only feedback” (or supported by a body-state/contact ablation). As written, the faculty vs. interface distinction is under-determined.
  2. [§3.2; Table 2; Table 4] Headline success requires both geometric criteria and subjective acknowledgment/active stop (FindSR/NavSR/InteractSR in Table 2 vs. Geo* in Table 4). That design choice is defensible for “action intelligence,” but it conflates spatial competence with the model’s willingness to emit a stop/claim in the harnessed format. Several large GeoNavSR–NavSR and GeoInteractSR–InteractSR gaps (e.g., GPT-5.5 27.7% vs. 13.9%; InternVL GeoNavSR 12.6% vs. NavSR 0.8%) show that a non-trivial fraction of “failures” are termination/reporting failures under the scaffold. The paper should report Geo* as co-primary in the main table and discuss how much of the embodied-self-awareness narrative is driven by stop/claim behavior rather than body–scene geometry alone.
  3. [Abstract; Finding 1; Table 3; §2.2] Table 3 shows the skill-specific verifier is decisive: removing it collapses NavSR 27%→2% and InteractSR 18.9%→0% on the mini-val. The benchmark therefore measures harnessed VLMs (high-to-mid-to-low scaffold + per-skill verifier; Fig. 3, Table 6), not raw off-the-shelf models. This is disclosed, but the Abstract and Finding 1 (“none solves”; “what current VLMs lack”) should state clearly that results are under this scaffold, and that without the verifier the loop essentially does not close. Otherwise readers will over-read the numbers as pure foundation-model competence.
minor comments (6)
  1. [§3.1; Abstract] Interaction is instantiated only as sit on a subset of categories (§3.1; 597 sit episodes). Richer manipulation is acknowledged as future work in §6; a sentence in the abstract limiting “interact” to sit would avoid over-generalization.
  2. [§3.1; Fig. 7] Difficulty tiers use fixed thresholds (distance ≤3.5/≤8 m, etc.; Fig. 7). Brief sensitivity of success rates to these cutoffs would strengthen the stratified claims.
  3. [§3.2; Table 2] Motion jerk is a useful coherence metric, but the denoising/timescale details (≈0.27 s / stride 8) are only lightly specified; a short formula or appendix note would aid reproduction.
  4. [Table 1; §4.1] Climb upstairs/downstairs achievement ratios in Table 1 are 0.74–0.79 (stable gain). Confirm in text that the VLM is not expected to compensate for this gain, or that parameters are calibrated so commanded ≈ achieved.
  5. [§5.1] Related work on object-goal nav and VLN is appropriate; a slightly sharper contrast with Habitat 3.0 humanoid setups (Puig et al., 2024) on what half-physics adds beyond avatar animation would help placement.
  6. [Tables 2–4; title page] Minor polish: “HumanClawBench” vs. “HumanCLAW-Bench” casing is inconsistent in table captions; arXiv date line says July 30, 2026.

Circularity Check

0 steps flagged

Empirical benchmark paper with geometric success criteria and rule-based failure labels; no derivation that reduces predictions to fitted inputs or self-definitional premises.

full rationale

HumanCLAW is a systems/evaluation paper, not a first-principles derivation. The load-bearing claims—none of nine off-the-shelf VLMs solves the bench; best InteractSR 16.8%; post-recognition failures concentrate in self-localization and body placement—are measured against external geometric ground truth (semantic pixel count, AABB distance, pelvis–mesh contact) plus logged action streams, with root causes assigned by fixed deterministic rules over rollouts (Appendix B, Table 7). The motion layer’s skill fidelity (Table 1) is a separate zero-shot check against commanded magnitudes, not a fit reused as a VLM ‘prediction.’ Self-citations (e.g. half-physics from Siyao et al. 2025) supply infrastructure and are not invoked as uniqueness theorems that force the central diagnosis. Framing failures as ‘embodied self-awareness’ is interpretive labeling of measured patterns, not a circular reduction of Eq. X to Eq. Y. No fitted-input-called-prediction, self-definitional loop, or load-bearing self-citation chain is present. Score 0 is appropriate.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 4 invented entities

Load-bearing commitments are methodological: that atomic skills + half-physics isolate decision-level action intelligence; that staged geometric criteria define success; that automated log rules correctly name root causes; and that frozen internet VLMs are the right probe of generalist action intelligence. No physical constants are fitted. Free choices are thresholds, skill set, and harness design.

free parameters (7)
  • Decision chunk duration / horizon = 0.5 s (15 frames @ 30 Hz)
    0.5 s chunks at 30 fps (15 future frames from 5 history) set the closed-loop timescale and what counts as a single decision.
  • FindSR pixel threshold = 100 px
    Target must occupy ≥100 semantic pixels in 512×512 ego view to count as found.
  • NavSR distance threshold = 20 cm (primary)
    Success when min distance to target AABB < 20 cm (with 1 m variant reported).
  • Difficulty tier cutoffs = dist ≤3.5/≤8 m; choice ≤2/≤5; obstacle ≤2.5/≤5
    Hand thresholds on geodesic distance, choice (turns+rooms), and obstacle density define easy/medium/hard splits.
  • Harness history defaults = hist 10, img 1
    Baseline text history length and single current image are design choices shown to matter in ablation.
  • Half-physics passive stiffness λ = λ = 1.0
    Joint stiffness in passive-body simulation affects contact compliance and collision outcomes.
  • Skill ControlNet training budgets = 5e5–1.5e6 steps; lr 3e-4; bs 2048
    Per-skill optimization steps and filtering of AMASS clips shape zero-shot skill reliability, which underpins decision-level attribution.
axioms (6)
  • domain assumption Half-physics preserves task-relevant physical consequences (collision, gravity, object displacement) while removing balance/motor-tracking failures as confounds for scoring action decisions.
    Core of §2.4 and Fig. 6; without it, failures cannot be cleanly read as decision errors.
  • ad hoc to paper A finite set of atomic parameterized skills (walk, side_step, step_back, turn, climb, sit, stop) is an adequate action vocabulary for measuring long-horizon action intelligence on find-navigate-sit.
    §2.3; paper admits finer/coarser vocabularies would redraw the decision/execution boundary (§6).
  • ad hoc to paper Generalizable action intelligence should be probed with frozen off-the-shelf VLMs (reasoning) rather than policies fitted on embodiment trajectories.
    Stated philosophical commitment in §1 and §6; not empirically compared to fitted baselines.
  • ad hoc to paper Staged success requires both geometric criteria and subjective acknowledgment/stop in the model’s outputs for headline metrics.
    §3.2; Geo* variants partially relax this, but main Table 2 uses the joint criterion.
  • domain assumption Automated top-to-bottom rules over navmesh distance, mesh contact, pixels, actions, and stated visible state assign a single true root cause per failed episode.
    Appendix B / Table 7; underpins the failure funnel percentages in Fig. 8.
  • domain assumption Standard simulators and assets (AI Habitat, Bullet, HSSD, AMASS) are faithful enough substrates for the claimed physical consequences and motion prior.
    §2.3–3.1 engineering base shared with prior embodied work.
invented entities (4)
  • Action intelligence (as operational closed-loop component of spatial intelligence) independent evidence
    purpose: Name the capacity being measured: moment-to-moment selection, parameterization, and sequencing of bodily skills under physical feedback.
    Defined in §1; useful framing, not a physical entity. Independent handle is performance on HumanCLAW-Bench itself.
  • Embodied self-awareness (online estimate of own body state, scene relation, and recent action consequences) no independent evidence
    purpose: Unify the dominant failure modes after perception (stop-while-far, unaware arrival/jam, sit-on-air, unnoticed collisions).
    Introduced as the missing capacity in Abstract/§4/§6. Partly operationalized via metrics, but risks conflating model deficit with missing proprioceptive inputs.
  • HumanCLAW harness (high-to-mid-to-low scaffold + skill-specific verifier) independent evidence
    purpose: Make open-ended VLMs emit reliable atomic skill calls without action finetuning.
    §2.2 system component; ablations show verifier/mid-level matter. Engineering construct, not a natural kind.
  • Plug-and-play skill ControlNet bag on a shared motion DiT prior independent evidence
    purpose: Zero-shot-reliable continuous realization of VLM skill strings with extensible skills.
    §2.3 method; fidelity table supports functional reliability vs generic text-to-motion.

pith-pipeline@v1.2.0-daily-grok45 · 39953 in / 4626 out tokens · 93745 ms · 2026-07-30T13:50:48.668101+00:00 · methodology

0 comments
read the original abstract

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 12 linked inside Pith

  1. [1]

    Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S \"u nderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  2. [2]

    _0 : A vision-language-action flow model for general robot control, 2024

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. _0 : A vis...

  3. [3]

    Turner, Eric Undersander, and Tsung-Yen Yang

    Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M. Turner, Eric Undersander, and Tsung-Yen Yang. PARTNR : A benchmark for planning...

  4. [4]

    SpatialVLM : Endowing vision-language models with spatial reasoning capabilities

    Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. SpatialVLM : Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14455--14465, 2024

  5. [5]

    LoTa-Bench : Benchmarking language-oriented task planners for embodied agents

    Jae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim, and Minsu Jang. LoTa-Bench : Benchmarking language-oriented task planners for embodied agents. In The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=ADSxCpCu9s

  6. [6]

    Umo: Unified in-context learning unlocks motion foundation model priors

    Xiaoyan Cong, Zekun Li, Zhiyang Dou, Hongyu Li, Omid Taheri, Chuan Guo, Abhay Mittal, Sizhe An, Taku Komura, Wojciech Matusik, et al. Umo: Unified in-context learning unlocks motion foundation model priors. arXiv preprint arXiv:2603.15975, 2026

  7. [7]

    Bullet physics simulation

    Erwin Coumans. Bullet physics simulation. In ACM SIGGRAPH 2015 Courses, page 7:1. ACM, 2015. doi:10.1145/2776880.2792704

  8. [8]

    Moving by looking: Towards vision-driven avatar motion generation

    Markos Diomataris, Berat Mert Albaba, Giorgio Becherini, Partha Ghosh, Omid Taheri, and Michael J Black. Moving by looking: Towards vision-driven avatar motion generation. arXiv preprint arXiv:2509.19259, 2025

  9. [9]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. PaLM-E : An embodied...

  10. [10]

    Manipulate-anything: Automating real-world robots using vision-language models

    Jiafei Duan, Wentao Yuan, Wilbert Pumacay, Yi Ru Wang, Kiana Ehsani, Dieter Fox, and Ranjay Krishna. Manipulate-anything: Automating real-world robots using vision-language models. In Conference on Robot Learning, 2024

  11. [11]

    Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine

    Dibya Ghosh, Homer Rich Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Quan Vuong, Ted Xiao, Pannag R. Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems...

  12. [12]

    Generating diverse and natural 3d human motions from text

    Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5152--5161, 2022. doi:10.1109/CVPR52688.2022.00509

  13. [13]

    MoMask : Generative masked modeling of 3d human motions

    Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. MoMask : Generative masked modeling of 3d human motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1900--1910, 2024. https://openaccess.thecvf.com/content/CVPR2024/html/Guo_MoMask_Generative_Masked_Modeling_of_3D_Human_Motions_CVPR_2024_paper.html

  14. [14]

    Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J. Black. Stochastic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11354--11364, 2021. doi:10.1109/ICCV48922.2021.01118

  15. [15]

    ESI-Bench : Towards embodied spatial intelligence that closes the perception-action loop

    Yining Hong, Jiageng Liu, Han Yin, Manling Li, Leonidas Guibas, Li Fei-Fei, Jiajun Wu, and Yejin Choi. ESI-Bench : Towards embodied spatial intelligence that closes the perception-action loop. arXiv preprint arXiv:2605.18746, 2026

  16. [16]

    VoxPoser : Composable 3d value maps for robotic manipulation with language models

    Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. VoxPoser : Composable 3d value maps for robotic manipulation with language models. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 540--562. PMLR, 2023 a . https://proceedings.mlr.press/v229/huang23b.html

  17. [17]

    Inner monologue: Embodied reasoning through planning with language models

    Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Tomas Jackson, Noah Brown, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. Inner monologue: Embodied reasoning through planning with language models. In Proceedings of the 6th Conference on Rob...

  18. [18]

    Como: Controllable motion generation through language guided pose code editing

    Yiming Huang, Weilin Wan, Yue Yang, Chris Callison-Burch, Mark Yatskar, and Lingjie Liu. Como: Controllable motion generation through language guided pose code editing. In European Conference on Computer Vision, pages 180--196. Springer, 2024

  19. [19]

    Do as i can, not as i say: Grounding language in robotic affordances

    Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Proceedings of the 6th Conference on Robot Learning, volume 205 of Proceedings of Machine Learning Research, pages 287--318. PMLR, 2023...

  20. [20]

    Iam: Identity-aware human motion and shape joint generation

    Wenqi Jia, Zekun Li, Abhay Mittal, Chengcheng Tang, Chuan Guo, Lezi Wang, James Matthew Rehg, Lingling Tao, and Size An. Iam: Identity-aware human motion and shape joint generation. arXiv preprint arXiv:2604.25164, 2026

  21. [21]

    Chang, and Manolis Savva

    Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X. Chang, and Manolis Savva. Habitat synthetic scenes dataset ( HSSD -200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...

  22. [22]

    Foster, Pannag R

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P. Foster, Pannag R. Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA : An open-source vision-language-action model. In Proceedings of the 8th Conference ...

  23. [23]

    MolmoAct : Action reasoning models that can reason in space

    Jason Lee, Jiafei Duan, Haoquan Fang, Yuquan Deng, Shuo Liu, Boyang Han, Mohammadreza Salehi, Jae Sung Hwang, et al. MolmoAct : Action reasoning models that can reason in space. arXiv preprint arXiv:2508.07917, 2025

  24. [24]

    BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation

    Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart \'i n-Mart \'i n, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation. In Proceedings of the 6th Conference on Robot Learning, volume 205 of Proceedings of Mac...

  25. [25]

    Unfolding spatial cognition: Evaluating multimodal models on visual simulations

    Linjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma, Yinuo Yang, Ziang Li, Yejin Choi, and Ranjay Krishna. Unfolding spatial cognition: Evaluating multimodal models on visual simulations. arXiv preprint arXiv:2506.04633, 2025

  26. [26]

    Embodied agent interface: Benchmarking LLMs for embodied decision making

    Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, Weiyu Liu, Percy Liang, Li Fei-Fei, Jiayuan Mao, and Jiajun Wu. Embodied agent interface: Benchmarking LLMs for embodied decision making. In Advances in Neural Information Processing Systems, volume 37, 2024. doi:10.52202/079017...

  27. [27]

    Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens

    Zekun Li, Sizhe An, Chengcheng Tang, Chuan Guo, Ivan Shugurov, Linguang Zhang, Amy Zhao, Srinath Sridhar, Lingling Tao, and Abhay Mittal. Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...

  28. [28]

    Genhsi: Controllable generation of human-scene interaction videos

    Zekun Li, Rui Zhou, Rahul Sajnani, Xiaoyan Cong, Daniel Ritchie, and Srinath Sridhar. Genhsi: Controllable generation of human-scene interaction videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 138--149, 2026 b

  29. [29]

    Code as policies: Language model programs for embodied control

    Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In IEEE International Conference on Robotics and Automation, 2023

  30. [30]

    Large model empowered embodied AI : A survey on decision-making and embodied learning, 2025

    Wenlong Liang, Rui Zhou, Yang Ma, Bing Zhang, Songlin Li, Yijia Liao, and Ping Kuang. Large model empowered embodied AI : A survey on decision-making and embodied learning, 2025. https://arxiv.org/abs/2508.10399

  31. [31]

    VisualAgentBench : Towards large multimodal models as visual foundation agents

    Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, et al. VisualAgentBench : Towards large multimodal models as visual foundation agents. arXiv preprint arXiv:2408.06327, 2024

  32. [32]

    A survey on vision-language-action models for embodied AI

    Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied AI . IEEE Transactions on Neural Networks and Learning Systems, 2026. doi:10.1109/TNNLS.2025.3650584

  33. [33]

    GR00T N1 : An open foundation model for generalist humanoid robots, 2025

    NVIDIA , Johan Bjorck, Fernando Casta \ n eda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, et al. GR00T N1 : An open foundation model for generalist humanoid robots, 2025. https://arxiv.org/abs/2503.14734

  34. [34]

    VirtualHome : Simulating household activities via programs

    Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. VirtualHome : Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8494--8502, 2018. https://openaccess.thecvf.com/content_cvpr_2018/html/Puig_VirtualHome_Simulating_Household_CVPR...

  35. [35]

    Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi

    Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, Vladimir Vondrus, Theophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara ...

  36. [36]

    Habitat: A platform for embodied AI research

    Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A platform for embodied AI research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9339--9347, 2019. doi:10.1109/ICCV.2019.00943

  37. [37]

    ALFRED : A benchmark for interpreting grounded instructions for everyday tasks

    Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. ALFRED : A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10740--10749, 2020. https://openaccess.thecvf.com/content_CV...

  38. [38]

    Bailando: 3 D dance generation by actor-critic GPT with choreographic memory

    Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando: 3 D dance generation by actor-critic GPT with choreographic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  39. [39]

    Bailando++: 3 D dance GPT with choreographic memory

    Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando++: 3 D dance GPT with choreographic memory. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 45 0 (12): 0 14192--14207, 2023

  40. [40]

    Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment

    Li Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin, Ziwei Liu, Henghui Ding, Lei Yang, and Chen Change Loy. Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment. In International Conference on Learning Representations (ICLR), 2024

  41. [41]

    Half-physics: Enabling kinematic 3d human model with physical interactions

    Li Siyao, Yao Feng, Omid Taheri, Chen Change Loy, and Michael J Black. Half-physics: Enabling kinematic 3d human model with physical interactions. arXiv preprint arXiv:2507.23778, 2025

  42. [42]

    Sadler, Wei-Lun Chao, and Yu Su

    Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler, Wei-Lun Chao, and Yu Su. LLM-Planner : Few-shot grounded planning for embodied agents with large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2998--3009, 2023. https://openaccess.thecvf.com/content/ICCV2023/html/Song_LLM-Planner_Few-Shot_Gr...

  43. [43]

    Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra

    Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel X. Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Trainin...

  44. [44]

    Cradle: Empowering foundation agents towards general computer control

    Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li, Bohan Zhou, Junpeng Yue, Jiechuan Jiang, Yewen Li, et al. Cradle: Empowering foundation agents towards general computer control. arXiv preprint arXiv:2403.03186, 2024

  45. [45]

    Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. https://openreview.net/forum?id=SJ1kSyO2jwu

  46. [46]

    Voyager: An open-ended embodied agent with large language models

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. https://openreview.net/forum?id=ehfRiF0R3a

  47. [47]

    Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents

    Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, and Yitao Liang. Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents. In Advances in Neural Information Processing Systems, volume 36, 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/6b8dfb8c0c12e6fafc6c256cb08a5ca7-Abstract-...

  48. [48]

    Text2interact: High-fidelity and diverse text-to-two-person interaction generation

    Qingxuan Wu, Zhiyang Dou, Chuan Guo, Yiming Huang, Qiao Feng, Bing Zhou, Jian Wang, and Lingjie Liu. Text2interact: High-fidelity and diverse text-to-two-person interaction generation. arXiv preprint arXiv:2510.06504, 2025

  49. [49]

    OmniControl : Control any joint at any time for human motion generation

    Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang. OmniControl : Control any joint at any time for human motion generation. In The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=gd0lAEtWso

  50. [50]

    Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

    Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces. arXiv preprint arXiv:2412.14171, 2024

  51. [51]

    EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents

    Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, Heng Ji, Huan Zhang, and Tong Zhang. EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. In Proceedings of the 42nd International Conference on Machine Lear...

  52. [52]

    PhysDiff : Physics-guided human motion diffusion model

    Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. PhysDiff : Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16010--16021, 2023. doi:10.1109/ICCV51070.2023.01467

  53. [53]

    Zhang, Thomas L

    Alex L. Zhang, Thomas L. Griffiths, Karthik R. Narasimhan, and Ofir Press. VideoGameBench : Can vision-language models complete popular video games? arXiv preprint arXiv:2505.18134, 2025 a

  54. [54]

    Egoreact: Egocentric video-driven 3d human reaction generation

    Libo Zhang, Zekun Li, Tianyu Li, Zeyu Cao, Rui Xu, Xiaoxiao Long, Wenjia Wang, Jingbo Wang, Yuan Liu, Wenping Wang, et al. Egoreact: Egocentric video-driven 3d human reaction generation. arXiv preprint arXiv:2512.22808, 2025 b

  55. [55]

    The wanderings of odysseus in 3d scenes

    Yan Zhang and Siyu Tang. The wanderings of odysseus in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20481--20491, 2022. https://openaccess.thecvf.com/content/CVPR2022/html/Zhang_The_Wanderings_of_Odysseus_in_3D_Scenes_CVPR_2022_paper.html

  56. [56]

    Yan Zhang, Yao Feng, Alp \'a r Cseke, Nitin Saini, Nathan Bajandas, Nicolas Heron, and Michael J. Black. PRIMAL : Physically reactive and interactive motor model for avatar learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12725--12736, 2025 c . https://openaccess.thecvf.com/content/ICCV2025/html/Zhang_PRIMAL_Phys...

  57. [57]

    Synthesizing diverse human motions in 3d indoor scenes

    Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, and Siyu Tang. Synthesizing diverse human motions in 3d indoor scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14738--14749, 2023. https://openaccess.thecvf.com/content/ICCV2023/html/Zhao_Synthesizing_Diverse_Human_Motions_in_3D_Indoor_Scenes_ICCV_2023_paper.html

  58. [58]

    RT-2 : Vision-language-action models transfer web knowledge to robotic control

    Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. RT-2 : Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 2165--2183. PMLR, 2023. http...