REVIEW 3 major objections 6 minor 58 references
Off-the-shelf vision-language models cannot yet act through a body: the best full success rate on a long-horizon humanoid bench is 16.8%, and the bottleneck is embodied self-awareness, not seeing the target.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 13:50 UTC pith:FTGAKXJR
load-bearing objection Clean middle-layer benchmark: frozen VLMs fail after recognition on body-relative tracking; the diagnosis is useful but interface-bound. the 3 major comments →
HumanCLAW: Can Vision-Language Models Act Through a Body?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
No current off-the-shelf VLM solves closed-loop whole-body find-navigate-interact under physical consequence: the best model completes the full progression on only 16.8% of episodes, and after target recognition the dominant failures are egocentric body and self-localization errors—stopping while far, not noticing arrival or jamming, and sitting into empty space—rather than perception or motor execution. The missing capacity is embodied self-awareness.
What carries the argument
HumanCLAW: a harnessed VLM issues one atomic parameterized skill each half-second; a skill-conditioned motion generator turns it into continuous full-body motion; a half-physics simulator applies contact, gravity, and collisions while factoring out balance and motor-tracking failures, so each failure attributes at the decision level.
Load-bearing premise
The setup assumes that with only egocentric video and text history—and no felt body pose or contact signal—failed episodes still mainly show a missing internal sense of the body, not an interface that never tells the model what its limbs are doing.
What would settle it
Add proprioceptive body state and contact feedback to the same frozen VLMs on the same episodes: if full find-navigate-interact success then jumps from the mid-teens toward reliable completion while collision and false-arrival errors collapse, the diagnosis of missing embodied self-awareness as an internal faculty is weakened; if the gap remains, it is strengthened.
If this is right
- Benchmarking embodied VLMs must score closed-loop body decisions under physical outcome, not only open-loop plans or camera navigation.
- Gains in image recognition alone will not close long-horizon humanoid task success if self-localization and termination stay broken.
- Atomic skill interfaces plus reusable motion priors let new skills and tasks be added without retraining the decision maker.
- Progress should target persistent body-state estimates, calibrated stop/arrive decisions, and anticipation of where the body will land after each skill.
Where Pith is reading between the lines
- If body awareness is the bottleneck, training or architectural work that forces models to predict their own limb occupancy and contact from egocentric history may transfer across embodiments more than collecting more action trajectories.
- The same self-localization failures would likely appear in any first-person agent that must stop next to and then precisely place a body relative to furniture, not only humanoids.
- A fair next experiment is a controlled ablation that restores proprioception and touch while holding the skill vocabulary fixed, to separate missing input from missing faculty.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HumanCLAW, a closed-loop evaluation framework that lets frozen off-the-shelf VLMs issue atomic whole-body skill commands (walk, turn, climb, sit, etc.), which a skill-conditioned motion DiT realizes as 0.5s full-body chunks executed in a half-physics Habitat/Bullet simulator that preserves contact, gravity, and collisions while factoring out balance and motor-tracking failures. On HumanCLAW-Bench (1,218 egocentric find–navigate–interact episodes over 41 HSSD houses), nine frontier VLMs are evaluated with staged objective and acknowledged success metrics, action-quality and disturbance measures, harness ablations, skill-fidelity checks, and automated root-cause attribution. None solves the benchmark; the best full InteractSR is 16.8% (Gemini-3.1). The authors argue that target recognition is largely intact once the object is rendered, and that post-recognition failures concentrate in egocentric self-localization, arrival/termination, and body placement—termed missing embodied self-awareness.
Significance. If the empirical picture holds, the work supplies a useful middle layer between symbolic agent benchmarks and end-to-end VLAs: continuous full-body motion with physical scene consequences, decision-level failure attribution, and a frozen generalist decision maker. Strengths include multi-model staged metrics with objective vs. acknowledged splits (Table 4), skill functional fidelity vs. MoMask (Table 1), harness ablations (Table 3), collision-by-body-part breakdowns (Table 5), and fully automated, reproducible root-cause rules (Appendix B, Table 7). The plug-and-play skill ControlNet design and half-physics decoupling are concrete systems contributions. The headroom (best full success 16.8%, NavSR peaking at 42.4%) and the focus on closed-loop action intelligence rather than open-loop spatial QA make the benchmark a credible stress test for the next generation of embodied foundation models.
major comments (3)
- [Abstract; Findings 4–6; §2.4; §6] Abstract and Findings 4–6 attribute post-recognition failures primarily to a missing internal faculty of “embodied self-awareness,” yet §2.4, Finding 6, and §6 state that the agent receives only egocentric RGB and text history, with no proprioceptive pose or contact/tactile channel, and that collisions “are never felt.” This interface choice is load-bearing for the diagnosis: many reported errors (unaware jammed, stop while far, sit on air, leg/arm collisions in Table 5 and Fig. 9) are exactly what an information-starved controller would produce. The empirical claim that none of nine VLMs solves the bench remains intact, but the abstract’s punchline should be scoped to “under egocentric vision-only feedback” (or supported by a body-state/contact ablation). As written, the faculty vs. interface distinction is under-determined.
- [§3.2; Table 2; Table 4] Headline success requires both geometric criteria and subjective acknowledgment/active stop (FindSR/NavSR/InteractSR in Table 2 vs. Geo* in Table 4). That design choice is defensible for “action intelligence,” but it conflates spatial competence with the model’s willingness to emit a stop/claim in the harnessed format. Several large GeoNavSR–NavSR and GeoInteractSR–InteractSR gaps (e.g., GPT-5.5 27.7% vs. 13.9%; InternVL GeoNavSR 12.6% vs. NavSR 0.8%) show that a non-trivial fraction of “failures” are termination/reporting failures under the scaffold. The paper should report Geo* as co-primary in the main table and discuss how much of the embodied-self-awareness narrative is driven by stop/claim behavior rather than body–scene geometry alone.
- [Abstract; Finding 1; Table 3; §2.2] Table 3 shows the skill-specific verifier is decisive: removing it collapses NavSR 27%→2% and InteractSR 18.9%→0% on the mini-val. The benchmark therefore measures harnessed VLMs (high-to-mid-to-low scaffold + per-skill verifier; Fig. 3, Table 6), not raw off-the-shelf models. This is disclosed, but the Abstract and Finding 1 (“none solves”; “what current VLMs lack”) should state clearly that results are under this scaffold, and that without the verifier the loop essentially does not close. Otherwise readers will over-read the numbers as pure foundation-model competence.
minor comments (6)
- [§3.1; Abstract] Interaction is instantiated only as sit on a subset of categories (§3.1; 597 sit episodes). Richer manipulation is acknowledged as future work in §6; a sentence in the abstract limiting “interact” to sit would avoid over-generalization.
- [§3.1; Fig. 7] Difficulty tiers use fixed thresholds (distance ≤3.5/≤8 m, etc.; Fig. 7). Brief sensitivity of success rates to these cutoffs would strengthen the stratified claims.
- [§3.2; Table 2] Motion jerk is a useful coherence metric, but the denoising/timescale details (≈0.27 s / stride 8) are only lightly specified; a short formula or appendix note would aid reproduction.
- [Table 1; §4.1] Climb upstairs/downstairs achievement ratios in Table 1 are 0.74–0.79 (stable gain). Confirm in text that the VLM is not expected to compensate for this gain, or that parameters are calibrated so commanded ≈ achieved.
- [§5.1] Related work on object-goal nav and VLN is appropriate; a slightly sharper contrast with Habitat 3.0 humanoid setups (Puig et al., 2024) on what half-physics adds beyond avatar animation would help placement.
- [Tables 2–4; title page] Minor polish: “HumanClawBench” vs. “HumanCLAW-Bench” casing is inconsistent in table captions; arXiv date line says July 30, 2026.
Circularity Check
Empirical benchmark paper with geometric success criteria and rule-based failure labels; no derivation that reduces predictions to fitted inputs or self-definitional premises.
full rationale
HumanCLAW is a systems/evaluation paper, not a first-principles derivation. The load-bearing claims—none of nine off-the-shelf VLMs solves the bench; best InteractSR 16.8%; post-recognition failures concentrate in self-localization and body placement—are measured against external geometric ground truth (semantic pixel count, AABB distance, pelvis–mesh contact) plus logged action streams, with root causes assigned by fixed deterministic rules over rollouts (Appendix B, Table 7). The motion layer’s skill fidelity (Table 1) is a separate zero-shot check against commanded magnitudes, not a fit reused as a VLM ‘prediction.’ Self-citations (e.g. half-physics from Siyao et al. 2025) supply infrastructure and are not invoked as uniqueness theorems that force the central diagnosis. Framing failures as ‘embodied self-awareness’ is interpretive labeling of measured patterns, not a circular reduction of Eq. X to Eq. Y. No fitted-input-called-prediction, self-definitional loop, or load-bearing self-citation chain is present. Score 0 is appropriate.
Axiom & Free-Parameter Ledger
free parameters (7)
- Decision chunk duration / horizon =
0.5 s (15 frames @ 30 Hz)
- FindSR pixel threshold =
100 px
- NavSR distance threshold =
20 cm (primary)
- Difficulty tier cutoffs =
dist ≤3.5/≤8 m; choice ≤2/≤5; obstacle ≤2.5/≤5
- Harness history defaults =
hist 10, img 1
- Half-physics passive stiffness λ =
λ = 1.0
- Skill ControlNet training budgets =
5e5–1.5e6 steps; lr 3e-4; bs 2048
axioms (6)
- domain assumption Half-physics preserves task-relevant physical consequences (collision, gravity, object displacement) while removing balance/motor-tracking failures as confounds for scoring action decisions.
- ad hoc to paper A finite set of atomic parameterized skills (walk, side_step, step_back, turn, climb, sit, stop) is an adequate action vocabulary for measuring long-horizon action intelligence on find-navigate-sit.
- ad hoc to paper Generalizable action intelligence should be probed with frozen off-the-shelf VLMs (reasoning) rather than policies fitted on embodiment trajectories.
- ad hoc to paper Staged success requires both geometric criteria and subjective acknowledgment/stop in the model’s outputs for headline metrics.
- domain assumption Automated top-to-bottom rules over navmesh distance, mesh contact, pixels, actions, and stated visible state assign a single true root cause per failed episode.
- domain assumption Standard simulators and assets (AI Habitat, Bullet, HSSD, AMASS) are faithful enough substrates for the claimed physical consequences and motion prior.
invented entities (4)
-
Action intelligence (as operational closed-loop component of spatial intelligence)
independent evidence
-
Embodied self-awareness (online estimate of own body state, scene relation, and recent action consequences)
no independent evidence
-
HumanCLAW harness (high-to-mid-to-low scaffold + skill-specific verifier)
independent evidence
-
Plug-and-play skill ControlNet bag on a shared motion DiT prior
independent evidence
read the original abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
Reference graph
Works this paper leans on
-
[1]
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S \"u nderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[2]
_0 : A vision-language-action flow model for general robot control, 2024
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. _0 : A vis...
Pith/arXiv arXiv 2024
-
[3]
Turner, Eric Undersander, and Tsung-Yen Yang
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M. Turner, Eric Undersander, and Tsung-Yen Yang. PARTNR : A benchmark for planning...
2025
-
[4]
SpatialVLM : Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. SpatialVLM : Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14455--14465, 2024
2024
-
[5]
LoTa-Bench : Benchmarking language-oriented task planners for embodied agents
Jae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim, and Minsu Jang. LoTa-Bench : Benchmarking language-oriented task planners for embodied agents. In The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=ADSxCpCu9s
2024
-
[6]
Umo: Unified in-context learning unlocks motion foundation model priors
Xiaoyan Cong, Zekun Li, Zhiyang Dou, Hongyu Li, Omid Taheri, Chuan Guo, Abhay Mittal, Sizhe An, Taku Komura, Wojciech Matusik, et al. Umo: Unified in-context learning unlocks motion foundation model priors. arXiv preprint arXiv:2603.15975, 2026
arXiv 2026
-
[7]
Erwin Coumans. Bullet physics simulation. In ACM SIGGRAPH 2015 Courses, page 7:1. ACM, 2015. doi:10.1145/2776880.2792704
arXiv 2015
-
[8]
Moving by looking: Towards vision-driven avatar motion generation
Markos Diomataris, Berat Mert Albaba, Giorgio Becherini, Partha Ghosh, Omid Taheri, and Michael J Black. Moving by looking: Towards vision-driven avatar motion generation. arXiv preprint arXiv:2509.19259, 2025
arXiv 2025
-
[9]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. PaLM-E : An embodied...
2023
-
[10]
Manipulate-anything: Automating real-world robots using vision-language models
Jiafei Duan, Wentao Yuan, Wilbert Pumacay, Yi Ru Wang, Kiana Ehsani, Dieter Fox, and Ranjay Krishna. Manipulate-anything: Automating real-world robots using vision-language models. In Conference on Robot Learning, 2024
2024
-
[11]
Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine
Dibya Ghosh, Homer Rich Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Quan Vuong, Ted Xiao, Pannag R. Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems...
-
[12]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5152--5161, 2022. doi:10.1109/CVPR52688.2022.00509
arXiv 2022
-
[13]
MoMask : Generative masked modeling of 3d human motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. MoMask : Generative masked modeling of 3d human motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1900--1910, 2024. https://openaccess.thecvf.com/content/CVPR2024/html/Guo_MoMask_Generative_Masked_Modeling_of_3D_Human_Motions_CVPR_2024_paper.html
1900
-
[14]
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J. Black. Stochastic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11354--11364, 2021. doi:10.1109/ICCV48922.2021.01118
arXiv 2021
-
[15]
ESI-Bench : Towards embodied spatial intelligence that closes the perception-action loop
Yining Hong, Jiageng Liu, Han Yin, Manling Li, Leonidas Guibas, Li Fei-Fei, Jiajun Wu, and Yejin Choi. ESI-Bench : Towards embodied spatial intelligence that closes the perception-action loop. arXiv preprint arXiv:2605.18746, 2026
Pith/arXiv arXiv 2026
-
[16]
VoxPoser : Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. VoxPoser : Composable 3d value maps for robotic manipulation with language models. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 540--562. PMLR, 2023 a . https://proceedings.mlr.press/v229/huang23b.html
2023
-
[17]
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Tomas Jackson, Noah Brown, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. Inner monologue: Embodied reasoning through planning with language models. In Proceedings of the 6th Conference on Rob...
2023
-
[18]
Como: Controllable motion generation through language guided pose code editing
Yiming Huang, Weilin Wan, Yue Yang, Chris Callison-Burch, Mark Yatskar, and Lingjie Liu. Como: Controllable motion generation through language guided pose code editing. In European Conference on Computer Vision, pages 180--196. Springer, 2024
2024
-
[19]
Do as i can, not as i say: Grounding language in robotic affordances
Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Proceedings of the 6th Conference on Robot Learning, volume 205 of Proceedings of Machine Learning Research, pages 287--318. PMLR, 2023...
2023
-
[20]
Iam: Identity-aware human motion and shape joint generation
Wenqi Jia, Zekun Li, Abhay Mittal, Chengcheng Tang, Chuan Guo, Lezi Wang, James Matthew Rehg, Lingling Tao, and Size An. Iam: Identity-aware human motion and shape joint generation. arXiv preprint arXiv:2604.25164, 2026
Pith/arXiv arXiv 2026
-
[21]
Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X. Chang, and Manolis Savva. Habitat synthetic scenes dataset ( HSSD -200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...
arXiv 2024
-
[22]
Foster, Pannag R
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P. Foster, Pannag R. Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA : An open-source vision-language-action model. In Proceedings of the 8th Conference ...
2025
-
[23]
MolmoAct : Action reasoning models that can reason in space
Jason Lee, Jiafei Duan, Haoquan Fang, Yuquan Deng, Shuo Liu, Boyang Han, Mohammadreza Salehi, Jae Sung Hwang, et al. MolmoAct : Action reasoning models that can reason in space. arXiv preprint arXiv:2508.07917, 2025
Pith/arXiv arXiv 2025
-
[24]
BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart \'i n-Mart \'i n, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation. In Proceedings of the 6th Conference on Robot Learning, volume 205 of Proceedings of Mac...
2023
-
[25]
Unfolding spatial cognition: Evaluating multimodal models on visual simulations
Linjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma, Yinuo Yang, Ziang Li, Yejin Choi, and Ranjay Krishna. Unfolding spatial cognition: Evaluating multimodal models on visual simulations. arXiv preprint arXiv:2506.04633, 2025
Pith/arXiv arXiv 2025
-
[26]
Embodied agent interface: Benchmarking LLMs for embodied decision making
Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, Weiyu Liu, Percy Liang, Li Fei-Fei, Jiayuan Mao, and Jiajun Wu. Embodied agent interface: Benchmarking LLMs for embodied decision making. In Advances in Neural Information Processing Systems, volume 37, 2024. doi:10.52202/079017...
-
[27]
Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens
Zekun Li, Sizhe An, Chengcheng Tang, Chuan Guo, Ivan Shugurov, Linguang Zhang, Amy Zhao, Srinath Sridhar, Lingling Tao, and Abhay Mittal. Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...
2026
-
[28]
Genhsi: Controllable generation of human-scene interaction videos
Zekun Li, Rui Zhou, Rahul Sajnani, Xiaoyan Cong, Daniel Ritchie, and Srinath Sridhar. Genhsi: Controllable generation of human-scene interaction videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 138--149, 2026 b
2026
-
[29]
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In IEEE International Conference on Robotics and Automation, 2023
2023
-
[30]
Large model empowered embodied AI : A survey on decision-making and embodied learning, 2025
Wenlong Liang, Rui Zhou, Yang Ma, Bing Zhang, Songlin Li, Yijia Liao, and Ping Kuang. Large model empowered embodied AI : A survey on decision-making and embodied learning, 2025. https://arxiv.org/abs/2508.10399
Pith/arXiv arXiv 2025
-
[31]
VisualAgentBench : Towards large multimodal models as visual foundation agents
Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, et al. VisualAgentBench : Towards large multimodal models as visual foundation agents. arXiv preprint arXiv:2408.06327, 2024
Pith/arXiv arXiv 2024
-
[32]
A survey on vision-language-action models for embodied AI
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied AI . IEEE Transactions on Neural Networks and Learning Systems, 2026. doi:10.1109/TNNLS.2025.3650584
arXiv 2026
-
[33]
GR00T N1 : An open foundation model for generalist humanoid robots, 2025
NVIDIA , Johan Bjorck, Fernando Casta \ n eda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, et al. GR00T N1 : An open foundation model for generalist humanoid robots, 2025. https://arxiv.org/abs/2503.14734
Pith/arXiv arXiv 2025
-
[34]
VirtualHome : Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. VirtualHome : Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8494--8502, 2018. https://openaccess.thecvf.com/content_cvpr_2018/html/Puig_VirtualHome_Simulating_Household_CVPR...
2018
-
[35]
Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi
Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, Vladimir Vondrus, Theophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara ...
2024
-
[36]
Habitat: A platform for embodied AI research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A platform for embodied AI research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9339--9347, 2019. doi:10.1109/ICCV.2019.00943
arXiv 2019
-
[37]
ALFRED : A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. ALFRED : A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10740--10749, 2020. https://openaccess.thecvf.com/content_CV...
2020
-
[38]
Bailando: 3 D dance generation by actor-critic GPT with choreographic memory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando: 3 D dance generation by actor-critic GPT with choreographic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[39]
Bailando++: 3 D dance GPT with choreographic memory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando++: 3 D dance GPT with choreographic memory. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 45 0 (12): 0 14192--14207, 2023
2023
-
[40]
Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment
Li Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin, Ziwei Liu, Henghui Ding, Lei Yang, and Chen Change Loy. Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment. In International Conference on Learning Representations (ICLR), 2024
2024
-
[41]
Half-physics: Enabling kinematic 3d human model with physical interactions
Li Siyao, Yao Feng, Omid Taheri, Chen Change Loy, and Michael J Black. Half-physics: Enabling kinematic 3d human model with physical interactions. arXiv preprint arXiv:2507.23778, 2025
Pith/arXiv arXiv 2025
-
[42]
Sadler, Wei-Lun Chao, and Yu Su
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler, Wei-Lun Chao, and Yu Su. LLM-Planner : Few-shot grounded planning for embodied agents with large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2998--3009, 2023. https://openaccess.thecvf.com/content/ICCV2023/html/Song_LLM-Planner_Few-Shot_Gr...
2023
-
[43]
Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra
Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel X. Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Trainin...
2021
-
[44]
Cradle: Empowering foundation agents towards general computer control
Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li, Bohan Zhou, Junpeng Yue, Jiechuan Jiang, Yewen Li, et al. Cradle: Empowering foundation agents towards general computer control. arXiv preprint arXiv:2403.03186, 2024
Pith/arXiv arXiv 2024
-
[45]
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. https://openreview.net/forum?id=SJ1kSyO2jwu
2023
-
[46]
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. https://openreview.net/forum?id=ehfRiF0R3a
2024
-
[47]
Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, and Yitao Liang. Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents. In Advances in Neural Information Processing Systems, volume 36, 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/6b8dfb8c0c12e6fafc6c256cb08a5ca7-Abstract-...
2023
-
[48]
Text2interact: High-fidelity and diverse text-to-two-person interaction generation
Qingxuan Wu, Zhiyang Dou, Chuan Guo, Yiming Huang, Qiao Feng, Bing Zhou, Jian Wang, and Lingjie Liu. Text2interact: High-fidelity and diverse text-to-two-person interaction generation. arXiv preprint arXiv:2510.06504, 2025
arXiv 2025
-
[49]
OmniControl : Control any joint at any time for human motion generation
Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang. OmniControl : Control any joint at any time for human motion generation. In The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=gd0lAEtWso
2024
-
[50]
Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces. arXiv preprint arXiv:2412.14171, 2024
Pith/arXiv arXiv 2024
-
[51]
EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents
Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, Heng Ji, Huan Zhang, and Tong Zhang. EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. In Proceedings of the 42nd International Conference on Machine Lear...
2025
-
[52]
PhysDiff : Physics-guided human motion diffusion model
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. PhysDiff : Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16010--16021, 2023. doi:10.1109/ICCV51070.2023.01467
arXiv 2023
-
[53]
Alex L. Zhang, Thomas L. Griffiths, Karthik R. Narasimhan, and Ofir Press. VideoGameBench : Can vision-language models complete popular video games? arXiv preprint arXiv:2505.18134, 2025 a
Pith/arXiv arXiv 2025
-
[54]
Egoreact: Egocentric video-driven 3d human reaction generation
Libo Zhang, Zekun Li, Tianyu Li, Zeyu Cao, Rui Xu, Xiaoxiao Long, Wenjia Wang, Jingbo Wang, Yuan Liu, Wenping Wang, et al. Egoreact: Egocentric video-driven 3d human reaction generation. arXiv preprint arXiv:2512.22808, 2025 b
arXiv 2025
-
[55]
The wanderings of odysseus in 3d scenes
Yan Zhang and Siyu Tang. The wanderings of odysseus in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20481--20491, 2022. https://openaccess.thecvf.com/content/CVPR2022/html/Zhang_The_Wanderings_of_Odysseus_in_3D_Scenes_CVPR_2022_paper.html
2022
-
[56]
Yan Zhang, Yao Feng, Alp \'a r Cseke, Nitin Saini, Nathan Bajandas, Nicolas Heron, and Michael J. Black. PRIMAL : Physically reactive and interactive motor model for avatar learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12725--12736, 2025 c . https://openaccess.thecvf.com/content/ICCV2025/html/Zhang_PRIMAL_Phys...
2025
-
[57]
Synthesizing diverse human motions in 3d indoor scenes
Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, and Siyu Tang. Synthesizing diverse human motions in 3d indoor scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14738--14749, 2023. https://openaccess.thecvf.com/content/ICCV2023/html/Zhao_Synthesizing_Diverse_Human_Motions_in_3D_Indoor_Scenes_ICCV_2023_paper.html
2023
-
[58]
RT-2 : Vision-language-action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. RT-2 : Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 2165--2183. PMLR, 2023. http...
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.