REVIEW 4 major objections 5 minor 70 references
A robot world model conditioned on rendered nominal robot geometry rather than raw action commands or logged future states learns only scene response and generalizes to unseen robots.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 04:23 UTC pith:P7BS6QCP
load-bearing objection A conceptually clean interface for robot video world models—render the deployment-available nominal trajectory as robot geometry—with experiments that support the design trends but don't yet isolate the claimed scene-response gains. the 4 major comments →
Robot-Factored World Models via Robot Rendering
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that robot world models should be conditioned not on action commands and not on logged robot states, but on a deployment-available nominal trajectory: the controller-realized motion produced from the command before scene interaction. This trajectory is rendered through the robot's geometric description into a camera-aligned RGB mesh video plus an end-effector-only depth channel, and paired with a static stream of scene RGB and depth along the same camera path. The learned model, a latent video diffusion model, is conditioned jointly on these streams and a scene-only text prompt, so its only task is to predict interaction-induced scene changes. On two manipulation
What carries the argument
The load-bearing objects are the nominal-trajectory operator ΦR (maps an action sequence into a robot-only trajectory via the robot's controller and kinematics) and the rendering operator ΠR (projects that trajectory through the robot's URDF geometry into a camera-aligned mesh RGB video and end-effector depth). A static context renderer ΠS supplies scene RGB and depth along the same camera trajectory, constructing the condition c = [E(Brgb), E(Mrgb), E(Dscene), E(Deef)] fed into a latent video diffusion model trained with a flow-matching loss. The pair ΦR and ΠR moves action realization out of the learned network; end-effector depth is what resolves whether an overlapping object is in front
Load-bearing premise
The nominal trajectory, produced by replaying commands in a scene-free robot simulation or collision-free shadow rollout, is assumed to match the controller-realized motion the actual robot would execute in free space before touching anything; if that replay diverges, the rendered interface is not the deployment-available signal it claims to be.
What would settle it
Take a physical robot and a fixed action sequence in free space (no scene objects), record its executed joint or end-effector trajectory, and compare it to the scene-free replayed nominal trajectory from the same start state; if the two differ by more than the robot's tracking error, the nominal trajectory is not deployment-available and the anti-leakage conditioning argument fails.
If this is right
- A world model trained this way no longer needs to learn how a specific robot's controller translates commands into motion; that factor is computed once and rendered.
- Because conditioning is visual and embodiment-agnostic, the same trained model accepts a new robot at inference simply by rendering its own geometry.
- Human demonstrations can be converted into robot-interaction rollouts by retargeting hand and arm motion and rendering it as robot geometry, without retraining.
- The depth pairing of end-effector and scene depth should make predicted contact, proximity, and occlusion ordering more reliable than image-plane overlap alone.
- Training and inference must both use nominal prompts; the paper's oracle diagnostic shows that training on logged future-state prompts degrades deployment-time performance when swapped for nominal prompts.
Where Pith is reading between the lines
- If this holds, robot world models could be pretrained once on a large corpus of mixed robot data and then applied to new hardware by only swapping the renderer, cutting the cost of embodied model adaptation.
- The anti-leakage property suggests a stress test: evaluate world models on contact-rich failure clips (missed grasps, slips); a model that leaks logged states would look artificially good on those clips, while the nominal interface reveals its true predictive limit.
- The interface's reliance on a known static scene suggests a natural extension to partially observed real scenes by coupling the rendered robot stream with online 3D reconstruction, an extension the paper notes is not yet fully closed.
- One could also render non-robot actors (other arms, tools, or articulated furniture) through the same geometry-plus-depth interface, turning the method into a general visual action prompt for interactive video generation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes robot-factored world models for action-conditioned video prediction. Instead of conditioning a video diffusion model directly on raw action commands or on logged future robot states, it first rolls each action through the robot's controller and kinematics to obtain a deployment-available 'nominal trajectory' q1:F (Eq. 1), and renders this trajectory through the URDF into camera-aligned mesh RGB and end-effector depth (Eq. 2). A static, camera-aware scene context supplies appearance, depth, and viewpoint (Eq. 3). The world model then predicts future video conditioned on this rendered geometry plus static context and a scene-only text prompt (Eq. 4). The authors argue this factorization keeps action realization and robot-geometry rendering outside the learned model, so the model only has to learn how the scene responds to the visible robot motion. Experiments on DROID and RoboCasa-GR1 compare against vector/pose-conditioned baselines, ablate nominal vs. raw-action rendering and depth, and provide qualitative demonstrations of prompt following, unseen embodiments, and human-demo retargeting.
Significance. If the central claim is supported, the paper offers a clean and practical interface idea: represent the action as explicit, camera-aligned robot geometry computed before scene interaction, thereby removing the need for a video model to learn embodiment-specific action realization. The method is concrete, the experiments use held-out clips, and the ablations in Table 2 directly separate nominal-vs-raw rendering and the depth pair. The paper also ships substantial implementation detail and an honest limitations section. However, the quantitative evidence is weakened by a confound: the rendered mesh condition is almost identical to the robot pixels in the target video, so full-frame reconstruction metrics do not isolate scene-response prediction. In addition, no error bars are reported anywhere, and the embodiment-generalization claims rest on qualitative still frames. These issues are fixable and do not invalidate the underlying idea, but they are load-bearing for the paper's strongest conclusions.
major comments (4)
- [Tables 1–2, §4.2] The main quantitative comparison is confounded. The conditioning stream contains M^rgb, the rendered robot mesh from the nominal trajectory, which is pixel-aligned with the robot region of the target video (exactly identical in RoboCasa-GR1 and very close in DROID). A model can achieve high full-frame PSNR/SSIM/LPIPS by copying the rendered robot pixels into its output while the AdaLN state-vector baseline must synthesize the entire robot from a low-dimensional state. Thus the reported gains may reflect robot-pixel reconstruction rather than better scene-response prediction. The paper should report metrics that exclude the rendered robot region, or quantitative interaction metrics such as object displacement, contact accuracy, or grasp outcome, before claiming the model 'learns how objects respond' (Eq. 4).
- [Tables 1–4, §4.2–4.3] No error bars or repeated-seed results are reported. Evaluation sets are small (256 and 128 clips) and video diffusion sampling is stochastic; some differences in Table 2 are small (e.g., SSIM 0.872 vs. 0.874, LPIPS 0.164 vs. 0.161). The paper should provide confidence intervals, multiple seeds, or a significance test. Without this, the current margins are difficult to interpret, and the claim that the rendered interface 'outperforms' baselines is not yet robust.
- [§4.3, Table 2, raw-action mesh row] The raw-action mesh comparison is potentially a strawman. Appendix A states that raw DROID joint targets can 'run ahead' of the physically trackable motion, so directly rendering raw actions produces visually implausible, impossible robot motion. The improvement of nominal mesh over raw-action mesh may then reflect rendering distortion rather than the value of realizing the action through the controller. Please quantify the raw-vs-nominal trajectory gap (e.g., tracking error or per-joint velocity statistics) or add an intermediate baseline that smooths raw actions, so the action-realization claim is isolated.
- [§4.4, Fig. 6, Appendix E] The third stated contribution—embodiment generalization to unseen robots—is supported only by qualitative still frames (HRDexDB, DexMimicGen). No quantitative metrics or evaluation protocol are given for these zero-shot settings. Since this is a headline contribution, provide either quantitative results on unseen embodiments or explicitly reframe this section as a qualitative proof-of-concept. The current evidence does not support a strong generalization claim.
minor comments (5)
- [§3.1, Eq. 4] The phrase 'learns how objects respond' is causal language. The model is a conditional generative model, not a causal estimator; suggest softening to 'predicts scene response' or adding a caveat that the learned association is observational.
- [§4.1] The exact definition of 'raw action mesh' in Table 2 is not stated in Section 4.1; Appendix A clarifies that raw DROID actions are logged joint/gripper targets, but the rendering procedure for raw actions should be described in the main text for reproducibility.
- [Figures 3, 5, 6] The qualitative figures are hard to evaluate from static stills. Consider adding video links or a supplementary video, and for Figure 5 report a quantitative prompt-following metric (e.g., trajectory tracking of the predicted robot region) rather than a single example.
- [§5] The limitations paragraph is candid, but the static-context assumption and the need for URDF/calibration directly constrain the claimed generality. These caveats should be reflected earlier, in the abstract or introduction, so readers do not overstate the method's applicability.
- [Appendix B, Table 3] The training hyperparameters table is useful, but the guidance scale 1.0 with a distilled LoRA may affect sample quality; please state whether the baselines use the same inference protocol and whether any tuning was performed per method.
Circularity Check
No load-bearing circularity; only a minor non-load-bearing self-citation.
full rationale
The derivation chain (Eqs. 1-4) is a constructive preprocessing interface, not a derivation that reduces to its inputs. Eq. 1 obtains the nominal trajectory by replaying actions through a controller; Eq. 2 renders it into mesh and end-effector depth; Eq. 4 conditions a diffusion model on the rendered streams plus static context. No fitted parameter is renamed as a prediction, and the main quantitative claims (Tables 1, 2, 4) are evaluated on held-out clips; the logged-state rows of Table 4 are explicitly an oracle diagnostic, not a claimed deployment prediction. The only self-citation is Dexterous World Models [44] (Sec. 3.5), which shares three authors with this paper and supplies the residual-dynamics/inpainting backbone and static-context design. This dependency is structural but not load-bearing for the paper's distinctive contribution: the nominal-trajectory rendering interface is implemented and validated on held-out data independently of any theorem or claim imported from [44]. The skeptical full-frame-metric confound (rendered mesh is visually close to the robot pixels in the target video) is a real external-validity concern about what the metrics isolate, but it is an evaluation confound rather than a circular derivation: the model's scene-response prediction is not logically forced by the conditioning streams. Accordingly no circular step is present; the score is 2 only to register the minor, non-load-bearing self-citation.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Residual-dynamics/video-inpainting formulation from Dexterous World Models [44] can learn scene response around a static context.
- domain assumption Scene-free replay (Isaac Lab / shadow rollout) yields a deployment-available nominal trajectory that matches real pre-interaction robot motion.
- domain assumption Static context stream (repeated first frame or robot-free simulated render) lets the model attribute all change to robot interaction.
- domain assumption Known URDF and camera-to-robot calibration are available for each embodiment.
- domain assumption Depth inputs (FoundationStereo estimate or simulator depth) are accurate enough to resolve contact-relevant proximity and occlusion.
- domain assumption PSNR/SSIM/LPIPS are meaningful proxies for world-model action-following quality.
read the original abstract
Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two distinct processes: they are first realized into robot motion by the robot body and controller, and the scene then responds through contact and object motion. Conditioning directly on action commands asks the world model to learn the realization process itself, while conditioning on logged future states leaks the interaction outcomes it is meant to predict. We propose robot-factored world models, which move two robot-specific factors outside the world model. First, action realization: each command is rolled through the robot's own controller and kinematics into a deployment-available nominal trajectory, a middle signal that avoids both action-realization learning and future-state leakage. Second, robot rendering: this nominal trajectory is rendered through the robot URDF, factoring the robot's geometry, kinematics, and appearance out of the model and into explicit rendered robot geometry. To resolve depth ambiguity, we pair end-effector depth with scene depth, giving geometric cues for contact and occlusion beyond image-plane overlap. Together, camera-aware static RGB/depth context and rendered robot geometry form a shared visual world-model interface that stays consistent across viewpoints and robot embodiments, so the model sees the action only as visible robot geometry and learns how objects respond to it. Our experiments show that the rendered interface outperforms vector-conditioned baselines and generalizes to unseen robot embodiments at inference. We further demonstrate that our model generates robot manipulation videos from human demonstrations by retargeting and rendering the hand motion as robot geometry.
Figures
Reference graph
Works this paper leans on
-
[1]
Ha and J
D. Ha and J. Schmidhuber. Recurrent world models facilitate policy evolution. InNeurIPS, 2018
2018
-
[2]
A. Bar, G. Zhou, D. Tran, T. Darrell, and Y . LeCun. Navigation world models. InCVPR, 2025
2025
-
[3]
F. Zhu, H. Wu, S. Guo, Y . Liu, C. Cheang, and T. Kong. Irasim: A fine-grained world model for robot manipulation. InICCV, 2025
2025
-
[4]
Y . Guo, L. X. Shi, J. Chen, and C. Finn. Ctrl-world: A controllable generative world model for robot manipulation. InICLR, 2026
2026
-
[5]
S. Gao, W. Liang, K. Zheng, A. Malik, S. Ye, S. Yu, W.-C. Tseng, Y . Dong, K. Mo, C.-H. Lin, Q. Ma, S. Nah, L. Magne, J. Xiang, Y . Xie, R. Zheng, D. Niu, Y . L. Tan, K. Zentner, G. Kurian, S. Indupuru, P. Jannaty, J. Gu, J. Zhang, J. Malik, P. Abbeel, M.-Y . Liu, Y . Zhu, J. Jang, and L. Fan. Dreamdojo: A generalist robot world model from large-scale hum...
Pith/arXiv arXiv 2026
-
[6]
A. K. Sharma, Y . Sun, N. Lu, Y . Zhang, J. Liu, and S. Yang. World-gymnast: Training robots with reinforcement learning in a world model.arXiv preprint arXiv:2602.02454, 2026
arXiv 2026
-
[7]
J. Quevedo, A. K. Sharma, Y . Sun, V . Suryavanshi, P. Liang, and S. Yang. Worldgym: World model as an environment for policy evaluation.arXiv preprint arXiv:2506.00613, 2025
arXiv 2025
-
[8]
Y . Wang, C. Wen, H. Guo, S. Peng, M. Qin, H. Bao, X. Zhou, and R. Hu. Precise action-to-video generation through visual action prompts. InICCV, 2025
2025
-
[9]
Hafner, T
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. InICML, 2019
2019
-
[10]
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023
Pith/arXiv arXiv 2023
-
[11]
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg. Daydreamer: World models for physical robot learning. InCoRL, 2023
2023
-
[12]
Hansen, H
N. Hansen, H. Su, and X. Wang. Td-mpc2: Scalable, robust world models for continuous control. InICLR, 2024
2024
-
[13]
N. Agarwal, A. Ali, M. Bala, Y . Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y . Chen, Y . Cui, Y . Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025
Pith/arXiv arXiv 2025
-
[14]
J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y . Fang, F. Hu, S. Huang, K. Kundalia, Y .-C. Lin, et al. Dreamgen: Unlocking generalization in robot learning through video world models.arXiv preprint arXiv:2505.12705, 2025
Pith/arXiv arXiv 2025
-
[15]
Bruce, M
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. InICML, 2024
2024
-
[16]
Alonso, A
E. Alonso, A. Jelley, V . Micheli, A. Kanervisto, A. Storkey, T. Pearce, and F. Fleuret. Diffusion for world modeling: Visual details matter in atari.NeurIPS, 2024
2024
-
[17]
Valevski, Y
D. Valevski, Y . Leviathan, M. Arar, and S. Fruchter. Diffusion models are real-time game engines. InICLR, 2025
2025
-
[18]
Z. Yang, Y . Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun. Unisim: A neural closed-loop sensor simulator. InCVPR, 2023. 9
2023
-
[19]
J. Wu, S. Yin, N. Feng, X. He, D. Li, J. Hao, and M. Long. ivideogpt: Interactive videogpts are scalable world models.NeurIPS, 2024
2024
-
[20]
H. Zhu, Y . Wang, J. Zhou, W. Chang, Y . Zhou, Z. Li, J. Chen, C. Shen, J. Pang, and T. He. Aether: Geometric-aware unified world modeling. InICCV, 2025
2025
-
[21]
J. Zhou, H. Gao, V . V oleti, A. Vasishta, C.-H. Yao, M. Boss, P. Torr, C. Rupprecht, and V . Jampani. Stable virtual camera: Generative view synthesis with diffusion models. InICCV, 2025
2025
-
[22]
L. Russell, A. Hu, L. Bertoni, G. Fedoseev, J. Shotton, E. Arani, and G. Corrado. Gaia-2: A controllable multi-view generative world model for autonomous driving.arXiv preprint arXiv:2503.20523, 2025
Pith/arXiv arXiv 2025
-
[23]
S. Gao, S. Zhou, Y . Du, J. Zhang, and C. Gan. Adaworld: Learning adaptable world models with latent actions.arXiv preprint arXiv:2503.18938, 2025
Pith/arXiv arXiv 2025
-
[24]
M. Huang, Y . Xiang, Z. Liang, J. Huang, J. Wang, Z. Xu, F. Tan, H. Zhou, M. Yang, and G. Che. Coworld-vla: Thinking in a multi-expert world model for autonomous driving.arXiv preprint arXiv:2605.10426, 2026
Pith/arXiv arXiv 2026
-
[25]
A. Liang, P. Czempin, M. Hong, Y . Zhou, E. Biyik, and S. Tu. Clam: Continuous latent action models for robot learning from unlabeled demonstrations.arXiv preprint arXiv:2505.04999, 2025
Pith/arXiv arXiv 2025
-
[26]
Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li. Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025
Pith/arXiv arXiv 2025
-
[27]
Q. Garrido, T. Nagarajan, B. Terver, N. Ballas, Y . LeCun, and M. Rabbat. Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026
arXiv 2026
-
[28]
Zhang, A
L. Zhang, A. Rao, and M. Agrawala. Adding conditional control to text-to-image diffusion models. InICCV, 2023
2023
-
[29]
Zhang, Y
Y . Zhang, Y . Wei, X. ZHANG, W. Zuo, Q. Tian, et al. Controlvideo: Training-free controllable text-to-video generation. InICLR, 2024
2024
-
[30]
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y . Zhang, Y . Shen, D. Zhao, and J. Zhou. Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 2023
2023
-
[31]
Y . Guo, C. Yang, A. Rao, Z. Liang, Y . Wang, Y . Qiao, M. Agrawala, D. Lin, and B. Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023
Pith/arXiv arXiv 2023
-
[32]
Z. Wang, Z. Yuan, X. Wang, Y . Li, T. Chen, M. Xia, P. Luo, and Y . Shan. Motionctrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH, 2024
2024
-
[33]
S. Yin, C. Wu, J. Liang, J. Shi, H. Li, G. Ming, and N. Duan. Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory.arXiv preprint arXiv:2308.08089, 2023
Pith/arXiv arXiv 2023
-
[34]
Zhang, J
Z. Zhang, J. Liao, M. Li, Z. Dai, B. Qiu, S. Zhu, L. Qin, and W. Wang. Tora: Trajectory-oriented diffusion transformer for video generation. InCVPR, 2025
2025
-
[35]
M. Niu, X. Cun, X. Wang, Y . Zhang, Y . Shan, and Y . Zheng. Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model. In ECCV, 2024
2024
-
[36]
H. Zhou, C. Wang, R. Nie, J. Liu, D. Yu, Q. Yu, and C. Wang. Trackgo: A flexible and efficient method for controllable video generation. InAAAI, 2025. 10
2025
-
[37]
Jiang, Z
Z. Jiang, Z. Han, C. Mao, J. Zhang, Y . Pan, and Y . Liu. Vace: All-in-one video creation and editing. InICCV, 2025
2025
-
[38]
D. Geng, C. Herrmann, J. Hur, F. Cole, S. Zhang, T. Pfaff, T. Lopez-Guevara, Y . Aytar, M. Rubinstein, C. Sun, et al. Motion prompting: Controlling video generation with motion trajectories. InCVPR, 2025
2025
-
[39]
J. Shin, Z. Li, R. Zhang, J.-Y . Zhu, J. Park, E. Shechtman, and X. Huang. Motionstream: Real-time video generation with interactive motion controls.arXiv preprint arXiv:2511.01266, 2025
arXiv 2025
-
[40]
H. Qiu, Z. Chen, Z. Wang, Y . He, M. Xia, and Z. Liu. Freetraj: Tuning-free trajectory control in video diffusion models.arXiv preprint arXiv:2406.16863, 2024
Pith/arXiv arXiv 2024
-
[41]
W.-D. K. Ma, J. P. Lewis, and W. B. Kleijn. Trailblazer: Trajectory control for diffusion-based video generation. InACM SIGGRAPH Asia, 2024
2024
-
[42]
Y . Jain, A. Nasery, V . Vineet, and H. Behl. Peekaboo: Interactive video generation via masked- diffusion. InCVPR, 2024
2024
-
[43]
Y . Li, X. Wang, Z. Zhang, Z. Wang, Z. Yuan, L. Xie, Y . Shan, and Y . Zou. Image conductor: Precision control for interactive video synthesis. InAAAI, 2025
2025
-
[44]
B. Kim, T. Kim, J. Lee, and H. Joo. Dexterous world models.arXiv preprint arXiv:2512.17907, 2025
arXiv 2025
-
[45]
Videox-fun: A video generation pipeline for diffusion transformer, 2026
aigc apps. Videox-fun: A video generation pipeline for diffusion transformer, 2026. URL https://github.com/aigc-apps/VideoX-Fun
2026
-
[46]
D. P. Kingma and M. Welling. Auto-encoding variational bayes. InICLR, 2014
2014
-
[47]
Team Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[48]
Lipman, R
Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. InICLR, 2023
2023
-
[49]
X. Liu, C. Gong, and Q. Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. InICLR, 2023
2023
-
[50]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024
2024
-
[51]
Nasiriany, A
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. InRSS, 2024
2024
-
[52]
T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026
arXiv 2026
-
[53]
J. Lim, T. Ha, M. Choi, J. Kim, B. Kim, S. Jeon, and H. Joo. Hrdexdb: A large-scale dataset of dexterous human and robotic hand grasps.arXiv preprint arXiv:2604.14944, 2026
Pith/arXiv arXiv 2026
-
[54]
Y .-W. Chao, W. Yang, Y . Xiang, P. Molchanov, A. Handa, J. Tremblay, Y . S. Narang, K. Van Wyk, U. Iqbal, S. Birchfield, J. Kautz, and D. Fox. DexYCB: A benchmark for capturing hand grasping of objects. InCVPR, 2021. 11
2021
-
[55]
M. Mittal, P. Roth, J. Tigue, A. Richard, O. Zhang, P. Du, A. Serrano-Muñoz, X. Yao, R. Zür- brügg, N. Rudin, et al. Isaac lab: A gpu-accelerated simulation framework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025
Pith/arXiv arXiv 2025
-
[56]
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y . Levi, Z. English, V . V oleti, A. Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023
Pith/arXiv arXiv 2023
-
[57]
Peebles and S
W. Peebles and S. Xie. Scalable diffusion models with transformers. InICCV, 2023
2023
-
[58]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. InCVPR, 2018
2018
-
[59]
B. Sundaralingam, S. K. S. Hari, A. Fishman, C. Garrett, K. Van Wyk, V . Blukis, A. Millane, H. Oleynikova, A. Handa, F. Ramos, N. Ratliff, and D. Fox. curobo: Parallelized collision-free minimum-jerk robot motion generation.arXiv preprint arXiv:2310.17274, 2023
Pith/arXiv arXiv 2023
-
[60]
Y . Qin, W. Yang, B. Huang, K. Van Wyk, H. Su, X. Wang, Y .-W. Chao, and D. Fox. Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system. InRSS, 2023
2023
-
[61]
C. M. Kim, B. Yi, H. Choi, Y . Ma, K. Goldberg, and A. Kanazawa. Pyroki: A modular toolkit for robot kinematic optimization. InIROS, 2025
2025
-
[62]
NVIDIA, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z. Yu, A....
Pith/arXiv arXiv 2025
-
[63]
B. Wen, M. Trepte, J. Aribido, J. Kautz, O. Gallo, and S. Birchfield. Foundationstereo: Zero-shot stereo matching. InCVPR, 2025
2025
-
[64]
S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Pith/arXiv arXiv 2025
-
[65]
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen. LoRA: Low-rank adaptation of large language models. InICLR, 2022
2022
-
[66]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InICLR, 2019
2019
-
[67]
Contributors
L. Contributors. Lightx2v: Light video generation inference framework, 2025. URL https: //github.com/ModelTC/lightx2v
2025
-
[68]
B. Zi, W. Peng, X. Qi, J. Wang, S. Zhao, R. Xiao, and K.-F. Wong. Minimax-remover: Taming bad noise helps video object removal.Advances in Neural Information Processing Systems, 38: 75518–75547, 2026
2026
-
[69]
Jiang, Y
Z. Jiang, Y . Xie, K. Lin, Z. Xu, W. Wan, A. Mandlekar, L. J. Fan, and Y . Zhu. DexMimicGen: Automated data generation for bimanual dexterous manipulation via imitation learning. In ICRA, pages 16923–16930, 2025. 12 A Interface Construction Details Dataset-specific nominalization.The main paper distinguishes raw actions, nominal robot-only trajectories, a...
2025
-
[70]
Four raw-frame states are stacked for each Wan latent frame, giving a 184D AdaLN input
For joint DROID+RoboCasa training, DROID fills the first 7 entries and RoboCasa fills the next 39 entries of a zero-padded 46D union vector. Four raw-frame states are stacked for each Wan latent frame, giving a 184D AdaLN input. Inference.Wan 2.1 14B variants are sampled with the LightX2V [ 67] CFG-and-step-distilled LoRA at 4 denoising steps and guidance...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.