{"total":14,"items":[{"citing_arxiv_id":"2607.00836","ref_index":18,"ref_count":2,"confidence":0.98,"is_internal_anchor":true,"paper_title":"From World Models to World Action Models: A Concise Tutorial for Robotics","primary_cat":"cs.RO","submitted_at":"2026-07-01T11:56:54+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.22136","ref_index":53,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data","primary_cat":"cs.RO","submitted_at":"2026-06-20T16:31:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Wh0 generates scalable egocentric human manipulation videos with world models and converts them to boost pretrained VLA models' zero-shot dexterous task success from 8.3% to 38.9% on 18 real-world tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20905","ref_index":90,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Vesta: A Generalist Embodied Reasoning Model","primary_cat":"cs.RO","submitted_at":"2026-06-18T20:01:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Vesta is a unified embodied generalist model that outperforms specialist baselines by over 20% on average and improves real-world robotic task success by over 35%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20781","ref_index":133,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"World Action Models: A Survey","primary_cat":"cs.RO","submitted_at":"2026-06-18T17:05:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"The shared lesson is that the rendered future can be a common planning currency even when the recovery module changes. Modular Render-and-Decode WAMs make that separation explicit. DreamGen [73] uses generated robot futures to obtain pseudo-actions before downstream policy training, so the coupling is offline rather than inference-time imagination. 4DGen [108], RIGVid [133], TC-IDM [119], and GraspDreamer [147] keep the world model and action recovery separate through pose tracking, VLM filtering, tool-centric point-cloud conversion, or human-grasp retargeting. VERA [94] makes the split even sharper by leaving the video planner action-free and training a Jacobian inverse-dynamics translator. These works are useful because the world"},{"citing_arxiv_id":"2606.12995","ref_index":13,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training","primary_cat":"cs.RO","submitted_at":"2026-06-11T07:31:05+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A humanoid robot can carry out diverse manipulation tasks in a zero-shot way by imitating one AI-generated video, using contact-aware trajectory optimization instead of task-specific policy training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12604","ref_index":44,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations","primary_cat":"cs.RO","submitted_at":"2026-06-10T19:01:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"EgoEngine transforms egocentric human videos into high-fidelity robot data enabling zero-shot visuomotor dexterous policy learning without real-robot demonstrations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.24934","ref_index":31,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos","primary_cat":"cs.RO","submitted_at":"2026-05-24T08:26:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"HumanEgo reports 92.5% average success on four real robot tasks using only 15-30 minutes of human video per task and zero robot data, with zero-shot transfer to new robots and cameras.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22272","ref_index":58,"ref_count":2,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors","primary_cat":"cs.RO","submitted_at":"2026-05-21T10:15:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Imagine2Real enables zero-shot humanoid-object interaction by unifying motions as 4D point trajectories, tracking only base/hands/object keypoints inside a BFM latent space, and training with progressive simple rewards for mocap deployment.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12090","ref_index":81,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"World Action Models: The Next Frontier in Embodied AI","primary_cat":"cs.RO","submitted_at":"2026-05-12T13:10:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Cascaded W AM Explicit UniPi [6], VLP [ 7], RoboEnvision [9], ThisThat [ 65], TesserAct [66], MVISTA-4D [67] Say ,Dream,and Act [10], Gen2Act [68], A VDC [8], Im2Flow2Act [69], 3DFlowAction [70] NovaFlow [71], Dream2Flow [72], Dreamitate [ 73], 4DGen [ 74], RIGVid [75], L VP [76] Vidar [77], Veo-Act [78], pi0.7 [ 79], V AG [80] Implicit VPP [11], VILP [ 81], Video Policy [13], ARDuP [ 82], mimic-video [ 12], LAP A [15], villa-X [ 83], S-V AM [14], OmniVTA [84], MWM [85] Joint W AM Autoregression GR1 [86], grmg [ 87], GR2 [88], Co TVLA [89], WorldVLA [90], rynnvla2 [91] VLA-JEP A [92], F1-VLA [93] Diffusion-based P AD [21], VideoVLA [94], UWM [20], DreamZero [ 17], CosmosPolicy [16], FLARE [95], UV A [96]"},{"citing_arxiv_id":"2605.02667","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AnchorD: Metric Grounding of Monocular Depth Using Factor Graphs","primary_cat":"cs.RO","submitted_at":"2026-05-04T14:48:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AnchorD anchors monocular depth priors in metric sensor data via patch-wise affine alignment using factor graph optimization, improving accuracy on non-Lambertian objects and introducing a new benchmark dataset with dense ground truth.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.04974","ref_index":74,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data","primary_cat":"cs.RO","submitted_at":"2026-04-04T15:37:11+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Video-to-robot control methods cluster into three interface families, and the field’s main bottleneck is grounding video-derived predictions into dependable closed-loop robot behavior.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.28489","ref_index":202,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms","primary_cat":"eess.IV","submitted_at":"2026-03-30T14:23:45+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"In twisted bilayer nodal d-wave superconductors, interlayer hopping creates nodes on the C2 axis and Bogoliubov flat bands when the single-layer Berry connection is parallel to that axis.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Drive-WM [182], Vista [183], MiLA [184], ADriver-I [185], [186], Drivedreamer [19], MagicDrive-V2 [40], DriveArena [187], MAD [188] Epona [189], GenAD [190], DriveLaW [191], DrivingGPT [192], VaV AM [193] Embodied AI Vidar [194], DreamGen [195], GenMimic [196], RBench [197], GigaWorld-0 [198], RIGVid [199], LuciBot [200], Gen2Act [201], Dreamitate [202] World-Env [203], EV AC [204], Ctrl-World [205], VideoAgent [206], VIPER [207], WorldEval [208], Genie Envisioner [209], World-Gymnast [210], DreamDojo [211] GR-1 [212], VILP [213], UV A [214], RoboEnvision [215], GEVRM [216], EnerVerse [217], LingBot-V A [218], Cosmos Policy [219],Fast-W AM [220],LeWorld- Model [221],DreamZero [222] Game & Interactive"},{"citing_arxiv_id":"2603.09030","ref_index":17,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"PlayWorld: Learning Robot World Models from Autonomous Play","primary_cat":"cs.RO","submitted_at":"2026-03-09T23:58:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"PlayWorld learns high-fidelity robot world models from unsupervised self-play, producing physically consistent video predictions that outperform models trained on human data and enabling 65% better real-world policy performance via model-based RL.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2508.13073","ref_index":199,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey","primary_cat":"cs.RO","submitted_at":"2025-08-18T16:45:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Learning from Human Videos Leverage human videos to adapt robot policies, enabling cross-domain transfer. Human-Robot Semantic Alignment [42], UniVLA [194], LAPA [195], VPDD [196], 3D-VLA [197], Humanoid-VLA [198] World Model-based VLA Integrate predictive world mod- els into VLA to model environ- ment dynamics. World-VLA [38], World4Omni [43], 3D- VLA [197], RIGVid [199], FoundationPose [200], V-JEPA 2-AC [201] TABLE 5: Representative RL approaches for VLA. ✓ stands for online and ✗ stands for offline. \"S\" stands for sparse reward, \"D\" for dense reward, \"RM\" denotes a pre-trained reward model, \"GPT\" denotes a reward given by GPT, and \"TC\" denotes a reward function based on task completion. Method RL Algo. Online Reward Formulation"}],"limit":50,"offset":0}