{"work":{"id":"59a1bdfa-635f-4266-a860-c30157fa91f2","openalex_id":"https://openalex.org/W7104607800","doi":"10.48550/arxiv.2511.04831","arxiv_id":"2511.04831","raw_key":null,"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","authors":null,"authors_text":"NVIDIA: Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du","year":2025,"venue":"cs.RO","abstract":"We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single extensible platform. We highlight its application to a diverse set of challenges, including whole-body control, cross-embodiment mobility, contact-rich and dexterous manipulation, and the integration of human demonstrations for skill acquisition. Finally, we discuss upcoming integration with the differentiable, GPU-accelerated Newton physics engine, which promises new opportunities for scalable, data-efficient, and gradient-based approaches to robot learning. We believe Isaac Lab's combination of advanced simulation capabilities, rich sensing, and data-center scale execution will help unlock the next generation of breakthroughs in robotics research.","external_url":"https://arxiv.org/abs/2511.04831","cited_by_count":2,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2511.04831","created_at":"2026-05-09T05:50:26.025143+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","render_title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning"},"hub":{"state":{"work_id":"59a1bdfa-635f-4266-a860-c30157fa91f2","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":136,"external_cited_by_count":2,"distinct_field_count":8,"first_pith_cited_at":"2025-05-25T09:17:22+00:00","last_pith_cited_at":"2026-07-09T17:50:47+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T05:09:22.940109+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":7},{"context_role":"method","n":4},{"context_role":"baseline","n":2},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":8},{"context_polarity":"use_method","n":3},{"context_polarity":"baseline","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","claims":[{"claim_text":"We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"whereq des and˙q des denote the desired joint position and velocity setpoints, respectively. Domain Randomization:To reduce the sim-to-real gap, we apply domain randomization [36], [37] in both driving and stunt modes. The full set of randomized parameters is given in Table II. VI. EXPERIMENTALRESULTS We train our policy in the high-throughput, GPU-based simulator IsaacLab [38], and then deploy it on real hardware for evaluation. Depending on the task, training takes 12-24 hours on an NVIDIA L40","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"We evaluateFlashSACon a diverse suite of robotic control tasks, measuring both asymptotic performance and wall-clock time (measured on a single RTX 5090 GPU). Our experiments span low- and high-dimensional state-based control, vision-based control, and sim-to-real humanoid locomotion. 5.1 State-Based RL on GPU-based Simulators Experimental Setup.We evaluate on 25 state-based control tasks drawn from four GPU-based simulators: IsaacLab [54], MuJoCo Playground [93], ManiSkill3 [81], and Genesis [4","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"velocity, joint states, contact indicators, and etc), andz t is the predicted explicit stair geometry representation. The policy outputs actiona t ∈ Aexecuted by a whole- body PD controller. The policyπ θ(at|ot)is optimized with PPO to maximize the expected discounted return: max θ Eπθ \" TX t=0 γtrt # ,(2) wherer t follows the default IsaacLab rough-locomotion weighted reward [17]. B. Explicit Stair Geometry Representation 1) Point Cloud Input:The raw perception input consists of local point clo","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"for isolated execution, we leverage the reasoning capabilities of vision-language models to synthesize and refine multi-agent actions collaboratively. 2.3 Simulation-based Robot Learning Simulation platforms, such as Isaac Sim [40] and MuJoCo [55], constitute essen- tial infrastructure for robotic learning, facilitating the safe exploration of com- plex control strategies without the risk of physical damage to hardware [41,43]. Conventional methodologies typically involve training control polici","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"TABLE I:Comparison of physical and perceptual capabilities across parallel robotics simulators. Simulators Physics Engine Batch Physics VRAM Usage Integrated Batch IK Batch Renderer Batch Render Fidelity 3DGS Env. Num. Dynamic 3DGS Scene 3DGS Render FPS Startup Speed Physics Cross Platform MuJoCo/MJX [46] Brax/MJX CPU/GPU⋆⋆×Madrona + - - - + L IsaacLab [39] PhysX5 GPU⋆ ⋆ ⋆ ⋆ ⋆✓omni.RTX ++ - - - ++ L ManiSkill [45] PhysX5 GPU⋆ ⋆ ⋆✓Vulkan SBR + - - - +++ W/L Genesis [61] Taichi GPU⋆⋆✓Madrona + - -","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"lower are the lower-body joint positions and velocities, anda t−1 lower is the previous lower-body action. The action outputq lower ∈R 15 is a 15-dimensional vector of target joint positions:2×6for the two legs and3for the waist motors. We adopt a teacher-student framework to train the lower- body policy. The teacher policy is first trained in simulation using PPO [53] with access to privileged information, and is subsequently distilled into a student policy via DAgger [54]. The student policy o","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (7 contexts).","role_counts":[{"n":7,"context_role":"background"},{"n":4,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T21:33:00.252125+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"cba9b0f9-e61e-43c7-a3ae-1cf07ccffeae","orcid":null,"display_name":"NVIDIA: Mayank Mittal"},{"id":"0732e2f2-1342-41d9-9ac2-a4d93fb1287e","orcid":null,"display_name":"Pascal Roth"},{"id":"ca928b39-525a-4810-b137-cf48e1cd66ab","orcid":null,"display_name":"James Tigue"},{"id":"4ca33c47-b4c9-4ae3-9017-9d62080f89ec","orcid":null,"display_name":"Antoine Richard"},{"id":"1cede019-1c72-4c96-a03a-0d8934eee352","orcid":null,"display_name":"Octi Zhang"},{"id":"f43f789b-33a3-4555-841d-325ed3164def","orcid":null,"display_name":"Peter Du"}]},"error":null,"updated_at":"2026-07-02T21:33:00.248661+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:46:19.534348+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":18},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":9},{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning","work_id":"a21210c8-5b8f-429a-accc-fb4ca1efd19d","shared_citers":6},{"title":"robosuite: A Modular Simulation Framework and Benchmark for Robot Learning","work_id":"d616d4ba-7713-4e3e-8c9e-dfebbb8f1abf","shared_citers":6},{"title":"ManiSkill3: GPU parallelized robotics simulation and rendering for generaliz- able embodied AI","work_id":"fd93f487-a562-4878-b9a6-d2fe843f20b4","shared_citers":5},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":5},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":5},{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":5},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":4},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":4},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","work_id":"6fe159e0-fa73-481a-88d4-4719c15140be","shared_citers":4},{"title":"Re3 Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation","work_id":"863259aa-bf61-4d7c-9998-47a69ffce75f","shared_citers":4},{"title":"RSL-RL: A learning library for robotics research","work_id":"715d23f4-4767-4719-956d-ce4e005f53c5","shared_citers":4},{"title":"Solving rubik’s cube with a robot hand","work_id":"81bf9cee-6de8-49f6-b967-3cb853e5ba67","shared_citers":4},{"title":"Available: https://arxiv.org/abs/2502.01143","work_id":"6bc3c45b-e07e-438a-aecb-4b7d8907a2c6","shared_citers":3},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":3},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","work_id":"13253de2-3d89-415c-8c2f-3adb25d4c337","shared_citers":3},{"title":"Gemini Robotics: Bringing AI into the Physical World","work_id":"f7c5ce10-8364-4fbe-964f-2802b81c3a98","shared_citers":3},{"title":"Karen Liu, Jiajun Wu, and Li Fei-Fei","work_id":"6b62f6a4-fd69-4dd7-b480-b27f0bf09632","shared_citers":3},{"title":"Much ado about noising: Dispelling the myths of generative robotic control","work_id":"cab94379-b09c-4084-8de9-d49329d27cd0","shared_citers":3},{"title":"Mujoco playground","work_id":"0325a224-ac67-49fb-a451-9fb130f682c9","shared_citers":3},{"title":"RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots","work_id":"11232b35-bd17-402a-9234-951c46015815","shared_citers":3},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":3},{"title":"World Action Models are Zero-shot Policies","work_id":"9a85fc69-74df-450e-94cd-69d186e9e830","shared_citers":3}],"time_series":[{"n":40,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:49:11.710245+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:46:25.044200+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","claims":[{"claim_text":"We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"whereq des and˙q des denote the desired joint position and velocity setpoints, respectively. Domain Randomization:To reduce the sim-to-real gap, we apply domain randomization [36], [37] in both driving and stunt modes. The full set of randomized parameters is given in Table II. VI. EXPERIMENTALRESULTS We train our policy in the high-throughput, GPU-based simulator IsaacLab [38], and then deploy it on real hardware for evaluation. Depending on the task, training takes 12-24 hours on an NVIDIA L40","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"We evaluateFlashSACon a diverse suite of robotic control tasks, measuring both asymptotic performance and wall-clock time (measured on a single RTX 5090 GPU). Our experiments span low- and high-dimensional state-based control, vision-based control, and sim-to-real humanoid locomotion. 5.1 State-Based RL on GPU-based Simulators Experimental Setup.We evaluate on 25 state-based control tasks drawn from four GPU-based simulators: IsaacLab [54], MuJoCo Playground [93], ManiSkill3 [81], and Genesis [4","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"velocity, joint states, contact indicators, and etc), andz t is the predicted explicit stair geometry representation. The policy outputs actiona t ∈ Aexecuted by a whole- body PD controller. The policyπ θ(at|ot)is optimized with PPO to maximize the expected discounted return: max θ Eπθ \" TX t=0 γtrt # ,(2) wherer t follows the default IsaacLab rough-locomotion weighted reward [17]. B. Explicit Stair Geometry Representation 1) Point Cloud Input:The raw perception input consists of local point clo","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"for isolated execution, we leverage the reasoning capabilities of vision-language models to synthesize and refine multi-agent actions collaboratively. 2.3 Simulation-based Robot Learning Simulation platforms, such as Isaac Sim [40] and MuJoCo [55], constitute essen- tial infrastructure for robotic learning, facilitating the safe exploration of com- plex control strategies without the risk of physical damage to hardware [41,43]. Conventional methodologies typically involve training control polici","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"TABLE I:Comparison of physical and perceptual capabilities across parallel robotics simulators. Simulators Physics Engine Batch Physics VRAM Usage Integrated Batch IK Batch Renderer Batch Render Fidelity 3DGS Env. Num. Dynamic 3DGS Scene 3DGS Render FPS Startup Speed Physics Cross Platform MuJoCo/MJX [46] Brax/MJX CPU/GPU⋆⋆×Madrona + - - - + L IsaacLab [39] PhysX5 GPU⋆ ⋆ ⋆ ⋆ ⋆✓omni.RTX ++ - - - ++ L ManiSkill [45] PhysX5 GPU⋆ ⋆ ⋆✓Vulkan SBR + - - - +++ W/L Genesis [61] Taichi GPU⋆⋆✓Madrona + - -","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"lower are the lower-body joint positions and velocities, anda t−1 lower is the previous lower-body action. The action outputq lower ∈R 15 is a 15-dimensional vector of target joint positions:2×6for the two legs and3for the waist motors. We adopt a teacher-student framework to train the lower- body policy. The teacher policy is first trained in simulation using PPO [53] with access to privileged information, and is subsequently distilled into a student policy via DAgger [54]. The student policy o","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (7 contexts).","role_counts":[{"n":7,"context_role":"background"},{"n":4,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T21:33:00.254829+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","claims":[{"claim_text":"We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:48:32.268268+00:00"}},"summary":{"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning","claims":[{"claim_text":"We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":18},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":9},{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning","work_id":"a21210c8-5b8f-429a-accc-fb4ca1efd19d","shared_citers":6},{"title":"robosuite: A Modular Simulation Framework and Benchmark for Robot Learning","work_id":"d616d4ba-7713-4e3e-8c9e-dfebbb8f1abf","shared_citers":6},{"title":"ManiSkill3: GPU parallelized robotics simulation and rendering for generaliz- able embodied AI","work_id":"fd93f487-a562-4878-b9a6-d2fe843f20b4","shared_citers":5},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":5},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":5},{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":5},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":4},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":4},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","work_id":"6fe159e0-fa73-481a-88d4-4719c15140be","shared_citers":4},{"title":"Re3 Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation","work_id":"863259aa-bf61-4d7c-9998-47a69ffce75f","shared_citers":4},{"title":"RSL-RL: A learning library for robotics research","work_id":"715d23f4-4767-4719-956d-ce4e005f53c5","shared_citers":4},{"title":"Solving rubik’s cube with a robot hand","work_id":"81bf9cee-6de8-49f6-b967-3cb853e5ba67","shared_citers":4},{"title":"Available: https://arxiv.org/abs/2502.01143","work_id":"6bc3c45b-e07e-438a-aecb-4b7d8907a2c6","shared_citers":3},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":3},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","work_id":"13253de2-3d89-415c-8c2f-3adb25d4c337","shared_citers":3},{"title":"Gemini Robotics: Bringing AI into the Physical World","work_id":"f7c5ce10-8364-4fbe-964f-2802b81c3a98","shared_citers":3},{"title":"Karen Liu, Jiajun Wu, and Li Fei-Fei","work_id":"6b62f6a4-fd69-4dd7-b480-b27f0bf09632","shared_citers":3},{"title":"Much ado about noising: Dispelling the myths of generative robotic control","work_id":"cab94379-b09c-4084-8de9-d49329d27cd0","shared_citers":3},{"title":"Mujoco playground","work_id":"0325a224-ac67-49fb-a451-9fb130f682c9","shared_citers":3},{"title":"RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots","work_id":"11232b35-bd17-402a-9234-951c46015815","shared_citers":3},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":3},{"title":"World Action Models are Zero-shot Policies","work_id":"9a85fc69-74df-450e-94cd-69d186e9e830","shared_citers":3}],"time_series":[{"n":40,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"4ca33c47-b4c9-4ae3-9017-9d62080f89ec","orcid":null,"display_name":"Antoine Richard","source":"manual","import_confidence":0.72},{"id":"ca928b39-525a-4810-b137-cf48e1cd66ab","orcid":null,"display_name":"James Tigue","source":"manual","import_confidence":0.72},{"id":"cba9b0f9-e61e-43c7-a3ae-1cf07ccffeae","orcid":null,"display_name":"NVIDIA: Mayank Mittal","source":"manual","import_confidence":0.72},{"id":"1cede019-1c72-4c96-a03a-0d8934eee352","orcid":null,"display_name":"Octi Zhang","source":"manual","import_confidence":0.72},{"id":"0732e2f2-1342-41d9-9ac2-a4d93fb1287e","orcid":null,"display_name":"Pascal Roth","source":"manual","import_confidence":0.72},{"id":"f43f789b-33a3-4555-841d-325ed3164def","orcid":null,"display_name":"Peter Du","source":"manual","import_confidence":0.72}]}}