TAVIS is a released benchmark showing active vision improves imitation learning in a task-dependent manner, multi-task policies struggle with shifts, and imitation produces human-like anticipatory gaze.
super hub Mixed citations
Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
Mixed citation behavior. Most common role is background (57%).
abstract
We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single extensible platform. We highlight its application to a diverse set of challenges, including whole-body control, cross-embodiment mobility, contact-rich and dexterous manipulation, and the integration of human demonstrations for skill acquisition. Finally, we discuss upcoming integration with the differentiable, GPU-accelerated Newton physics engine, which promises new opportunities for scalable, data-efficient, and gradient-based approaches to robot learning. We believe Isaac Lab's combination of advanced simulation capabilities, rich sensing, and data-center scale execution will help unlock the next generation of breakthroughs in robotics research.
hub tools
citation-role summary
citation-polarity summary
claims ledger
- abstract We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single
authors
co-cited works
representative citing papers
Reinforcement learning produces a policy for passive inline skating on a humanoid robot that achieves up to 50% lower cost of transport than walking and transfers zero-shot to physical hardware.
A hybrid simulation enables an RL policy that throws cable-suspended payloads from a quadrotor with up to 50% lower landing error and 30% shorter duration than model-based baselines, including a vision-only variant.
jaxipm is the first GPU-batched IPOPT solver in JAX using heterogeneous iteration fusion and iteration-level batching, delivering up to 32.85x higher throughput than standard IPOPT on quadrotor NMPC benchmarks.
WireCraft is a new configurable simulation benchmark for industrial DLO manipulation with three task families, dual physics models, and shared evaluation of RL, IL, and VLA policies showing high success under privileged state but bottlenecks for vision-based methods.
HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.
HARBOR is a new agentic harness framework that automates robot RL workflows end-to-end across 16 tasks in manipulation, locomotion, and dexterous control, matching or exceeding default configurations while enabling sim-to-real transfer.
VoLoAgent uses a VLM to steer heterogeneous robot capabilities as interruptible tools for long-horizon manipulation and introduces the RoboVoLo benchmark, claiming substantial outperformance over single VLA/VLM or tool-based systems with real-robot validation.
CoP tactile representation with differentiable calibration enables zero-shot sim-to-real transfer and outperforms binary and raw-taxel baselines on peg-in-hole insertion and ball balancing with a multi-fingered hand.
Risk-aware domain randomization in contact-rich sampling-based predictive control reshapes the basin of attraction around contact-producing actions in the optimizer's effective cost landscape.
VUDA enables spatial sharing between CUDA and Vulkan on GPUs via channel redirection and page-table grafting, achieving up to 85% higher throughput than temporal baselines in embodied AI tasks.
An acoustic dimer of two subwavelength scatterers can achieve unidirectional transverse scattering (transverse Kerker effect) while maintaining strong overall scattering.
BRRL derives an analytic optimal policy for regularized constrained RL that guarantees monotonic improvement and yields the BPO algorithm that matches or exceeds PPO.
A model-free system uses 2D point trackers to achieve causal 6D pose tracking and incremental 3D reconstruction for multiple unseen rigid objects from RGB-D video, with recovery from complete occlusions.
Controller gains affect learnability differently for behavior cloning, RL from scratch, and sim-to-real transfer, so optimal gains depend on the learning paradigm rather than desired task behavior.
A feed-forward feature-Gaussian plus one-step geometry-aware pixel-flow simulator converts large image collections into 20K interactive scenes and 10M+ navigation samples that improve zero-shot Habitat and real-robot performance.
Finger-conserving resource-aware grasps, selected by curriculum RL, raise sequential dexterous second-subtask success over greedy stability-first baselines on HANDFUL-Bench and a real LEAP hand.
A modular benchmark of 100 dexterous manipulation tasks across 3 arms and 6 hands with 3,180 demonstrations reveals that current policies (Diffusion Policy, DP3, OpenVLA, π0.5) achieve only 34% mean success, exposing unsolved challenges in contact-rich and precise manipulation.
A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.
A proprioceptive humanoid policy trained with slope-adaptive ZMP regularization plus biomechanical reward gating traverses outdoor grass slopes to 32.1° without online exteroception.
Calf-integrated dual arms on a Go2 enable ground-level bimanual loco-manipulation with four-foot stance and VLM skill sequencing, demonstrated on three simulation tasks.
A guided VAE trained on pro StarCraft replays enables four latent-space traversal strategies to produce counterfactual improvement trajectories for amateur players.
Foot-mounted proximity sensors provide pre-contact feedback that, when integrated into RL, improves quadruped traversal robustness on discrete terrain with reliable sim-to-real transfer.
Generates 48,000 synthetic VLK trajectories in 3D-reconstructed scenes to train a policy for egocentric perception-based humanoid navigation and object transport, shown on physical Unitree G1 robot.
citing papers explorer
-
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
TAVIS is a released benchmark showing active vision improves imitation learning in a task-dependent manner, multi-task policies struggle with shifts, and imitation produces human-like anticipatory gaze.
-
Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot
Reinforcement learning produces a policy for passive inline skating on a humanoid robot that achieves up to 50% lower cost of transport than walking and transfers zero-shot to physical hardware.
-
Learning to Throw: Agile and Accurate Cable-Suspended Payload Delivery with a Quadrotor
A hybrid simulation enables an RL policy that throws cable-suspended payloads from a quadrotor with up to 50% lower landing error and 30% shorter duration than model-based baselines, including a vision-only variant.
-
Scaling Nonlinear Optimization: Many Problems One GPU
jaxipm is the first GPU-batched IPOPT solver in JAX using heterogeneous iteration fusion and iteration-level batching, delivering up to 32.85x higher throughput than standard IPOPT on quadrotor NMPC benchmarks.
-
WireCraft: A Simulation Benchmark for Industrial DLO Manipulation
WireCraft is a new configurable simulation benchmark for industrial DLO manipulation with three task families, dual physics models, and shared evaluation of RL, IL, and VLA policies showing high success under privileged state but bottlenecks for vision-based methods.
-
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.
-
HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning
HARBOR is a new agentic harness framework that automates robot RL workflows end-to-end across 16 tasks in manipulation, locomotion, and dexterous control, matching or exceeding default configurations while enabling sim-to-real transfer.
-
VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation
VoLoAgent uses a VLM to steer heterogeneous robot capabilities as interruptible tools for long-horizon manipulation and introduces the RoboVoLo benchmark, claiming substantial outperformance over single VLA/VLM or tool-based systems with real-robot validation.
-
Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation
CoP tactile representation with differentiable calibration enables zero-shot sim-to-real transfer and outperforms binary and raw-taxel baselines on peg-in-hole insertion and ball balancing with a multi-fingered hand.
-
On Surprising Effects of Risk-Aware Domain Randomization for Contact-Rich Sampling-based Predictive Control
Risk-aware domain randomization in contact-rich sampling-based predictive control reshapes the basin of attraction around contact-producing actions in the optimizer's effective cost landscape.
-
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
VUDA enables spatial sharing between CUDA and Vulkan on GPUs via channel redirection and page-table grafting, achieving up to 85% higher throughput than temporal baselines in embodied AI tasks.
-
FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing
An acoustic dimer of two subwavelength scatterers can achieve unidirectional transverse scattering (transverse Kerker effect) while maintaining strong overall scattering.
-
Bounded Ratio Reinforcement Learning
BRRL derives an analytic optimal policy for regularized constrained RL that guarantees monotonic improvement and yields the BPO algorithm that matches or exceeds PPO.
-
Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
A model-free system uses 2D point trackers to achieve causal 6D pose tracking and incremental 3D reconstruction for multiple unseen rigid objects from RGB-D video, with recovery from complete occlusions.
-
Tune to Learn: How Controller Gains Shape Robot Policy Learning
Controller gains affect learnability differently for behavior cloning, RL from scratch, and sim-to-real transfer, so optimal gains depend on the learning paradigm rather than desired task behavior.
-
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
A feed-forward feature-Gaussian plus one-step geometry-aware pixel-flow simulator converts large image collections into 20K interactive scenes and 10M+ navigation samples that improve zero-shot Habitat and real-robot performance.
-
HANDFUL: Sequential Grasp-Conditioned Dexterous Manipulation with Resource Awareness
Finger-conserving resource-aware grasps, selected by curriculum RL, raise sequential dexterous second-subtask success over greedy stability-first baselines on HANDFUL-Bench and a real LEAP hand.
-
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
A modular benchmark of 100 dexterous manipulation tasks across 3 arms and 6 hands with 3,180 demonstrations reveals that current policies (Diffusion Policy, DP3, OpenVLA, π0.5) achieve only 34% mean success, exposing unsolved challenges in contact-rich and precise manipulation.
-
ContactMimic: Humanoid Object Interaction via Contact Control
A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.
-
Physics-Guided Biomechanical Gait Adaptation for Humanoid Locomotion on Extreme Sloped Terrains
A proprioceptive humanoid policy trained with slope-adaptive ZMP regularization plus biomechanical reward gating traverses outdoor grass slopes to 32.1° without online exteroception.
-
Calf-Integrated Arms for Bimanual Quadruped Loco-Manipulation
Calf-integrated dual arms on a Go2 enable ground-level bimanual loco-manipulation with four-foot stance and VLM skill sequencing, demonstrated on three simulation tasks.
-
Play Like Champions: Counterfactual Feedback Generation in Latent Space
A guided VAE trained on pro StarCraft replays enables four latent-space traversal strategies to produce counterfactual improvement trajectories for amateur players.
-
Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing
Foot-mounted proximity sensors provide pre-contact feedback that, when integrated into RL, improves quadruped traversal robustness on discrete terrain with reliable sim-to-real transfer.
-
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Generates 48,000 synthetic VLK trajectories in 3D-reconstructed scenes to train a policy for egocentric perception-based humanoid navigation and object transport, shown on physical Unitree G1 robot.
-
Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field
A graspability field learned from synthesized grasps provides a dense reward signal for an RL policy that performs closed-loop non-prehensile manipulation leading to successful grasps.
-
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
An auto-regressive diffusion planner trained with scheduled prefix sampling, coupled asynchronously to a pretrained universal tracker, enables closed-loop humanoid whole-body control with replanning under disturbances and zero-shot moving-target reaching.
-
AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance
AnyBody distills a privileged teacher tracker into a latent unit-sphere representation and uses a masked transformer to drive humanoid control from arbitrary keypoint subsets.
-
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
An automated real-to-sim pipeline builds digital twins and affordance-preserving cousins from video, yielding sim evaluations that correlate with real robot policy success and zero-shot sim-to-real gains.
-
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
SceneBot conditions a humanoid tracking policy on motion references and contact labels, using reconstructed scene-interaction data to unify free-space locomotion with contact-rich manipulation and terrain tasks.
-
OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation
OmniContact introduces contact flow as a shared representation of body trajectories and contact signals to learn and chain loco-manipulation meta-skills, reporting 98.7% success on box carrying and 76.5% on push-stack tasks.
-
CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation
CoorDex distills privileged body and hand motion teachers into proprioceptive latent priors and composes them via shared-context residual RL heads to enable continuous high-DoF dexterous loco-manipulation.
-
Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
Wh0 generates scalable egocentric human manipulation videos with world models and converts them to boost pretrained VLA models' zero-shot dexterous task success from 8.3% to 38.9% on 18 real-world tasks.
-
Rotation-Aware Point-Cloud Embeddings for Vision-Based In-Hand Reorientation
Learns rotation-aware point-cloud embeddings calibrated to SO(3) geodesic error, enabling model-free RL for vision-based in-hand reorientation without pose or flow inputs.
-
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
FAST uses termination-rate-triggered virtual continuation plus masked normalized PPO loss to cut parallel RL sampling latency by ≥1.78× without biasing autonomous-driving policies.
-
Inductive Generalization for Robotic Manipulation
The paper introduces an inductive generalization evaluation protocol for manipulation policies and shows that SOTA vision-language-action models fail on progressively harder task variants.
-
CRAX: Fast Safe Reinforcement Learning Benchmarking
CRAX is a new fast benchmark suite for constrained RL built on MJX, with six environment suites and tasks across difficulty levels, showing no single safe RL method dominates and benefits from curriculum learning.
-
VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents
VOiLA combines distilled diffusion-model POMDP samplers with the vectorized VOPP planner, achieving data-efficient planning, generalization to unseen layouts, and simulation-to-real transfer on a quadruped.
-
TactSpace: Learning a Physics-enriched Shared Latent Space for Tactile Sim-to-Real Transfer
A multi-modal learning framework aligns simulated and real tactile data in a shared latent space for zero-shot sim-to-real transfer, with reported gains from multi-physics modalities and a released simulation implementation.
-
APT: Atomic Physical Transitions for Causal Video-Language Understanding
Introduces APT chains as ordered causal transition sequences and APT-Tune to improve VLM transition detection while preserving event-level performance.
-
AnnotateAnything: Automatic Annotation of 3D Assets for Robot Manipulation
AnnotateAnything converts passive 3D assets into manipulation-ready assets by combining vision-language reasoning for semantics with parallel physics pipelines for executable action annotations such as grasps and articulations.
-
Mana: Dexterous Manipulation of Articulated Tools
Mana framework achieves zero-shot sim-to-real transfer for grasping and in-hand manipulation of four articulated tools using a coarse-to-fine animation-inspired pipeline.
-
From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation
FlowPilot combines anchored flow matching for multimodal action pre-training with human-in-the-loop preference learning to improve long-horizon monocular sidewalk navigation, reporting 42% success in simulation and reduced interruptions in real-world tests.
-
SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation
SIMPLE is a new large-scale simulation benchmark for humanoid loco-manipulation that integrates accurate dynamics and photorealistic rendering and demonstrates policy transfer from simulation to physical robots.
-
Continual Quadruped Robots Coordination via Semantic Skill Discovery
Conquer is a retrieve-adapt-update semantic skill-library framework with a SAG backbone for continual multi-quadruped coordination, reporting 95.6% average success in simulations and real-world validation.
-
Rapid co-design of Buoyancy-assisted robots for Challenging Locomotion using Gaussian Evolutionary Specialists
GES framework uses Gaussian-partitioned specialist policies to co-optimize morphology and control for buoyancy-assisted legged robots, reporting 5-25% performance gains, 3x hardware obstacle improvement, and 37% faster design search versus baselines.
-
TAGA: Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion
TAGA learns terrain-aware active gaze behaviors for humanoid robots via RL alone, enabling generalizable locomotion with 1.2m real-world gap traversal.
-
HORIZON: Recoverability-Governed Curriculum for Physical-Domain Scaling
HORIZON is a recoverability-governed checkpointed frontier curriculum for on-policy physical-domain scaling on quadruped locomotion that identifies three regularities: uneven widening, non-monotonic composition, and the necessity of joint on-policy interaction.
-
S2M-Trek: From Single to Multi-Sphere Transport via Per-Frame Deep Sets on a Wheel-Legged Robot
Per-Frame Deep Sets enables scaling single-sphere to five-sphere transport on a quadruped by performing permutation-invariant pooling within each history frame, reaching 100% no-drop success in simulation where standard encoders plateau.
-
Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
Dynamic scene graphs serve as explicit memory to improve imitation learning policies for spatial-temporal reasoning under partial observability in mobile and tabletop manipulation.
-
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
UniLab is a CPU/GPU heterogeneous system for robot RL training using MuJoCoUni and MotrixSim backends that reports 3-10x end-to-end efficiency improvements and cross-platform compatibility beyond CUDA.