VLA-RL applies online RL to pretrained VLAs, yielding a 4.5% gain over strong baselines on 40 LIBERO manipulation tasks and matching commercial models like π₀-FAST.
Bootstrap- ping reinforcement learning with imitation 10 for vision-based agile flight
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
RL framework for agile drone racing combines task-aware switching and physically informed procedural track generation to achieve 7.4x better zero-shot generalization to unseen tracks while maintaining competitive speeds.
PerchRL applies two-stage RL with randomized trajectories, temporal augmentation, and visibility-aware rewards to achieve vision-based perching on irregularly moving inclined platforms.
Unsupervised behavioral mode discovery combined with mutual information rewards enables RL fine-tuning of multimodal generative policies that achieves higher success rates without losing action diversity.
SCAL aligns source and target latent features conditioned on system state, reducing target imitation loss to a source loss plus a conditional-KL term, and reports strong sample efficiency in BARC-CARLA.
citing papers explorer
-
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
VLA-RL applies online RL to pretrained VLAs, yielding a 4.5% gain over strong baselines on 40 LIBERO manipulation tasks and matching commercial models like π₀-FAST.
-
Bridging Performance and Generalization in Reinforcement Learning for Agile Flight
RL framework for agile drone racing combines task-aware switching and physically informed procedural track generation to achieve 7.4x better zero-shot generalization to unseen tracks while maintaining competitive speeds.
-
PerchRL: Vision-Based Agile Perching on Inclined Platforms under Rapid and Irregular Motion
PerchRL applies two-stage RL with randomized trajectories, temporal augmentation, and visibility-aware rewards to achieve vision-based perching on irregularly moving inclined platforms.
-
Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
Unsupervised behavioral mode discovery combined with mutual information rewards enables RL fine-tuning of multimodal generative policies that achieves higher success rates without losing action diversity.
-
State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning
SCAL aligns source and target latent features conditioned on system state, reducing target imitation loss to a source loss plus a conditional-KL term, and reports strong sample efficiency in BARC-CARLA.