Presents MT-Libero, a GPU-parallel multi-task RL benchmark in Isaac Lab, and DGPO, an on-policy method combining importance-weighted PPO with adaptive behavior cloning from demonstrations.
Shikun Liu, Edward Johns, and Andrew J
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.
TOPPO reformulates PPO with critic balancing to address gradient ill-conditioning in multi-task RL and reports stronger mean and tail performance than SAC baselines on Meta-World+ using fewer parameters and steps.
citing papers explorer
-
GPU-Parallel Multi-Task Reinforcement Learning with Demonstration Guided Policy Optimization
Presents MT-Libero, a GPU-parallel multi-task RL benchmark in Isaac Lab, and DGPO, an on-policy method combining importance-weighted PPO with adaptive behavior cloning from demonstrations.
-
Exploring Line Bundle Standard Models with Transformers
A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.
-
TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing
TOPPO reformulates PPO with critic balancing to address gradient ill-conditioning in multi-task RL and reports stronger mean and tail performance than SAC baselines on Meta-World+ using fewer parameters and steps.