A multi-agent RL high-level planner outputs task-space velocities that a GPU-parallel QP low-level controller converts to joint velocities while enforcing limits and collisions, yielding robust sim-to-real dexterous grasping with zero-shot steerability.
Counterfactual multi-agent policy gradients
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
CONDITIONAL 2representative citing papers
A shared consensus vector, generated before any action, lets cooperative agents act simultaneously and lets the whole joint policy be trained with single-agent PPO.
citing papers explorer
-
Learning Reactive Dexterous Grasping via Hierarchical Task-Space RL Planning and Joint-Space QP Control
A multi-agent RL high-level planner outputs task-space velocities that a GPU-parallel QP low-level controller converts to joint velocities while enforcing limits and collisions, yielding robust sim-to-real dexterous grasping with zero-shot steerability.
-
Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
A shared consensus vector, generated before any action, lets cooperative agents act simultaneously and lets the whole joint policy be trained with single-agent PPO.