A single shared PPO policy with action masking and a potential reward based on the robots' centroid is trained to solve dynamic multi-robot rendezvous navigation in factory grids.
Interaction-aware multi-agent reinforcement learning for mobile agents with individual goals,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing
A single shared PPO policy with action masking and a potential reward based on the robots' centroid is trained to solve dynamic multi-robot rendezvous navigation in factory grids.