Pith. sign in

Safe rlhf: Safe reinforcement learning from human feedback

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.MA 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

cs.MA · 2025-01-23 · reject · novelty 5.0

SS-MARL combines graph-neural-network message passing with a trust-region constrained update and a new multi-constraint recovery step, reporting better reward and safety trade-offs than MACPO, MAPPO, and InforMARL in cooperative navigation with up to 96 agents.

citing papers explorer

Showing 1 of 1 citing paper.

  • Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System cs.MA · 2025-01-23 · reject · none · ref 5

    SS-MARL combines graph-neural-network message passing with a trust-region constrained update and a new multi-constraint recovery step, reporting better reward and safety trade-offs than MACPO, MAPPO, and InforMARL in cooperative navigation with up to 96 agents.