Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.
Christiano, Jan Leike, Tom B
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
Logit averaging inside GRPO yields higher or comparable benchmark accuracy to KL-regularized GRPO without using KL terms or a critic.
citing papers explorer
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.
-
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
Logit averaging inside GRPO yields higher or comparable benchmark accuracy to KL-regularized GRPO without using KL terms or a critic.