DEMAR combines dual ensembled Q-value targets with an L1 regularizer on mixing-network weights to curb multiagent Q-value overestimation in value-mixing Q-learning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MA 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
DEMAR combines dual ensembled Q-value targets with an L1 regularizer on mixing-network weights to curb multiagent Q-value overestimation in value-mixing Q-learning.