M3DT combines a Decision Transformer with grouped, separately trained expert modules and a learned router, achieving better normalized scores than baselines across 10 to 160 multi-task RL tasks.
• Ant-dir (Rothfuss et al., 2018): We also use 40 tasks in this domain, each with a goal direction uniformly sampled in a two-dimensional plane
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
M3DT combines a Decision Transformer with grouped, separately trained expert modules and a learned router, achieving better normalized scores than baselines across 10 to 160 multi-task RL tasks.