MoRSE trains role- and subtask-specific LoRA experts with a prototype router and hierarchical GRPO, improving LLM code-generation benchmarks and held-out task generalization.
Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.MA 1years
2026 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
MoRSE trains role- and subtask-specific LoRA experts with a prototype router and hierarchical GRPO, improving LLM code-generation benchmarks and held-out task generalization.