AgentCDM uses two-stage RL training, first with ACH reasoning scaffolding then with the scaffold gradually removed, to make a Qwen-7B decision agent outperform voting, dictatorial, and prompted-reasoning baselines on MMLU, MMLU-Pro, and ARC-Challenge.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning
AgentCDM uses two-stage RL training, first with ACH reasoning scaffolding then with the scaffold gradually removed, to make a Qwen-7B decision agent outperform voting, dictatorial, and prompted-reasoning baselines on MMLU, MMLU-Pro, and ARC-Challenge.