REVIEW 3 major objections 5 minor 46 references
Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning
T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Four learned gates can make a multi-agent LLM system more accurate than fixed ensembles while using roughly half the active compute.
desk verdict Solid systems paper: four jointly trained gates plus hot-swap calibration deliver real efficiency gains on three of four benchmarks; CoGRPO credit is the softest link but not a load-bearing collapse. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
GRADE’s coordination harness: four lightweight MLP gates plus CoGRPO, which samples a group of joint coordination rollouts, normalizes their dense rewards into a shared advantage, and applies that advantage to every gate and surviving agent that participated, without a learned critic.
What would settle it
Retrain or re-evaluate the same architecture with a method that gives distinct counterfactual credit to each gate (or that removes std-normalization and clipping) and check whether the accuracy and compute advantages over Puppeteer-style fixed ensembles disappear on MMLUPro and GPQA.
Extended reading notes
Core claim
A hierarchical multi-agent system controlled by four jointly trained gates (agent selection, hierarchy depth, selective inter-agent communication, and branch pruning), optimized with a shared group-relative advantage under CoGRPO, can outperform fixed multi-agent ensembles and monolithic baselines on standard reasoning benchmarks while activating only about 17B parameters on average and supporting runtime expert substitution via per-agent calibration.
Load-bearing premise
That giving every gate and surviving agent the same shared group-relative advantage from a rollout is stable and fair enough credit assignment for a hierarchy of discrete routing, depth, communication, and pruning decisions.
Editorial extensions
If this is right
- Per-query adaptive depth and agent count can replace always-on multi-agent stacks for many reasoning workloads, cutting active parameters roughly in half without accuracy loss on multi-domain benchmarks.
- Selective, gated communication between specialists is more useful than full broadcast; forcing all pairs to communicate hurts accuracy.
- Runtime expert swaps become practical if each backend keeps a cheap affine calibration map refreshed on a few dozen anchor queries.
- Harder tasks naturally recruit more agents and deeper hierarchy levels once the gates are trained, so difficulty-aware compute allocation can be learned rather than hand-tuned.
Reading between the lines
- The same gate stack could be attached to stronger math specialists to close the remaining AIME gap without permanently paying for those specialists on easy queries.
- Shared-advantage training may scale poorly if the number of discrete gate decisions grows much larger; sparse counterfactual baselines might then become necessary.
- Calibration maps suggest a practical pattern for any multi-model system that mixes open-weight and API backends under one learned router.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GRADE, a hierarchical multi-agent LLM system whose four learned gates (Assignment, Depth, Cross-Read, Prune) jointly decide which experts to activate, how deep a query should go, which agent pairs may communicate, and which branches to drop before combination. Training uses CoGRPO, a critic-free multi-agent adaptation of GRPO that samples groups of joint coordination rollouts and assigns a single group-relative advantage to every participating gate and surviving agent. Experts live in a hot-swappable registry with per-agent affine calibration maps so that a backend can be replaced at inference with only 64 anchor queries. Empirically, at ~17B average active parameters GRADE reports 94.8/78.2/57.4 on GSM8K/MMLUPro/GPQA, beating the strongest baseline (Puppeteer, ~28B) by 4.8 and 3.3 points on the latter two while remaining competitive on AIME-2025; ablations attribute the largest accuracy drops to hierarchy and masked cross-attention, and show that calibration is required for safe hot-swapping.
Significance. If the results hold under broader scrutiny, the work is a concrete step toward compute-adaptive multi-agent reasoning: it jointly learns depth, routing, selective communication and pruning inside one hierarchical policy, rather than fixing topology by prompt or always running a full ensemble. The Expert Registry plus per-agent calibration is a practical contribution for production systems that must swap models without full retraining. Strengths include three-seed reporting with standard deviations, leave-one-out component ablations (Table 9), RL-objective and fixed-k sweeps (Table 4), hot-swap recovery curves (Table 3, Figure 2), cross-read frequency and token-usage analyses (Figures 3–4), and a significance table against Puppeteer (Table 11). The central efficiency claim is therefore empirically well-supported even if the precise credit-assignment story remains partly architectural.
major comments (3)
- Section 2.4 (Collaborative Credit Assignment) and Eqs. (11)–(12) assign one shared group-normalized advantage Â^(i) uniformly to every gate decision and every surviving agent. Table 4 shows CoGRPO ahead of MAGRPO-obj / Shared Reward / MAPoRL-obj / MAPPO, but those variants still operate on the same joint hierarchical action space; they do not isolate whether the shared scalar is low-variance enough for the discrete Depth / Assignment / Cross-Read / Prune factors. Because Table 9 already attributes the largest accuracy drops to hierarchy and masked cross-attention, the causal link between CoGRPO’s credit assignment and the ~17B efficiency numbers in Table 1 is not fully established. A targeted ablation (e.g., per-gate counterfactual baselines on a subset of queries, or gate-wise advantage decomposition) would strengthen the claim that the joint RL recipe, not only the architecture, drives
- Active-parameter accounting (Section 3.1, Tables 1–2) defines “~17B average active parameters” as the expected total parameters evaluated across forward passes, ranging from 0.5B (depth 0) to ~32B (all six sub-agents). The comparison to Puppeteer’s fixed ~28B is central to the efficiency claim, yet the paper does not report the empirical distribution of depth / k / cross-read decisions on the evaluation sets, nor a FLOPs or wall-clock breakdown that would let a reader recompute the average under alternative definitions (e.g., counting only unique model weights resident in VRAM). Without that distribution, the half-compute claim is harder to audit and may be sensitive to the particular agent pool (three Qwen-7B copies + Phi-3-mini + Llama-3.2-3B).
- On AIME-2025 (Table 1), GRADE trails Puppeteer by 1.9 points (25.3 vs 27.2) and the authors correctly note that per-agent capacity dominates. The discussion (Section 3.3) suggests adding math-specialist backends, but the manuscript does not test whether the Depth / Assignment gates still allocate compute efficiently when the registry contains stronger specialists, nor whether CoGRPO’s cost term continues to prevent collapse onto the strongest model. A short specialist-swap experiment would clarify whether the adaptive-depth story generalizes beyond the current heterogeneous but relatively weak pool.
minor comments (5)
- Figure 1 caption and Section 2.1: managers are described as “trainable two-layer MLP” modules that “generate no tokens”; a one-sentence clarification that they only supply the grouping term in Eq. (1) would prevent readers from expecting manager-level generation.
- Eq. (6) and the surrounding text: the Cross-Read mask R is applied as an element-wise product inside the softmax argument; state explicitly whether invalid pairs are set to −∞ or to 0 before the softmax, as the two conventions differ in gradient behavior.
- Table 11 reports a pooled two-sample t-test with df=4 (three seeds). Note that multiple-comparison correction is not applied; a brief remark that the MMLUPro result remains significant under Bonferroni would help cautious readers.
- Appendix A.1 lists many free hyperparameters (λbal, λD, λR, τP, α, etc.). A short sensitivity paragraph or a pointer to which of these were tuned vs. fixed a priori would improve reproducibility.
- Typographical: “affinity” / “efficient” appear with ligature artifacts in several places; normalize to ASCII “affinity” / “efficient” for arXiv rendering consistency.
Circularity Check
No significant circularity: empirical multi-agent RL system evaluated on external public benchmarks with standard group-relative advantages.
full rationale
GRADE’s load-bearing claims are empirical accuracy and active-parameter comparisons on GSM8K, MMLUPro, GPQA, and AIME-2025 against external baselines (Table 1–2). CoGRPO (Eqs. 9–14) is a critic-free adaptation of GRPO that assigns a shared group-normalized advantage Â^(i) to participating gates and agents; the baseline is computed from the same G rollouts that produce the gradient, which is standard policy-gradient practice and does not redefine the target metrics. Rewards combine terminal correctness, margin/agent log-prob terms, and a cost penalty (Eq. 10)—none of these are fitted constants later re-presented as predictions. Ablations (Tables 4, 9) and hot-swap calibration (Table 3, Eq. 15) are leave-one-out or controlled substitutions, not self-definitional identities. Citations (GRPO, MAPoRL, MAGRPO, Puppeteer, etc.) are to independent prior work; there is no load-bearing self-citation uniqueness theorem or ansatz smuggled in as a forced derivation. The paper is self-contained against external benchmarks and does not reduce its central results to inputs by construction.
Assumptions & free parameters
free parameters (9)
- group size G =
8
- clip radius η =
0.2
- KL weight β =
0.02
- cost weight λc =
0.05
- margin / agent-align weights λm, λa =
0.10 / 0.05
- load-balance / depth / read-sparsity λbal, λD, λR =
0.01 / 0.02 / 0.01
- prune threshold τP =
0.35
- calibration momentum α and anchor count =
0.05 / 64
- model dimension d, Adam LR, training steps, batch size =
64 / 1e-4 / 200 / 48
assumptions (4)
- domain assumption Group-relative advantage (mean/std within G rollouts) is an unbiased and sufficiently low-variance baseline for hierarchical multi-gate policies without a learned critic.
- domain assumption Frozen pretrained experts plus lightweight adapters and answer heads can be coordinated by small gates trained only on the last two encoder layers and the gates themselves.
- ad hoc to paper A single shared scalar advantage can be assigned to every participating gate and agent (collaborative credit assignment) without per-gate counterfactual baselines.
- domain assumption Standard correctness + cost reward plus load-balance / depth / sparsity regularizers prevent collapse of the discrete gates.
invented entities (4)
-
Four coordination gates (Assignment, Depth, Cross-Read, Prune)
-
CoGRPO (Collaborative Group-Relative Policy Optimization)
-
Expert Registry with per-agent calibration maps κa
-
Manager modules as pure grouping signals (no token generation)
Cite this review
Pith. "Pith review of Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning." pith.science (2026). https://pith.science/paper/JJHWDBOO
@misc{pith2026260710836,
author = {Pith},
title = {Pith review of: Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/JJHWDBOO}},
note = {Machine review of arXiv:2607.10836}
}
abstract
Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We present GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning), a hierarchical multi-agent system in which four lightweight learned gates jointly govern agent selection, hierarchy depth, inter-agent communication, and branch pruning. Training uses CoGRPO (Collaborative Group-Relative Policy Optimization), a novel critic-free recipe that adapts GRPO to multi-agent hierarchies and assigns a shared advantage signal to every gate and agent that participated in a rollout. Agent models are drawn from a hot-swappable Expert Registry; per-agent calibration maps allow experts to be replaced at inference time without retraining. At $\sim$17B average active parameters, GRADE outperforms all baselines on GSM8K, MMLUPro, and GPQA, surpassing the strongest baseline by 4.8 points on MMLUPro at half the active compute. On AIME-2025, where model depth dominates, GRADE remains competitive to existing frameworks. Ablations isolate the hierarchy and masked cross-attention as the largest contributors to accuracy, and show that per-agent calibration is necessary for safe hot-swapping.
Reference graph
Works this paper leans on
-
[1]
Emergent Abilities of Large Language Models,
J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Y ogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P . Liang, J. Dean, and W. Fedus, “Emergent Abilities of Large Language Models,” Transactions on Machine Learning Research, 2022, survey Certification. [Online]. Available: https://openreview.net/forum?id=yzkSU5zdwD
2022
-
[2]
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations,
Q. Wu, G. Bansal, J. Zhang, Y . Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang, “AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations,” in First Conference on Language Modeling, 2024. [Online]. Available: https://openreview.net/forum?id=BAakY1hNKS
2024
-
[3]
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework,
S. Hong, M. Zhuge, J. Chen, X. Zheng, Y . Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Y au, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber, “MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=VtmBAGCN7o
2024
-
[4]
CAMEL: Communicative Agents for
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “CAMEL: Communicative Agents for ”Mind” Exploration of Large Language Model Society,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Asso- ciates, Inc., 2023, pp. 51 991–52 008. [Online]. Available: ...
2023
-
[5]
ChatDev: Communicative Agents for Software Development,
C. Qian, W. Liu, H. Liu, N. Chen, Y . Dang, J. Li, C. Y ang, W. Chen, Y . Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun, “ChatDev: Communicative Agents for Software Development,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W . Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Thaila...
2024
-
[6]
Improving Factuality and Reasoning in Language Models through Multiagent Debate,
Y . Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving Factuality and Reasoning in Language Models through Multiagent Debate,” in Proceedings of the 41st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkam...
2024
-
[7]
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate,
T. Liang, Z. He, W. Jiao, X. Wang, Y . Wang, R. Wang, Y . Y ang, S. Shi, and Z. Tu, “Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computati...
2024
-
[8]
Self-Refine: Iterative Refinement with Self-Feedback,
A. Madaan, N. Tandon, P . Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Y ang, S. Gupta, B. P . Majumder, K. Hermann, S. Welleck, A. Y azdanbakhsh, and P . Clark, “Self-Refine: Iterative Refinement with Self-Feedback,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Ha...
2023
Show all 46 references
-
[9]
ReAct: Synergizing Reasoning and Acting in Language Models,
S. Y ao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y . Cao, “ReAct: Synergizing Reasoning and Acting in Language Models,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=WE_vluYUL-X
2023
-
[10]
LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion,
D. Jiang, X. Ren, and B. Y . Lin, “LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , A. Rogers, J. Boyd-Graber, and N....
2023
-
[11]
Self-Consistency Improves Chain of Thought Reasoning in Language Models,
X. Wang, J. Wei, D. Schuurmans, Q. V . Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” in The Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openr...
2023
-
[12]
Mixture-of-Agents Enhances Large Language Model Capabilities,
J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou, “Mixture-of-Agents Enhances Large Language Model Capabilities,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=h0ZfDIrj7T
2025
-
[13]
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms,
S. Yuan, K. Song, J. Chen, X. Tan, D. Li, and D. Y ang, “EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Langua...
2025
-
[14]
Multi-Agent Collaboration via Evolving Orchestration,
Y . Dang, C. Qian, X. Luo, J. Fan, Z. Xie, R. Shi, W. Chen, C. Y ang, X. Che, Y . Tian, X. Xiong, L. Han, Z. Liu, and M. Sun, “Multi-Agent Collaboration via Evolving Orchestration,” in Advances in Neural Information Processing Systems , D. Belgrave, C. Zhang, H. Lin, R. Pascan...
2025
-
[15]
MAPoRL: Multi-agent post-co-training for collaborative large language models with reinforcement learning,
C. Park, S. Han, X. Guo, A. E. Ozdaglar, K. Zhang, and J.-K. Kim, “MAPoRL: Multi-agent post-co-training for collaborative large language models with reinforcement learning,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: L...
2025
-
[16]
LLM Collaboration with Multi-Agent Reinforcement Learning,
S. Liu, Z. Liang, X. Lyu, and C. Amato, “LLM Collaboration with Multi-Agent Reinforcement Learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 40, no. 38. Association for the Advancement of Artificial Intelligence, Mar. 2026, pp. 32 150–32 158. [O...
2026 doi
-
[17]
Adaptive computation time for recurrent neural networks,
A. Graves, “Adaptive computation time for recurrent neural networks,” arXiv preprint arXiv:1603.08983v6 , 2016. [Online]. Available: https://arxiv.org/abs/1603.08983v6
2016 arXiv
-
[18]
Universal Transformers,
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser, “Universal Transformers,” in International Conference on Learning Representations, 2019. [Online]. Available: https://openreview.net/forum?id=HyzdRiR9Y7
2019
-
[19]
Depth-Adaptive Transformer,
M. Elbayad, J. Gu, E. Grave, and M. Auli, “Depth-Adaptive Transformer,” in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=SJg7KhVKPH
2020
-
[20]
Confident Adaptive Language Modeling,
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V . Tran, Y . Tay, and D. Metzler, “Confident Adaptive Language Modeling,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran As...
2022
-
[21]
Mixture of depths: Dynamically allocating compute in transformer language models,
D. Raposo, S. Ritter, B. Richards, T. Lillicrap, P . C. Humphreys, and A. Santoro, “Mixture of depths: Dynamically allocating compute in transformer language models,” arXiv preprint arXiv:2404.02258v1 , 2024. [Online]. Available: https://arxiv.org/abs/2404.02258v1
2024 arXiv
-
[22]
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,
L. Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,” Transactions on Machine Learning Research , 2024. [Online]. Available: https://openreview.net/forum? id=cSimKw5p6R
2024
-
[23]
RouteLLM: Learning to Route LLMs from Preference Data,
I. Ong, A. Almahairi, V . Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to Route LLMs from Preference Data,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/f...
2025
-
[24]
Learning Multiagent Communication with Backpropagation,
S. Sukhbaatar, a. szlam, and R. Fergus, “Learning Multiagent Communication with Backpropagation,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available...
2016
-
[25]
Learning to Communicate with Deep Multi-Agent Reinforcement Learning,
J. Foerster, I. A. Assael, N. de Freitas, and S. Whiteson, “Learning to Communicate with Deep Multi-Agent Reinforcement Learning,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29, 2016. [Online]. A...
2016
-
[26]
Learning Attentional Communication for Multi-Agent Cooperation,
J. Jiang and Z. Lu, “Learning Attentional Communication for Multi-Agent Cooperation,” in Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31, 2018. [Online]. Available: https://pro...
2018
-
[27]
Individualized Controlled Continuous Communication Model for Multiagent Cooperative and Competitive Tasks,
A. Singh, T. Jain, and S. Sukhbaatar, “Individualized Controlled Continuous Communication Model for Multiagent Cooperative and Competitive Tasks,” inInternational Conference on Learning Representations, 2019. [Online]. Available: https://openreview.net/forum?id=rye7knCqK7
2019
-
[28]
TarMAC: Targeted Multi-Agent Communication,
A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, and J. Pineau, “TarMAC: Targeted Multi-Agent Communication,” in Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, ser. Proceedings of Machi...
2019
-
[29]
DeepSeekMath: Pushing the limits of mathematical reasoning in open language models,
Z. Shao, P . Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . K. Li, Y . Wu, and D. Guo, “DeepSeekMath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300v3 , 2024. [Online]. Available: https://arxiv.org/abs/2402.03300v3
2024 arXiv
-
[30]
Categorical Reparameterization with Gumbel-Softmax,
E. Jang, S. Gu, and B. Poole, “Categorical Reparameterization with Gumbel-Softmax,” in International Conference on Learning Representations, 2017. [Online]. Available: https://openreview.net/forum?id=rkE3y85ee
2017
-
[31]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P . Welinder, P . F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructi...
2022
-
[32]
Proximal policy optimization algorithms,
J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347v2, 2017. [Online]. Available: https://arxiv.org/abs/1707.06347v2
2017 arXiv
-
[33]
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,
D. Guo, D. Y ang, H. Zhang, J. Song, P . Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y . Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, ...
2025 arXiv
-
[34]
The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games,
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y . Wang, A. Bayen, and Y . WU, “The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35....
2022
-
[35]
On Calibration of Modern Neural Networks,
C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y . W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017,...
2017
-
[36]
Calibration of Pre-trained Transformers,
S. Desai and G. Durrett, “Calibration of Pre-trained Transformers,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp...
2020
-
[37]
Training verifiers to solve math word problems,
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. T worek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168v2, 2021. [Online]. Available: https://arxiv.org/abs/2110.14168v2
2021 arXiv
-
[38]
MMLU-Pro: A More Robust and Chal- lenging Multi-Task Language Understanding Benchmark,
Y . Wang, X. Ma, G. Zhang, Y . Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen, “MMLU-Pro: A More Robust and Chal- lenging Multi-Task Language Understanding Benchmark,” in Advances in Neural Information...
2024
-
[39]
GPQA: A Graduate-Level Google-Proof Q&A Benchmark,
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, and S. R. Bowman, “GPQA: A Graduate-Level Google-Proof Q&A Benchmark,” in First Conference on Language Modeling, 2024. [Online]. Available: https://openreview.net/forum?id=Ti67584b98
2024
-
[40]
Beyond benchmarks: Matharena as an evaluation platform for mathematics with LLMs,
J. Dekoninck, N. Jovanović, T. Gehrunger, K. Rögnvaldsson, I. Petrov, C. Sun, and M. Vechev, “Beyond benchmarks: Matharena as an evaluation platform for mathematics with LLMs,” in 3rd AI for Math Workshop: Toward Self-Evolving Scientific Agents at ICML 2026, 2026. [Online]. Av...
2026
-
[41]
Chain-of- Thought Prompting Elicits Reasoning in Large Language Models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V . Le, and D. Zhou, “Chain-of- Thought Prompting Elicits Reasoning in Large Language Models,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, ...
2022
-
[42]
Phi-3 Technical Report: A Highly Capable Language Model Locally on Y our Phone,
M. I. Abdin, S. Ade Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Hassan Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, A. Benhaim, M. Bilenko, J. Bjorck, S. Bubeck, M. Cai, C. C. T. Mendes, W. Chen, V . Chaudhary, P . Chopra, A. D. Giorno, G. de Rosa, M. Dixon, R. Elda...
2024
-
[43]
The Llama 3 herd of models,
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Y ang, A. Fan, A. Goyal, A. Hartshorn, A. Y ang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spatar...
2024 arXiv
-
[44]
Mixtral of experts,
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P . Stock, S. Subramanian, S. Y ang, S. Antoniak, T. L. Scao, T. Gervet,...
2024 arXiv
-
[45]
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” in International Conference on Learning Representations, 2017. [Online]. Available: https://openreview.net/forum?i...
2017
-
[46]
Switch transformers: scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, Jan. 2022. [Online]. Available: http://jmlr.org/papers/v23/21-0998.html 17 Gated...
2022
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.