REVIEW 3 major objections 4 minor 49 references
Maximum Score Routing For Mixture-of-Experts
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read MaxScore treats MoE routing as a min-cost maximum-flow problem, and claims this beats both capacity-constrained and unconstrained baselines at equal FLOPs.
desk verdict Plausible MoE routing idea, but the body is unreadable mojibake; there is nothing to audit until the authors provide a clean manuscript and code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Minimum-cost maximum-flow (MCMF) routing combined with a SoftTopk operator. The flow problem assigns tokens to expert capacity so that the maximum possible number of tokens is routed at minimum cost; the SoftTopk relaxation gives the discrete assignment a differentiable training signal. The flow constraints do the load balancing and elimination of token dropping, while SoftTopk keeps gradients flowing to the router.
What would settle it
Train one MoE with MaxScore and one with a standard top-k baseline at matched FLOPs, then evaluate both using the hard minimum-cost maximum-flow assignment at inference. If the MaxScore advantage vanishes, or if it only appears when the flow solver's overhead is excluded from the FLOP count, the central claim is falsified.
Extended reading notes
Core claim
MaxScore's core discovery is that token-to-expert routing can be formulated as a global assignment problem: a minimum-cost maximum-flow problem on a graph where tokens are sources and experts are sinks with capacities, and that this discrete assignment can be trained with a SoftTopk relaxation. The paper argues this combination eliminates the two failure modes of existing routers: capacity-saturated experts force token dropping, underutilized experts create padding waste, and unconstrained routers drift into load imbalance. By routing through maximum flow, MaxScore is claimed to keep all tokens, avoid padding, and maintain balanced expert load by construction, and the reported experiments sh
Load-bearing premise
That the loss gradients from the SoftTopk relaxation stay aligned with the hard minimum-cost maximum-flow assignment that is actually used to route tokens, so optimizing the surrogate also improves the true discrete router.
Editorial extensions
If this is right
- Capacity-limited MoE layers can be trained without token dropping or padding waste, so training and inference see the same routing behavior.
- Load balancing becomes a constraint of the assignment problem rather than a separately tuned auxiliary objective.
- The equivalent-FLOPs claim means MaxScore can replace existing top-k routers without adding a compute premium at the tested scales.
- Because every token is assigned, experts receive balanced utilization, which should improve hardware efficiency.
- Both constrained baselines that drop tokens and unconstrained baselines that risk imbalance are reportedly beaten on training loss and evaluation score.
Reading between the lines
- If the gains are real, they suggest the improvement comes from global assignment rather than per-token greedy selection; a direct test would compare MCMF routing with a top-k router that sends the same number of tokens per expert.
- One untested extension is annealing or sharpening the SoftTopk temperature during training to move the surrogate closer to the hard assignment; the paper does not report such a schedule.
- The same flow-plus-relaxation pattern could extend to other capacity-constrained allocation problems, such as expert choice, device placement, or attention sparsification, where a differentiable proxy for a discrete optimizer is needed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes MaxScore, a routing method for sparse Mixture-of-Experts models that formulates token-to-expert assignment as a minimum-cost maximum-flow problem and augments it with a SoftTopk operator. The abstract claims that this resolves limitations of iterative rerouting and optimal-transport formulations, and that it achieves lower training losses and higher evaluation scores at equivalent FLOPs relative to both capacity-constrained and unconstrained baselines. The full text provided for review is unreadable mojibake: equations, algorithm descriptions, experimental tables, and references cannot be inspected. The header garbles a second arXiv identifier (arXiv:2508.12802v1 [cs.CV]) and the bibliography consists of unresolved placeholder labels. As a result, neither the technical derivation nor the empirical evidence behind the central comparative claim can be audited.
Significance. If the central claim is true, MaxScore would be a useful contribution to MoE routing: it promises to avoid token dropping and padding inefficiency while maintaining load balance, at no additional FLOPs cost. The conceptual combination of min-cost max-flow with a differentiable top-k relaxation is plausible and worth exploring. However, the manuscript as submitted provides no auditable support for these claims. The abstract contains no dataset names, model scales, baseline identities, numerical results, or ablation details, and the body is unreadable. There are no machine-checked proofs, no reproducible code artifacts, and no derivations that can be checked. The significance is therefore entirely conditional on evidence that is not present in the submission.
major comments (3)
- [Full text (unreadable)] The complete body of the submission is mojibake. No equation, algorithm pseudocode, experimental table, or proof is readable. For instance, the header line 'arXiv:2508.12802v1 [cs.CV] 18 Aug 2025' appears inside the text, and the bibliography consists of unresolved placeholder labels. It is consequently impossible to audit the min-cost maximum-flow formulation, the SoftTopk construction, the capacity/load-balancing constraints, the training procedure, or the evaluation methodology. This is load-bearing because the paper's central claim is an empirical comparison.
- [Abstract] The claim of 'lower training losses and higher evaluation scores at equivalent FLOPs' is asserted without any numbers, dataset names, model scales, baseline identities, or ablation descriptions. The 'equivalent FLOPs' comparison is not defined: it is unclear whether the cost of solving the min-cost flow at every training step is included, and whether any load-balancing or capacity-related terms in the loss are counted. If the solver overhead or balancing constraints are excluded, the comparison would not be apples-to-apples. These details must appear in the paper, not only in a linked repository.
- [Abstract / SoftTopk alignment] The method trains with a SoftTopk operator but presumably routes by the hard min-cost maximum-flow assignment at inference. The manuscript gives no argument or evidence that the differentiable surrogate tracks the discrete objective, nor that gradients propagate correctly through the flow solver. Without such an analysis or experiment, lower training loss on the surrogate need not translate into better evaluation performance under the hard assignment. This is a second load-bearing uncertainty in the main claim.
minor comments (4)
- [Header] The header contains a second arXiv identifier, 'arXiv:2508.12802v1 [cs.CV] 18 Aug 2025', which appears to be a copy/paste error or an incorrect arXiv metadata line. This should be corrected.
- [References] The bibliography appears as unresolved placeholder labels; no complete references are readable. The submission must include a proper reference list.
- [General] The abstract points to a GitHub repository for implementation details. While a repository is useful, the paper itself must contain the full experimental configuration, hyperparameters, number of runs, variance measures, and the exact definition of 'equivalent FLOPs' to be verifiable.
- [Title] Minor typographical point: the title in the abstract block is 'Maximum Score Routing For Mixture-of-Experts' with 'For' capitalized; ensure consistent capitalization with the official title.
Circularity Check
No demonstrable circularity; the MaxScore claim is an empirical comparison and no equation or fitted-input reduction is auditable in this rendering.
full rationale
The paper's central claim is empirical and comparative: MaxScore, defined as routing modeled as a minimum-cost maximum-flow problem integrated with a SoftTopk operator, achieves lower training losses and higher evaluation scores at equivalent FLOPs versus constrained and unconstrained baselines. No equations, derivations, or fitted-parameter descriptions are readable in the provided full text, which is mojibake. I therefore cannot exhibit the specific reduction required by the circularity standard: there is no quotable equation showing that a prediction is identical by construction to an input, no fitted parameter renamed as a prediction, and no load-bearing self-citation chain. The abstract's SoftTopk-vs-hard-assignment concern and the FLOPs-accounting concern are verification or correctness risks, not demonstrated circularity. Under the hard evidence rule, the absence of readable derivation text means no circular step can be established, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption Gradients through the SoftTopk relaxation approximate the gradients of the hard min-cost maximum-flow assignment closely enough for end-to-end training to improve the true routing objective.
- domain assumption The minimum-cost maximum-flow solver's computational overhead is fairly captured in the 'equivalent FLOPs' comparison against baselines.
Cite this review
Pith. "Pith review of Maximum Score Routing For Mixture-of-Experts." pith.science (2026). https://pith.science/paper/NYYM5AKO
@misc{pith2026250812801,
author = {Pith},
title = {Pith review of: Maximum Score Routing For Mixture-of-Experts},
year = {2026},
howpublished = {\url{https://pith.science/paper/NYYM5AKO}},
note = {Machine review of arXiv:2508.12801}
}
abstract
Routing networks in sparsely activated mixture-of-experts (MoE) dynamically allocate input tokens to top-k experts through differentiable sparse transformations, enabling scalable model capacity while preserving computational efficiency. Traditional MoE networks impose an expert capacity constraint to ensure GPU-friendly computation. However, this leads to token dropping when capacity is saturated and results in low hardware efficiency due to padding in underutilized experts. Removing the capacity constraint, in turn, compromises load balancing and computational efficiency. To address these issues, we propose Maximum Score Routing ($\mathbf{MaxScore}$), a novel MoE routing paradigm that models routing as a minimum-cost maximum-flow problem and integrates a SoftTopk operator. MaxScore resolves the fundamental limitations of iterative rerouting and optimal transport formulations, achieving lower training losses and higher evaluation scores at equivalent FLOPs compared to both constrained and unconstrained baselines. Implementation details and experimental configurations can be obtained from $\href{https://github.com/dongbw18/MaxScore.git}{MaxScore}$.
Reference graph
Works this paper leans on
-
[1]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. https://arxiv.org/abs/2305.13245 Gqa: Training generalized multi-query transformer models from multi-head checkpoints . Preprint, arXiv:2305.13245
arXiv 2023
-
[2]
Richard Bellman. 1958. On a routing problem. Quarterly of applied mathematics, 16(1):87--90
work page 1958
-
[3]
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup. 2016. https://arxiv.org/abs/1511.06297 Conditional computation in neural networks for faster models . Preprint, arXiv:1511.06297
arXiv 2016
-
[4]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 Piqa: Reasoning about physical commonsense in natural language . Preprint, arXiv:1911.11641
arXiv 2019
-
[5]
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. https://arxiv.org/abs/1604.06174 Training deep nets with sublinear memory cost . Preprint, arXiv:1604.06174
arXiv 2016
-
[6]
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew Johnson, Katie Millican, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, ...
arXiv 2022
-
[7]
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://arxiv.org/abs/1905.10044 Boolq: Exploring the surprising difficulty of natural yes/no questions . Preprint, arXiv:1905.10044
arXiv 2019
-
[8]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457
arXiv 2018
Show all 49 references
-
[9]
Marco Cuturi. 2013. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in neural information processing systems, 26
2013
-
[10]
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. 2024. https://arxiv.org/abs/2401.06066 Deepseekmoe: Towards ultimate...
2024 arXiv
-
[11]
Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Yang, Haowei...
2024 arXiv
-
[12]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2024 arXiv
-
[13]
David Eigen, Marc'Aurelio Ranzato, and Ilya Sutskever. 2014. https://arxiv.org/abs/1312.4314 Learning factored representations in a deep mixture of experts . Preprint, arXiv:1312.4314
2014 arXiv
-
[14]
William Fedus, Barret Zoph, and Noam Shazeer. 2022. https://arxiv.org/abs/2101.03961 Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity . Preprint, arXiv:2101.03961
2022 arXiv
-
[15]
Lester Randolph Ford. 1956. Network flow theory. Rand Corporation Paper, Santa Monica, 1956
1956
-
[16]
Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. 2022. https://arxiv.org/abs/2211.15841 Megablocks: Efficient sparse training with mixture-of-experts . Preprint, arXiv:2211.15841
2022 arXiv
-
[17]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[18]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[19]
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zheng Leng Thai, Kaihuo Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai L...
2024 arXiv
-
[20]
Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong. 2023. https://arxiv.org/abs/2206.03382 Tutel: Adaptive mixture-of-experts at scale . Preprint, a...
2023 arXiv
-
[21]
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, Marek Cygan, and Sebastian Jaszczur. 2024. https://arxiv.org/abs/2402.07871 Scaling laws for fine-grained mixt...
2024 arXiv
-
[22]
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://arxiv.org/abs/1704.04683 Race: Large-scale reading comprehension dataset from examinations . Preprint, arXiv:1704.04683
2017 arXiv
-
[23]
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. https://arxiv.org/abs/2006.16668 Gshard: Scaling giant models with conditional computation and automatic sharding . Preprint, arXiv:2006.16668
2020 arXiv
-
[24]
Ilya Loshchilov and Frank Hutter. 2019. https://arxiv.org/abs/1711.05101 Decoupled weight decay regularization . Preprint, arXiv:1711.05101
2019 arXiv
-
[25]
André F. T. Martins and Ramón Fernandez Astudillo. 2016. https://arxiv.org/abs/1602.02068 From softmax to sparsemax: A sparse model of attention and multi-label classification . Preprint, arXiv:1602.02068
2016 arXiv
-
[26]
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. https://arxiv.org/abs/1809.02789 Can a suit of armor conduct electricity? a new dataset for open book question answering . Preprint, arXiv:1809.02789
2018 arXiv
-
[27]
Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali F...
2024 arXiv
-
[28]
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. https://arxiv.org/abs/1606.06031 The lambada dataset: Word prediction requiring a broad discourse context . Prepri...
2016 arXiv
-
[29]
Ben Peters, Vlad Niculae, and André F. T. Martins. 2019. https://arxiv.org/abs/1905.05702 Sparse sequence-to-sequence models . Preprint, arXiv:1905.05702
2019 arXiv
-
[30]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. https://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer . arXiv e-prints
2019 arXiv
-
[31]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. https://arxiv.org/abs/1910.02054 Zero: Memory optimizations toward training trillion parameter models . Preprint, arXiv:1910.02054
2020 arXiv
-
[32]
Noam Shazeer. 2020. https://arxiv.org/abs/2002.05202 Glu variants improve transformer . Preprint, arXiv:2002.05202
2020 arXiv
-
[33]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. https://arxiv.org/abs/1701.06538 Outrageously large neural networks: The sparsely-gated mixture-of-experts layer . Preprint, arXiv:1701.06538
2017 arXiv
-
[34]
Jianlin Su. 2024. https://spaces.ac.cn/archives/10373 After softmax: Finding a smooth approximation for top-k
2024
-
[35]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. https://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding . Preprint, arXiv:2104.09864
2023 arXiv
-
[36]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 Lla...
2023 arXiv
-
[37]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[38]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . CoRR, abs/1706.03762
2017 arXiv
-
[39]
Gary R Waissi. 1994. Network flows: Theory, algorithms, and applications
1994
-
[40]
Ziteng Wang, Jun Zhu, and Jianfei Chen. 2025. https://arxiv.org/abs/2412.14711 Remoe: Fully differentiable mixture-of-experts with relu routing . Preprint, arXiv:2412.14711
2025 arXiv
-
[41]
Liu, and Matt Gardner
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. https://arxiv.org/abs/1707.06209 Crowdsourcing multiple choice science questions . Preprint, arXiv:1707.06209
2017 arXiv
-
[42]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020 arXiv
-
[43]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://arxiv.org/abs/1905.07830 Hellaswag: Can a machine really finish your sentence? Preprint, arXiv:1905.07830
2019 arXiv
-
[44]
Biao Zhang and Rico Sennrich. 2019. https://arxiv.org/abs/1910.07467 Root mean square layer normalization . Preprint, arXiv:1910.07467
2019 arXiv
-
[45]
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. https://arxiv.org/abs/1810.12885 Record: Bridging the gap between human and machine commonsense reading comprehension . Preprint, arXiv:1810.12885
2018 arXiv
-
[46]
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon. 2022. https://arxiv.org/abs/2202.09368 Mixture-of-experts with expert choice routing . Preprint, arXiv:2202.09368
2022 arXiv
-
[47]
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022. https://arxiv.org/abs/2202.08906 St-moe: Designing stable and transferable sparse expert models . Preprint, arXiv:2202.08906
2022 arXiv
-
[48]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[49]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.