Pith. sign in

REVIEW 3 major objections 4 minor 49 references

Maximum Score Routing For Mixture-of-Experts

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read MaxScore treats MoE routing as a min-cost maximum-flow problem, and claims this beats both capacity-constrained and unconstrained baselines at equal FLOPs.

desk verdict Plausible MoE routing idea, but the body is unreadable mojibake; there is nothing to audit until the authors provide a clean manuscript and code. read the letter →

arxiv 2508.12801 v1 pith:NYYM5AKO submitted 2025-08-18 cs.LG cs.CL

classification cs.LGcs.CL
keywords mixture-of-expertstokenroutingminimum-costmaximum-flowSoftTopkloadbalancingdroppingsparseMoEdifferentiable
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes Maximum Score Routing (MaxScore), a routing scheme for sparsely activated mixture-of-experts models in which every token is assigned to experts by solving a minimum-cost maximum-flow problem, made differentiable through a SoftTopk operator. The paper's central claim is that this resolves the trade-off that forces other MoE routers either to drop tokens when experts hit capacity or to sacrifice load balance when capacity is removed. MaxScore is reported to achieve lower training losses and higher evaluation scores than both constrained and unconstrained baselines at equivalent FLOPs. If that holds, MoE training can keep hardware-friendly capacity limits without losing tokens or padding, making the routing objective and the actual assignment closer.

What carries the argument

Minimum-cost maximum-flow (MCMF) routing combined with a SoftTopk operator. The flow problem assigns tokens to expert capacity so that the maximum possible number of tokens is routed at minimum cost; the SoftTopk relaxation gives the discrete assignment a differentiable training signal. The flow constraints do the load balancing and elimination of token dropping, while SoftTopk keeps gradients flowing to the router.

What would settle it

Train one MoE with MaxScore and one with a standard top-k baseline at matched FLOPs, then evaluate both using the hard minimum-cost maximum-flow assignment at inference. If the MaxScore advantage vanishes, or if it only appears when the flow solver's overhead is excluded from the FLOP count, the central claim is falsified.

Watch

Extended reading notes

Core claim

MaxScore's core discovery is that token-to-expert routing can be formulated as a global assignment problem: a minimum-cost maximum-flow problem on a graph where tokens are sources and experts are sinks with capacities, and that this discrete assignment can be trained with a SoftTopk relaxation. The paper argues this combination eliminates the two failure modes of existing routers: capacity-saturated experts force token dropping, underutilized experts create padding waste, and unconstrained routers drift into load imbalance. By routing through maximum flow, MaxScore is claimed to keep all tokens, avoid padding, and maintain balanced expert load by construction, and the reported experiments sh

Load-bearing premise

That the loss gradients from the SoftTopk relaxation stay aligned with the hard minimum-cost maximum-flow assignment that is actually used to route tokens, so optimizing the surrogate also improves the true discrete router.

Editorial extensions

If this is right

  • Capacity-limited MoE layers can be trained without token dropping or padding waste, so training and inference see the same routing behavior.
  • Load balancing becomes a constraint of the assignment problem rather than a separately tuned auxiliary objective.
  • The equivalent-FLOPs claim means MaxScore can replace existing top-k routers without adding a compute premium at the tested scales.
  • Because every token is assigned, experts receive balanced utilization, which should improve hardware efficiency.
  • Both constrained baselines that drop tokens and unconstrained baselines that risk imbalance are reportedly beaten on training loss and evaluation score.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the gains are real, they suggest the improvement comes from global assignment rather than per-token greedy selection; a direct test would compare MCMF routing with a top-k router that sends the same number of tokens per expert.
  • One untested extension is annealing or sharpening the SoftTopk temperature during training to move the surrogate closer to the hard assignment; the paper does not report such a schedule.
  • The same flow-plus-relaxation pattern could extend to other capacity-constrained allocation problems, such as expert choice, device placement, or attention sparsification, where a differentiable proxy for a discrete optimizer is needed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript proposes MaxScore, a routing method for sparse Mixture-of-Experts models that formulates token-to-expert assignment as a minimum-cost maximum-flow problem and augments it with a SoftTopk operator. The abstract claims that this resolves limitations of iterative rerouting and optimal-transport formulations, and that it achieves lower training losses and higher evaluation scores at equivalent FLOPs relative to both capacity-constrained and unconstrained baselines. The full text provided for review is unreadable mojibake: equations, algorithm descriptions, experimental tables, and references cannot be inspected. The header garbles a second arXiv identifier (arXiv:2508.12802v1 [cs.CV]) and the bibliography consists of unresolved placeholder labels. As a result, neither the technical derivation nor the empirical evidence behind the central comparative claim can be audited.

Significance. If the central claim is true, MaxScore would be a useful contribution to MoE routing: it promises to avoid token dropping and padding inefficiency while maintaining load balance, at no additional FLOPs cost. The conceptual combination of min-cost max-flow with a differentiable top-k relaxation is plausible and worth exploring. However, the manuscript as submitted provides no auditable support for these claims. The abstract contains no dataset names, model scales, baseline identities, numerical results, or ablation details, and the body is unreadable. There are no machine-checked proofs, no reproducible code artifacts, and no derivations that can be checked. The significance is therefore entirely conditional on evidence that is not present in the submission.

major comments (3)
  1. [Full text (unreadable)] The complete body of the submission is mojibake. No equation, algorithm pseudocode, experimental table, or proof is readable. For instance, the header line 'arXiv:2508.12802v1 [cs.CV] 18 Aug 2025' appears inside the text, and the bibliography consists of unresolved placeholder labels. It is consequently impossible to audit the min-cost maximum-flow formulation, the SoftTopk construction, the capacity/load-balancing constraints, the training procedure, or the evaluation methodology. This is load-bearing because the paper's central claim is an empirical comparison.
  2. [Abstract] The claim of 'lower training losses and higher evaluation scores at equivalent FLOPs' is asserted without any numbers, dataset names, model scales, baseline identities, or ablation descriptions. The 'equivalent FLOPs' comparison is not defined: it is unclear whether the cost of solving the min-cost flow at every training step is included, and whether any load-balancing or capacity-related terms in the loss are counted. If the solver overhead or balancing constraints are excluded, the comparison would not be apples-to-apples. These details must appear in the paper, not only in a linked repository.
  3. [Abstract / SoftTopk alignment] The method trains with a SoftTopk operator but presumably routes by the hard min-cost maximum-flow assignment at inference. The manuscript gives no argument or evidence that the differentiable surrogate tracks the discrete objective, nor that gradients propagate correctly through the flow solver. Without such an analysis or experiment, lower training loss on the surrogate need not translate into better evaluation performance under the hard assignment. This is a second load-bearing uncertainty in the main claim.
minor comments (4)
  1. [Header] The header contains a second arXiv identifier, 'arXiv:2508.12802v1 [cs.CV] 18 Aug 2025', which appears to be a copy/paste error or an incorrect arXiv metadata line. This should be corrected.
  2. [References] The bibliography appears as unresolved placeholder labels; no complete references are readable. The submission must include a proper reference list.
  3. [General] The abstract points to a GitHub repository for implementation details. While a repository is useful, the paper itself must contain the full experimental configuration, hyperparameters, number of runs, variance measures, and the exact definition of 'equivalent FLOPs' to be verifiable.
  4. [Title] Minor typographical point: the title in the abstract block is 'Maximum Score Routing For Mixture-of-Experts' with 'For' capitalized; ensure consistent capitalization with the official title.

Circularity Check

0 steps flagged · score 0.0 of 10

No demonstrable circularity; the MaxScore claim is an empirical comparison and no equation or fitted-input reduction is auditable in this rendering.

full rationale

The paper's central claim is empirical and comparative: MaxScore, defined as routing modeled as a minimum-cost maximum-flow problem integrated with a SoftTopk operator, achieves lower training losses and higher evaluation scores at equivalent FLOPs versus constrained and unconstrained baselines. No equations, derivations, or fitted-parameter descriptions are readable in the provided full text, which is mojibake. I therefore cannot exhibit the specific reduction required by the circularity standard: there is no quotable equation showing that a prediction is identical by construction to an input, no fitted parameter renamed as a prediction, and no load-bearing self-citation chain. The abstract's SoftTopk-vs-hard-assignment concern and the FLOPs-accounting concern are verification or correctness risks, not demonstrated circularity. Under the hard evidence rule, the absence of readable derivation text means no circular step can be established, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities are identifiable from the abstract; the method-level hyperparameters (SoftTopk temperature, flow-cost structure, any balance terms) presumably exist in the method section but the text is unreadable, so nothing further can be extracted.

assumptions (2)
  • domain assumption Gradients through the SoftTopk relaxation approximate the gradients of the hard min-cost maximum-flow assignment closely enough for end-to-end training to improve the true routing objective.
    Invoked in the abstract via 'integrates a SoftTopk operator'; this is the core trainability assumption of the method and is unverified at abstract level.
  • domain assumption The minimum-cost maximum-flow solver's computational overhead is fairly captured in the 'equivalent FLOPs' comparison against baselines.
    The abstract claims FLOPs parity versus baselines; whether solver cost and capacity padding are accounted symmetrically cannot be checked from the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Maximum Score Routing For Mixture-of-Experts." pith.science (2026). https://pith.science/paper/NYYM5AKO

@misc{pith2026250812801,
  author       = {Pith},
  title        = {Pith review of: Maximum Score Routing For Mixture-of-Experts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NYYM5AKO}},
  note         = {Machine review of arXiv:2508.12801}
}
abstract

Routing networks in sparsely activated mixture-of-experts (MoE) dynamically allocate input tokens to top-k experts through differentiable sparse transformations, enabling scalable model capacity while preserving computational efficiency. Traditional MoE networks impose an expert capacity constraint to ensure GPU-friendly computation. However, this leads to token dropping when capacity is saturated and results in low hardware efficiency due to padding in underutilized experts. Removing the capacity constraint, in turn, compromises load balancing and computational efficiency. To address these issues, we propose Maximum Score Routing ($\mathbf{MaxScore}$), a novel MoE routing paradigm that models routing as a minimum-cost maximum-flow problem and integrates a SoftTopk operator. MaxScore resolves the fundamental limitations of iterative rerouting and optimal transport formulations, achieving lower training losses and higher evaluation scores at equivalent FLOPs compared to both constrained and unconstrained baselines. Implementation details and experimental configurations can be obtained from $\href{https://github.com/dongbw18/MaxScore.git}{MaxScore}$.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

49 extracted references · 5 canonical work pages

  1. [1]

    Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. https://arxiv.org/abs/2305.13245 Gqa: Training generalized multi-query transformer models from multi-head checkpoints . Preprint, arXiv:2305.13245

  2. [2]

    Richard Bellman. 1958. On a routing problem. Quarterly of applied mathematics, 16(1):87--90

  3. [3]

    Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup. 2016. https://arxiv.org/abs/1511.06297 Conditional computation in neural networks for faster models . Preprint, arXiv:1511.06297

  4. [4]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 Piqa: Reasoning about physical commonsense in natural language . Preprint, arXiv:1911.11641

  5. [5]

    Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. https://arxiv.org/abs/1604.06174 Training deep nets with sublinear memory cost . Preprint, arXiv:1604.06174

  6. [6]

    Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew Johnson, Katie Millican, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, ...

  7. [7]

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://arxiv.org/abs/1905.10044 Boolq: Exploring the surprising difficulty of natural yes/no questions . Preprint, arXiv:1905.10044

  8. [8]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457

Show all 49 references
  1. [9]

    Marco Cuturi. 2013. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in neural information processing systems, 26

  2. [10]

    Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. 2024. https://arxiv.org/abs/2401.06066 Deepseekmoe: Towards ultimate...

  3. [11]

    Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Yang, Haowei...

  4. [12]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  5. [13]

    David Eigen, Marc'Aurelio Ranzato, and Ilya Sutskever. 2014. https://arxiv.org/abs/1312.4314 Learning factored representations in a deep mixture of experts . Preprint, arXiv:1312.4314

  6. [14]

    William Fedus, Barret Zoph, and Noam Shazeer. 2022. https://arxiv.org/abs/2101.03961 Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity . Preprint, arXiv:2101.03961

  7. [15]

    Lester Randolph Ford. 1956. Network flow theory. Rand Corporation Paper, Santa Monica, 1956

  8. [16]

    Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. 2022. https://arxiv.org/abs/2211.15841 Megablocks: Efficient sparse training with mixture-of-experts . Preprint, arXiv:2211.15841

  9. [17]

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...

  10. [18]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  11. [19]

    Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zheng Leng Thai, Kaihuo Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai L...

  12. [20]

    Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong. 2023. https://arxiv.org/abs/2206.03382 Tutel: Adaptive mixture-of-experts at scale . Preprint, a...

  13. [21]

    Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, Marek Cygan, and Sebastian Jaszczur. 2024. https://arxiv.org/abs/2402.07871 Scaling laws for fine-grained mixt...

  14. [22]

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://arxiv.org/abs/1704.04683 Race: Large-scale reading comprehension dataset from examinations . Preprint, arXiv:1704.04683

  15. [23]

    Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. https://arxiv.org/abs/2006.16668 Gshard: Scaling giant models with conditional computation and automatic sharding . Preprint, arXiv:2006.16668

  16. [24]

    Ilya Loshchilov and Frank Hutter. 2019. https://arxiv.org/abs/1711.05101 Decoupled weight decay regularization . Preprint, arXiv:1711.05101

  17. [25]

    André F. T. Martins and Ramón Fernandez Astudillo. 2016. https://arxiv.org/abs/1602.02068 From softmax to sparsemax: A sparse model of attention and multi-label classification . Preprint, arXiv:1602.02068

  18. [26]

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. https://arxiv.org/abs/1809.02789 Can a suit of armor conduct electricity? a new dataset for open book question answering . Preprint, arXiv:1809.02789

  19. [27]

    Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi

    Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali F...

  20. [28]

    Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. https://arxiv.org/abs/1606.06031 The lambada dataset: Word prediction requiring a broad discourse context . Prepri...

  21. [29]

    Ben Peters, Vlad Niculae, and André F. T. Martins. 2019. https://arxiv.org/abs/1905.05702 Sparse sequence-to-sequence models . Preprint, arXiv:1905.05702

  22. [30]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. https://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer . arXiv e-prints

  23. [31]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. https://arxiv.org/abs/1910.02054 Zero: Memory optimizations toward training trillion parameter models . Preprint, arXiv:1910.02054

  24. [32]

    Noam Shazeer. 2020. https://arxiv.org/abs/2002.05202 Glu variants improve transformer . Preprint, arXiv:2002.05202

  25. [33]

    Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. https://arxiv.org/abs/1701.06538 Outrageously large neural networks: The sparsely-gated mixture-of-experts layer . Preprint, arXiv:1701.06538

  26. [34]

    Jianlin Su. 2024. https://spaces.ac.cn/archives/10373 After softmax: Finding a smooth approximation for top-k

  27. [35]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. https://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding . Preprint, arXiv:2104.09864

  28. [36]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 Lla...

  29. [37]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  30. [38]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . CoRR, abs/1706.03762

  31. [39]

    Gary R Waissi. 1994. Network flows: Theory, algorithms, and applications

  32. [40]

    Ziteng Wang, Jun Zhu, and Jianfei Chen. 2025. https://arxiv.org/abs/2412.14711 Remoe: Fully differentiable mixture-of-experts with relu routing . Preprint, arXiv:2412.14711

  33. [41]

    Liu, and Matt Gardner

    Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. https://arxiv.org/abs/1707.06209 Crowdsourcing multiple choice science questions . Preprint, arXiv:1707.06209

  34. [42]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  35. [43]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://arxiv.org/abs/1905.07830 Hellaswag: Can a machine really finish your sentence? Preprint, arXiv:1905.07830

  36. [44]

    Biao Zhang and Rico Sennrich. 2019. https://arxiv.org/abs/1910.07467 Root mean square layer normalization . Preprint, arXiv:1910.07467

  37. [45]

    Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. https://arxiv.org/abs/1810.12885 Record: Bridging the gap between human and machine commonsense reading comprehension . Preprint, arXiv:1810.12885

  38. [46]

    Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon. 2022. https://arxiv.org/abs/2202.09368 Mixture-of-experts with expert choice routing . Preprint, arXiv:2202.09368

  39. [47]

    Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022. https://arxiv.org/abs/2202.08906 St-moe: Designing stable and transferable sparse expert models . Preprint, arXiv:2202.08906

  40. [48]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  41. [49]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.