REVIEW 3 major objections 6 minor 1 cited by
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Sigmoid self-attention needs fewer samples than softmax attention.
desk verdict The sigmoid-gating MoE analysis is a real technical contribution, but the headline claim about self-attention sample complexity does not follow from it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the representation of one row of the attention matrix as a mixture of experts with quadratic affinity scores: $[\mathrm{SigmoidAttn}(X)]_{i,:} = \sum_j \sigma(x_i B x_j^\top)\, x_j W_V$, with $B = W_Q W_K^\top / \sqrt{d_k}$. This rewrites learning attention weights as estimating the parameters of a sigmoid-gating mixture-of-experts regression. The proofs are carried by Voronoi loss functions, which measure parameter discrepancies cell by cell, together with strong and weak identifiability conditions expressed as linear independence of partial derivatives; those conditions determine whether Taylor-expanded parameter differences can be separated. The sigmoid's element-wise, unnormalized structure removes the softmax normalization constraint and, in the dense regime, lets first-order Taylor terms dominate, yielding the fast $O_P((\log n/n)^{1/2})$ expert rate.
What would settle it
Fit a sigmoid-gated mixture of experts with polynomial experts to synthetic data from the same model in the dense over-specified regime and plot the Voronoi loss against $n$; the paper predicts decay of order $n^{-1/2}$, so observing decay of order $n^{-1/4}$ or slower, the softmax baseline rate, would refute the central claim.
Extended reading notes
Core claim
The central claim is a concrete sample-complexity separation between sigmoid and softmax attention. Using the equivalence that each row of the attention output is a mixture of experts, with gate $\sigma(x_i B x_j^\top)$ for sigmoid and value rows $x_j W_V$ as experts, the authors analyze the sigmoid-gating mixture-of-experts regression model. They show that in the dense regime, weakly identifiable experts such as ReLU, GELU, and polynomial experts are estimated at rate $O_P((\log n/n)^{1/2})$, so only $O(\epsilon^{-2})$ samples are needed for approximation error $\epsilon$. The comparable softmax analysis from the baseline they compare against yields $O(\epsilon^{-4})$ for strongly identifiable experts and exponential $O(\exp(\epsilon^{-1/\tau}))$ for polynomial experts. The paper therefore claims sigmoid attention is more sample-efficient than softmax attention in the dense regime and equally efficient in the sparse regime.
Load-bearing premise
The analysis assumes the sample complexity of estimating a sigmoid-gated mixture-of-experts regression transfers directly to the sample complexity of learning self-attention parameters in a Transformer, even though in attention the experts are random input tokens shared across rows and the fitted parameters are $W_Q$, $W_K$, and $W_V$.
Editorial extensions
If this is right
- In the dense regime, sigmoid self-attention reaches the same approximation error as softmax with quadratically fewer samples, $O(\epsilon^{-2})$ versus $O(\epsilon^{-4})$.
- Polynomial experts, which are exponentially hard under softmax attention, become polynomially easy under sigmoid attention in the dense regime.
- Under the sparse regime, sigmoid attention is not worse than softmax: both require $O(\epsilon^{-4})$ samples for strongly identifiable experts.
- The same separation holds under partially quadratic affinity scores, where linear experts also move from exponential to polynomial sample complexity.
- The result provides a statistical justification for the empirical success of sigmoid attention: removing token competition also removes a statistical bottleneck in the dense regime.
Reading between the lines
- If the transfer assumption holds, the paper implies sigmoid attention is preferable in small-sample settings such as few-shot learning or low-resource modeling, where sample efficiency matters more than raw capacity.
- The mixture-of-experts representation suggests a testable architectural prediction: attention heads whose value projections are well approximated by low-degree polynomials should benefit most from switching to sigmoid gating.
- A rigorous extension to multi-head attention via hierarchical mixtures of experts, which the authors flag as future work, would likely preserve the dense-regime separation if the hierarchy inherits weak identifiability.
- Because input-dependent gating weights are typical in trained models, the paper's argument predicts that practical gains from sigmoid attention should be widespread rather than confined to specially constructed cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to prove that sigmoid self-attention is more sample-efficient than softmax self-attention. It does so by representing each row of a self-attention matrix as a mixture of experts (MoE) with quadratic affinity scores, then analyzing the sample complexity of estimating the parameters of a sigmoid-gating MoE regression model (Eq. 5). The authors derive regression-function convergence rates under sparse and dense regimes for the gating parameters, translate these into parameter and expert convergence rates via Voronoi losses, and compare them with rates imported from a prior softmax-MoE analysis [1]. In the dense regime they claim a polynomial O(epsilon^-2) sample complexity for sigmoid-gating experts, versus O(epsilon^-4) or exponential rates for softmax, concluding that sigmoid self-attention is more sample-efficient. The paper also presents numerical experiments on the MoE regression models.
Significance. If the central claim were established, this would be a notable theoretical result giving a rigorous statistical justification for the empirically observed advantages of sigmoid self-attention. The technical machinery developed for sigmoid-gating MoE convergence—bracketing-entropy regression bounds, Voronoi-loss lower bounds, weak identifiability conditions, and a minimax lower bound—contains interesting components that could be of independent value. However, the paper does not provide a theorem connecting the MoE regression rates to the estimation of actual self-attention parameters, and the dense-regime analysis is carried out against a proxy target rather than the ground-truth experts. As a result, the headline claim about sigmoid versus softmax self-attention is not supported by the results as they stand.
major comments (3)
- [Section 2 and Section 4] No formal result connects the sample complexity of the MoE regression model in Eq. (5) to the sample complexity of learning self-attention parameters. In Eq. (5) the unknown quantities are fixed expert and gating parameters estimated from i.i.d. pairs (X_i, Y_i), whereas in the self-attention representation of Section 2 the 'experts' are the input token projections x_j W_V, which are random and shared across rows, and the learnable parameters are W_Q, W_K, W_V. The algebraic rewriting of one attention row as an MoE does not map one estimation problem onto the other, and no theorem in the paper states that the rates for the regression model transfer to attention parameter estimation. Therefore the abstract claim that 'sigmoid self-attention has lower sample complexity than softmax self-attention' does not follow from the MoE analysis.
- [Section 4.2, Theorem 3, Corollary 1, and Proposition 3] The dense-regime rate is for an over-parameterized proxy, not for the true experts. Corollary 1 states that inf_{G in M_N(Theta)\M_{N*}(Theta)} ||f_{\hat G_n} - f_G||_{L2(mu)} = O_P(sqrt(log n/n)), and Theorem 3 bounds L3(\hat G_n, \bar G) where \bar G is the minimizer of ||f_G - f_{G*}|| over that same excluded set. The text then concludes that 'it takes those experts only a polynomial number of samples O(epsilon^-2) to achieve an approximation error of epsilon'. This conclusion is not established because the error is measured against \bar G, not against the ground-truth expert parameters G*. In fact, Proposition 3 in Appendix B.5 shows that a dense over-specified sum of sigmoid gates cannot converge to the true single-sigmoid gating function, so there is no reason to expect \bar eta_i = eta*_i or \bar f_G = f_{G*}. The softmax rates imported from [1] are for estimation of the ground-truth experts, so the comparison in Table 1 compares distances to two different targets and does not support the claimed sample-efficiency advantage.
- [Section 4.1.2, Theorem 2] The exponential sample-complexity claim for polynomial experts is not supported by Theorem 2. That theorem gives inf_{\hat G_n} sup_{G} E[L_{2,r}(\hat G_n, G)] \gtrsim n^{-1/2} for every r \ge 1. This is a minimax lower bound on the r-th-power Voronoi loss, and it implies at most that the corresponding parameter discrepancies cannot be estimated faster than a polynomial rate of order n^{-1/(2r)} for fixed r. The subsequent text claims that the parameter convergence rates are 'slower than any polynomial rates O_P(n^{-1/2r}) for any r \ge 1, potentially as slow as O_P(1/log^tau(n))'. This is internally inconsistent because n^{-1/2r} is itself a polynomial rate, and the 'potentially as slow as n^{-1/log^tau(n)}' assertion is not a consequence of any proved statement. Consequently, the 'exponential number of data O(exp(epsilon^{-1/tau}))' for sigmoid-gating polynomial experts in the sparse regime, as listed in Table 1, is not proven, and the comparison with the softmax exponential rate from [1] is not established.
minor comments (6)
- [Section 5, Setup] The phrase 'ynthetic data' should be 'synthetic data'.
- [Section 6, Conclusion] The sentence 'Our results show that sigmoid self-attention has a higher sample complexity than the softmax version in the more common dense regime' contradicts the abstract, the title, and Table 1; it should read 'lower sample complexity'.
- [Appendix A, Related Works] The sentence 'Le et al. Furthermore, Akbarian et al. [1]...' contains an incomplete citation 'Le et al.' with no reference or title; please complete it or remove it.
- [Section 4.1.1, Theorem 1] The statement 'If the expert function ... then the lower bound ... holds true ... then L1( bGn, G∗) = OP(...)' has a double-'then' construction that should be rephrased for clarity.
- [Appendix B.1, Step 4] The covering numbers |\Delta_\tau| and |\Omega_\tau| are deterministic quantities but are written with OP(...); they should use O(...) notation.
- [Section 5 and Figure 1] The Voronoi loss L3(\hat G_n, G) is plotted for the sigmoid model fitted to data generated from a softmax-gating MoE, but the target measure G for the sigmoid fit is never defined; please specify how G is chosen in that setting.
Circularity Check
No significant circularity: the sigmoid MoE rates are independently derived, though the dense-regime sample-complexity claim targets a proxy rather than the true experts.
full rationale
No circular step is established. The sigmoid-gating MoE rates in Sections 3 and 4 are derived from a least-squares estimator defined in equation (6) via bracketing-entropy bounds (Appendix B.1) and Voronoi-loss lower bounds (Theorems 1-3), so they do not presuppose the softmax rates or the attention claim. The Section 2 attention-to-MoE identity is an algebraic rewriting of one attention row, not a fit. The softmax baseline is imported from reference [1], whose authors overlap with the present paper; this is load-bearing self-citation, but [1] is a separate stated-assumption analysis of softmax quadratic gating, so it does not make the sigmoid derivation circular under the hard rules. The main concern is a target mismatch, not circularity: in the dense regime, Theorem 3 gives inf_G L3(hat G_n, G) = O_P(sqrt(log n/n)) for G defined as the argmin over over-specified mixtures, and Section 4.2 then concludes that 'it takes those experts only a polynomial number of samples O(epsilon^-2) to achieve an approximation error of epsilon' without showing that the proxy parameters equal the true experts or that the proxy regression function equals f_{G*}; this is a transfer gap that affects the correctness of the comparison, but it is not an equation-level reduction of the prediction to its inputs. The Section 6 single-head limitation is a scope statement, not a circular step. The score is set to 2 to reflect the low-level self-citation burden while noting that no circular derivation was found.
Assumptions & free parameters
assumptions (5)
- domain assumption The data follows the regression model Yi = f_G*(Xi) + eps_i with i.i.d. Gaussian noise and bounded input domain (Eq. 4).
- domain assumption Strong or weak identifiability conditions (Definitions 1-4) hold for the expert functions, including ReLU/GELU networks and polynomial experts.
- standard math Standard empirical process theory results (bracketing entropy, Le Cam's lemma, Fatou's lemma) apply as used in the proofs.
- ad hoc to paper The self-attention row representation as an MoE with quadratic affinity scores (Section 2) preserves the statistical estimation problem of attention.
- ad hoc to paper The dense regime is more common in practice than the sparse regime.
Cite this review
Pith. "Pith review of Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective." pith.science (2026). https://pith.science/paper/WWV7XGQO
@misc{pith2026250200281,
author = {Pith},
title = {Pith review of: Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/WWV7XGQO}},
note = {Machine review of arXiv:2502.00281}
}
read the original abstract
At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on the most salient information. However, the softmax structure slows down the attention computation due to its row-wise nature, and it inherently introduces competition among tokens: as the weight assigned to one token increases, the weights of others decrease. This competitive dynamic may narrow the focus of self-attention to a limited set of features, potentially overlooking other informative characteristics. Recent experimental studies have shown that using the element-wise sigmoid function helps eliminate token competition and reduce the computational overhead. Despite these promising empirical results, a rigorous comparison between sigmoid and softmax self-attention mechanisms remains absent in the literature. This paper closes this gap by theoretically demonstrating that sigmoid self-attention is more sample-efficient than its softmax counterpart. Toward that goal, we represent the self-attention matrix as a mixture of experts and show that ``experts'' in sigmoid self-attention require significantly less data to achieve the same approximation error as those in softmax self-attention.
Figures
Forward citations
Cited by 1 Pith paper
-
SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
A CNN-Transformer U-Net with hierarchical expert routing and a shape-adapting hub reports state-of-the-art colon histopathology segmentation Dice of 95.57% on EBHI, 95.16% on DigestPath, and 94.17% on GlaS.
Reference graph
Works this paper leans on
-
[1]
P. Akbarian, H. Nguyen, X. Han, and N. Ho. Quadratic gating functions in mixture of experts: A statistical insight.arXiv preprint arXiv:2410.11222, 2024. (Cited on pages 2, 5, 9, 10, 12, and 33.)
arXiv 2024
- [2]
- [3]
-
[4]
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021. (Cited on page 1.)
work page 2021
-
[5]
Z. Chen, Y. Deng, Y. Wu, Q. Gu, and Y. Li. Towards understanding the mixture-of-experts layer in deep learning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 23049–23062. Curran Associates, Inc., 2022.(Cited on page 3.)
work page 2022
-
[6]
Z. Chen, Y. Shen, M. Ding, Z. Chen, H. Zhao, E. G. Learned-Miller, and C. Gan. Mod-squad: Designing mixtures of experts as modular multi-task learners. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11828–11837, 2023.(Cited on page 3.) 45
work page 2023
-
[7]
Z. Chi, L. Dong, S. Huang, D. Dai, S. Ma, B. Patra, S. Singhal, P. Bajaj, X. Song, X.-L. Mao, H. Huang, and F. Wei. On the representation collapse of sparse mixture of experts. In A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, editors,Advances in Neural Information Processing Systems, 2022. (Cited on page 3.)
work page 2022
-
[8]
R. Csordás, K. Irie, and J. Schmidhuber. Approximating two-layer feedforward networks for efficient transformers. In H. Bouamor, J. Pino, and K. Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, pages 674–692, Singapore, Dec. 2023. Association for Computational Linguistics. (Cited on page 3.)
work page 2023
Show all 51 references
-
[9]
Csordás, P
R. Csordás, P. Pi´kekos, K. Irie, and J. Schmidhuber. Switchhead: Accelerating transformers with mixture-of-experts attention. arXiv preprint arXiv:2312.07987, 2023. (Cited on pages 2 and 12.)
2023 arXiv
-
[10]
T. Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, 2024. (Cited on page 2.)
2024
-
[11]
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Re. Flashattention: Fast and memory-efficient exact attention with IO-awareness. InAdvances in Neural Information Processing Systems,
-
[12]
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, ...
2024 arXiv
-
[13]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics, 2019. (Cited on page 1.)
2019
-
[14]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 46 words: Transformers for image recognition at scale.ArXiv, abs/2010.11929, 2020. (Cited on page 1.)
2010 arXiv
-
[15]
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. S. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. ...
2021
-
[16]
Faria and G
S. Faria and G. Soromenho. Fitting mixtures of linear regressions. Journal of Statistical Computation and Simulation, 80(2):201–225, 2010. (Cited on page 3.)
2010
-
[17]
Fedus, B
W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23:1–39, 2022. (Cited on page 3.)
2022
-
[18]
N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu. Mixture of informed experts for multilingual speech recognition. InICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6234–6238....
2021
-
[19]
X. Gu, T. Pang, C. Du, Q. Liu, F. Zhang, C. Du, Y. Wang, and M. Lin. When attention sink emerges in language models: An empirical view.arXiv preprint arXiv:2410.10781, 2024. (Cited on page 1.)
2024 arXiv
-
[20]
Hazimeh, Z
H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. Chi. Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning. Advances in Neural Information Processing Systems, 34:29335–29347, 2021. (Cit...
2021
-
[21]
Ho, C.-Y
N. Ho, C.-Y. Yang, and M. I. Jordan. Convergence rates for Gaussian mixtures of experts. Journal of Machine Learning Research, 23(323):1–81, 2022. (Cited on page 13.)
2022
-
[22]
E. S. Hu, K. Ahn, Q. Liu, H. Xu, M. Tomar, A. Langford, D. Jayaraman, A. Lamb, and J. Langford. Learning to achieve goals with belief state transformers.ArXiv, abs/2410.23506,
-
[23]
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3, 1991. (Cited on pages 2 and 3.)
1991
-
[24]
A. Jain, V. P. Singh, and S. P. Rath. A multi-accent acoustic model using mixture of experts for speech recognition. InInterspeech, pages 779–783, 2019.(Cited on page 3.)
2019
-
[25]
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T...
2024 arXiv
-
[26]
P. Jin, B. Zhu, L. Yuan, and S. Yan. Moh: Multi-head attention as mixture-of-head attention. arXiv preprint arXiv:2410.11842, 2024. (Cited on page 12.)
2024
-
[27]
M. I. Jordan and R. A. Jacobs. Hierarchical mixtures of experts and the EM algorithm.Neural Computation, 6:181–214, 1994. (Cited on pages 3 and 12.)
1994
-
[28]
C. Kim, J. Park, J. Shin, H. Lee, P. Abbeel, and K. Lee. Preference transformer: Modeling human preferences using transformers for rl.arXiv preprint arXiv:2303.00957, 2023. (Cited on page 1.)
2023 arXiv
-
[29]
Kwon and S.-W
Y. Kwon and S.-W. Chung. Mole: Mixture of language experts for multi-lingual automatic speech recognition. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.(Cited on page 3.)
2023
-
[30]
B. Lindsay. Mixture models: Theory, geometry and applications. In NSF-CBMS Regional Conference Series in Probability and Statistics. IMS, Hayward, CA., 1995.(Cited on page 3.)
1995
-
[31]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.(Cited on page 1.)
2021
-
[32]
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi. Modeling task relationships in multi- task learning with multi-gate mixture-of-experts. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1930–1939, 2018.(Cited on page 3.)
1930
-
[33]
Manole and N
T. Manole and N. Ho. Refined convergence rates for maximum likelihood estimation under finite mixture models. InProceedings of the 39th International Conference on Machine Learning, volume 162 ofProceedings of Machine Learning Research, pages 14979–15006. PMLR, 17–23 Jul 2022....
2022
-
[34]
E. F. Mendes and W. Jiang. Convergence rates for mixture-of-experts.arXiv preprint arxiv 1110.2058, 2011. (Cited on page 13.)
2011 arXiv
-
[35]
Muennighoff, L
N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. D. Morrison, S. Min, W. Shi, P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Hajishirzi. Ol...
2024 arXiv
-
[36]
Nguyen, P
H. Nguyen, P. Akbarian, F. Yan, and N. Ho. Statistical perspective of top-k sparse softmax gating mixture of experts. InInternational Conference on Learning Representations, 2024. (Cited on page 13.)
2024
-
[37]
Nguyen, X
H. Nguyen, X. Han, C. W. Harris, S. Saria, and N. Ho. On expert estimation in hierarchical mixture of experts: Beyond softmax gating functions.arXiv preprint arXiv:2410.02935, 2024. (Cited on page 12.)
2024 arXiv
-
[38]
Nguyen, N
H. Nguyen, N. Ho, and A. Rinaldo. On least square estimation in softmax gating mixture of experts. In Proceedings of the ICML, 2024. (Cited on pages 6 and 12.) 48
2024
-
[39]
Nguyen, T
H. Nguyen, T. Nguyen, and N. Ho. Demystifying softmax gating function in Gaussian mixture of experts. InAdvances in Neural Information Processing Systems, 2023. (Cited on page 13.)
2023
-
[40]
Puigcerver, C
J. Puigcerver, C. Riquelme, B. Mustafa, and N. Houlsby. From sparse to soft mixtures of experts. In The Twelfth International Conference on Learning Representations, 2024. (Cited on page 3.)
2024
-
[41]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.(Cited on page 1.)
2021
-
[42]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020. (Cited on page 1.)
2020
-
[43]
Ramapuram, F
J. Ramapuram, F. Danieli, E. Dhekane, F. Weers, D. Busbridge, P. Ablin, T. Likhomanenko, J. Digani, Z. Gu, A. Shidani, and R. Webb. Theory, analysis, and best practices for sigmoid self-attention. arXiv preprint arXiv:2409.04431, 2024. (Cited on pages 1, 2, and 12.)
2024 arXiv
-
[44]
Shazeer, A
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer. InIn International Conference on Learning Representations, 2017. (Cited on page 3.)
2017
-
[45]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023. (Cited on page 1.)
2023 arXiv
-
[46]
van de Geer.Empirical Processes in M-estimation
S. van de Geer.Empirical Processes in M-estimation. Cambridge University Press, 2000.(Cited on pages 5 and 13.)
2000
-
[47]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.(Cited on pages 1, 3, and 12.)
2017
-
[48]
X. Wu, S. Huang, W. Wang, and F. Wei. Multi-head mixture-of-experts.arXiv preprint arXiv:2404.15045, 2024. (Cited on pages 2 and 12.)
2024 arXiv
-
[49]
F. Yan, H. Nguyen, D. Le, P. Akbarian, and N. Ho. Understanding expert structures on minimax parameter estimation in contaminated mixture of experts. InProceedings of The 28th International Conference on Artificial Intelligence and Statistics, 2025. (Cited on page 13.)
2025
-
[50]
Z. You, S. Feng, D. Su, and D. Yu. Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts.arXiv preprint arXiv:2105.03036, 2021. (Cited on page 3.)
2021 arXiv
-
[51]
Y. Zhou, N. Du, Y. Huang, D. Peng, C. Lan, D. Huang, S. Shakeri, D. So, A. Dai, Y. Lu, Z. Chen, Q. Le, C. Cui, J. Laudon, and J. Dean. Brainformers: Trading simplicity for efficiency. In International Conference on Machine Learning, pages 42531–42542. PMLR, 2023.(Cited on page 3.) 49
2023
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.