REVIEW 5 major objections 7 minor 1 cited by
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
T0 review · 5 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read More Transformer layers improve performance only when those layers produce diverse outputs, and each extra layer yields a smaller gain—matching how performance scales logarithmically with parameters.
desk verdict Solid application of known ensemble-diversity theory to Transformer layers, but the submodularity headline rests on an assumption the authors themselves admit is unrealistic. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is information-theoretic diversity, defined as conditional total correlation $I(U_{1:|L|}\mid Y)$ minus total correlation $I(U_{1:|L|})$ for the per-layer logit contributions $U_{1:|L|}$ in the residual stream. Theorems 1-3 establish a bias-diversity decomposition of mean squared logit error, and Theorems 4-8 bound prediction error using $H(Y)-I(U_{1:|L|};Y)$ and show that, when layer outputs are conditionally independent given the label, this mutual information is submodular and non-decreasing. That submodularity is the mechanism producing diminishing returns from added layers and therefore the link to parameter scaling laws.
What would settle it
On a trained Transformer, compute the marginal accuracy gain of successively adding the next layer's output to the residual sum, or compare accuracy between depth-$n$ and depth-$(n-1)$ prefixes. If these marginal gains do not tend to shrink with depth across several tasks, the submodularity claim is contradicted; a complementary check is to estimate the conditional total correlation $I(U_{1:|L|}\mid Y)$ from hidden states, since a large positive value shows the theorem's key premise fails in practice.
Extended reading notes
Core claim
The paper's central claim is that a Transformer's prediction error decomposes into a bias term (each layer's own error) minus a diversity term (how much layers disagree), so error falls as layer outputs become mutually diverse. At the information-theoretic level the error is governed by $H(Y)-I(U_{1:|L|};Y)$, and joint mutual information splits into relevancy plus conditional redundancy minus redundancy; the paper identifies conditional redundancy minus redundancy as information-theoretic diversity. Under the condition that layer outputs are independent given the true label, $I(U_{1:|L|};Y)$ and information-theoretic diversity are submodular and non-decreasing, so each added layer brings a smaller gain than the last. Assuming layer count is proportional to parameter count, this explains why scaling laws show logarithmic performance gains with parameters.
Load-bearing premise
The result that extra layers give diminishing returns assumes that different layers' outputs are statistically independent once the true label is known; real layers are heavily dependent, and the paper itself notes that nonzero conditional redundancy can make performance non-monotonic.
Editorial extensions
If this is right
- Adding a layer that simply repeats the behavior of existing layers can fail to improve accuracy; the paper's layer-sharing experiments show lower diversity and limited gains.
- When layers are conditionally independent given the label, each additional layer yields a smaller accuracy gain, giving a layer-wise explanation of logarithmic parameter scaling.
- Performance can fall when extra layers are added if conditional redundancy dominates, so depth increases are not guaranteed improvements; the paper observes this non-monotonicity.
- Because embedding layers contribute little by weight of layer count, most bias and diversity come from attention and MLP blocks, and attention diversity is relatively low, pointing to redundant attention heads as a source of waste.
Reading between the lines
- If layer diversity is the real driver of scaling gains, architectures that explicitly diversify layer outputs, such as by discouraging redundant attention heads or using diverse per-layer objectives, might capture depth-like gains with fewer parameters.
- The conditional-independence assumption underpinning the submodularity theorem is strong; real layers share information about the label, so the diminishing-returns prediction likely holds only as a smoothed trend, not pointwise.
- The information-theoretic diversity measure could serve as a monitoring metric during training or pruning: layers or heads that add little conditional redundancy relative to redundancy may be safe to remove.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a theoretical account of Transformer layer scaling by applying bias-diversity (ambiguity) decomposition to the residual stream. The authors show that the MSE between the ensemble logits and the true target decomposes into a bias term and a diversity term (Theorems 1-3), and then introduce an information-theoretic decomposition of an error bound, where relevancy, conditional redundancy, and redundancy determine the joint mutual information between layer outputs and the label (Theorem 4). The main theoretical findings are that adding layers improves performance only if layers are diverse, and that under a conditional-independence assumption, the mutual information and the information-theoretic diversity are submodular, which the authors interpret as a mechanism behind diminishing returns in parameter scaling laws (Theorems 5-8). Experiments on eight NLP tasks with several open-weight LLMs report correlations between the MSE-based bias/diversity metrics and accuracy, module-level contributions, and a decreasing trend of performance with layer count across models.
Significance. If the central claim were fully established, the paper would provide a mechanistic interpretation of scaling laws linking layer diversity to diminishing returns. The paper's positive contributions include a clean application of the Krogh-Vedelsby decomposition to the residual stream with explicit module-level bias/diversity terms, and an interesting empirical observation that layer-shared MobileLLM models exhibit much lower diversity than non-shared models (Table 1). However, the main submodularity result is conditional on an independence assumption that the paper itself concedes is violated in practice, and the empirical evidence is indirect and confounded. With appropriate qualification of the claims and a corrected proof/analysis, the paper could be a useful contribution to the interpretability literature.
major comments (5)
- [Abstract and §3.4] The abstract's statement that 'performance gains from increasing the number of layers exhibit submodularity' is not supported by the theorems as qualified. Theorem 7 proves submodularity of I(U_{1:|L|};Y) only under the assumption that the layer outputs are independent given Y. The paper explicitly notes in §3.4 that 'in the actual situation, we can expect performance improvement by increased Conditional Redundancy, and it makes the performance improvement non-monotonic' and in §5.5 that performance degradation indicates Conditional Redundancy > 0. The abstract and conclusion (§6) present submodularity as an empirical property without these caveats; the claims should be rephrased as a conditional theoretical result about an idealized case, not a property of actual Transformers.
- [§5.5 and Figure 6] The empirical evidence for diminishing returns is not a test of the theory. Figure 6 compares different models with different layer sizes (33, 57, 65, 81) that also differ in parameter count, hidden dimension, architecture, and training data; this is a cross-model comparison, not a controlled manipulation of layer count. Moreover, the diversity measured in Figures 3, 5, and 7 is the MSE ambiguity of Eq. 8, not the information-theoretic diversity of Eq. 13 that appears in Theorems 4 and 7. Consequently, the observed decreasing trend and the interpretation that performance degradation 'indicates that Conditional Redundancy becomes larger than zero' do not directly confirm the submodularity theorem. The authors should either use within-model layer ablations (e.g., truncating a single model at different depths) or explicitly label the evidence as suggestive rather than confirmatory.
- [§3.4, Eq. (13)] The correspondence between the bias-diversity decomposition and the information-theoretic decomposition is asserted rather than proved. The sentence 'the bias term corresponds to the relevancy term and the diversity term corresponds to the information-theoretic diversity' is not accompanied by a derivation connecting Eq. (8) to Eq. (13). The two decompositions operate on different quantities (MSE of logits vs. mutual information of layer-output random variables), and the paper does not define the probabilistic model under which a layer output u^{(i)} is a random variable, nor how the expectation in Eq. (8) relates to the entropies in Eq. (13). This gap matters because the experimental diversity (Fig. 3) is the former, while Theorems 5-8 concern the latter.
- [§3.4, Theorem 7] Even granting the conditional-independence assumption, submodularity of I(U_{1:|L|};Y) does not imply submodularity of the prediction error or of performance. Equation (12) gives upper and lower bounds on p(Y≠g(U)), not an equality. The inference 'the effect of adding layers on performance decreases as the number of layers increases' requires an argument that the error probability tracks I(U_{1:|L|};Y) monotonically and with diminishing differences; the paper does not provide this. The supermodularity of the bounds (Theorem 8) is a statement about the bounds, not about the actual error.
- [Appendix B.7, Eqs. (34)-(39)] The proof that I(U_{1:|L|};Y) is submodular is not self-contained. In Eq. (37), the term H(u|Y,U_{1:|L|}) - H(u|Y,U_{1:|L|+1}) is set equal to H(u|Y)-H(u|Y), but this equality requires that the candidate element u and the newly added element are conditionally independent of each other and of the entire set given Y; the theorem's assumption is stated for 'the variables in U_{1:|L|}' and does not clearly extend to the enlarged sets used in the marginal-gain comparison. The proof should state this extension explicitly and correct the notation (the same symbol u is used for the candidate added to two different sets). Without this, the submodularity claim is not rigorously established.
minor comments (7)
- [Throughout] There are several typos in technical terms: 'diveristy' in §3.4, 'Information Theoretic Diveristy' in Theorem 5 and Eq. (13), 'Divergence' for 'Diversity' in B.7, and 'Relevancy' for 'Relevance'.
- [§5.2] The sentence 'the fact that the Spearman correlation is lower than the rank correlation' is self-contradictory because Spearman's rho is a rank correlation. Based on the values in Figure 3, Spearman correlations are generally higher than Pearson correlations, so the sentence likely intended the opposite.
- [§4.2] The citation of [Touvron et al., 2023] (Llama-2 report) for the 'trend to repurpose such models as small-scale LLMs' appears unrelated to layer reuse; a more relevant reference should be provided.
- [Theorem 2] Theorem 2's notation 'Diversity→0 (Bias→0)' is ambiguous; it should be written as 'Bias→0 implies Diversity→0' (and similarly for the module-level statements) to avoid confusion about the direction of the implication.
- [Appendix B.5, Eqs. (21)-(26)] The proof of the monotonicity of Redundancy contains an algebraic error: H(U_{1:|L|+1}) is written as H(u^{(|L|+1)}) + H(u^{(|L|+1)}|U_{1:|L|}), which is dimensionally inconsistent; the second term should be H(U_{1:|L|}|u^{(|L|+1)}) or the first term should be H(U_{1:|L|}). Although the monotonicity conclusion is a standard result, the derivation as written is incorrect and should be fixed.
- [§5.5, Figure 6] The x-axis of Figure 6 shows layer sizes, but the models also differ in parameter count, hidden dimension, and architecture. A note should be added to clarify the confound, or a controlled comparison (e.g., pruning layers from a single model) should be reported.
- [Appendix A] The limitations paragraph mentions approximation issues for probability distributions but does not mention the conditional-independence assumption behind Theorems 7-8; stating this limitation explicitly would help readers calibrate the scope of the submodularity result.
Circularity Check
No significant circularity: the paper imports external bias-diversity and information-theoretic decompositions and does not fit parameters that are then renamed as predictions; Theorem 7's conditional-independence caveat is a scope limitation, not circularity.
full rationale
The derivation chain is self-contained with respect to circularity. The bias-diversity decomposition (Theorem 1, Eqs. 7-10) is explicitly attributed to Krogh and Vedelsby 1994 and Wood et al. 2024, and is an algebraic identity rather than a fitted quantity; the claim that higher diversity lowers MSE follows mathematically from that identity without any parameter being fit to the target claim. The information-theoretic results (Theorems 4-8, Eqs. 12-13) are imported from external sources (Brown 2009; Zhou and Li 2010; Krause and Guestrin 2005; Fujishige 1978; Studeny 2006) and are not self-citations. No load-bearing premise is justified solely by the authors' own prior work, and no uniqueness theorem is invoked to forbid alternatives. The main abstract claim that performance gains exhibit submodularity is conditional on the assumption that U_{1:|L|} are independent given Y; the paper itself acknowledges in Sections 3.4 and 5.5 that in actual Transformers Conditional Redundancy is non-zero and performance can be non-monotonic. That is a legitimate scope/over-claim concern about whether the theorem applies to real models, but it is not a circular reduction: Theorem 7 remains an external mathematical result under a stated assumption, and the empirical section does not pretend to fit that assumption into the conclusion. The paper's own limitations note in Appendix A also flags the gap between logits and softmax probabilities and the difficulty of computing the information-theoretic diversity without approximation. These caveats weaken the practical force of the conclusions but do not make the derivation equivalent to its inputs by construction. Hence no circularity step can be quoted from the text, and the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- number of layers as proxy for number of parameters
- p-value threshold and significance test choices =
0.05
- z-transformation for averaging correlations
assumptions (5)
- domain assumption RMSNorm scaling s can be moved inside the sum and combined with W_out to define u^(i)
- domain assumption Logits are treated as a sum of independent layer contributions u^(i)
- domain assumption The true distribution of logits u_hat exists and can be approximated by a one-hot vector in Eq. 14
- ad hoc to paper Layer outputs are independent given the label Y (Theorem 7)
- domain assumption Mutual information bounds apply to Transformer layers
Cite this review
Pith. "Pith review of Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws." pith.science (2026). https://pith.science/paper/PBOIYGPA
@misc{pith2026250524009,
author = {Pith},
title = {Pith review of: Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws},
year = {2026},
howpublished = {\url{https://pith.science/paper/PBOIYGPA}},
note = {Machine review of arXiv:2505.24009}
}
read the original abstract
Transformers deliver outstanding performance across a wide range of tasks and are now a dominant backbone architecture for large language models (LLMs). Their task-solving performance is improved by increasing parameter size, as shown in the recent studies on parameter scaling laws. Although recent mechanistic-interpretability studies have deepened our understanding of the internal behavior of Transformers by analyzing their residual stream, the relationship between these internal mechanisms and the parameter scaling laws remains unclear. To bridge this gap, we focus on layers and their size, which mainly decide the parameter size of Transformers. For this purpose, we first theoretically investigate the layers within the residual stream through a bias-diversity decomposition. The decomposition separates (i) bias, the error of each layer's output from the ground truth, and (ii) diversity, which indicates how much the outputs of each layer differ from each other. Analyzing Transformers under this theory reveals that performance improves when individual layers make predictions close to the correct answer and remain mutually diverse. We show that diversity becomes especially critical when individual layers' outputs are far from the ground truth. Finally, we introduce an information-theoretic diversity and show our main findings that adding layers enhances performance only when those layers behave differently, i.e., are diverse. We also reveal the performance gains from increasing the number of layers exhibit submodularity: marginal improvements diminish as additional layers increase, mirroring the logarithmic convergence predicted by the parameter scaling laws. Experiments on multiple semantic-understanding tasks with various LLMs empirically confirm the theoretical properties derived in this study.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
When Does Sparsity Mitigate the Curse of Depth in LLMs
Implicit and explicit sparsity reduce residual-stream variance and improve layer effectiveness metrics, enabling a depth-scaling recipe with about 4.6 points higher downstream accuracy.
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
arXiv 2024
-
[2]
Hewett, Mojan Javaheripi, Piero Kauffmann, James R
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu,...
arXiv 2024
-
[3]
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, ...
2022
-
[4]
An information theoretic perspective on multiple classifier systems
Gavin Brown. An information theoretic perspective on multiple classifier systems. In Proceedings of the 8th International Workshop on Multiple Classifier Systems, MCS '09, page 344–353, Berlin, Heidelberg, 2009. Springer-Verlag. ISBN 9783642023255. doi:10.1007/978-3-642-02326-2_35. URL https://doi.org/10.1007/978-3-642-02326-2_35
-
[5]
Gavin Brown, Jeremy L. Wyatt, and Peter TiŇo. Managing diversity in regression ensembles. Journal of Machine Learning Research, 6 0 (55): 0 1621--1650, 2005. URL http://jmlr.org/papers/v6/brown05a.html
work page 2005
-
[6]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
1901
-
[7]
When parts are greater than sums: Individual LLM components can outperform full models
Ting-Yun Chang, Jesse Thomason, and Robin Jia. When parts are greater than sums: Individual LLM components can outperform full models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10280--10299, Miami, Florida, USA, November 2024. Association for...
-
[8]
B ool Q : Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. B ool Q : Exploring the surprising difficulty of natural yes/no questions. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics:...
2019
Show all 39 references
-
[9]
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. URL https://arxiv.org/abs/1803.05457
2018 arXiv
-
[10]
Averaging correlations: Expected values and bias in combined pearson rs and fisher's z transformations
David M Corey, William P Dunlap, and Michael J Burke. Averaging correlations: Expected values and bias in combined pearson rs and fisher's z transformations. The Journal of general psychology, 125 0 (3): 0 245--261, 1998
1998
-
[11]
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyzdRiR9Y7
2019
-
[12]
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1 0 (1): 0 12, 2021
2021
-
[13]
Polymatroidal dependence structure of a set of random variables
Satoru Fujishige. Polymatroidal dependence structure of a set of random variables. Information and Control, 39 0 (1): 0 55--72, 1978. ISSN 0019-9958. doi:https://doi.org/10.1016/S0019-9958(78)91063-X. URL https://www.sciencedirect.com/science/article/pii/S001999587891063X
1978 doi
-
[14]
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language P...
2021 doi
-
[15]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[16]
What matters in transformers? not all attention is needed, 2024
Shwai He, Guoheng Sun, Zheyu Shen, and Ang Li. What matters in transformers? not all attention is needed, 2024. URL https://arxiv.org/abs/2406.15786
2024 arXiv
-
[17]
Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically, 2017. URL https://arxiv.org/abs/1712.00409
2017 arXiv
-
[18]
Submodular combinatorial information measures with applications in machine learning
Rishabh Iyer, Ninad Khargoankar, Jeff Bilmes, and Himanshu Asanani. Submodular combinatorial information measures with applications in machine learning. In Vitaly Feldman, Katrina Ligett, and Sivan Sabato, editors, Proceedings of the 32nd International Conference on Algorithmi...
2021
-
[19]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[20]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. URL https://arxiv.org/abs/2001.08361
2020 arXiv
-
[21]
Near-optimal nonmyopic value of information in graphical models
Andreas Krause and Carlos Guestrin. Near-optimal nonmyopic value of information in graphical models. In Proceedings of the Twenty-First Conference on Uncertainty in Artificial Intelligence, UAI'05, page 324–331, Arlington, Virginia, USA, 2005. AUAI Press. ISBN 0974903914
2005
-
[22]
Neural network ensembles, cross validation, and active learning
Anders Krogh and Jesper Vedelsby. Neural network ensembles, cross validation, and active learning. In G. Tesauro, D. Touretzky, and T. Leen, editors, Advances in Neural Information Processing Systems, volume 7. MIT Press, 1994. URL https://proceedings.neurips.cc/paper_files/pa...
1994
-
[23]
Mobilellm: optimizing sub-billion parameter language models for on-device use cases
Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra. Mobilellm: optimizing sub-billion parameter language models for on-device use cases. In P...
2024
-
[24]
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kam...
2022 arXiv
-
[25]
Probable error of a correlation coefficient
Student. Probable error of a correlation coefficient. Biometrika, pages 302--310, 1908
1908
-
[26]
Probabilistic conditional independence structures
Milan Studeny. Probabilistic conditional independence structures. Springer Science & Business Media, 2006
2006
-
[27]
Lessons on parameter sharing across layers in transformers
Sho Takase and Shun Kiyono. Lessons on parameter sharing across layers in transformers. In Nafise Sadat Moosavi, Iryna Gurevych, Yufang Hou, Gyuwan Kim, Young Jin Kim, Tal Schuster, and Ameeta Agrawal, editors, Proceedings of the Fourth Workshop on Simple and Efficient Natural...
2023 doi
-
[28]
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[29]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural In...
2017
-
[30]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE : A multi-task benchmark and analysis platform for natural language understanding. 2019. In the Proceedings of ICLR
2019
-
[31]
Bilateral multi-perspective matching for natural language sentences
Zhiguo Wang, Wael Hamza, and Radu Florian. Bilateral multi-perspective matching for natural language sentences. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI-17 , pages 4144--4150, 2017. doi:10.24963/ijcai.2017/579. URL htt...
2017 doi
-
[32]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. T...
2022
-
[33]
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Marilyn Walker, Heng Ji, and Amanda Stent, editors, Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Com...
2018 doi
-
[34]
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[35]
Webb, Henry W
Danny Wood, Tingting Mu, Andrew M. Webb, Henry W. J. Reeve, Mikel Luj\' a n, and Gavin Brown. A unified theory of diversity in ensemble learning. J. Mach. Learn. Res., 24 0 (1), mar 2024. ISSN 1532-4435
2024
-
[36]
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML'20. JMLR.org, 2020
2020
-
[37]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proce...
2019
-
[38]
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. Character-level convolutional networks for text classification. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL ht...
2015
-
[39]
Multi-information ensemble diversity
Zhi-Hua Zhou and Nan Li. Multi-information ensemble diversity. In Neamat El Gayar, Josef Kittler, and Fabio Roli, editors, Multiple Classifier Systems, pages 134--144, Berlin, Heidelberg, 2010. Springer Berlin Heidelberg. ISBN 978-3-642-12127-2
2010
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.