REVIEW 3 major objections 2 minor 95 references
On the Fitness Landscape in the $NK$ Model
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The dense NK fitness landscape is exactly solvable: its free energy and maximum fitness have explicit thermodynamic limits, and near-optimal peaks form an exponentially large orthogonal set.
desk verdict The abstract promises a real result for the dense NK regime, but the supplied full text is a different paper; verdict must wait on the actual PDF. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the NK random field viewed as a spin-glass Hamiltonian on the $N$-locus hypercube, where each locus interacts with $K$ randomly selected others; in the dense regime, the analysis is carried by the spin-glass free-energy variational principle. The load-bearing mechanism is the overlap-gap property: near-maximum genomes have asymptotically zero pairwise overlap, which converts the geometric claim about many peaks into a quantitative statement about the overlap distribution of high-fitness configurations.
What would settle it
Run the NK model for increasing $N$ with $K/N$ fixed at several values of $\alpha$; if the empirical free energy or maximum fitness per locus does not approach the paper's predicted limits, the central claim is wrong. A more direct check is exhaustive enumeration for small $N$: the number of near-fittest mutually orthogonal genomes should grow exponentially in $N$ at a rate depending on $\alpha$.
Extended reading notes
Core claim
In the dense regime $K/N \to \alpha \in (0,1]$, the NK random field is treated as a disordered spin system on the hypercube $\{0,1\}^N$. The main result gives the exact limits, per locus, of the quenched free energy at all temperatures and of the maximum fitness, with the limits depending only on $\alpha$. The geometric companion is an overlap-gap statement: any two genomes whose fitness is within a small window of the global maximum are asymptotically orthogonal, i.e. they differ on almost all loci, and the family of such near-fittest genomes contains exponentially many mutually orthogonal members. Consequently, near-fittest evolutionary paths become impossible when fitness levels approach
Load-bearing premise
The derivation assumes the dense NK random field converges to a spin-glass system for which a rigorous free-energy variational principle holds; if that convergence fails, the claimed exact limits do not follow from the method.
Editorial extensions
If this is right
- For any $\alpha \in (0,1]$, the free energy per locus and the maximum fitness per locus have definite large-$N$ limits, so numerical simulations of the NK model have analytic benchmarks to match.
- The landscape is provably rugged: exponentially many near-fittest genomes are pairwise almost completely different, so no single narrow region of genome space contains the top of the landscape.
- Adaptation to the global maximum is blocked in a precise sense: paths connecting near-fittest genomes do not exist as the fitness threshold approaches the maximum.
- For small interaction density $\alpha$, evolution can still proceed without fitness loss: a fitness-maintaining path exists with high probability.
Reading between the lines
- The supplied body text under this title is a different manuscript (an evaluation of language models), so this extraction rests on the title, abstract, and subject class; the quantitative theorems should be verified against the original source before citation.
- If the variational-principle route is correct, the dense NK model should sit in the same universality class as other Gaussian spin glasses, which would imply that locating the global maximum is computationally hard for the same reason it is hard in those models.
- The rate at which the number of near-fittest orthogonal genomes grows in $N$ (the exponential rate as a function of $\alpha$) is not stated in the abstract; measuring that rate numerically would be a direct test of the geometric picture.
- The small-$\alpha$ path result suggests a sharp threshold in $\alpha$ between path-connected and path-blocked fitness landscapes; locating that threshold is a natural next problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper, as represented by its arXiv metadata and abstract (arXiv:2508.12464, math.PR), claims to study the NK model in the regime K/N -> alpha in (0,1], using spin-glass methods. The abstract announces exact limits for the free energy at any temperature and for the maximum fitness, an exponentially large collection of near-fittest pairwise asymptotically orthogonal genomes, overlap gap properties, and quantitative statements about the possibility/impossibility of evolutionary paths near the global maximum. However, the supplied full text is not an NK-model paper at all; it is an unrelated evaluation study of OpenAI's GPT-OSS language models (arXiv:2508.12461v3). No NK model is defined, no main theorem is stated, and no proof, lemma, or derivation supporting any of the announced results appears anywhere in the manuscript body. Consequently, none of the paper's central claims can be checked from the submitted material.
Significance. If the results announced in the abstract were properly proved, they would constitute a substantial contribution to the mathematical theory of the NK model: exact solution of the dense regime (K/N -> alpha), a rigorous variational formula for the free energy, and a quantitative picture of the fitness landscape's multiple-peak geometry. These would be of interest to probabilists, statistical physicists, and theoretical evolutionary biologists. I cannot assess the significance of the actual contribution, however, because the manuscript contains no statement of the claimed theorems beyond the abstract and no supporting argument. The only verifiable content is an unrelated LLM evaluation, which has no bearing on the NK model. There is no evidence of machine-checked proofs, reproducible derivations, or even precise definitions on which to judge the mathematical claims.
major comments (3)
- [Full text (whole body)] The body of the submitted manuscript is 'Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models' (arXiv:2508.12461v3), not an NK-model paper. There is no definition of the NK model, no random-field construction, no theorem statements, no proofs, and no overlap-gap definitions. The abstract's central claims—exact free-energy limits, maximum fitness, exponential near-fittest orthogonal genomes, and evolutionary-path impossibilities—are therefore entirely unsupported by the supplied text. This is not a presentation issue but a complete absence of the mathematical content that the abstract announces.
- [Abstract] The abstract asserts that the NK model is treated 'via the spin glass methodologies' and that this yields 'exact limits.' No precise conditions are stated: the distribution of the NK random field, the nature of the K/N -> alpha limit, the convergence of the field to a spin-glass-type model, or the variational principle being invoked are all unspecified. If the intended paper contains a proof, these elements are load-bearing; in the present manuscript they appear only as assertions. The claimed exact limits do not follow from the stated method without these details.
- [Title and submission metadata] The manuscript's title, subject class (math.PR), and abstract concern the NK model, but the actual text is an empirical evaluation of language models. This internal inconsistency means the submission does not contain the research it claims to present. Even as a 'wrong file' scenario, the submitted manuscript is not refereeable as a probability paper, and the issue cannot be repaired by minor local edits within the current text.
minor comments (2)
- [References] The reference list is the bibliography of the GPT-OSS evaluation paper (e.g., Radford et al., Hendrycks et al., Kaplan et al.). No references to the NK model literature—Kauffman, Levin, Weinberger, or subsequent mathematical works on NK landscapes—appear, which is inconsistent with the abstract's claims.
- [Figures and tables] All tables and figures (e.g., Table I on model parameters, Figures 1–5 on benchmark results) concern large language models. None provide any numerical, exact, or illustrative support for the NK-model results described in the abstract.
Circularity Check
No circularity demonstrable: the supplied full text is a different arXiv paper (2508.12461v3), so the NK-model derivation chain is absent and no reduction-to-input can be exhibited.
full rationale
The abstract of arXiv:2508.12464 claims exact free-energy limits, maximum fitness, and multiple-peak structure in the NK model obtained 'via the spin glass methodologies.' The provided full text, however, is the manuscript of arXiv:2508.12461v3 ('Is GPT-OSS Good? ...'), an LLM evaluation paper with no NK-model content. Consequently there are no equations, coupling constructions, interpolation arguments, or overlap-gap derivations from the claimed NK paper to audit. Under the hard rules, circularity can only be claimed when the specific reduction is exhibited (Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction). No such reduction is present. The mismatch is a verification blocker, not evidence of circularity: the absence of proof cannot itself show that the exact limits are equivalent to the inputs. No self-citation chain, no ansatz smuggled by citation, and no fitted-input-called-prediction pattern can be identified from the abstract alone. The abstract's invocation of 'spin glass methodologies' is load-bearing but unverified in the supplied text; that is a correctness or completeness concern, not a demonstrated circular step. The honest finding is therefore a non-finding: no significant circularity can be established from the available material.
Assumptions & free parameters
Cite this review
Pith. "Pith review of On the Fitness Landscape in the $NK$ Model." pith.science (2026). https://pith.science/paper/76WY6YDF
@misc{pith2026250812464,
author = {Pith},
title = {Pith review of: On the Fitness Landscape in the $NK$ Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/76WY6YDF}},
note = {Machine review of arXiv:2508.12464}
}
abstract
The $NK$ model, introduced by Kauffman, Levin, and Weinberger, is a random field used to describe the fitness landscape of certain species with $N$ genetic loci, each interacting with $K$ others. The model has wide applications in understanding evolutionary and natural selection as it captures ruggedness feature of the fitness landscape. Earlier literature has been focused on the case $K$ being a fixed positive integer and used tools from Ergodic and Markov theory. In this paper, by viewing it as a statistical physics object, we investigate the $NK$ model in the regime $K/N\to\alpha \in(0,1]$ via the spin glass methodologies. Our main result identifies the exact limits for the free energy at any temperature and the maximum fitness. Moreover, we show that the $NK$ model exhibits a multiple-peak structure, namely, the number of near-fittest genomes that are asymptotically orthogonal to each other is exponentially large. Based on establishing the overlap gap properties, we obtain quantitative descriptions for the geometry of the fitness landscape and deduce that, in particular, near-fittest evolutionary paths become impossible as the fitness levels of the genomes approach the global maximum for any $\alpha\in (0,1]$. Nevertheless, we also show that by choosing $\alpha$ sufficiently small, an evolutionary path maintained at a given fitness level can be constructed with high probability.
Reference graph
Works this paper leans on
-
[1]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” inInternational Conference on Learning Representations, 2017
2017
-
[2]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”The Journal of Machine Learning Research, vol. 23, no. 1, pp. 5232–5270, 2022
2022
-
[3]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,”arXiv preprint arXiv:2001.08361, 2020
arXiv 2001
-
[4]
Scaling laws for autoregressive generative modeling,
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Grayet al., “Scaling laws for autoregressive generative modeling,”arXiv preprint arXiv:2010.14701, 2020
arXiv 2010
-
[5]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., “Training compute-optimal large language models,”arXiv preprint arXiv:2203.15556, 2022
arXiv 2022
-
[6]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskeveret al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[7]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azharet al., “Llama: Open and efficient foundation language models,”arXiv preprint arXiv:2302.13971, 2023
arXiv 2023
-
[8]
Gemma: Open models based on gemini research and technology,
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Loveet al., “Gemma: Open models based on gemini research and technology,”arXiv preprint arXiv:2403.08295, 2024
arXiv 2024
Show all 95 references
-
[9]
Deepseek llm: Scaling open-source language models with longtermism,
D.-A. and:, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Donget al., “Deepseek llm: Scaling open-source language models with longtermism,”arXiv preprint arXiv:2401.02954, 2024
2024 arXiv
-
[10]
Qwen technical report,
J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huanget al., “Qwen technical report,”arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[11]
Qwen2.5 technical report,
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...
2024 arXiv
-
[12]
Phi- 3 technical report: A highly capable language model locally on your phone,
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behlet al., “Phi- 3 technical report: A highly capable language model locally on your phone,”arXiv preprint arXiv:2404.14219, 2024
2024 arXiv
-
[13]
Emergent abilities of large language models,
J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzleret al., “Emergent abilities of large language models,”arXiv preprint arXiv:2206.07682, 2022
2022 arXiv
-
[14]
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?”Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[15]
Measuring massive multitask language understanding,
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2021
2009 arXiv
-
[16]
Training verifiers to solve math word problems,
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakanoet al., “Training verifiers to solve math word problems,”arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[17]
Solv- ing quantitative reasoning problems with language models,
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., “Solv- ing quantitative reasoning problems with language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022
2022
-
[18]
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,”arXiv preprint arXiv:2107.03374, 2021
2021 arXiv
-
[19]
Program synthesis with large language models,
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Leet al., “Program synthesis with large language models,”arXiv preprint arXiv:2108.07732, 2021
2021 arXiv
-
[20]
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chimet al., “Starcoder: may the source be with you!”arXiv preprint arXiv:2305.06161, 2023
2023 arXiv
-
[21]
Codegen: An open large language model for code with multi-turn program synthesis,
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y . Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,”arXiv preprint arXiv:2203.13474, 2022
2022 arXiv
-
[22]
Incoder: A generative model for code infilling and synthesis,
D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis, “Incoder: A generative model for code infilling and synthesis,”arXiv preprint arXiv:2204.05999, 2022
2022 arXiv
-
[23]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,
Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fuet al., “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 991–63 010, 2023
2023
-
[24]
Language models are few-shot multilingual learners,
G. I. Winata, A. Madotto, Z. Lin, R. Liu, J. Yosinski, and P. Fung, “Language models are few-shot multilingual learners,”arXiv preprint arXiv:2109.07684, 2021
2021 arXiv
-
[25]
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”arXiv preprint arXiv:2306.05685, 2023
2023 arXiv
-
[26]
Lamda: Language models for dialog applications,
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,”arXiv preprint arXiv:2201.08239, 2022
2022 arXiv
-
[27]
Recipes for building an open-domain chatbot,
S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y . Liu, J. Xu, M. Ott, E. M. Smith, Y .-L. Boureauet al., “Recipes for building an open-domain chatbot,”arXiv preprint arXiv:2004.13637, 2021
2004 arXiv
-
[28]
Glam: Efficient scaling of language models with mixture-of-experts,
N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firatet al., “Glam: Efficient scaling of language models with mixture-of-experts,”International Conference on Machine Learning, pp. 5547–5569, 2022
2022
-
[29]
Retrieval- augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschelet al., “Retrieval- augmented generation for knowledge-intensive nlp tasks,”Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020
2020
-
[30]
Efficient large scale language modeling with mixtures of experts,
M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V . Lin, J. Du, S. Iyer, R. Pasunuruet al., “Efficient large scale language modeling with mixtures of experts,”arXiv preprint arXiv:2112.10684, 2021
2021 arXiv
-
[31]
Mixture-of-experts with expert choice routing,
Y . Zhou, T. Lei, H. Liu, N. Du, Y . Huang, V . Zhao, A. M. Daiet al., “Mixture-of-experts with expert choice routing,”Advances in Neural Information Processing Systems, vol. 35, pp. 7103–7114, 2022
2022
-
[32]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosaleet al., “Llama 2: Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[33]
The falcon series of open language models,
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al., “The falcon series of open language models,”arXiv preprint arXiv:2311.16867, 2023
2023 arXiv
-
[34]
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,
T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”Advances in Neural Information Processing Systems, vol. 35, pp. 30 318–30 332, 2022
2022
-
[35]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[36]
Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,
G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosenet al., “Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,”arXiv preprint arXiv...
2025 arXiv
-
[37]
(2025) Claude opus 4.1
Anthropic. (2025) Claude opus 4.1. Accessed: 2025-08-26. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-1
2025
-
[38]
Qwen3 technical report,
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025
2025 arXiv
-
[39]
D. Team. (2025) Deepseek-v3.1 release report. Accessed: 2025- 08-26. [Online]. Available: https://api-docs.deepseek.com/zh-cn/news/ news250821
2025
-
[40]
Eight things to know about large language models,
S. R. Bowman, “Eight things to know about large language models,” arXiv preprint arXiv:2304.00612, 2023
2023 arXiv
-
[41]
Ai and the everything in the whole wide world benchmark,
I. D. Raji, E. Denton, E. M. Bender, A. Hanna, and A. Paullada, “Ai and the everything in the whole wide world benchmark,”arXiv preprint arXiv:2111.15366, 2021
2021 arXiv
-
[42]
Holistic evaluation of language models,
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,” inTransactions on Machine Learning Research, 2023
2023
-
[43]
With little power comes great responsibility,
D. Card, P. Henderson, U. Khandelwal, R. Jia, K. Mahowald, and D. Ju- rafsky, “With little power comes great responsibility,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 9263–9274
2020
-
[44]
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,
O. Sainz, J. A. Campos, I. Garc ´ıa-Ferrero, J. Etxaniz, O. L. de La- calle, and E. Agirre, “Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,”arXiv preprint arXiv:2310.18018, 2023
2023 arXiv
-
[45]
A survey of large language models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Donget al., “A survey of large language models,”arXiv preprint arXiv:2303.18223, 2023
2023 arXiv
-
[46]
Large language models: A survey,
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Am- atriain, and J. Gao, “Large language models: A survey,”arXiv preprint arXiv:2402.06196, 2024
2024 arXiv
-
[47]
llama.cpp: Inference of llama model in pure c/c++,
G. Gerganov, “llama.cpp: Inference of llama model in pure c/c++,” https: //github.com/ggerganov/llama.cpp, 2023
2023
-
[48]
Gshard: Scaling giant models with conditional computation and automatic sharding,
D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,”arXiv preprint arXiv:2006.16668, 2021
2006 arXiv
-
[49]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[50]
Scaling language models: Methods, analysis & insights from training gopher,
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Younget al., “Scaling language models: Methods, analysis & insights from training gopher,”arXiv preprint arXiv:2112.11446, 2021
2021 arXiv
-
[51]
Palm: Scaling language modeling with pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmannet al., “Palm: Scaling language modeling with pathways,”arXiv preprint arXiv:2204.02311, 2022
2022 arXiv
-
[52]
Finqa: A dataset of numerical reasoning over financial data,
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y . Wang, “Finqa: A dataset of numerical reasoning over financial data,” 2022. [Online]. Available: https://arxiv.org/abs/2109.00122
2022 arXiv
-
[53]
Medical exam question answering with large-scale reading comprehension,
X. Zhang, J. Wu, Z. He, X. Liu, and Y . Su, “Medical exam question answering with large-scale reading comprehension,” 2018. [Online]. Available: https://arxiv.org/abs/1802.10279
2018 arXiv
-
[54]
Experimenting with legal ai solutions: The case of question-answering for access to justice,
J. Li, R. Bhambhoria, S. Dahan, and X. Zhu, “Experimenting with legal ai solutions: The case of question-answering for access to justice,”arXiv preprint arXiv:2409.07713, 2024
2024 arXiv
-
[55]
Crowdsourcing multiple choice science questions,
J. Welbl, N. F. Liu, and M. Gardner, “Crowdsourcing multiple choice science questions,” 2017. [Online]. Available: https://arxiv.org/abs/1707. 06209
2017
-
[56]
Piqa: Reasoning about physical commonsense in natural language,
Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” 2019. [Online]. Available: https://arxiv.org/abs/1911.11641
2019 arXiv
-
[57]
Dialogsum: A real-life scenario dialogue summarization dataset,
Y . Chen, Y . Liu, L. Chen, and Y . Zhang, “Dialogsum: A real-life scenario dialogue summarization dataset,” 2021. [Online]. Available: https://arxiv.org/abs/2105.06762
2021 arXiv
-
[58]
Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,”International Conference on Machine Learning, pp. 4411–4421, 2020
2020
-
[59]
mt5: A massively multilingual pre-trained text-to-text transformer,
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,”arXiv preprint arXiv:2010.11934, 2021
2010 arXiv
-
[60]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”arXiv preprint arXiv:2206.04615, 2022
2022 arXiv
-
[61]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language mod- els are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[62]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 24 824–24 837, 2022
2022
-
[63]
Large lan- guage models are zero-shot reasoners,
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large lan- guage models are zero-shot reasoners,”Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022
2022
-
[64]
Large language models are not fair evaluators,
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” arXiv preprint arXiv:2305.17926, 2023
2023 arXiv
-
[65]
The curious case of neural text degeneration,
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y . Choi, “The curious case of neural text degeneration,”arXiv preprint arXiv:1904.09751, 2019
1904 arXiv
-
[66]
Neural text generation with unlikelihood training,
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston, “Neural text generation with unlikelihood training,”arXiv preprint arXiv:1908.04319, 2019
1908 arXiv
-
[67]
The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,
L. Ruis, A. K. Kemfert, D. Hupkes, L. Kaiser, and B. M. Lake, “The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,”Advances in Neural Information Processing Systems, vol. 37, 2024
2024
-
[68]
Locally typical sampling,
C. Meister, T. Pimentel, G. Wiher, and R. Cotterell, “Locally typical sampling,”Transactions of the Association for Computational Linguis- tics, vol. 11, pp. 102–121, 2023
2023
-
[69]
The hitchhiker’s guide to testing statistical significance in natural language processing,
R. Dror, G. Baumer, S. Shlomov, and R. Reichart, “The hitchhiker’s guide to testing statistical significance in natural language processing,” Proceedings of the 56th Annual Meeting of the Association for Compu- tational Linguistics, vol. 1, pp. 1383–1392, 2018
2018
-
[70]
Note on the sampling error of the difference between correlated proportions or percentages,
Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,”Psychometrika, vol. 12, no. 2, pp. 153–157, 1947
1947
-
[71]
Controlling the false discovery rate: a practical and powerful approach to multiple testing,
Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,”Journal of the Royal statistical society: series B (Methodological), vol. 57, no. 1, pp. 289–300, 1995
1995
-
[72]
Cohen,Statistical power analysis for the behavioral sciences
J. Cohen,Statistical power analysis for the behavioral sciences. Lawrence Erlbaum Associates, 1988
1988
-
[73]
Efron and R
B. Efron and R. J. Tibshirani,An introduction to the bootstrap. CRC press, 1994
1994
-
[74]
Statistical significance tests for machine translation evalua- tion,
P. Koehn, “Statistical significance tests for machine translation evalua- tion,” inProceedings of the 2004 conference on empirical methods in natural language processing, 2004, pp. 388–395
2004
-
[75]
Information-theoretic probing for linguistic structure,
T. Pimentel, J. Valvoda, R. Hall Maudslay, R. Zmigrod, A. Williams, and R. Cotterell, “Information-theoretic probing for linguistic structure,” arXiv preprint arXiv:2004.03061, 2020
2004 arXiv
-
[76]
A statistical analysis of summariza- tion evaluation metrics using resampling methods,
D. Deutsch, R. Dror, and D. Roth, “A statistical analysis of summariza- tion evaluation metrics using resampling methods,”Transactions of the Association for Computational Linguistics, vol. 9, pp. 1132–1146, 2021
2021
-
[77]
Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,
J. Pineau, P. Vincent-Lamarre, K. Sinha, V . Larivi `ere, A. Beygelzimer, F. d’Alch´e Buc, E. Fox, and H. Larochelle, “Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,”The Journal of Machine Learning Research, vo...
2019
-
[78]
Show your work: Improved reporting of experimental results,
J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show your work: Improved reporting of experimental results,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Languag...
2019
-
[79]
Data contamination: From memorization to exploitation,
I. Magar and R. Schwartz, “Data contamination: From memorization to exploitation,”arXiv preprint arXiv:2203.08242, 2022
2022 arXiv
-
[80]
Evading data contamination detection: Exploring test-time preprocessing methods,
J. Dekoninck, M. Whitehouse, M. Suppanen, S. Cullen, M. Kuchnik, and J. Mitchell, “Evading data contamination detection: Exploring test-time preprocessing methods,”arXiv preprint arXiv:2411.02643, 2024
2024 arXiv
-
[81]
Inverse scaling can become u-shaped,
J. Wei, N. Hou, A. Lampinen, X. Chen, D. Huang, Y . Tay, X. Chen, Y . Lu, D. Zhou, T. Maet al., “Inverse scaling can become u-shaped,” arXiv preprint arXiv:2211.02011, 2023
2023 arXiv
-
[82]
Inverse scaling: When bigger isn’t better,
I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liuet al., “Inverse scaling: When bigger isn’t better,”arXiv preprint arXiv:2306.09479, 2023
2023 arXiv
-
[83]
Predictability and surprise in large generative models,
D. Ganguli, D. Hernandez, L. Lovitt, N. DasSarma, T. Henighan, A. Jones, N. Joseph, J. Kernion, B. Mann, A. Askellet al., “Predictability and surprise in large generative models,”Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1747– 1764, 2022
2022
-
[84]
Large language models struggle to learn long-tail knowledge,
N. Kandpal, H. Deng, A. Roberts, E. Wallace, and C. Raffel, “Large language models struggle to learn long-tail knowledge,”International Conference on Machine Learning, pp. 15 696–15 707, 2023
2023
-
[85]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdh- ery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,”arXiv preprint arXiv:2203.11171, 2023
2023 arXiv
-
[86]
Rethink reporting of evaluation results in ai,
R. Burnell, W. Schellaert, J. Burden, T. D. Ullman, F. Martinez-Plumed, J. B. Tenenbaum, D. Rutar, L. G. Cheke, J. Sohl-Dickstein, M. Mitchell et al., “Rethink reporting of evaluation results in ai,”Science, vol. 380, no. 6641, pp. 136–138, 2023
2023
-
[87]
Impact of pretraining term frequencies on few-shot numerical reasoning,
Y . Razeghi, R. L. Logan IV , M. Gardner, and S. Singh, “Impact of pretraining term frequencies on few-shot numerical reasoning,”arXiv preprint arXiv:2202.07206, 2022
2022 arXiv
-
[88]
A systematic evaluation of large language models of code,
F. F. Xu, U. Alon, G. Neubig, and V . J. Hellendoorn, “A systematic evaluation of large language models of code,”Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming, pp. 1–10, 2022
2022
-
[89]
Energy and policy consider- ations for deep learning in nlp,
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy consider- ations for deep learning in nlp,”Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 3645–3650, 2019
2019
-
[90]
Carbon emissions and large neural network training,
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,”arXiv preprint arXiv:2104.10350, 2021
2021 arXiv
-
[91]
Scale efficiently: Insights from pretraining and finetuning transformers,
Y . Tay, M. Dehghani, J. Rao, W. Fedus, S. Abnar, H. W. Chung, S. Narang, D. Yogatama, A. Vaswani, and D. Metzler, “Scale efficiently: Insights from pretraining and finetuning transformers,”arXiv preprint arXiv:2109.10686, 2022
2022 arXiv
-
[92]
Teoria statistica delle classi e calcolo delle probabilita,
C. Bonferroni, “Teoria statistica delle classi e calcolo delle probabilita,” Pubblicazioni del R Istituto Superiore di Scienze Economiche e Com- mericiali di Firenze, vol. 8, pp. 3–62, 1936
1936
-
[93]
New effect size rules of thumb,
S. S. Sawilowsky, “New effect size rules of thumb,”Journal of modern applied statistical methods, vol. 8, no. 2, p. 26, 2009
2009
-
[94]
The measurement of observer agreement for categorical data,
J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,”biometrics, pp. 159–174, 1977
1977
-
[95]
Prior and posterior networks: A survey on evidential deep learning methods for uncertainty estima- tion,
D. Ulmer, C. Hardmeier, and J. Frellsen, “Prior and posterior networks: A survey on evidential deep learning methods for uncertainty estima- tion,”arXiv preprint arXiv:2110.03051, 2022
2022 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.