REVIEW 3 major objections 2 minor 95 references
The dense NK fitness landscape is exactly solvable: its free energy and maximum fitness have explicit thermodynamic limits, and near-optimal peaks form an exponentially large orthogonal set.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection The abstract promises a real result for the dense NK regime, but the supplied full text is a different paper; verdict must wait on the actual PDF. the 3 major comments →
On the Fitness Landscape in the $NK$ Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
In the dense regime $K/N \to \alpha \in (0,1]$, the NK random field is treated as a disordered spin system on the hypercube $\{0,1\}^N$. The main result gives the exact limits, per locus, of the quenched free energy at all temperatures and of the maximum fitness, with the limits depending only on $\alpha$. The geometric companion is an overlap-gap statement: any two genomes whose fitness is within a small window of the global maximum are asymptotically orthogonal, i.e. they differ on almost all loci, and the family of such near-fittest genomes contains exponentially many mutually orthogonal members. Consequently, near-fittest evolutionary paths become impossible when fitness levels approach
What carries the argument
The central object is the NK random field viewed as a spin-glass Hamiltonian on the $N$-locus hypercube, where each locus interacts with $K$ randomly selected others; in the dense regime, the analysis is carried by the spin-glass free-energy variational principle. The load-bearing mechanism is the overlap-gap property: near-maximum genomes have asymptotically zero pairwise overlap, which converts the geometric claim about many peaks into a quantitative statement about the overlap distribution of high-fitness configurations.
Load-bearing premise
The derivation assumes the dense NK random field converges to a spin-glass system for which a rigorous free-energy variational principle holds; if that convergence fails, the claimed exact limits do not follow from the method.
What would settle it
Run the NK model for increasing $N$ with $K/N$ fixed at several values of $\alpha$; if the empirical free energy or maximum fitness per locus does not approach the paper's predicted limits, the central claim is wrong. A more direct check is exhaustive enumeration for small $N$: the number of near-fittest mutually orthogonal genomes should grow exponentially in $N$ at a rate depending on $\alpha$.
If this is right
- For any $\alpha \in (0,1]$, the free energy per locus and the maximum fitness per locus have definite large-$N$ limits, so numerical simulations of the NK model have analytic benchmarks to match.
- The landscape is provably rugged: exponentially many near-fittest genomes are pairwise almost completely different, so no single narrow region of genome space contains the top of the landscape.
- Adaptation to the global maximum is blocked in a precise sense: paths connecting near-fittest genomes do not exist as the fitness threshold approaches the maximum.
- For small interaction density $\alpha$, evolution can still proceed without fitness loss: a fitness-maintaining path exists with high probability.
Where Pith is reading between the lines
- The supplied body text under this title is a different manuscript (an evaluation of language models), so this extraction rests on the title, abstract, and subject class; the quantitative theorems should be verified against the original source before citation.
- If the variational-principle route is correct, the dense NK model should sit in the same universality class as other Gaussian spin glasses, which would imply that locating the global maximum is computationally hard for the same reason it is hard in those models.
- The rate at which the number of near-fittest orthogonal genomes grows in $N$ (the exponential rate as a function of $\alpha$) is not stated in the abstract; measuring that rate numerically would be a direct test of the geometric picture.
- The small-$\alpha$ path result suggests a sharp threshold in $\alpha$ between path-connected and path-blocked fitness landscapes; locating that threshold is a natural next problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper, as represented by its arXiv metadata and abstract (arXiv:2508.12464, math.PR), claims to study the NK model in the regime K/N -> alpha in (0,1], using spin-glass methods. The abstract announces exact limits for the free energy at any temperature and for the maximum fitness, an exponentially large collection of near-fittest pairwise asymptotically orthogonal genomes, overlap gap properties, and quantitative statements about the possibility/impossibility of evolutionary paths near the global maximum. However, the supplied full text is not an NK-model paper at all; it is an unrelated evaluation study of OpenAI's GPT-OSS language models (arXiv:2508.12461v3). No NK model is defined, no main theorem is stated, and no proof, lemma, or derivation supporting any of the announced results appears anywhere in the manuscript body. Consequently, none of the paper's central claims can be checked from the submitted material.
Significance. If the results announced in the abstract were properly proved, they would constitute a substantial contribution to the mathematical theory of the NK model: exact solution of the dense regime (K/N -> alpha), a rigorous variational formula for the free energy, and a quantitative picture of the fitness landscape's multiple-peak geometry. These would be of interest to probabilists, statistical physicists, and theoretical evolutionary biologists. I cannot assess the significance of the actual contribution, however, because the manuscript contains no statement of the claimed theorems beyond the abstract and no supporting argument. The only verifiable content is an unrelated LLM evaluation, which has no bearing on the NK model. There is no evidence of machine-checked proofs, reproducible derivations, or even precise definitions on which to judge the mathematical claims.
major comments (3)
- [Full text (whole body)] The body of the submitted manuscript is 'Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models' (arXiv:2508.12461v3), not an NK-model paper. There is no definition of the NK model, no random-field construction, no theorem statements, no proofs, and no overlap-gap definitions. The abstract's central claims—exact free-energy limits, maximum fitness, exponential near-fittest orthogonal genomes, and evolutionary-path impossibilities—are therefore entirely unsupported by the supplied text. This is not a presentation issue but a complete absence of the mathematical content that the abstract announces.
- [Abstract] The abstract asserts that the NK model is treated 'via the spin glass methodologies' and that this yields 'exact limits.' No precise conditions are stated: the distribution of the NK random field, the nature of the K/N -> alpha limit, the convergence of the field to a spin-glass-type model, or the variational principle being invoked are all unspecified. If the intended paper contains a proof, these elements are load-bearing; in the present manuscript they appear only as assertions. The claimed exact limits do not follow from the stated method without these details.
- [Title and submission metadata] The manuscript's title, subject class (math.PR), and abstract concern the NK model, but the actual text is an empirical evaluation of language models. This internal inconsistency means the submission does not contain the research it claims to present. Even as a 'wrong file' scenario, the submitted manuscript is not refereeable as a probability paper, and the issue cannot be repaired by minor local edits within the current text.
minor comments (2)
- [References] The reference list is the bibliography of the GPT-OSS evaluation paper (e.g., Radford et al., Hendrycks et al., Kaplan et al.). No references to the NK model literature—Kauffman, Levin, Weinberger, or subsequent mathematical works on NK landscapes—appear, which is inconsistent with the abstract's claims.
- [Figures and tables] All tables and figures (e.g., Table I on model parameters, Figures 1–5 on benchmark results) concern large language models. None provide any numerical, exact, or illustrative support for the NK-model results described in the abstract.
Circularity Check
No circularity demonstrable: the supplied full text is a different arXiv paper (2508.12461v3), so the NK-model derivation chain is absent and no reduction-to-input can be exhibited.
full rationale
The abstract of arXiv:2508.12464 claims exact free-energy limits, maximum fitness, and multiple-peak structure in the NK model obtained 'via the spin glass methodologies.' The provided full text, however, is the manuscript of arXiv:2508.12461v3 ('Is GPT-OSS Good? ...'), an LLM evaluation paper with no NK-model content. Consequently there are no equations, coupling constructions, interpolation arguments, or overlap-gap derivations from the claimed NK paper to audit. Under the hard rules, circularity can only be claimed when the specific reduction is exhibited (Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction). No such reduction is present. The mismatch is a verification blocker, not evidence of circularity: the absence of proof cannot itself show that the exact limits are equivalent to the inputs. No self-citation chain, no ansatz smuggled by citation, and no fitted-input-called-prediction pattern can be identified from the abstract alone. The abstract's invocation of 'spin glass methodologies' is load-bearing but unverified in the supplied text; that is a correctness or completeness concern, not a demonstrated circular step. The honest finding is therefore a non-finding: no significant circularity can be established from the available material.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of On the Fitness Landscape in the $NK$ Model." pith.science (2026). https://pith.science/paper/76WY6YDF
@misc{pith2026250812464,
author = {Pith},
title = {Pith review of: On the Fitness Landscape in the $NK$ Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/76WY6YDF}},
note = {Machine review of arXiv:2508.12464}
}
abstract
The $NK$ model, introduced by Kauffman, Levin, and Weinberger, is a random field used to describe the fitness landscape of certain species with $N$ genetic loci, each interacting with $K$ others. The model has wide applications in understanding evolutionary and natural selection as it captures ruggedness feature of the fitness landscape. Earlier literature has been focused on the case $K$ being a fixed positive integer and used tools from Ergodic and Markov theory. In this paper, by viewing it as a statistical physics object, we investigate the $NK$ model in the regime $K/N\to\alpha \in(0,1]$ via the spin glass methodologies. Our main result identifies the exact limits for the free energy at any temperature and the maximum fitness. Moreover, we show that the $NK$ model exhibits a multiple-peak structure, namely, the number of near-fittest genomes that are asymptotically orthogonal to each other is exponentially large. Based on establishing the overlap gap properties, we obtain quantitative descriptions for the geometry of the fitness landscape and deduce that, in particular, near-fittest evolutionary paths become impossible as the fitness levels of the genomes approach the global maximum for any $\alpha\in (0,1]$. Nevertheless, we also show that by choosing $\alpha$ sufficiently small, an evolutionary path maintained at a given fitness level can be constructed with high probability.
Reference graph
Works this paper leans on
-
[1]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” inInternational Conference on Learning Representations, 2017
2017
-
[2]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”The Journal of Machine Learning Research, vol. 23, no. 1, pp. 5232–5270, 2022
2022
-
[3]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,”arXiv preprint arXiv:2001.08361, 2020
Pith/arXiv arXiv 2001
-
[4]
Scaling laws for autoregressive generative modeling,
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Grayet al., “Scaling laws for autoregressive generative modeling,”arXiv preprint arXiv:2010.14701, 2020
Pith/arXiv arXiv 2010
-
[5]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., “Training compute-optimal large language models,”arXiv preprint arXiv:2203.15556, 2022
Pith/arXiv arXiv 2022
-
[6]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskeveret al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[7]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azharet al., “Llama: Open and efficient foundation language models,”arXiv preprint arXiv:2302.13971, 2023
Pith/arXiv arXiv 2023
-
[8]
Gemma: Open models based on gemini research and technology,
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Loveet al., “Gemma: Open models based on gemini research and technology,”arXiv preprint arXiv:2403.08295, 2024
Pith/arXiv arXiv 2024
-
[9]
Deepseek llm: Scaling open-source language models with longtermism,
D.-A. and:, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Donget al., “Deepseek llm: Scaling open-source language models with longtermism,”arXiv preprint arXiv:2401.02954, 2024
Pith/arXiv arXiv 2024
-
[10]
J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huanget al., “Qwen technical report,”arXiv preprint arXiv:2309.16609, 2023
Pith/arXiv arXiv 2023
-
[11]
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y . Fan, Y . Su, Y . Zhang, Y . Wan, Y . Liu, Z. Cui, Z. Zhang, ...
Pith/arXiv arXiv 2024
-
[12]
Phi- 3 technical report: A highly capable language model locally on your phone,
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behlet al., “Phi- 3 technical report: A highly capable language model locally on your phone,”arXiv preprint arXiv:2404.14219, 2024
Pith/arXiv arXiv 2024
-
[13]
Emergent abilities of large language models,
J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzleret al., “Emergent abilities of large language models,”arXiv preprint arXiv:2206.07682, 2022
Pith/arXiv arXiv 2022
-
[14]
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?”Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[15]
Measuring massive multitask language understanding,
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2021
Pith/arXiv arXiv 2009
-
[16]
Training verifiers to solve math word problems,
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakanoet al., “Training verifiers to solve math word problems,”arXiv preprint arXiv:2110.14168, 2021
Pith/arXiv arXiv 2021
-
[17]
Solv- ing quantitative reasoning problems with language models,
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., “Solv- ing quantitative reasoning problems with language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022
2022
-
[18]
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,”arXiv preprint arXiv:2107.03374, 2021
Pith/arXiv arXiv 2021
-
[19]
Program synthesis with large language models,
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Leet al., “Program synthesis with large language models,”arXiv preprint arXiv:2108.07732, 2021
Pith/arXiv arXiv 2021
-
[20]
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chimet al., “Starcoder: may the source be with you!”arXiv preprint arXiv:2305.06161, 2023
Pith/arXiv arXiv 2023
-
[21]
Codegen: An open large language model for code with multi-turn program synthesis,
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y . Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,”arXiv preprint arXiv:2203.13474, 2022
Pith/arXiv arXiv 2022
-
[22]
Incoder: A generative model for code infilling and synthesis,
D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis, “Incoder: A generative model for code infilling and synthesis,”arXiv preprint arXiv:2204.05999, 2022
Pith/arXiv arXiv 2022
-
[23]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,
Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fuet al., “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 991–63 010, 2023
2023
-
[24]
Language models are few-shot multilingual learners,
G. I. Winata, A. Madotto, Z. Lin, R. Liu, J. Yosinski, and P. Fung, “Language models are few-shot multilingual learners,”arXiv preprint arXiv:2109.07684, 2021
Pith/arXiv arXiv 2021
-
[25]
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”arXiv preprint arXiv:2306.05685, 2023
Pith/arXiv arXiv 2023
-
[26]
Lamda: Language models for dialog applications,
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,”arXiv preprint arXiv:2201.08239, 2022
Pith/arXiv arXiv 2022
-
[27]
Recipes for building an open-domain chatbot,
S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y . Liu, J. Xu, M. Ott, E. M. Smith, Y .-L. Boureauet al., “Recipes for building an open-domain chatbot,”arXiv preprint arXiv:2004.13637, 2021
Pith/arXiv arXiv 2004
-
[28]
Glam: Efficient scaling of language models with mixture-of-experts,
N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firatet al., “Glam: Efficient scaling of language models with mixture-of-experts,”International Conference on Machine Learning, pp. 5547–5569, 2022
2022
-
[29]
Retrieval- augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschelet al., “Retrieval- augmented generation for knowledge-intensive nlp tasks,”Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020
2020
-
[30]
Efficient large scale language modeling with mixtures of experts,
M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V . Lin, J. Du, S. Iyer, R. Pasunuruet al., “Efficient large scale language modeling with mixtures of experts,”arXiv preprint arXiv:2112.10684, 2021
Pith/arXiv arXiv 2021
-
[31]
Mixture-of-experts with expert choice routing,
Y . Zhou, T. Lei, H. Liu, N. Du, Y . Huang, V . Zhao, A. M. Daiet al., “Mixture-of-experts with expert choice routing,”Advances in Neural Information Processing Systems, vol. 35, pp. 7103–7114, 2022
2022
-
[32]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosaleet al., “Llama 2: Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023
Pith/arXiv arXiv 2023
-
[33]
The falcon series of open language models,
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al., “The falcon series of open language models,”arXiv preprint arXiv:2311.16867, 2023
Pith/arXiv arXiv 2023
-
[34]
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,
T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”Advances in Neural Information Processing Systems, vol. 35, pp. 30 318–30 332, 2022
2022
-
[35]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021
Pith/arXiv arXiv 2021
-
[36]
G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosenet al., “Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,”arXiv preprint arXiv:2507.06261, 2025
Pith/arXiv arXiv 2025
-
[37]
(2025) Claude opus 4.1
Anthropic. (2025) Claude opus 4.1. Accessed: 2025-08-26. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-1
2025
-
[38]
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[39]
D. Team. (2025) Deepseek-v3.1 release report. Accessed: 2025- 08-26. [Online]. Available: https://api-docs.deepseek.com/zh-cn/news/ news250821
2025
-
[40]
Eight things to know about large language models,
S. R. Bowman, “Eight things to know about large language models,” arXiv preprint arXiv:2304.00612, 2023
Pith/arXiv arXiv 2023
-
[41]
Ai and the everything in the whole wide world benchmark,
I. D. Raji, E. Denton, E. M. Bender, A. Hanna, and A. Paullada, “Ai and the everything in the whole wide world benchmark,”arXiv preprint arXiv:2111.15366, 2021
Pith/arXiv arXiv 2021
-
[42]
Holistic evaluation of language models,
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,” inTransactions on Machine Learning Research, 2023
2023
-
[43]
With little power comes great responsibility,
D. Card, P. Henderson, U. Khandelwal, R. Jia, K. Mahowald, and D. Ju- rafsky, “With little power comes great responsibility,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 9263–9274
2020
-
[44]
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,
O. Sainz, J. A. Campos, I. Garc ´ıa-Ferrero, J. Etxaniz, O. L. de La- calle, and E. Agirre, “Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,”arXiv preprint arXiv:2310.18018, 2023
Pith/arXiv arXiv 2023
-
[45]
A survey of large language models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Donget al., “A survey of large language models,”arXiv preprint arXiv:2303.18223, 2023
Pith/arXiv arXiv 2023
-
[46]
Large language models: A survey,
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Am- atriain, and J. Gao, “Large language models: A survey,”arXiv preprint arXiv:2402.06196, 2024
Pith/arXiv arXiv 2024
-
[47]
llama.cpp: Inference of llama model in pure c/c++,
G. Gerganov, “llama.cpp: Inference of llama model in pure c/c++,” https: //github.com/ggerganov/llama.cpp, 2023
2023
-
[48]
Gshard: Scaling giant models with conditional computation and automatic sharding,
D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,”arXiv preprint arXiv:2006.16668, 2021
Pith/arXiv arXiv 2006
-
[49]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[50]
Scaling language models: Methods, analysis & insights from training gopher,
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Younget al., “Scaling language models: Methods, analysis & insights from training gopher,”arXiv preprint arXiv:2112.11446, 2021
Pith/arXiv arXiv 2021
-
[51]
Palm: Scaling language modeling with pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmannet al., “Palm: Scaling language modeling with pathways,”arXiv preprint arXiv:2204.02311, 2022
Pith/arXiv arXiv 2022
-
[52]
Finqa: A dataset of numerical reasoning over financial data,
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y . Wang, “Finqa: A dataset of numerical reasoning over financial data,” 2022. [Online]. Available: https://arxiv.org/abs/2109.00122
Pith/arXiv arXiv 2022
-
[53]
Medical Exam Question Answering with Large-scale Reading Comprehension
X. Zhang, J. Wu, Z. He, X. Liu, and Y . Su, “Medical exam question answering with large-scale reading comprehension,” 2018. [Online]. Available: https://arxiv.org/abs/1802.10279
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[54]
Experimenting with Legal AI Solutions: The Case of Question-Answering for Access to Justice
J. Li, R. Bhambhoria, S. Dahan, and X. Zhu, “Experimenting with legal ai solutions: The case of question-answering for access to justice,”arXiv preprint arXiv:2409.07713, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[55]
Crowdsourcing multiple choice science questions,
J. Welbl, N. F. Liu, and M. Gardner, “Crowdsourcing multiple choice science questions,” 2017. [Online]. Available: https://arxiv.org/abs/1707. 06209
work page 2017
-
[56]
Piqa: Reasoning about physical commonsense in natural language,
Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” 2019. [Online]. Available: https://arxiv.org/abs/1911.11641
Pith/arXiv arXiv 2019
-
[57]
Dialogsum: A real-life scenario dialogue summarization dataset,
Y . Chen, Y . Liu, L. Chen, and Y . Zhang, “Dialogsum: A real-life scenario dialogue summarization dataset,” 2021. [Online]. Available: https://arxiv.org/abs/2105.06762
Pith/arXiv arXiv 2021
-
[58]
Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,”International Conference on Machine Learning, pp. 4411–4421, 2020
work page 2020
-
[59]
mt5: A massively multilingual pre-trained text-to-text transformer,
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,”arXiv preprint arXiv:2010.11934, 2021
Pith/arXiv arXiv 2010
-
[60]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”arXiv preprint arXiv:2206.04615, 2022
Pith/arXiv arXiv 2022
-
[61]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language mod- els are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[62]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 24 824–24 837, 2022
2022
-
[63]
Large lan- guage models are zero-shot reasoners,
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large lan- guage models are zero-shot reasoners,”Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022
2022
-
[64]
Large language models are not fair evaluators,
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” arXiv preprint arXiv:2305.17926, 2023
Pith/arXiv arXiv 2023
-
[65]
The curious case of neural text degeneration,
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y . Choi, “The curious case of neural text degeneration,”arXiv preprint arXiv:1904.09751, 2019
Pith/arXiv arXiv 1904
-
[66]
Neural text generation with unlikelihood training,
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston, “Neural text generation with unlikelihood training,”arXiv preprint arXiv:1908.04319, 2019
Pith/arXiv arXiv 1908
-
[67]
L. Ruis, A. K. Kemfert, D. Hupkes, L. Kaiser, and B. M. Lake, “The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,”Advances in Neural Information Processing Systems, vol. 37, 2024
work page 2024
-
[68]
C. Meister, T. Pimentel, G. Wiher, and R. Cotterell, “Locally typical sampling,”Transactions of the Association for Computational Linguis- tics, vol. 11, pp. 102–121, 2023
work page 2023
-
[69]
The hitchhiker’s guide to testing statistical significance in natural language processing,
R. Dror, G. Baumer, S. Shlomov, and R. Reichart, “The hitchhiker’s guide to testing statistical significance in natural language processing,” Proceedings of the 56th Annual Meeting of the Association for Compu- tational Linguistics, vol. 1, pp. 1383–1392, 2018
work page 2018
-
[70]
Note on the sampling error of the difference between correlated proportions or percentages,
Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,”Psychometrika, vol. 12, no. 2, pp. 153–157, 1947
1947
-
[71]
Controlling the false discovery rate: a practical and powerful approach to multiple testing,
Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,”Journal of the Royal statistical society: series B (Methodological), vol. 57, no. 1, pp. 289–300, 1995
1995
-
[72]
Cohen,Statistical power analysis for the behavioral sciences
J. Cohen,Statistical power analysis for the behavioral sciences. Lawrence Erlbaum Associates, 1988
work page 1988
-
[73]
B. Efron and R. J. Tibshirani,An introduction to the bootstrap. CRC press, 1994
work page 1994
-
[74]
Statistical significance tests for machine translation evalua- tion,
P. Koehn, “Statistical significance tests for machine translation evalua- tion,” inProceedings of the 2004 conference on empirical methods in natural language processing, 2004, pp. 388–395
work page 2004
-
[75]
Information-theoretic probing for linguistic structure,
T. Pimentel, J. Valvoda, R. Hall Maudslay, R. Zmigrod, A. Williams, and R. Cotterell, “Information-theoretic probing for linguistic structure,” arXiv preprint arXiv:2004.03061, 2020
Pith/arXiv arXiv 2004
-
[76]
A statistical analysis of summariza- tion evaluation metrics using resampling methods,
D. Deutsch, R. Dror, and D. Roth, “A statistical analysis of summariza- tion evaluation metrics using resampling methods,”Transactions of the Association for Computational Linguistics, vol. 9, pp. 1132–1146, 2021
work page 2021
-
[77]
J. Pineau, P. Vincent-Lamarre, K. Sinha, V . Larivi `ere, A. Beygelzimer, F. d’Alch´e Buc, E. Fox, and H. Larochelle, “Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,”The Journal of Machine Learning Research, vol. 22, no. 1, pp. 7459–7478, 2021
work page 2019
-
[78]
Show your work: Improved reporting of experimental results,
J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show your work: Improved reporting of experimental results,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 2185–2194
work page 2019
-
[79]
Data contamination: From memorization to exploitation,
I. Magar and R. Schwartz, “Data contamination: From memorization to exploitation,”arXiv preprint arXiv:2203.08242, 2022
Pith/arXiv arXiv 2022
-
[80]
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
J. Dekoninck, M. Whitehouse, M. Suppanen, S. Cullen, M. Kuchnik, and J. Mitchell, “Evading data contamination detection: Exploring test-time preprocessing methods,”arXiv preprint arXiv:2411.02643, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.