Pith. sign in

REVIEW 3 major objections 2 minor 95 references

On the Fitness Landscape in the $NK$ Model

T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The dense NK fitness landscape is exactly solvable: its free energy and maximum fitness have explicit thermodynamic limits, and near-optimal peaks form an exponentially large orthogonal set.

desk verdict The abstract promises a real result for the dense NK regime, but the supplied full text is a different paper; verdict must wait on the actual PDF. read the letter →

arxiv 2508.12464 v1 pith:76WY6YDF submitted 2025-08-17 math.PR

classification math.PR MSC 60K3582B44
keywords NKmodelfitnesslandscapespinglassfreeenergymaximumoverlapgappropertyevolutionarypaths
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that when $K/N \to \alpha \in (0,1]$, the NK model of fitness landscapes can be solved exactly in the thermodynamic limit. The authors derive the limiting free energy at any temperature and the limiting maximum fitness, and they show that near-fittest genomes are asymptotically orthogonal to each other and exponentially numerous. This matters because the NK model is a standard picture of rugged evolution; exact limits turn heuristic intuition about ruggedness into quantitative statements about whether adaptive paths can reach the global maximum. The paper also proves that near-fittest evolutionary paths disappear as fitness approaches the global maximum for every $\alpha \in (0,1]$, while for small $\alpha$ a path held at a fixed fitness level exists with high probability.

What carries the argument

The central object is the NK random field viewed as a spin-glass Hamiltonian on the $N$-locus hypercube, where each locus interacts with $K$ randomly selected others; in the dense regime, the analysis is carried by the spin-glass free-energy variational principle. The load-bearing mechanism is the overlap-gap property: near-maximum genomes have asymptotically zero pairwise overlap, which converts the geometric claim about many peaks into a quantitative statement about the overlap distribution of high-fitness configurations.

What would settle it

Run the NK model for increasing $N$ with $K/N$ fixed at several values of $\alpha$; if the empirical free energy or maximum fitness per locus does not approach the paper's predicted limits, the central claim is wrong. A more direct check is exhaustive enumeration for small $N$: the number of near-fittest mutually orthogonal genomes should grow exponentially in $N$ at a rate depending on $\alpha$.

Watch

Extended reading notes

Core claim

In the dense regime $K/N \to \alpha \in (0,1]$, the NK random field is treated as a disordered spin system on the hypercube $\{0,1\}^N$. The main result gives the exact limits, per locus, of the quenched free energy at all temperatures and of the maximum fitness, with the limits depending only on $\alpha$. The geometric companion is an overlap-gap statement: any two genomes whose fitness is within a small window of the global maximum are asymptotically orthogonal, i.e. they differ on almost all loci, and the family of such near-fittest genomes contains exponentially many mutually orthogonal members. Consequently, near-fittest evolutionary paths become impossible when fitness levels approach

Load-bearing premise

The derivation assumes the dense NK random field converges to a spin-glass system for which a rigorous free-energy variational principle holds; if that convergence fails, the claimed exact limits do not follow from the method.

Editorial extensions

If this is right

  • For any $\alpha \in (0,1]$, the free energy per locus and the maximum fitness per locus have definite large-$N$ limits, so numerical simulations of the NK model have analytic benchmarks to match.
  • The landscape is provably rugged: exponentially many near-fittest genomes are pairwise almost completely different, so no single narrow region of genome space contains the top of the landscape.
  • Adaptation to the global maximum is blocked in a precise sense: paths connecting near-fittest genomes do not exist as the fitness threshold approaches the maximum.
  • For small interaction density $\alpha$, evolution can still proceed without fitness loss: a fitness-maintaining path exists with high probability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The supplied body text under this title is a different manuscript (an evaluation of language models), so this extraction rests on the title, abstract, and subject class; the quantitative theorems should be verified against the original source before citation.
  • If the variational-principle route is correct, the dense NK model should sit in the same universality class as other Gaussian spin glasses, which would imply that locating the global maximum is computationally hard for the same reason it is hard in those models.
  • The rate at which the number of near-fittest orthogonal genomes grows in $N$ (the exponential rate as a function of $\alpha$) is not stated in the abstract; measuring that rate numerically would be a direct test of the geometric picture.
  • The small-$\alpha$ path result suggests a sharp threshold in $\alpha$ between path-connected and path-blocked fitness landscapes; locating that threshold is a natural next problem.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper, as represented by its arXiv metadata and abstract (arXiv:2508.12464, math.PR), claims to study the NK model in the regime K/N -> alpha in (0,1], using spin-glass methods. The abstract announces exact limits for the free energy at any temperature and for the maximum fitness, an exponentially large collection of near-fittest pairwise asymptotically orthogonal genomes, overlap gap properties, and quantitative statements about the possibility/impossibility of evolutionary paths near the global maximum. However, the supplied full text is not an NK-model paper at all; it is an unrelated evaluation study of OpenAI's GPT-OSS language models (arXiv:2508.12461v3). No NK model is defined, no main theorem is stated, and no proof, lemma, or derivation supporting any of the announced results appears anywhere in the manuscript body. Consequently, none of the paper's central claims can be checked from the submitted material.

Significance. If the results announced in the abstract were properly proved, they would constitute a substantial contribution to the mathematical theory of the NK model: exact solution of the dense regime (K/N -> alpha), a rigorous variational formula for the free energy, and a quantitative picture of the fitness landscape's multiple-peak geometry. These would be of interest to probabilists, statistical physicists, and theoretical evolutionary biologists. I cannot assess the significance of the actual contribution, however, because the manuscript contains no statement of the claimed theorems beyond the abstract and no supporting argument. The only verifiable content is an unrelated LLM evaluation, which has no bearing on the NK model. There is no evidence of machine-checked proofs, reproducible derivations, or even precise definitions on which to judge the mathematical claims.

major comments (3)
  1. [Full text (whole body)] The body of the submitted manuscript is 'Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models' (arXiv:2508.12461v3), not an NK-model paper. There is no definition of the NK model, no random-field construction, no theorem statements, no proofs, and no overlap-gap definitions. The abstract's central claims—exact free-energy limits, maximum fitness, exponential near-fittest orthogonal genomes, and evolutionary-path impossibilities—are therefore entirely unsupported by the supplied text. This is not a presentation issue but a complete absence of the mathematical content that the abstract announces.
  2. [Abstract] The abstract asserts that the NK model is treated 'via the spin glass methodologies' and that this yields 'exact limits.' No precise conditions are stated: the distribution of the NK random field, the nature of the K/N -> alpha limit, the convergence of the field to a spin-glass-type model, or the variational principle being invoked are all unspecified. If the intended paper contains a proof, these elements are load-bearing; in the present manuscript they appear only as assertions. The claimed exact limits do not follow from the stated method without these details.
  3. [Title and submission metadata] The manuscript's title, subject class (math.PR), and abstract concern the NK model, but the actual text is an empirical evaluation of language models. This internal inconsistency means the submission does not contain the research it claims to present. Even as a 'wrong file' scenario, the submitted manuscript is not refereeable as a probability paper, and the issue cannot be repaired by minor local edits within the current text.
minor comments (2)
  1. [References] The reference list is the bibliography of the GPT-OSS evaluation paper (e.g., Radford et al., Hendrycks et al., Kaplan et al.). No references to the NK model literature—Kauffman, Levin, Weinberger, or subsequent mathematical works on NK landscapes—appear, which is inconsistent with the abstract's claims.
  2. [Figures and tables] All tables and figures (e.g., Table I on model parameters, Figures 1–5 on benchmark results) concern large language models. None provide any numerical, exact, or illustrative support for the NK-model results described in the abstract.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity demonstrable: the supplied full text is a different arXiv paper (2508.12461v3), so the NK-model derivation chain is absent and no reduction-to-input can be exhibited.

full rationale

The abstract of arXiv:2508.12464 claims exact free-energy limits, maximum fitness, and multiple-peak structure in the NK model obtained 'via the spin glass methodologies.' The provided full text, however, is the manuscript of arXiv:2508.12461v3 ('Is GPT-OSS Good? ...'), an LLM evaluation paper with no NK-model content. Consequently there are no equations, coupling constructions, interpolation arguments, or overlap-gap derivations from the claimed NK paper to audit. Under the hard rules, circularity can only be claimed when the specific reduction is exhibited (Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction). No such reduction is present. The mismatch is a verification blocker, not evidence of circularity: the absence of proof cannot itself show that the exact limits are equivalent to the inputs. No self-citation chain, no ansatz smuggled by citation, and no fitted-input-called-prediction pattern can be identified from the abstract alone. The abstract's invocation of 'spin glass methodologies' is load-bearing but unverified in the supplied text; that is a correctness or completeness concern, not a demonstrated circular step. The honest finding is therefore a non-finding: no significant circularity can be established from the available material.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The actual NK manuscript was not available, so no axiology could be performed. The abstract alone gives no fitted parameters or invented entities; alpha is the fixed asymptotic ratio, not a fitted quantity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Fitness Landscape in the $NK$ Model." pith.science (2026). https://pith.science/paper/76WY6YDF

@misc{pith2026250812464,
  author       = {Pith},
  title        = {Pith review of: On the Fitness Landscape in the $NK$ Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/76WY6YDF}},
  note         = {Machine review of arXiv:2508.12464}
}
abstract

The $NK$ model, introduced by Kauffman, Levin, and Weinberger, is a random field used to describe the fitness landscape of certain species with $N$ genetic loci, each interacting with $K$ others. The model has wide applications in understanding evolutionary and natural selection as it captures ruggedness feature of the fitness landscape. Earlier literature has been focused on the case $K$ being a fixed positive integer and used tools from Ergodic and Markov theory. In this paper, by viewing it as a statistical physics object, we investigate the $NK$ model in the regime $K/N\to\alpha \in(0,1]$ via the spin glass methodologies. Our main result identifies the exact limits for the free energy at any temperature and the maximum fitness. Moreover, we show that the $NK$ model exhibits a multiple-peak structure, namely, the number of near-fittest genomes that are asymptotically orthogonal to each other is exponentially large. Based on establishing the overlap gap properties, we obtain quantitative descriptions for the geometry of the fitness landscape and deduce that, in particular, near-fittest evolutionary paths become impossible as the fitness levels of the genomes approach the global maximum for any $\alpha\in (0,1]$. Nevertheless, we also show that by choosing $\alpha$ sufficiently small, an evolutionary path maintained at a given fitness level can be constructed with high probability.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

95 extracted references · 21 canonical work pages

  1. [1]

    Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” inInternational Conference on Learning Representations, 2017

  2. [2]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”The Journal of Machine Learning Research, vol. 23, no. 1, pp. 5232–5270, 2022

  3. [3]

    Scaling laws for neural language models,

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,”arXiv preprint arXiv:2001.08361, 2020

  4. [4]

    Scaling laws for autoregressive generative modeling,

    T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Grayet al., “Scaling laws for autoregressive generative modeling,”arXiv preprint arXiv:2010.14701, 2020

  5. [5]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., “Training compute-optimal large language models,”arXiv preprint arXiv:2203.15556, 2022

  6. [6]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskeveret al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019

  7. [7]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azharet al., “Llama: Open and efficient foundation language models,”arXiv preprint arXiv:2302.13971, 2023

  8. [8]

    Gemma: Open models based on gemini research and technology,

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Loveet al., “Gemma: Open models based on gemini research and technology,”arXiv preprint arXiv:2403.08295, 2024

Show all 95 references
  1. [9]

    Deepseek llm: Scaling open-source language models with longtermism,

    D.-A. and:, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Donget al., “Deepseek llm: Scaling open-source language models with longtermism,”arXiv preprint arXiv:2401.02954, 2024

  2. [10]

    Qwen technical report,

    J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huanget al., “Qwen technical report,”arXiv preprint arXiv:2309.16609, 2023

  3. [11]

    Qwen2.5 technical report,

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...

  4. [12]

    Phi- 3 technical report: A highly capable language model locally on your phone,

    M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behlet al., “Phi- 3 technical report: A highly capable language model locally on your phone,”arXiv preprint arXiv:2404.14219, 2024

  5. [13]

    Emergent abilities of large language models,

    J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzleret al., “Emergent abilities of large language models,”arXiv preprint arXiv:2206.07682, 2022

  6. [14]

    Are emergent abilities of large language models a mirage?

    R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?”Advances in Neural Information Processing Systems, vol. 36, 2024

  7. [15]

    Measuring massive multitask language understanding,

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2021

  8. [16]

    Training verifiers to solve math word problems,

    K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakanoet al., “Training verifiers to solve math word problems,”arXiv preprint arXiv:2110.14168, 2021

  9. [17]

    Solv- ing quantitative reasoning problems with language models,

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., “Solv- ing quantitative reasoning problems with language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022

  10. [18]

    Evaluating large language models trained on code,

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,”arXiv preprint arXiv:2107.03374, 2021

  11. [19]

    Program synthesis with large language models,

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Leet al., “Program synthesis with large language models,”arXiv preprint arXiv:2108.07732, 2021

  12. [20]

    Starcoder: may the source be with you!

    R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chimet al., “Starcoder: may the source be with you!”arXiv preprint arXiv:2305.06161, 2023

  13. [21]

    Codegen: An open large language model for code with multi-turn program synthesis,

    E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y . Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,”arXiv preprint arXiv:2203.13474, 2022

  14. [22]

    Incoder: A generative model for code infilling and synthesis,

    D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis, “Incoder: A generative model for code infilling and synthesis,”arXiv preprint arXiv:2204.05999, 2022

  15. [23]

    C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,

    Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fuet al., “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 991–63 010, 2023

  16. [24]

    Language models are few-shot multilingual learners,

    G. I. Winata, A. Madotto, Z. Lin, R. Liu, J. Yosinski, and P. Fung, “Language models are few-shot multilingual learners,”arXiv preprint arXiv:2109.07684, 2021

  17. [25]

    Judging llm-as-a-judge with mt-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”arXiv preprint arXiv:2306.05685, 2023

  18. [26]

    Lamda: Language models for dialog applications,

    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,”arXiv preprint arXiv:2201.08239, 2022

  19. [27]

    Recipes for building an open-domain chatbot,

    S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y . Liu, J. Xu, M. Ott, E. M. Smith, Y .-L. Boureauet al., “Recipes for building an open-domain chatbot,”arXiv preprint arXiv:2004.13637, 2021

  20. [28]

    Glam: Efficient scaling of language models with mixture-of-experts,

    N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firatet al., “Glam: Efficient scaling of language models with mixture-of-experts,”International Conference on Machine Learning, pp. 5547–5569, 2022

  21. [29]

    Retrieval- augmented generation for knowledge-intensive nlp tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschelet al., “Retrieval- augmented generation for knowledge-intensive nlp tasks,”Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020

  22. [30]

    Efficient large scale language modeling with mixtures of experts,

    M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V . Lin, J. Du, S. Iyer, R. Pasunuruet al., “Efficient large scale language modeling with mixtures of experts,”arXiv preprint arXiv:2112.10684, 2021

  23. [31]

    Mixture-of-experts with expert choice routing,

    Y . Zhou, T. Lei, H. Liu, N. Du, Y . Huang, V . Zhao, A. M. Daiet al., “Mixture-of-experts with expert choice routing,”Advances in Neural Information Processing Systems, vol. 35, pp. 7103–7114, 2022

  24. [32]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosaleet al., “Llama 2: Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023

  25. [33]

    The falcon series of open language models,

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al., “The falcon series of open language models,”arXiv preprint arXiv:2311.16867, 2023

  26. [34]

    Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,

    T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”Advances in Neural Information Processing Systems, vol. 35, pp. 30 318–30 332, 2022

  27. [35]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021

  28. [36]

    Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

    G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosenet al., “Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,”arXiv preprint arXiv...

  29. [37]

    (2025) Claude opus 4.1

    Anthropic. (2025) Claude opus 4.1. Accessed: 2025-08-26. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-1

  30. [38]

    Qwen3 technical report,

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025

  31. [39]

    D. Team. (2025) Deepseek-v3.1 release report. Accessed: 2025- 08-26. [Online]. Available: https://api-docs.deepseek.com/zh-cn/news/ news250821

  32. [40]

    Eight things to know about large language models,

    S. R. Bowman, “Eight things to know about large language models,” arXiv preprint arXiv:2304.00612, 2023

  33. [41]

    Ai and the everything in the whole wide world benchmark,

    I. D. Raji, E. Denton, E. M. Bender, A. Hanna, and A. Paullada, “Ai and the everything in the whole wide world benchmark,”arXiv preprint arXiv:2111.15366, 2021

  34. [42]

    Holistic evaluation of language models,

    P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,” inTransactions on Machine Learning Research, 2023

  35. [43]

    With little power comes great responsibility,

    D. Card, P. Henderson, U. Khandelwal, R. Jia, K. Mahowald, and D. Ju- rafsky, “With little power comes great responsibility,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 9263–9274

  36. [44]

    Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,

    O. Sainz, J. A. Campos, I. Garc ´ıa-Ferrero, J. Etxaniz, O. L. de La- calle, and E. Agirre, “Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,”arXiv preprint arXiv:2310.18018, 2023

  37. [45]

    A survey of large language models,

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Donget al., “A survey of large language models,”arXiv preprint arXiv:2303.18223, 2023

  38. [46]

    Large language models: A survey,

    S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Am- atriain, and J. Gao, “Large language models: A survey,”arXiv preprint arXiv:2402.06196, 2024

  39. [47]

    llama.cpp: Inference of llama model in pure c/c++,

    G. Gerganov, “llama.cpp: Inference of llama model in pure c/c++,” https: //github.com/ggerganov/llama.cpp, 2023

  40. [48]

    Gshard: Scaling giant models with conditional computation and automatic sharding,

    D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,”arXiv preprint arXiv:2006.16668, 2021

  41. [49]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  42. [50]

    Scaling language models: Methods, analysis & insights from training gopher,

    J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Younget al., “Scaling language models: Methods, analysis & insights from training gopher,”arXiv preprint arXiv:2112.11446, 2021

  43. [51]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmannet al., “Palm: Scaling language modeling with pathways,”arXiv preprint arXiv:2204.02311, 2022

  44. [52]

    Finqa: A dataset of numerical reasoning over financial data,

    Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y . Wang, “Finqa: A dataset of numerical reasoning over financial data,” 2022. [Online]. Available: https://arxiv.org/abs/2109.00122

  45. [53]

    Medical exam question answering with large-scale reading comprehension,

    X. Zhang, J. Wu, Z. He, X. Liu, and Y . Su, “Medical exam question answering with large-scale reading comprehension,” 2018. [Online]. Available: https://arxiv.org/abs/1802.10279

  46. [54]

    Experimenting with legal ai solutions: The case of question-answering for access to justice,

    J. Li, R. Bhambhoria, S. Dahan, and X. Zhu, “Experimenting with legal ai solutions: The case of question-answering for access to justice,”arXiv preprint arXiv:2409.07713, 2024

  47. [55]

    Crowdsourcing multiple choice science questions,

    J. Welbl, N. F. Liu, and M. Gardner, “Crowdsourcing multiple choice science questions,” 2017. [Online]. Available: https://arxiv.org/abs/1707. 06209

  48. [56]

    Piqa: Reasoning about physical commonsense in natural language,

    Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” 2019. [Online]. Available: https://arxiv.org/abs/1911.11641

  49. [57]

    Dialogsum: A real-life scenario dialogue summarization dataset,

    Y . Chen, Y . Liu, L. Chen, and Y . Zhang, “Dialogsum: A real-life scenario dialogue summarization dataset,” 2021. [Online]. Available: https://arxiv.org/abs/2105.06762

  50. [58]

    Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,

    J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,”International Conference on Machine Learning, pp. 4411–4421, 2020

  51. [59]

    mt5: A massively multilingual pre-trained text-to-text transformer,

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,”arXiv preprint arXiv:2010.11934, 2021

  52. [60]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

    A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”arXiv preprint arXiv:2206.04615, 2022

  53. [61]

    Language mod- els are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language mod- els are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  54. [62]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 24 824–24 837, 2022

  55. [63]

    Large lan- guage models are zero-shot reasoners,

    T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large lan- guage models are zero-shot reasoners,”Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022

  56. [64]

    Large language models are not fair evaluators,

    P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” arXiv preprint arXiv:2305.17926, 2023

  57. [65]

    The curious case of neural text degeneration,

    A. Holtzman, J. Buys, L. Du, M. Forbes, and Y . Choi, “The curious case of neural text degeneration,”arXiv preprint arXiv:1904.09751, 2019

  58. [66]

    Neural text generation with unlikelihood training,

    S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston, “Neural text generation with unlikelihood training,”arXiv preprint arXiv:1908.04319, 2019

  59. [67]

    The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,

    L. Ruis, A. K. Kemfert, D. Hupkes, L. Kaiser, and B. M. Lake, “The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,”Advances in Neural Information Processing Systems, vol. 37, 2024

  60. [68]

    Locally typical sampling,

    C. Meister, T. Pimentel, G. Wiher, and R. Cotterell, “Locally typical sampling,”Transactions of the Association for Computational Linguis- tics, vol. 11, pp. 102–121, 2023

  61. [69]

    The hitchhiker’s guide to testing statistical significance in natural language processing,

    R. Dror, G. Baumer, S. Shlomov, and R. Reichart, “The hitchhiker’s guide to testing statistical significance in natural language processing,” Proceedings of the 56th Annual Meeting of the Association for Compu- tational Linguistics, vol. 1, pp. 1383–1392, 2018

  62. [70]

    Note on the sampling error of the difference between correlated proportions or percentages,

    Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,”Psychometrika, vol. 12, no. 2, pp. 153–157, 1947

  63. [71]

    Controlling the false discovery rate: a practical and powerful approach to multiple testing,

    Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,”Journal of the Royal statistical society: series B (Methodological), vol. 57, no. 1, pp. 289–300, 1995

  64. [72]

    Cohen,Statistical power analysis for the behavioral sciences

    J. Cohen,Statistical power analysis for the behavioral sciences. Lawrence Erlbaum Associates, 1988

  65. [73]

    Efron and R

    B. Efron and R. J. Tibshirani,An introduction to the bootstrap. CRC press, 1994

  66. [74]

    Statistical significance tests for machine translation evalua- tion,

    P. Koehn, “Statistical significance tests for machine translation evalua- tion,” inProceedings of the 2004 conference on empirical methods in natural language processing, 2004, pp. 388–395

  67. [75]

    Information-theoretic probing for linguistic structure,

    T. Pimentel, J. Valvoda, R. Hall Maudslay, R. Zmigrod, A. Williams, and R. Cotterell, “Information-theoretic probing for linguistic structure,” arXiv preprint arXiv:2004.03061, 2020

  68. [76]

    A statistical analysis of summariza- tion evaluation metrics using resampling methods,

    D. Deutsch, R. Dror, and D. Roth, “A statistical analysis of summariza- tion evaluation metrics using resampling methods,”Transactions of the Association for Computational Linguistics, vol. 9, pp. 1132–1146, 2021

  69. [77]

    Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,

    J. Pineau, P. Vincent-Lamarre, K. Sinha, V . Larivi `ere, A. Beygelzimer, F. d’Alch´e Buc, E. Fox, and H. Larochelle, “Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,”The Journal of Machine Learning Research, vo...

  70. [78]

    Show your work: Improved reporting of experimental results,

    J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show your work: Improved reporting of experimental results,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Languag...

  71. [79]

    Data contamination: From memorization to exploitation,

    I. Magar and R. Schwartz, “Data contamination: From memorization to exploitation,”arXiv preprint arXiv:2203.08242, 2022

  72. [80]

    Evading data contamination detection: Exploring test-time preprocessing methods,

    J. Dekoninck, M. Whitehouse, M. Suppanen, S. Cullen, M. Kuchnik, and J. Mitchell, “Evading data contamination detection: Exploring test-time preprocessing methods,”arXiv preprint arXiv:2411.02643, 2024

  73. [81]

    Inverse scaling can become u-shaped,

    J. Wei, N. Hou, A. Lampinen, X. Chen, D. Huang, Y . Tay, X. Chen, Y . Lu, D. Zhou, T. Maet al., “Inverse scaling can become u-shaped,” arXiv preprint arXiv:2211.02011, 2023

  74. [82]

    Inverse scaling: When bigger isn’t better,

    I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liuet al., “Inverse scaling: When bigger isn’t better,”arXiv preprint arXiv:2306.09479, 2023

  75. [83]

    Predictability and surprise in large generative models,

    D. Ganguli, D. Hernandez, L. Lovitt, N. DasSarma, T. Henighan, A. Jones, N. Joseph, J. Kernion, B. Mann, A. Askellet al., “Predictability and surprise in large generative models,”Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1747– 1764, 2022

  76. [84]

    Large language models struggle to learn long-tail knowledge,

    N. Kandpal, H. Deng, A. Roberts, E. Wallace, and C. Raffel, “Large language models struggle to learn long-tail knowledge,”International Conference on Machine Learning, pp. 15 696–15 707, 2023

  77. [85]

    Self-consistency improves chain of thought reasoning in language models,

    X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdh- ery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,”arXiv preprint arXiv:2203.11171, 2023

  78. [86]

    Rethink reporting of evaluation results in ai,

    R. Burnell, W. Schellaert, J. Burden, T. D. Ullman, F. Martinez-Plumed, J. B. Tenenbaum, D. Rutar, L. G. Cheke, J. Sohl-Dickstein, M. Mitchell et al., “Rethink reporting of evaluation results in ai,”Science, vol. 380, no. 6641, pp. 136–138, 2023

  79. [87]

    Impact of pretraining term frequencies on few-shot numerical reasoning,

    Y . Razeghi, R. L. Logan IV , M. Gardner, and S. Singh, “Impact of pretraining term frequencies on few-shot numerical reasoning,”arXiv preprint arXiv:2202.07206, 2022

  80. [88]

    A systematic evaluation of large language models of code,

    F. F. Xu, U. Alon, G. Neubig, and V . J. Hellendoorn, “A systematic evaluation of large language models of code,”Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming, pp. 1–10, 2022

  81. [89]

    Energy and policy consider- ations for deep learning in nlp,

    E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy consider- ations for deep learning in nlp,”Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 3645–3650, 2019

  82. [90]

    Carbon emissions and large neural network training,

    D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,”arXiv preprint arXiv:2104.10350, 2021

  83. [91]

    Scale efficiently: Insights from pretraining and finetuning transformers,

    Y . Tay, M. Dehghani, J. Rao, W. Fedus, S. Abnar, H. W. Chung, S. Narang, D. Yogatama, A. Vaswani, and D. Metzler, “Scale efficiently: Insights from pretraining and finetuning transformers,”arXiv preprint arXiv:2109.10686, 2022

  84. [92]

    Teoria statistica delle classi e calcolo delle probabilita,

    C. Bonferroni, “Teoria statistica delle classi e calcolo delle probabilita,” Pubblicazioni del R Istituto Superiore di Scienze Economiche e Com- mericiali di Firenze, vol. 8, pp. 3–62, 1936

  85. [93]

    New effect size rules of thumb,

    S. S. Sawilowsky, “New effect size rules of thumb,”Journal of modern applied statistical methods, vol. 8, no. 2, p. 26, 2009

  86. [94]

    The measurement of observer agreement for categorical data,

    J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,”biometrics, pp. 159–174, 1977

  87. [95]

    Prior and posterior networks: A survey on evidential deep learning methods for uncertainty estima- tion,

    D. Ulmer, C. Hardmeier, and J. Frellsen, “Prior and posterior networks: A survey on evidential deep learning methods for uncertainty estima- tion,”arXiv preprint arXiv:2110.03051, 2022

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.