Pith. sign in

REVIEW 3 major objections 2 minor 95 references

The dense NK fitness landscape is exactly solvable: its free energy and maximum fitness have explicit thermodynamic limits, and near-optimal peaks form an exponentially large orthogonal set.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection The abstract promises a real result for the dense NK regime, but the supplied full text is a different paper; verdict must wait on the actual PDF. the 3 major comments →

arxiv 2508.12464 v1 pith:76WY6YDF submitted 2025-08-17 math.PR

On the Fitness Landscape in the $NK$ Model

classification math.PR MSC 60K3582B44
keywords NK modelfitness landscapespin glassfree energymaximum fitnessoverlap gap propertyevolutionary paths
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that when $K/N \to \alpha \in (0,1]$, the NK model of fitness landscapes can be solved exactly in the thermodynamic limit. The authors derive the limiting free energy at any temperature and the limiting maximum fitness, and they show that near-fittest genomes are asymptotically orthogonal to each other and exponentially numerous. This matters because the NK model is a standard picture of rugged evolution; exact limits turn heuristic intuition about ruggedness into quantitative statements about whether adaptive paths can reach the global maximum. The paper also proves that near-fittest evolutionary paths disappear as fitness approaches the global maximum for every $\alpha \in (0,1]$, while for small $\alpha$ a path held at a fixed fitness level exists with high probability.

Core claim

In the dense regime $K/N \to \alpha \in (0,1]$, the NK random field is treated as a disordered spin system on the hypercube $\{0,1\}^N$. The main result gives the exact limits, per locus, of the quenched free energy at all temperatures and of the maximum fitness, with the limits depending only on $\alpha$. The geometric companion is an overlap-gap statement: any two genomes whose fitness is within a small window of the global maximum are asymptotically orthogonal, i.e. they differ on almost all loci, and the family of such near-fittest genomes contains exponentially many mutually orthogonal members. Consequently, near-fittest evolutionary paths become impossible when fitness levels approach

What carries the argument

The central object is the NK random field viewed as a spin-glass Hamiltonian on the $N$-locus hypercube, where each locus interacts with $K$ randomly selected others; in the dense regime, the analysis is carried by the spin-glass free-energy variational principle. The load-bearing mechanism is the overlap-gap property: near-maximum genomes have asymptotically zero pairwise overlap, which converts the geometric claim about many peaks into a quantitative statement about the overlap distribution of high-fitness configurations.

Load-bearing premise

The derivation assumes the dense NK random field converges to a spin-glass system for which a rigorous free-energy variational principle holds; if that convergence fails, the claimed exact limits do not follow from the method.

What would settle it

Run the NK model for increasing $N$ with $K/N$ fixed at several values of $\alpha$; if the empirical free energy or maximum fitness per locus does not approach the paper's predicted limits, the central claim is wrong. A more direct check is exhaustive enumeration for small $N$: the number of near-fittest mutually orthogonal genomes should grow exponentially in $N$ at a rate depending on $\alpha$.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • For any $\alpha \in (0,1]$, the free energy per locus and the maximum fitness per locus have definite large-$N$ limits, so numerical simulations of the NK model have analytic benchmarks to match.
  • The landscape is provably rugged: exponentially many near-fittest genomes are pairwise almost completely different, so no single narrow region of genome space contains the top of the landscape.
  • Adaptation to the global maximum is blocked in a precise sense: paths connecting near-fittest genomes do not exist as the fitness threshold approaches the maximum.
  • For small interaction density $\alpha$, evolution can still proceed without fitness loss: a fitness-maintaining path exists with high probability.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The supplied body text under this title is a different manuscript (an evaluation of language models), so this extraction rests on the title, abstract, and subject class; the quantitative theorems should be verified against the original source before citation.
  • If the variational-principle route is correct, the dense NK model should sit in the same universality class as other Gaussian spin glasses, which would imply that locating the global maximum is computationally hard for the same reason it is hard in those models.
  • The rate at which the number of near-fittest orthogonal genomes grows in $N$ (the exponential rate as a function of $\alpha$) is not stated in the abstract; measuring that rate numerically would be a direct test of the geometric picture.
  • The small-$\alpha$ path result suggests a sharp threshold in $\alpha$ between path-connected and path-blocked fitness landscapes; locating that threshold is a natural next problem.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper, as represented by its arXiv metadata and abstract (arXiv:2508.12464, math.PR), claims to study the NK model in the regime K/N -> alpha in (0,1], using spin-glass methods. The abstract announces exact limits for the free energy at any temperature and for the maximum fitness, an exponentially large collection of near-fittest pairwise asymptotically orthogonal genomes, overlap gap properties, and quantitative statements about the possibility/impossibility of evolutionary paths near the global maximum. However, the supplied full text is not an NK-model paper at all; it is an unrelated evaluation study of OpenAI's GPT-OSS language models (arXiv:2508.12461v3). No NK model is defined, no main theorem is stated, and no proof, lemma, or derivation supporting any of the announced results appears anywhere in the manuscript body. Consequently, none of the paper's central claims can be checked from the submitted material.

Significance. If the results announced in the abstract were properly proved, they would constitute a substantial contribution to the mathematical theory of the NK model: exact solution of the dense regime (K/N -> alpha), a rigorous variational formula for the free energy, and a quantitative picture of the fitness landscape's multiple-peak geometry. These would be of interest to probabilists, statistical physicists, and theoretical evolutionary biologists. I cannot assess the significance of the actual contribution, however, because the manuscript contains no statement of the claimed theorems beyond the abstract and no supporting argument. The only verifiable content is an unrelated LLM evaluation, which has no bearing on the NK model. There is no evidence of machine-checked proofs, reproducible derivations, or even precise definitions on which to judge the mathematical claims.

major comments (3)
  1. [Full text (whole body)] The body of the submitted manuscript is 'Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models' (arXiv:2508.12461v3), not an NK-model paper. There is no definition of the NK model, no random-field construction, no theorem statements, no proofs, and no overlap-gap definitions. The abstract's central claims—exact free-energy limits, maximum fitness, exponential near-fittest orthogonal genomes, and evolutionary-path impossibilities—are therefore entirely unsupported by the supplied text. This is not a presentation issue but a complete absence of the mathematical content that the abstract announces.
  2. [Abstract] The abstract asserts that the NK model is treated 'via the spin glass methodologies' and that this yields 'exact limits.' No precise conditions are stated: the distribution of the NK random field, the nature of the K/N -> alpha limit, the convergence of the field to a spin-glass-type model, or the variational principle being invoked are all unspecified. If the intended paper contains a proof, these elements are load-bearing; in the present manuscript they appear only as assertions. The claimed exact limits do not follow from the stated method without these details.
  3. [Title and submission metadata] The manuscript's title, subject class (math.PR), and abstract concern the NK model, but the actual text is an empirical evaluation of language models. This internal inconsistency means the submission does not contain the research it claims to present. Even as a 'wrong file' scenario, the submitted manuscript is not refereeable as a probability paper, and the issue cannot be repaired by minor local edits within the current text.
minor comments (2)
  1. [References] The reference list is the bibliography of the GPT-OSS evaluation paper (e.g., Radford et al., Hendrycks et al., Kaplan et al.). No references to the NK model literature—Kauffman, Levin, Weinberger, or subsequent mathematical works on NK landscapes—appear, which is inconsistent with the abstract's claims.
  2. [Figures and tables] All tables and figures (e.g., Table I on model parameters, Figures 1–5 on benchmark results) concern large language models. None provide any numerical, exact, or illustrative support for the NK-model results described in the abstract.

Circularity Check

0 steps flagged

No circularity demonstrable: the supplied full text is a different arXiv paper (2508.12461v3), so the NK-model derivation chain is absent and no reduction-to-input can be exhibited.

full rationale

The abstract of arXiv:2508.12464 claims exact free-energy limits, maximum fitness, and multiple-peak structure in the NK model obtained 'via the spin glass methodologies.' The provided full text, however, is the manuscript of arXiv:2508.12461v3 ('Is GPT-OSS Good? ...'), an LLM evaluation paper with no NK-model content. Consequently there are no equations, coupling constructions, interpolation arguments, or overlap-gap derivations from the claimed NK paper to audit. Under the hard rules, circularity can only be claimed when the specific reduction is exhibited (Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction). No such reduction is present. The mismatch is a verification blocker, not evidence of circularity: the absence of proof cannot itself show that the exact limits are equivalent to the inputs. No self-citation chain, no ansatz smuggled by citation, and no fitted-input-called-prediction pattern can be identified from the abstract alone. The abstract's invocation of 'spin glass methodologies' is load-bearing but unverified in the supplied text; that is a correctness or completeness concern, not a demonstrated circular step. The honest finding is therefore a non-finding: no significant circularity can be established from the available material.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

The actual NK manuscript was not available, so no axiology could be performed. The abstract alone gives no fitted parameters or invented entities; alpha is the fixed asymptotic ratio, not a fitted quantity.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Fitness Landscape in the $NK$ Model." pith.science (2026). https://pith.science/paper/76WY6YDF

@misc{pith2026250812464,
  author       = {Pith},
  title        = {Pith review of: On the Fitness Landscape in the $NK$ Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/76WY6YDF}},
  note         = {Machine review of arXiv:2508.12464}
}
Share X Bluesky LinkedIn Reddit HN
abstract

The $NK$ model, introduced by Kauffman, Levin, and Weinberger, is a random field used to describe the fitness landscape of certain species with $N$ genetic loci, each interacting with $K$ others. The model has wide applications in understanding evolutionary and natural selection as it captures ruggedness feature of the fitness landscape. Earlier literature has been focused on the case $K$ being a fixed positive integer and used tools from Ergodic and Markov theory. In this paper, by viewing it as a statistical physics object, we investigate the $NK$ model in the regime $K/N\to\alpha \in(0,1]$ via the spin glass methodologies. Our main result identifies the exact limits for the free energy at any temperature and the maximum fitness. Moreover, we show that the $NK$ model exhibits a multiple-peak structure, namely, the number of near-fittest genomes that are asymptotically orthogonal to each other is exponentially large. Based on establishing the overlap gap properties, we obtain quantitative descriptions for the geometry of the fitness landscape and deduce that, in particular, near-fittest evolutionary paths become impossible as the fitness levels of the genomes approach the global maximum for any $\alpha\in (0,1]$. Nevertheless, we also show that by choosing $\alpha$ sufficiently small, an evolutionary path maintained at a given fitness level can be constructed with high probability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

95 extracted references · 21 canonical work pages · 3 internal anchors

  1. [1]

    Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” inInternational Conference on Learning Representations, 2017

  2. [2]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”The Journal of Machine Learning Research, vol. 23, no. 1, pp. 5232–5270, 2022

  3. [3]

    Scaling laws for neural language models,

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,”arXiv preprint arXiv:2001.08361, 2020

  4. [4]

    Scaling laws for autoregressive generative modeling,

    T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Grayet al., “Scaling laws for autoregressive generative modeling,”arXiv preprint arXiv:2010.14701, 2020

  5. [5]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., “Training compute-optimal large language models,”arXiv preprint arXiv:2203.15556, 2022

  6. [6]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskeveret al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019

  7. [7]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azharet al., “Llama: Open and efficient foundation language models,”arXiv preprint arXiv:2302.13971, 2023

  8. [8]

    Gemma: Open models based on gemini research and technology,

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Loveet al., “Gemma: Open models based on gemini research and technology,”arXiv preprint arXiv:2403.08295, 2024

  9. [9]

    Deepseek llm: Scaling open-source language models with longtermism,

    D.-A. and:, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Donget al., “Deepseek llm: Scaling open-source language models with longtermism,”arXiv preprint arXiv:2401.02954, 2024

  10. [10]

    Qwen technical report,

    J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huanget al., “Qwen technical report,”arXiv preprint arXiv:2309.16609, 2023

  11. [11]

    Qwen2.5 technical report,

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y . Fan, Y . Su, Y . Zhang, Y . Wan, Y . Liu, Z. Cui, Z. Zhang, ...

  12. [12]

    Phi- 3 technical report: A highly capable language model locally on your phone,

    M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behlet al., “Phi- 3 technical report: A highly capable language model locally on your phone,”arXiv preprint arXiv:2404.14219, 2024

  13. [13]

    Emergent abilities of large language models,

    J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzleret al., “Emergent abilities of large language models,”arXiv preprint arXiv:2206.07682, 2022

  14. [14]

    Are emergent abilities of large language models a mirage?

    R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?”Advances in Neural Information Processing Systems, vol. 36, 2024

  15. [15]

    Measuring massive multitask language understanding,

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2021

  16. [16]

    Training verifiers to solve math word problems,

    K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakanoet al., “Training verifiers to solve math word problems,”arXiv preprint arXiv:2110.14168, 2021

  17. [17]

    Solv- ing quantitative reasoning problems with language models,

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., “Solv- ing quantitative reasoning problems with language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022

  18. [18]

    Evaluating large language models trained on code,

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,”arXiv preprint arXiv:2107.03374, 2021

  19. [19]

    Program synthesis with large language models,

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Leet al., “Program synthesis with large language models,”arXiv preprint arXiv:2108.07732, 2021

  20. [20]

    Starcoder: may the source be with you!

    R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chimet al., “Starcoder: may the source be with you!”arXiv preprint arXiv:2305.06161, 2023

  21. [21]

    Codegen: An open large language model for code with multi-turn program synthesis,

    E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y . Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,”arXiv preprint arXiv:2203.13474, 2022

  22. [22]

    Incoder: A generative model for code infilling and synthesis,

    D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis, “Incoder: A generative model for code infilling and synthesis,”arXiv preprint arXiv:2204.05999, 2022

  23. [23]

    C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,

    Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fuet al., “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 991–63 010, 2023

  24. [24]

    Language models are few-shot multilingual learners,

    G. I. Winata, A. Madotto, Z. Lin, R. Liu, J. Yosinski, and P. Fung, “Language models are few-shot multilingual learners,”arXiv preprint arXiv:2109.07684, 2021

  25. [25]

    Judging llm-as-a-judge with mt-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”arXiv preprint arXiv:2306.05685, 2023

  26. [26]

    Lamda: Language models for dialog applications,

    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,”arXiv preprint arXiv:2201.08239, 2022

  27. [27]

    Recipes for building an open-domain chatbot,

    S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y . Liu, J. Xu, M. Ott, E. M. Smith, Y .-L. Boureauet al., “Recipes for building an open-domain chatbot,”arXiv preprint arXiv:2004.13637, 2021

  28. [28]

    Glam: Efficient scaling of language models with mixture-of-experts,

    N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firatet al., “Glam: Efficient scaling of language models with mixture-of-experts,”International Conference on Machine Learning, pp. 5547–5569, 2022

  29. [29]

    Retrieval- augmented generation for knowledge-intensive nlp tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschelet al., “Retrieval- augmented generation for knowledge-intensive nlp tasks,”Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020

  30. [30]

    Efficient large scale language modeling with mixtures of experts,

    M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V . Lin, J. Du, S. Iyer, R. Pasunuruet al., “Efficient large scale language modeling with mixtures of experts,”arXiv preprint arXiv:2112.10684, 2021

  31. [31]

    Mixture-of-experts with expert choice routing,

    Y . Zhou, T. Lei, H. Liu, N. Du, Y . Huang, V . Zhao, A. M. Daiet al., “Mixture-of-experts with expert choice routing,”Advances in Neural Information Processing Systems, vol. 35, pp. 7103–7114, 2022

  32. [32]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosaleet al., “Llama 2: Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023

  33. [33]

    The falcon series of open language models,

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al., “The falcon series of open language models,”arXiv preprint arXiv:2311.16867, 2023

  34. [34]

    Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,

    T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”Advances in Neural Information Processing Systems, vol. 35, pp. 30 318–30 332, 2022

  35. [35]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021

  36. [36]

    Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

    G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosenet al., “Gem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,”arXiv preprint arXiv:2507.06261, 2025

  37. [37]

    (2025) Claude opus 4.1

    Anthropic. (2025) Claude opus 4.1. Accessed: 2025-08-26. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-1

  38. [38]

    Qwen3 technical report,

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025

  39. [39]

    D. Team. (2025) Deepseek-v3.1 release report. Accessed: 2025- 08-26. [Online]. Available: https://api-docs.deepseek.com/zh-cn/news/ news250821

  40. [40]

    Eight things to know about large language models,

    S. R. Bowman, “Eight things to know about large language models,” arXiv preprint arXiv:2304.00612, 2023

  41. [41]

    Ai and the everything in the whole wide world benchmark,

    I. D. Raji, E. Denton, E. M. Bender, A. Hanna, and A. Paullada, “Ai and the everything in the whole wide world benchmark,”arXiv preprint arXiv:2111.15366, 2021

  42. [42]

    Holistic evaluation of language models,

    P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,” inTransactions on Machine Learning Research, 2023

  43. [43]

    With little power comes great responsibility,

    D. Card, P. Henderson, U. Khandelwal, R. Jia, K. Mahowald, and D. Ju- rafsky, “With little power comes great responsibility,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 9263–9274

  44. [44]

    Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,

    O. Sainz, J. A. Campos, I. Garc ´ıa-Ferrero, J. Etxaniz, O. L. de La- calle, and E. Agirre, “Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark,”arXiv preprint arXiv:2310.18018, 2023

  45. [45]

    A survey of large language models,

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Donget al., “A survey of large language models,”arXiv preprint arXiv:2303.18223, 2023

  46. [46]

    Large language models: A survey,

    S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Am- atriain, and J. Gao, “Large language models: A survey,”arXiv preprint arXiv:2402.06196, 2024

  47. [47]

    llama.cpp: Inference of llama model in pure c/c++,

    G. Gerganov, “llama.cpp: Inference of llama model in pure c/c++,” https: //github.com/ggerganov/llama.cpp, 2023

  48. [48]

    Gshard: Scaling giant models with conditional computation and automatic sharding,

    D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,”arXiv preprint arXiv:2006.16668, 2021

  49. [49]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  50. [50]

    Scaling language models: Methods, analysis & insights from training gopher,

    J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Younget al., “Scaling language models: Methods, analysis & insights from training gopher,”arXiv preprint arXiv:2112.11446, 2021

  51. [51]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmannet al., “Palm: Scaling language modeling with pathways,”arXiv preprint arXiv:2204.02311, 2022

  52. [52]

    Finqa: A dataset of numerical reasoning over financial data,

    Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y . Wang, “Finqa: A dataset of numerical reasoning over financial data,” 2022. [Online]. Available: https://arxiv.org/abs/2109.00122

  53. [53]

    Medical Exam Question Answering with Large-scale Reading Comprehension

    X. Zhang, J. Wu, Z. He, X. Liu, and Y . Su, “Medical exam question answering with large-scale reading comprehension,” 2018. [Online]. Available: https://arxiv.org/abs/1802.10279

  54. [54]

    Experimenting with Legal AI Solutions: The Case of Question-Answering for Access to Justice

    J. Li, R. Bhambhoria, S. Dahan, and X. Zhu, “Experimenting with legal ai solutions: The case of question-answering for access to justice,”arXiv preprint arXiv:2409.07713, 2024

  55. [55]

    Crowdsourcing multiple choice science questions,

    J. Welbl, N. F. Liu, and M. Gardner, “Crowdsourcing multiple choice science questions,” 2017. [Online]. Available: https://arxiv.org/abs/1707. 06209

  56. [56]

    Piqa: Reasoning about physical commonsense in natural language,

    Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” 2019. [Online]. Available: https://arxiv.org/abs/1911.11641

  57. [57]

    Dialogsum: A real-life scenario dialogue summarization dataset,

    Y . Chen, Y . Liu, L. Chen, and Y . Zhang, “Dialogsum: A real-life scenario dialogue summarization dataset,” 2021. [Online]. Available: https://arxiv.org/abs/2105.06762

  58. [58]

    Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,

    J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “Xtreme: A massively multilingual multi-task benchmark for evaluat- ing cross-lingual generalisation,”International Conference on Machine Learning, pp. 4411–4421, 2020

  59. [59]

    mt5: A massively multilingual pre-trained text-to-text transformer,

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,”arXiv preprint arXiv:2010.11934, 2021

  60. [60]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

    A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”arXiv preprint arXiv:2206.04615, 2022

  61. [61]

    Language mod- els are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language mod- els are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  62. [62]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 24 824–24 837, 2022

  63. [63]

    Large lan- guage models are zero-shot reasoners,

    T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large lan- guage models are zero-shot reasoners,”Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022

  64. [64]

    Large language models are not fair evaluators,

    P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” arXiv preprint arXiv:2305.17926, 2023

  65. [65]

    The curious case of neural text degeneration,

    A. Holtzman, J. Buys, L. Du, M. Forbes, and Y . Choi, “The curious case of neural text degeneration,”arXiv preprint arXiv:1904.09751, 2019

  66. [66]

    Neural text generation with unlikelihood training,

    S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston, “Neural text generation with unlikelihood training,”arXiv preprint arXiv:1908.04319, 2019

  67. [67]

    The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,

    L. Ruis, A. K. Kemfert, D. Hupkes, L. Kaiser, and B. M. Lake, “The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,”Advances in Neural Information Processing Systems, vol. 37, 2024

  68. [68]

    Locally typical sampling,

    C. Meister, T. Pimentel, G. Wiher, and R. Cotterell, “Locally typical sampling,”Transactions of the Association for Computational Linguis- tics, vol. 11, pp. 102–121, 2023

  69. [69]

    The hitchhiker’s guide to testing statistical significance in natural language processing,

    R. Dror, G. Baumer, S. Shlomov, and R. Reichart, “The hitchhiker’s guide to testing statistical significance in natural language processing,” Proceedings of the 56th Annual Meeting of the Association for Compu- tational Linguistics, vol. 1, pp. 1383–1392, 2018

  70. [70]

    Note on the sampling error of the difference between correlated proportions or percentages,

    Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,”Psychometrika, vol. 12, no. 2, pp. 153–157, 1947

  71. [71]

    Controlling the false discovery rate: a practical and powerful approach to multiple testing,

    Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,”Journal of the Royal statistical society: series B (Methodological), vol. 57, no. 1, pp. 289–300, 1995

  72. [72]

    Cohen,Statistical power analysis for the behavioral sciences

    J. Cohen,Statistical power analysis for the behavioral sciences. Lawrence Erlbaum Associates, 1988

  73. [73]

    Efron and R

    B. Efron and R. J. Tibshirani,An introduction to the bootstrap. CRC press, 1994

  74. [74]

    Statistical significance tests for machine translation evalua- tion,

    P. Koehn, “Statistical significance tests for machine translation evalua- tion,” inProceedings of the 2004 conference on empirical methods in natural language processing, 2004, pp. 388–395

  75. [75]

    Information-theoretic probing for linguistic structure,

    T. Pimentel, J. Valvoda, R. Hall Maudslay, R. Zmigrod, A. Williams, and R. Cotterell, “Information-theoretic probing for linguistic structure,” arXiv preprint arXiv:2004.03061, 2020

  76. [76]

    A statistical analysis of summariza- tion evaluation metrics using resampling methods,

    D. Deutsch, R. Dror, and D. Roth, “A statistical analysis of summariza- tion evaluation metrics using resampling methods,”Transactions of the Association for Computational Linguistics, vol. 9, pp. 1132–1146, 2021

  77. [77]

    Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,

    J. Pineau, P. Vincent-Lamarre, K. Sinha, V . Larivi `ere, A. Beygelzimer, F. d’Alch´e Buc, E. Fox, and H. Larochelle, “Improving reproducibility in machine learning research: a report from the neurips 2019 reproducibil- ity program,”The Journal of Machine Learning Research, vol. 22, no. 1, pp. 7459–7478, 2021

  78. [78]

    Show your work: Improved reporting of experimental results,

    J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show your work: Improved reporting of experimental results,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 2185–2194

  79. [79]

    Data contamination: From memorization to exploitation,

    I. Magar and R. Schwartz, “Data contamination: From memorization to exploitation,”arXiv preprint arXiv:2203.08242, 2022

  80. [80]

    A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers

    J. Dekoninck, M. Whitehouse, M. Suppanen, S. Cullen, M. Kuchnik, and J. Mitchell, “Evading data contamination detection: Exploring test-time preprocessing methods,”arXiv preprint arXiv:2411.02643, 2024

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.