REVIEW 4 major objections 6 minor 217 references
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that LLM reasoning is a directed stochastic walk over a latent skill graph, yielding closed-form accuracy and cost curves that unify training and inference compute scaling.
desk verdict A clean stylized theory of inference compute, but the training-inference bridge in Eq. (51) conflates skill acquisition with per-step execution success and needs major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is DS3, directed stochastic skill search: inference is a stochastic walk on a latent skill graph with distinguished STOP, IDLE, and BRANCH nodes, and in the simplified instantiation the walk moves over a fully connected relevant skill set with directionality $\hat{\iota}$ and fixed skill-success probability $p$. The load-bearing identity is that task completion is a sum of $m$ independent geometric waiting times, whose cumulative distribution function is a regularized incomplete $\beta$ function; every strategy comparison reduces to substituting the appropriate effective success probability and budget into that curve. On the training side, the paper extends a hierarchical skill-text tripartite graph to allow upward and downward skill traversal, and adds a two-step sigmoid mapping from pretraining loss to per-skill success probability to connect training and inference.
What would settle it
Run chain-of-thought with increasing token budgets on a fixed model and controlled tasks, fit $\hat{\iota}p$, and test whether accuracy follows the predicted incomplete-$\beta$ curve $I_{\hat{\iota}p}(m,\,T_{\max}-m+1)$; a systematic misfit, such as continued linear-in-log-token growth beyond the predicted saturation, would refute the idealized walk at the core of the claim.
Extended reading notes
Core claim
The central claim is that a random-walk model of reasoning suffices to reproduce the observed inference-scaling phenomenology. In the idealized instantiation, a task requires $m$ skills applied in a fixed order; at each step the model moves toward the next required skill with probability $\hat{\iota}$ and executes it successfully with probability $p$, so per-step progress is a Bernoulli trial with success probability $\hat{\iota}p$. The number of steps to finish the task is therefore negative binomial, and task success under a step budget is the regularized incomplete $\beta$ function $\psi_{\mathrm{CoT}}=I_{\hat{\iota}p}(m,\,T_{\max}-m+1)$, with matching cost expressions for chain-of-thought, tree-of-thought, best-of-N, and majority voting. Coupled to a hierarchical skill-text training graph, the framework predicts linear accuracy gains with log training and inference compute, strategy preference switching, and reasoning-elicited emergence; coupled to empirical loss-to-accuracy sigmoid fits, it reproduces pass@k coverage and majority-vote saturation.
Load-bearing premise
The load-bearing premise is that pretraining loss translates to per-skill success probability through a fixed monotone sigmoid and that reasoning at inference is a memoryless random walk with a constant per-step success probability; if real models deviate from either, the unified scaling predictions do not follow.
Editorial extensions
If this is right
- For chain-of-thought, task success is a closed-form incomplete-beta curve, so accuracy becomes predictable from roughly two fitted quantities: per-step progress and the number of required skills.
- Sequential reasoning is preferred for strong models and easy tasks, while branching tree-of-thought wins for weak models and hard tasks, with the crossover set by task length and branching factor.
- Jointly optimizing model size, training tokens, and inference tokens dominates fixing a training-compute-optimal ratio, so inference-aware design should favor smaller, overtrained models with explicit token budgets.
- Deeper reasoning steepens the accuracy-versus-model-size slope, which explains why chain-of-thought can elicit emergence on tasks where parameter scaling has plateaued.
- Best-of-N coverage keeps improving with more samples while majority voting saturates when the correct answer is not the majority, and both behaviors arise from one Beta-mixture model over task success rates.
Reading between the lines
- A testable extension is to recover a model's internal directionality coefficient $\hat{\iota}$ from its pass@k curve; if DS3 is right, that single latent number should predict performance across unrelated tasks without refitting.
- The framework suggests adaptive inference routing: a deployment system could estimate task difficulty and switch between sequential and parallel reasoning on the fly, turning strategy switching from an observed trend into a design rule.
- If inference scaling is as structured as claimed, capability forecasting and compute-oversight metrics should include inference compute in log form alongside training compute rather than treating training FLOPs alone.
- The same machinery may apply to latent-space or multi-agent reasoning by reinterpreting the transition kernel, although the paper only analyzes token-level autoregressive traces.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes directed stochastic skill search (DS3), a framework in which inference is modeled as a stochastic walk over a latent skill graph, and derives closed-form expressions for task success probability and compute cost under chain-of-thought (CoT), tree-of-thought (ToT(1)), best-of-N (BoN), and majority voting (MV). The central technical object is the CoT success formula psi_CoT = I_{iota-hat p}(m, T_max - m + 1) (Eq. 20), together with cost expressions (Eqs. 28-37). The paper then extends a prior tripartite skill-text graph model of training (Section III-A) to connect training-time skill acquisition to inference-time success, and separately links the framework to empirical two-step loss-to-accuracy mappings (Section III-B) and to Beta-mixture models of BoN and MV behavior (Section III-D). The claimed contributions are theoretical recoveries of empirically observed inference-scaling phenomena: linear accuracy gains with logarithmic compute, strategy switching with task difficulty and model capability, reasoning-elicited emergence, and BoN/MV behavior in a unified framework.
Significance. If the central derivation were sound, the paper would provide a useful unifying vocabulary for training-inference compute tradeoffs, and the closed-form negative-binomial formulas in Section II are clean and internally consistent under the stated idealizations. The paper is also honest in explicitly labeling the sigmoid loss-to-accuracy mapping as a functional ansatz (Section III-B) and in acknowledging that the two-mode Gaussian task prior is not yet robustly validated (Section III-D). The BoN/MV analysis is strengthened by fitting against public pass@k data from Brown et al. The significance is tempered, however, by two load-bearing problems: the training-inference bridge in Eq. (51) conflates skill acquisition with per-step execution success, and the MV saturation behavior is calibrated rather than predicted. No code is provided, and the simulation-based empirical alignment in Section III-A1 relies on many hand-picked parameters, so the claimed first-principles recoveries are not yet established.
major comments (4)
- [Section III-A, Eq. (51)] The success probability in Eq. (51) substitutes the training-time skill-acquisition probability p_s^(ell) from Eq. (48) into the per-step success slot of the negative-binomial CDF introduced for DS3 in Eq. (20). These are different quantities: p_s is the probability that a skill is acquired during pretraining, a one-time Bernoulli event, whereas the DS3 parameter p in Eqs. (16)-(17) is a per-attempt execution reliability that supports independent geometric retries. If a required skill was not acquired, no number of inference-time retries can make it succeed; if it was acquired, the per-attempt success probability is a separate reliability parameter that the tripartite model does not provide. Using p_s in the negative binomial therefore lets repeated attempts effectively 'acquire' the skill, which the training model does not permit. Since Eq. (51) is the mathematical bridge that produces the claimed first-principles recovery of linear-in-log-compute scaling and reasoning-elicited emergence, this conflation must be resolved, for example by introducing an explicit acquisition-then-execution model and showing that the qualitative trends survive.
- [Section II-F, Theorem 1 setup] The setup of Theorem 1 states that the strategy comparison is performed 'for total token budget Omega and per-step token cost omega' and then uses Omega/(sqrt(b) omega) as the effective number of sequential steps in the ToT(1) success formula. This is inconsistent with Eq. (29), which uses floor(T_max / b) = Omega/(b omega) under a shared token budget. The factor sqrt(b) comes from the attention-cost equivalence in Eq. (39), which equates compute budgets, not token budgets. As written, the proof that branching dominates in the low-capability regime rests on an unjustified change of budget metric; the theorem should either be stated under a fixed compute budget with the appropriate token rescaling or use the token-budget formula.
- [Section III-D, Eq. (62) and Fig. 10] The MV saturation behavior is not an independent prediction of the framework: Eq. (62) calibrates the dominant incorrect-output rate c_1 to the empirically observed MV plateau P_inf^MV, and the Beta parameters (alpha, beta, A) are fitted to the BoN pass@k curves before being inserted into Eq. (61). Consequently, the agreement in Fig. 10 demonstrates that the parametric family is flexible enough to fit the saturation point, not that the DS3 framework predicts MV behavior from first principles. The authors should state explicitly which components of Fig. 10 are fitted (the BoN parameters, the saturation point) and which are predicted (for example, the finite-k shape of the MV curve given those fitted inputs), and should temper the abstract's claim that MV behavior is 'captured within a unified analytical framework.'
- [Section III-A1 and Appendix B] The claimed recovery of linear accuracy scaling with log compute in Figs. 7.2 and 7.3 rests on simulations whose parameters are described as 'rounded and interpretable' and 'chosen to reflect plausible values without overfitting' (Section III-A1), with a large free-parameter set listed in Appendix B (skill pool sizes, tokens per skill, beta, delta, lambda, prior parameters, etc.). Without a sensitivity analysis or a fitting procedure with uncertainty quantification, the match to the Claude 3.7 and o1 curves is not evidence of predictive power of the model; it may simply reflect the flexibility of the parameterization. The authors should provide a sensitivity analysis over the main parameters, or fit the parameters to one subset of the data and show that the qualitative scaling trends persist on a held-out subset.
minor comments (6)
- [References] Reference [80] contains a URL placeholder '[insert-slug-if-available]' and is incomplete; it should be replaced with a proper citation or removed.
- [Section II-B, Eq. (26)] The piecewise linear-Gaussian approximation defines the central-region cutoff as z* = 2 without any derivation or justification; a sentence explaining this choice and its sensitivity would improve reproducibility.
- [Section III-A, Eq. (42)] The notation for m_{ell'} in Eq. (42) uses the same symbol m both for the task's required skills and inside the floor brackets, which is confusing; please use a distinct floor-bracket notation or a worded definition.
- [Table I, MV row] The MV success-probability entry in Table I is an expectation whose denominator is not clearly defined; the tie-breaking rule should be specified in the table caption or in the text before the table.
- [Appendix C-A, Fig. 9] The caption of Fig. 9 states that L_0 ranges over [3.0, 1.5] and that 'most difficult 100 represents L_0 = 1.5', which is confusing because the interval is written in descending order; please reorder or clarify.
- [Section III-B] The same symbols L and L_0 are used for pretraining loss and task-difficulty thresholds; to avoid confusion, consider renaming the task-difficulty parameter (e.g., to tau) throughout Sections III-B and III-C.
Circularity Check
The training–inference bridge is defined by substituting acquisition probability into the per-step execution slot p, while MV saturation and Chinchilla-deviation patterns are calibrated inputs rather than independent predictions.
-
self definitional
[Section III-A, Eq. (51)]
"Then the success probability for a task at level ℓ requiring m skills, under a CoT inference strategy with step budget Tmax = Ω_budget/ω, is given by: ψ^{CoT}_{ℓ,m} = max_{ℓ′∈{1,...,L}} γ^{m_{ℓ′}}_{ℓ′} I_{ι̂ p_s(ℓ′)}(m_{ℓ′}, Ω_budget/ω − m_{ℓ′} + 1), (51)."
In the idealized model, p is defined as 'the fixed probability of successfully applying any visited skill' and enters ψ_CoT = I_{ι̂p}(m, T_max - m + 1) through independent geometric retries. In Section III-A, p_s^(ℓ) is defined as the probability that a skill is acquired during training, i.e., that at least η of its associated concepts are learned. Eq. (51) inserts p_s into the same slot p of the negative-binomial CDF. Acquisition is a one-time binary training-time event; retrying an unacquired skill at inference cannot make it acquired. By construction, the training-to-inference link therefore treats task success as the CDF of acquisition events, so the claimed recovery of training-inference scaling patterns is the definition of the bridge, not an independent derivation.
-
fitted input called prediction
[Section III-D, Eq. (62) and surrounding text]
"Since the asymptotic MV accuracy is governed by the most frequent incorrect response, we calibrate the c_j values to match this limit. In particular, we set the dominant incorrect output rate to reproduce the empirically observed MV saturation point. This leads to the constraint: max_j c_j = I^{-1}(1 - P^MV_∞/A; α, β) / (1 - I^{-1}(1 - P^MV_∞/A; α, β)), (62), where P^MV_∞ measures MV saturation accuracy."
The error spectrum c_j is not derived from the training or inference model; it is fixed by inverting the observed MV saturation plateau P^MV_∞. The same c_j then enters the finite-k formula Ψ_MV(k,Y) in Eq. (61), producing the MV curves shown in Fig. 10. Thus the MV saturation behavior that is presented as captured by the unified framework is an input to the model: Eq. (62) is exactly the condition that the model's asymptotic MV limit equals the empirical saturation value.
1 more flagged steps
-
fitted input called prediction
[Section III-A, paragraph after Eq. (43)]
"To address this, we modify the mechanism of skill acquisition to depend directly on the number of learned concepts, without requiring co-occurrence between skills through shared intermediaries. This adjustment ensures that the framework reflects empirical results: for a fixed compute budget, performance declines consistently as one moves away from the Chinchilla-optimal κ (see Fig. 14)."
The training-side skill-acquisition mechanism is structurally altered for the explicit purpose of reproducing the observed Chinchilla-deviation pattern; the resulting Fig. 14 behavior is then described as the framework reflecting empirical results rather than as a falsifiable prediction. Any later 'theoretical recovery' of training-compute scaling that uses this modified acquisition mechanism inherits the target empirical pattern as an input.
full rationale
The Section II idealized DS3 formulas are internally consistent mathematical identities: given the stated memoryless random-walk assumptions, ψ_CoT = I_{ι̂p}(m, T_max - m + 1) and the cost expressions follow by direct probability calculus. Those formulas, by themselves, are not circular. The circularity appears at the training-inference interface. Eq. (51) replaces the per-attempt execution success probability p, which the negative-binomial retry model requires, with the training-time acquisition probability p_s. The paper defines these as different quantities, so the bridge is made by construction rather than derivation: unacquired skills are implicitly allowed to be 'acquired' through repeated inference attempts, which the training model does not permit. This substitution is the load-bearing step behind the claimed first-principles recovery of linear-in-log-compute scaling and reasoning-elicited emergence. Separate empirical inputs are also built into the claimed outputs. The MV analysis calibrates the dominant incorrect-output rate c_j to the observed saturation plateau, then generates MV curves from that same calibrated parameter; the Chinchilla-deviation behavior is obtained by explicitly modifying the skill-acquisition mechanism so the framework reflects the empirical pattern. These are transparent fits, but they mean the abstract's claims that these behaviors are 'theoretically recovered' and 'captured' overstate what is independently predicted. The paper does disclose the sigmoid loss-to-accuracy map as a 'functional ansatz' and acknowledges the push-forward construction is 'inherently sensitive to the assumed task distribution,' so those items do not count as hidden circularity. Self-citation to the authors' prior tripartite framework [35] is present and load-bearing for the training graph, but the immediate formulas are re-derived in this paper and no uniqueness theorem is imported, so it does not independently raise the score. Overall, the idealized inference theory is self-contained, but the central training-inference unification and two of the headline empirical reproductions reduce, at least partially, to calibrated inputs.
Assumptions & free parameters
free parameters (9)
- directionality delta / effective directionality iota-hat =
delta=0.5 in simulations; iota-hat = 1/2 + 1/(2M)
- per-skill success probability p =
set through sigmoid ansatz in empirical fits; implicit in simulations
- tokens per skill step omega =
25
- Beta mixture (alpha, beta, A) =
fit per model to pass@k curves in Appendix E-B
- incorrect-output spectrum c_j (with decay lambda) =
lambda = 0.592 (Llama-3-70B-Instruct), lambda = 0.752 (Llama-3-8B-Instruct)
- two-mode Gaussian task-prior (mu_m, sigma_m) =
jointly optimized with iota-hat in Fig. 12
- prerequisite scaling sigma_l = (1/2) ln(l) =
(1/2) ln(l)
- relevant skill set inflation beta =
5
- sigmoid loss-to-accuracy parameters a, b, d =
a=1, b=5, d=0 in AIME fits; a=1, b=20, d=0 in Appendix C-A
assumptions (7)
- domain assumption LLM knowledge forms a hierarchical skill-text tripartite graph with concepts and skills.
- domain assumption Inference is a directed stochastic walk on the skill graph with fixed, history-independent per-step success iota-hat p.
- domain assumption Tasks require strictly sequential execution of m required skills with no out-of-order completion.
- ad hoc to paper The pretraining-loss-to-skill-success mapping is a monotone sigmoid (functional ansatz).
- domain assumption Composability of skills is governed by an Erdos-Renyi giant component inside each hierarchy level.
- domain assumption Oracle verifiers select the best trace with no verification cost.
- domain assumption The relevant skill graph is restricted to the giant connected component.
invented entities (1)
-
latent skill graph with required, relevant-but-not-required, and irrelevant skills plus control nodes (BRANCH, IDLE, STOP)
Cite this review
Pith. "Pith review of A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search." pith.science (2026). https://pith.science/paper/S6ECUKFY
@misc{pith2026250700004,
author = {Pith},
title = {Pith review of: A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/S6ECUKFY}},
note = {Machine review of arXiv:2507.00004}
}
read the original abstract
Large language models (LLMs) demand considerable computational, energy, and financial resources during both training and deployment. While scaling laws for training have guided much of the field's recent progress, inference costs now represent a significant and growing component of the overall resource burden, particularly for reasoning-focused models. Existing characterizations of compute-optimality that consider model size, dataset size, and inference tokens in isolation or in fixed combinations risk overlooking more efficient operating points. We introduce directed stochastic skill search (DS3), a general framework that represents inference as stochastic traversal over a learned skill graph. From a simplified yet expressive instantiation, we derive closed-form expressions for task success and compute cost across a wide range of inference strategies -- including chain-of-thought (CoT) and tree-of-thought (ToT) -- enabling comparative analysis as a function of task difficulty and model capability. To that end, we extend a prior first-principles tripartite graph framework of LLM training to incorporate inference, and separately bridge DS3 with empirical methods that characterize LLM scaling behavior. We theoretically recover empirically observed patterns, including: linear accuracy scaling with logarithmic compute; variation in preferred inference strategies as a function of task difficulty and model capability; emergent behavior elicited by reasoning even when performance plateaus under parameter scaling; and both best-of-N (BoN) and majority voting behavior captured within a unified analytical framework. By explicitly characterizing training-inference interdependencies, our framework deepens theoretical understanding and supports principled algorithmic design and resource allocation.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv:2001.08361 [cs.LG], 2020
arXiv 2001
-
[2]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre, “Training compute-optimal large language models,” arXiv:2203.15556 [cs.CL], 2022
arXiv 2022
-
[3]
Compute-optimal LLMs provably generalize better with scale,
M. Finzi, S. Kapoor, D. Granziol, A. Gu, C. D. Sa, J. Z. Kolter, and A. G. Wilson, “Compute-optimal LLMs provably generalize better with scale,” arXiv:2504.15208 [cs.LG], 2025
arXiv 2025
-
[4]
Training compute of frontier AI models grows by 4-5x per year,
J. Sevilla and E. Rold ´an, “Training compute of frontier AI models grows by 4-5x per year,” 2024, accessed: 2025-04-26. [Online]. Available: https://epoch.ai/blog/training-compute-of-frontier-ai-models-grows-by-4-5x-per-year
2024
-
[5]
Increased compute efficiency and the diffusion of AI capabilities,
K. Pilz, L. Heim, and N. Brown, “Increased compute efficiency and the diffusion of AI capabilities,” arXiv:2311.15377 [cs.CY], 2024
arXiv 2024
-
[6]
Measuring the algorithmic efficiency of neural networks,
D. Hernandez and T. B. Brown, “Measuring the algorithmic efficiency of neural networks,” arXiv:2005.04305 [cs.LG], 2020
arXiv 2005
-
[7]
Algorithmic progress in language models,
A. Ho, T. Besiroglu, E. Erdil, D. Owen, R. Rahman, Z. C. Guo, D. Atkinson, N. Thompson, and J. Sevilla, “Algorithmic progress in language models,” arXiv:2403.05812 [cs.CL], 2024
arXiv 2024
-
[8]
Claude’s extended thinking,
Anthropic, “Claude’s extended thinking,” https://www.anthropic.com/news/visible-extended-thinking, Feb. 2025, Anthropic blog post. 20This kind of visibility would not be directly available in models that rely primarily on latent-space inference [181], [182]. 30
2025
Show all 217 references
-
[9]
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,
DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y . Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. L...
2025 arXiv
-
[10]
Gemini 2.5: Our most intelligent AI model,
K. Kavukcuoglu, “Gemini 2.5: Our most intelligent AI model,” https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, Mar. 2025, Google DeepMind blog post
2025
-
[11]
IBM Granite 3.2: Reasoning, vision, forecasting and more,
K. Soule and D. Bergmann, “IBM Granite 3.2: Reasoning, vision, forecasting and more,” https://www.ibm.com/new/announcements/ ibm-granite-3-2-open-source-reasoning-and-vision, Feb. 2025, accessed: 2025-05-06
2025
-
[12]
Phi-4- reasoning technical report,
M. Abdin, S. Agarwal, A. Awadallah, V . Balachandran, H. Behl, L. Chen, G. de Rosa, S. Gunasekar, M. Javaheripi, N. Joshi, P. Kauffmann, Y . Lara, C. C. T. Mendes, A. Mitra, B. Nushi, D. Papailiopoulos, O. Saarikivi, S. Shah, V . Shrivastava, V . Vineet, Y . Wu, S. Yousefi, an...
2025 arXiv
-
[13]
OpenAI o1 System Card,
OpenAI, “OpenAI o1 System Card,” OpenAI, Tech. Rep., December 2024. [Online]. Available: https://openai.com/index/openai-o1-system-card/
2024
-
[14]
Introducing OpenAI o3 and o4-mini,
——, “Introducing OpenAI o3 and o4-mini,” https://openai.com/index/introducing-o3-and-o4-mini/, Apr. 2025, openAI blog post
2025
-
[15]
Grok 3 Beta — The Age of Reasoning Agents,
xAI, “Grok 3 Beta — The Age of Reasoning Agents,” https://x.ai/news/grok-3, Feb. 2025, xAI news post
2025
-
[16]
The growing energy footprint of artificial intelligence,
A. De Vries, “The growing energy footprint of artificial intelligence,”Joule, vol. 7, no. 10, pp. 2191–2194, 2023
2023
-
[17]
Estimating the carbon footprint of BLOOM, a 176B parameter language model,
A. S. Luccioni, S. Viguier, and A.-L. Ligozat, “Estimating the carbon footprint of BLOOM, a 176B parameter language model,”Journal of Machine Learning Research, vol. 24, no. 253, pp. 1–15, 2023
2023
-
[18]
The carbon footprint of machine learning training will plateau, then shrink,
D. Patterson, J. Gonzalez, U. H ¨olzle, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “The carbon footprint of machine learning training will plateau, then shrink,” arXiv:2204.05149 [cs.LG], 2022
2022 arXiv
-
[19]
Sustainable AI: Environmental implications, challenges and opportunities,
C.-J. Wu, R. Raghavendra, U. Gupta, B. Acun, N. Ardalani, K. Maeng, G. Chang, F. Aga, J. Huang, C. Bai, M. Gschwind, A. Gupta, M. Ott, A. Melnikov, S. Candido, D. Brooks, G. Chauhan, B. Lee, H.-H. Lee, B. Akyildiz, M. Balandat, J. Spisak, R. Jain, M. Rabbat, and K. Hazelwood, ...
2022
-
[20]
The next wave of AI: Demand and adoption,
T. O’Malley, R. Sandler, and T. Young, “The next wave of AI: Demand and adoption,” https://www.ib.barclays/our-insights/3-point-perspective/ the-next-wave-of-AI-demand-and-adoption.html, 2024, accessed: 2025-04-25
2024
-
[21]
From efficiency gains to rebound effects: The problem of Jevons’ Paradox in AI’s polarized environmental debate,
A. S. Luccioni, E. Strubell, and K. Crawford, “From efficiency gains to rebound effects: The problem of Jevons’ Paradox in AI’s polarized environmental debate,” arXiv:2501.16548 [cs.CY], 2025
2025 arXiv
-
[22]
W. S. Jevons,The Coal Question; An Inquiry concerning the Progress of the Nation, and the Probable Exhaustion of our Coal-mines. Macmillan, 1866
-
[23]
1B user messages sent on ChatGPT every day,
E. Roth, “1B user messages sent on ChatGPT every day,” The Verge, 2024, accessed: 2025-04-22. [Online]. Available: https: //www.theverge.com/2024/12/4/24313097/chatgpt-300-million-weekly-users
2024
-
[24]
ChatGPT added one million users in the last hour,
K. Robison, “ChatGPT added one million users in the last hour,” The Verge, 2025, accessed: 2025-04-22. [Online]. Available: https://www.theverge.com/openai/639960/chatgpt-added-one-million-users-in-the-last-hour
2025
-
[25]
ChatGPT statistics and user trends (2025),
DemandSage, “ChatGPT statistics and user trends (2025),” 2025, accessed: 2025-04-22. [Online]. Available: https://www.demandsage.com/ chatgpt-statistics
2025
-
[26]
A systematic review of Green AI,
R. Verdecchia, J. Sallou, and L. Cruz, “A systematic review of Green AI,”Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, vol. 13, no. 4, p. e1507, 2023
2023
-
[27]
Deep Blue,
M. Campbell, A. J. Hoane, Jr., and F.-H. Hsu, “Deep Blue,”Artificial Intelligence, vol. 134, no. 1, pp. 57–83, 2002
2002
-
[28]
Mastering the game of Go with deep neural networks and tree search,
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V . Panneershelvam, M. Lanctotet al., “Mastering the game of Go with deep neural networks and tree search,”Nature, vol. 529, no. 7587, pp. 484–489, 2016
2016
-
[29]
Mastering the game of Go without human knowledge,
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Boltonet al., “Mastering the game of Go without human knowledge,”Nature, vol. 550, no. 7676, pp. 354–359, 2017
2017
-
[30]
Mastering Chess and Shogi by self-play with a general reinforcement learning algorithm,
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, “Mastering Chess and Shogi by self-play with a general reinforcement learning algorithm,” arXiv:1712.01815 [cs.AI], 2017
2017 arXiv
-
[31]
Scaling scaling laws with board games,
A. L. Jones, “Scaling scaling laws with board games,” arXiv:2104.03113 [cs.LG], 2021
2021 arXiv
-
[32]
Safe and nested subgame solving for imperfect-information games,
N. Brown and T. Sandholm, “Safe and nested subgame solving for imperfect-information games,” arXiv:1705.02955 [cs.AI], 2017
2017 arXiv
-
[33]
Human-level play in the game of Diplomacy by combining language models with strategic reasoning,
Meta Fundamental AI Research Diplomacy Team (FAIR), A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Huet al., “Human-level play in the game of Diplomacy by combining language models with strategic reasoning,”Science, vol. 378, no. 6624, p...
2022
-
[34]
Emergent abilities of large language models,
J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,” arXiv:2206.07682 [cs.CL], 2022
2022 arXiv
-
[35]
An information theory of compute-optimal size scaling, emergence, and plateaus in language models,
A. K. Nayak and L. R. Varshney, “An information theory of compute-optimal size scaling, emergence, and plateaus in language models,” inCompression Workshop @ NeurIPS 2024, 2024
2024
-
[36]
Multi-task Language Understanding on MMLU Leaderboard,
Papers with Code, “Multi-task Language Understanding on MMLU Leaderboard,” 2025, accessed: 2025-04-28. [Online]. Available: https://paperswithcode.com/sota/multi-task-language-understanding-on-mmlu
2025
-
[37]
Measuring massive multitask language understanding,
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[38]
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?” inAdvances in Neural Information Processing Systems, 2023, vol. 36, pp. 55 565–55 581
2023
-
[39]
The quantization model of neural scaling,
E. Michaud, Z. Liu, U. Girit, and M. Tegmark, “The quantization model of neural scaling,” inAdvances in Neural Information Processing Systems, 2023, vol. 36, pp. 28 699–28 722
2023
-
[40]
Circuit tracing: Revealing computational graphs in language models,
E. Ameisen, J. Lindsey, A. Pearce, W. Gurnee, N. L. Turner, B. Chen, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. Marcus, M. Sklar, A. Templeton, T. Bricken, C. McDougall, H. Cunningham, T. Henighan, A. Jermyn, A. Jones, A. Persic, Z. Qi, T. Ben Thompson, S. Zimmerman, K. R...
-
[41]
Curriculum learning,
Y . Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” inProceedings of the International Conference on Machine Learning (ICML), 2009, pp. 41–48
2009
-
[42]
A theory for emergence of complex skills in language models,
S. Arora and A. Goyal, “A theory for emergence of complex skills in language models,” arXiv:2307.15936 [cs.LG], 2023
2023 arXiv
-
[43]
A mathematical theory for learning semantic languages by abstract learners,
K.-Y . Liao, C.-S. Chang, and Y .-W. P. Hong, “A mathematical theory for learning semantic languages by abstract learners,”IEEE Journal on Selected Areas in Communications, 2025
2025
-
[44]
Skill-Mix: a flexible and expandable family of evaluations for AI models,
D. Yu, S. Kaur, A. Gupta, J. Brown-Cohen, A. Goyal, and S. Arora, “Skill-Mix: a flexible and expandable family of evaluations for AI models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024
2024
-
[45]
The learning curve: implications of a quantitative analysis,
C. R. Gallistel, S. Fairhurst, and P. Balsam, “The learning curve: implications of a quantitative analysis,”Proceedings of the National Academy of Sciences, vol. 101, no. 36, pp. 13 124–13 131, 2004
2004
-
[46]
Plateaus, dips, and leaps: Where to look for inventions and discoveries during skilled performance,
W. D. Gray and J. K. Lindstedt, “Plateaus, dips, and leaps: Where to look for inventions and discoveries during skilled performance,”Cognitive Science, vol. 41, no. 7, pp. 1838–1870, 2017
2017
-
[47]
A first-principles mathematical model integrates the disparate timescales of human learning,
M. Lu, T. Marghetis, and V . C. Yang, “A first-principles mathematical model integrates the disparate timescales of human learning,”npj Complexity, vol. 2, no. 1, p. 15, 2025
2025
-
[48]
Spin-glass models as error-correcting codes,
N. Sourlas, “Spin-glass models as error-correcting codes,”Nature, vol. 339, no. 6227, pp. 693–695, 1989
1989
-
[49]
Newell,Unified Theories of Cognition
A. Newell,Unified Theories of Cognition. Harvard University Press, 1994
1994
-
[50]
Barab ´asi,Network Science
A.-L. Barab ´asi,Network Science. Cambridge University Press, 2016
2016
-
[51]
Learning curves: Asymptotic values and rate of convergence,
C. Cortes, L. D. Jackel, S. Solla, V . Vapnik, and J. Denker, “Learning curves: Asymptotic values and rate of convergence,” inAdvances in Neural Information Processing Systems, 1993, vol. 6, pp. 327–334
1993
-
[52]
Deep learning scaling is predictable, empirically,
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y . Yang, and Y . Zhou, “Deep learning scaling is predictable, empirically,” arXiv:1712.00409 [cs.LG], 2017
2017 arXiv
-
[53]
A constructive prediction of the generalization error across scales,
J. S. Rosenfeld, A. Rosenfeld, Y . Belinkov, and N. Shavit, “A constructive prediction of the generalization error across scales,” arXiv:1909.12673 [cs.LG], 2019
1909 arXiv
-
[54]
Language models are few-shot learners,
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B....
2005 arXiv
-
[55]
Prediction and entropy of printed English,
C. E. Shannon, “Prediction and entropy of printed English,”Bell System Technical Journal, vol. 30, no. 1, pp. 50–64, 1951
1951
-
[56]
Explaining neural scaling laws,
Y . Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma, “Explaining neural scaling laws,”Proceedings of the National Academy of Sciences, vol. 121, no. 27, p. e2311878121, Jun. 2024
2024
-
[57]
Towards a universal scaling law of LLM training and inference,
C. Wu and R. Tang, “Towards a universal scaling law of LLM training and inference,” ScienceOpen Preprints, 2024
2024
-
[58]
The Llama 3 herd of models,
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...
2024 arXiv
-
[59]
Densing law of LLMs,
C. Xiao, J. Cai, W. Zhao, G. Zeng, B. Lin, J. Zhou, Z. Zheng, X. Han, Z. Liu, and M. Sun, “Densing law of LLMs,” arXiv:2412.04315 [cs.AI], 2024
2024 arXiv
-
[60]
How predictable is language model benchmark performance?
D. Owen, “How predictable is language model benchmark performance?” arXiv:2401.04757 [cs.LG], 2024. 32
2024 arXiv
-
[61]
Observational scaling laws and the predictability of language model performance,
Y . Ruan, C. J. Maddison, and T. Hashimoto, “Observational scaling laws and the predictability of language model performance,” arXiv:2405.10938 [cs.LG], 2024
2024 arXiv
-
[62]
Language models scale reliably with over-training and on downstream tasks,
S. Y . Gadre, G. Smyrnis, V . Shankar, S. Gururangan, M. Wortsman, R. Shao, J. Mercat, A. Fang, J. Li, S. Keh, R. Xin, M. Nezhurina, I. Vasiljevic, J. Jitsev, L. Soldaini, A. G. Dimakis, G. Ilharco, P. W. Koh, S. Song, T. Kollar, Y . Carmon, A. Dave, R. Heckel, N. Muennighoff,...
2024 arXiv
-
[63]
Broken neural scaling laws,
E. Caballero, K. Gupta, I. Rish, and D. Krueger, “Broken neural scaling laws,” arXiv:2210.14891 [cs.LG], 2023
2023 arXiv
-
[64]
Scaling laws for downstream task performance in machine translation,
B. Isik, N. Ponomareva, H. Hazimeh, D. Paparas, S. Vassilvitskii, and S. Koyejo, “Scaling laws for downstream task performance in machine translation,” arXiv:2402.04177 [cs.CL], 2025
2025
-
[65]
Exploring the limits of large scale pre-training,
S. Abnar, M. Dehghani, B. Neyshabur, and H. Sedghi, “Exploring the limits of large scale pre-training,” arXiv:2110.02095 [cs.LG], 2021
2021 arXiv
-
[66]
Scaling laws do not scale,
F. Diaz and M. Madaio, “Scaling laws do not scale,” arXiv:2307.03201 [cs.LG], 2024
2024 arXiv
-
[67]
Not-just-scaling laws: Towards a better understanding of the downstream impact of language model design decisions,
E. Liu, A. Bertsch, L. Sutawika, L. Tjuatja, P. Fernandes, L. Marinov, M. Chen, S. Singhal, C. Lawrence, A. Raghunathan, K. Gashteovski, and G. Neubig, “Not-just-scaling laws: Towards a better understanding of the downstream impact of language model design decisions,” arXiv:25...
2025
-
[68]
Same pre-training loss, better downstream: Implicit bias matters for language models,
H. Liu, S. M. Xie, Z. Li, and T. Ma, “Same pre-training loss, better downstream: Implicit bias matters for language models,” arXiv:2210.14199 [cs.LG], 2022
2022 arXiv
-
[69]
Overtrained language models are harder to fine-tune,
J. M. Springer, S. Goyal, K. Wen, T. Kumar, X. Yue, S. Malladi, G. Neubig, and A. Raghunathan, “Overtrained language models are harder to fine-tune,” arXiv:2503.19206 [cs.CL], 2025
2025 arXiv
-
[70]
Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning,
F. Chen, A. Raventos, N. Cheng, S. Ganguli, and S. Druckmann, “Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning,” arXiv:2502.07154 [cs.LG], 2025
2025
-
[71]
TruthfulQA: Measuring how models mimic human falsehoods,
S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” arXiv:2109.07958 [cs.CL], 2022
2022 arXiv
-
[72]
BBQ: A hand-built bias benchmark for question answering,
A. Parrish, A. Chen, N. Nangia, V . Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. R. Bowman, “BBQ: A hand-built bias benchmark for question answering,” arXiv:2110.08193 [cs.CL], 2022
2022 arXiv
-
[73]
Inverse scaling can become U-shaped,
J. Wei, N. Kim, Y . Tay, and Q. V . Le, “Inverse scaling can become U-shaped,” arXiv:2211.02011 [cs.CL], 2023
2023 arXiv
-
[74]
s1: Simple test-time scaling,
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Cand `es, and T. Hashimoto, “s1: Simple test-time scaling,” arXiv:2501.19393 [cs.CL], 2025
2025 arXiv
-
[75]
Reinforcement learning for reasoning in large language models with one training example,
Y . Wang, Q. Yang, Z. Zeng, L. Ren, L. Liu, B. Peng, H. Cheng, X. He, K. Wang, J. Gao, W. Chen, S. Wang, S. S. Du, and Y . Shen, “Reinforcement learning for reasoning in large language models with one training example,” arXiv:2504.20571 [cs.LG], 2025
2025 arXiv
-
[76]
Large language models are zero-shot reasoners,
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa, “Large language models are zero-shot reasoners,” arXiv:2205.11916 [cs.CL], 2023
2023 arXiv
-
[77]
Challenging BIG-Bench tasks and whether chain-of-thought can solve them,
M. Suzgun, N. Scales, N. Sch ¨arli, S. Gehrmann, Y . Tay, H. W. Chung, A. Chowdhery, Q. V . Le, E. H. Chi, D. Zhou, and J. Wei, “Challenging BIG-Bench tasks and whether chain-of-thought can solve them,” arXiv:2210.09261 [cs.CL], 2022
-
[78]
Understanding reasoning ability of language models from the perspective of reasoning paths aggregation,
X. Wang, A. Amayuelas, K. Zhang, L. Pan, W. Chen, and W. Y . Wang, “Understanding reasoning ability of language models from the perspective of reasoning paths aggregation,” arXiv:2402.03268 [cs.LG], 2024
2024 arXiv
-
[79]
Sequence to sequence learning with neural networks: What a decade,
I. Sutskever, “Sequence to sequence learning with neural networks: What a decade,” NeurIPS 2024 Test of Time Award Talk, Dec. 2024. [Online]. Available: https://www.youtube.com/watch?v=HlGi4OOuZyw
2024
-
[80]
AI doom from an LLM-plateau-ist perspective,
S. Byrnes, “AI doom from an LLM-plateau-ist perspective,” https://www.lesswrong.com/posts/[insert-slug-if-available], 2023, online essay
2023
-
[81]
The first wave of AI innovation is over. here’s what comes next,
G. Ritter and W. Lu, “The first wave of AI innovation is over. here’s what comes next,”Fast Company, Jul. 2024
2024
-
[82]
AI won’t plateau — if we give it time to think,
N. Brown, “AI won’t plateau — if we give it time to think,” TED Talk, Dec. 2024. [Online]. Available: https://www.youtube.com/watch?v=MG9oqntiJKg
2024
-
[83]
Scaling data-constrained language models,
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel, “Scaling data-constrained language models,” inAdvances in Neural Information Processing Systems, 2023, vol. 36, pp. 50 358–50 376
2023
-
[84]
Scaling laws for data filtering – data curation cannot be compute agnostic,
S. Goyal, P. Maini, Z. C. Lipton, A. Raghunathan, and J. Z. Kolter, “Scaling laws for data filtering – data curation cannot be compute agnostic,” arXiv:2404.07177 [cs.LG], 2024
2024 arXiv
-
[85]
The rising costs of training frontier AI models,
B. Cottier, R. Rahman, L. Fattorini, N. Maslej, and D. Owen, “The rising costs of training frontier AI models,” arXiv:2405.21015 [cs.CY], 2024
2024 arXiv
-
[86]
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting,
M. Turpin, J. Michael, E. Perez, and S. R. Bowman, “Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting,” arXiv:2305.04388 [cs.CL], 2023
2023 arXiv
-
[87]
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought,
A. Saparov and H. He, “Language models are greedy reasoners: A systematic formal analysis of chain-of-thought,” arXiv:2210.01240 [cs.CL], 2023
2023 arXiv
-
[88]
Efficient streaming language models with attention sinks,
G. Xiao, Y . Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” arXiv:2309.17453 [cs.CL], 2024
2024 arXiv
-
[89]
Don’t take things out of context: Attention intervention for enhancing chain-of-thought reasoning in large language models,
S. Yan, C. Shen, W. Wang, L. Xie, J. Liu, and J. Ye, “Don’t take things out of context: Attention intervention for enhancing chain-of-thought reasoning in large language models,” arXiv:2503.11154 [cs.CL], 2025
2025 arXiv
-
[90]
The curse of CoT: On the limitations of chain-of-thought in in-context learning,
T. Zheng, Y . Chen, C. Li, C. Li, Q. Zong, H. Shi, B. Xu, Y . Song, G. Y . Wong, and S. See, “The curse of CoT: On the limitations of chain-of-thought in in-context learning,” arXiv:2504.05081 [cs.CL], 2025
2025
-
[91]
When more is less: Understanding chain-of-thought length in LLMs,
Y . Wu, Y . Wang, T. Du, S. Jegelka, and Y . Wang, “When more is less: Understanding chain-of-thought length in LLMs,” arXiv:2502.07266 [cs.AI], 2025
2025 arXiv
-
[92]
Let’s verify step by step,
H. Lightman, V . Kosaraju, Y . Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” arXiv:2305.20050 [cs.LG], 2023
2023 arXiv
-
[93]
Training verifiers to solve math word problems,
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” arXiv:2110.14168 [cs.LG], 2021
2021 arXiv
-
[94]
When to solve, when to verify: Compute-optimal problem solving and generative verification for LLM reasoning,
N. Singhi, H. Bansal, A. Hosseini, A. Grover, K.-W. Chang, M. Rohrbach, and A. Rohrbach, “When to solve, when to verify: Compute-optimal problem solving and generative verification for LLM reasoning,” arXiv:2504.01005 [cs.CL], 2025
2025
-
[95]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv:2203.11171 [cs.CL], 2022
2022 arXiv
-
[96]
Solving quantitative reasoning problems with language models,
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y . Wu, B. Neyshabur, G. Gur-Ari, and V . Misra, “Solving quantitative reasoning problems with language models,” inAdvances in Neural Information Process...
2022
-
[97]
Self-Refine: Iterative refinement with self-feedback,
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-Refine: Iterative refinement with self-feedback,” arXiv:2303.17651 [cs.CL], 2023
2023 arXiv
-
[98]
Cost-of-Pass: An economic framework for evaluating language models,
M. H. Erol, B. El, M. Suzgun, M. Yuksekgonul, and J. Zou, “Cost-of-Pass: An economic framework for evaluating language models,” arXiv:2504.13359 [cs.AI], 2025
2025
-
[99]
Scaling laws for transfer,
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish, “Scaling laws for transfer,” arXiv:2102.01293 [cs.LG], 2021
2021 arXiv
-
[100]
Reproducible scaling laws for contrastive language-image learning,
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” inProceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...
2023
-
[101]
Distillation scaling laws,
D. Busbridge, A. Shidani, F. Weers, J. Ramapuram, E. Littwin, and R. Webb, “Distillation scaling laws,” arXiv:2502.08606 [cs.LG], 2025
2025 arXiv
-
[102]
Scaling laws for fine-grained mixture of experts,
J. Krajewski, J. Ludziejewski, K. Adamczewski, M. Pi ´oro, M. Krutul, S. Antoniak, K. Ciebiera, K. Kr ´ol, T. Odrzyg ´o´zd´z, P. Sankowski, M. Cygan, and S. Jaszczur, “Scaling laws for fine-grained mixture of experts,” arXiv:2402.07871 [cs.LG], 2024. 33
2024 arXiv
-
[103]
D. W. Thompson,On Growth and Form. Cambridge University Press, 1917
1917
-
[104]
On being the right size,
J. B. S. Haldane, “On being the right size,”Harper’s Magazine, vol. 152, pp. 424–427, 1926
1926
-
[105]
The evolutions of large brain size in mammals: the ‘over-700-gram club quartet’,
P. R. Manger, M. A. Spocter, and N. Patzke, “The evolutions of large brain size in mammals: the ‘over-700-gram club quartet’,”Brain Behavior and Evolution, vol. 82, no. 1, pp. 68–78, 2013
2013
-
[106]
Beyond Chinchilla-optimal: Accounting for inference in language model scaling laws,
N. Sardana, J. Portes, S. Doubov, and J. Frankle, “Beyond Chinchilla-optimal: Accounting for inference in language model scaling laws,” arXiv:2401.00448 [cs.LG], 2024
2024 arXiv
-
[107]
DeepSeek LLM: Scaling open-source language models with longtermism,
DeepSeek-AI, X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, H. Gao, K. Gao, W. Gao, R. Ge, K. Guan, D. Guo, J. Guo, G. Hao, Z. Hao, Y . He, W. Hu, P. Huang, E. Li, G. Li, J. Li, Y . Li, Y . K. Li, W. Liang, F. Lin, A. X. Liu, B. Liu, W. Liu,...
2024 arXiv
-
[108]
Cost-optimal grouped-query attention for long-context LLMs,
Y . Chen, Y . Wu, X. Han, Z. Liu, and M. Sun, “Cost-optimal grouped-query attention for long-context LLMs,” arXiv:2503.09579 [cs.CL], 2025
2025
-
[109]
PaLM 2 technical report,
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chenet al., “PaLM 2 technical report,” arXiv:2305.10403 [cs.CL], 2023
2023 arXiv
-
[110]
Chinchilla scaling: A replication attempt,
T. Besiroglu, E. Erdil, M. Barnett, and J. You, “Chinchilla scaling: A replication attempt,” arXiv:2404.10102 [cs.AI], 2024
2024 arXiv
-
[111]
Scaling LLM test-time compute optimally can be more effective than scaling model parameters,
C. Snell, J. Lee, K. Xu, and A. Kumar, “Scaling LLM test-time compute optimally can be more effective than scaling model parameters,” arXiv:2408.03314 [cs.LG], 2024
2024 arXiv
-
[112]
Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models,
Y . Wu, Z. Sun, S. Li, S. Welleck, and Y . Yang, “Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models,” arXiv:2408.00724 [cs.AI], 2025
2025 arXiv
-
[113]
Go smol or go home,
H. De Vries, “Go smol or go home,” 2023. [Online]. Available: https://www.harmdevries.com/post/model-size-vs-compute-overhead/
2023
-
[114]
Learning to reason with LLMs,
OpenAI, “Learning to reason with LLMs,” Sep. 2024. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2024
-
[115]
Trading off compute in training and inference,
P. Villalobos and D. Atkinson, “Trading off compute in training and inference,” 2023. [Online]. Available: https: //epoch.ai/blog/trading-off-compute-in-training-and-inference
2023
-
[116]
Large language monkeys: Scaling inference compute with repeated sampling,
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V . Le, C. R ´e, and A. Mirhoseini, “Large language monkeys: Scaling inference compute with repeated sampling,” arXiv:2407.21787 [cs.LG], 2024
2024 arXiv
-
[117]
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families,
F. M. Polo, S. Somerstep, L. Choshen, Y . Sun, and M. Yurochkin, “Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families,” arXiv:2412.06540 [cs.LG], 2025
2025
-
[118]
A simple model of inference scaling laws,
N. Levi, “A simple model of inference scaling laws,” arXiv:2410.16377 [stat.ML], 2024
2024 arXiv
-
[119]
Architectural styles of curiosity in global Wikipedia mobile app readership,
D. Zhou, S. Patankar, D. M. Lydon-Staley, P. Zurn, M. Gerlach, and D. S. Bassett, “Architectural styles of curiosity in global Wikipedia mobile app readership,”Science Advances, vol. 10, no. 43, p. eadn3268, 2024
2024
-
[120]
A big data approach to computational creativity: The curious case of Chef Watson,
L. R. Varshney, F. Pinel, K. R. Varshney, D. Bhattacharjya, A. Schoergendorfer, and Y .-M. Chee, “A big data approach to computational creativity: The curious case of Chef Watson,”IBM Journal of Research and Development, vol. 63, no. 1, pp. 7:1–7:18, 2019
2019
-
[121]
A coupon-collector model of machine-aided discovery,
A. Vempaty, L. R. Varshney, and P. K. Varshney, “A coupon-collector model of machine-aided discovery,” inKDD Workshop on Data-Driven Discovery, 2017. [Online]. Available: https://arxiv.org/abs/1708.03833
2017 arXiv
-
[122]
How well do LLMs compress their own chain-of-thought? a token complexity approach,
A. Lee, E. Che, and T. Peng, “How well do LLMs compress their own chain-of-thought? a token complexity approach,” arXiv:2503.01141 [cs.CL], 2025
2025 arXiv
-
[123]
Motivated numeracy and enlightened self-government,
D. M. Kahan, E. Peters, E. C. Dawson, and P. Slovic, “Motivated numeracy and enlightened self-government,”Behavioural Public Policy, vol. 1, no. 1, pp. 54–86, 2017
2017
-
[124]
Language (technology) is power: A critical survey of “bias
S. L. Blodgett, S. Barocas, H. Daum ´e III, and H. Wallach, “Language (technology) is power: A critical survey of “bias” in NLP,” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds...
2020
-
[125]
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” inProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021, pp. 610–623
2021
-
[126]
A survey on bias and fairness in machine learning,
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” arXiv:1908.09635 [cs.LG], 2022
1908 arXiv
-
[127]
Political biases and inconsistencies in bilingual GPT models—the cases of the US and China,
D. Zhou and Y . Zhang, “Political biases and inconsistencies in bilingual GPT models—the cases of the US and China,”Scientific Reports, vol. 14, no. 1, p. 25048, 2024
2024
-
[128]
A theory of fads, fashion, custom, and cultural change as informational cascades,
S. Bikhchandani, D. Hirshleifer, and I. Welch, “A theory of fads, fashion, custom, and cultural change as informational cascades,”Journal of Political Economy, vol. 100, no. 5, pp. 992–1026, 1992
1992
-
[129]
Data on Notable AI Models,
Epoch AI, “Data on Notable AI Models,” 6 2024, accessed: 2025-04-11. [Online]. Available: https://epoch.ai/data/notable-ai-models
2024
-
[130]
Measuring mathematical problem solving with the MATH dataset,
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” arXiv:2103.03874 [cs.LG], 2021
2021 arXiv
-
[131]
Pythia: A suite for analyzing large language models across training and scaling,
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. Van Der Wal, “Pythia: A suite for analyzing large language models across training and scaling,” inProceedings of th...
2023
-
[132]
Gemma: Open models based on Gemini research and technology,
Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. H ´eliou, A. Tacchetti, A. Bulanova, A. Paterson...
2024 arXiv
-
[133]
The long-run efficiency of real-time electricity pricing,
S. Borenstein, “The long-run efficiency of real-time electricity pricing,”The Energy Journal, vol. 26, no. 3, pp. 93–116, 2005
2005
-
[134]
Electricity Explained: Prices and Factors Affecting Prices,
U.S. Energy Information Administration, “Electricity Explained: Prices and Factors Affecting Prices,” 2023, accessed: 2025-04-23. [Online]. Available: https://www.eia.gov/energyexplained/electricity/prices-and-factors-affecting-prices.php
2023
-
[135]
Handa, D
K. Handa, D. Bent, A. Tamkin, M. McCain, E. Durmus, M. Stern, M. Schiraldi, S. Huang, S. Ritchie, S. Syverud, K. Jagadish, M. V o, M. Bell, and D. Ganguli. (2025) Anthropic education report: How university students use Claude. [Online]. Available: https://www.anthropic.com/new...
2025
-
[136]
Information batteries: storing opportunity power with speculative execution,
J. Switzer and B. Raghavan, “Information batteries: storing opportunity power with speculative execution,”SIGENERGY Energy Inform. Rev., vol. 1, no. 1, pp. 1–11, Nov. 2021. 34
2021
-
[137]
Risk-limited dispatch of knowledge work,
S. Agarwal, Y .-M. Chee, J. Lee, R. R. Sindhgatta, and L. R. Varshney, “Risk-limited dispatch of knowledge work,” Oct. 2014, US Patent App. 13/870,422
2014
-
[138]
Smart operation of smart grid: Risk-limiting dispatch,
P. P. Varaiya, F. F. Wu, and J. W. Bialek, “Smart operation of smart grid: Risk-limiting dispatch,”Proceedings of the IEEE, vol. 99, no. 1, pp. 40–57, 2011
2011
-
[139]
Sustainable supercomputing for AI: GPU power capping at HPC scale,
D. Zhao, S. Samsi, J. McDonald, B. Li, D. Bestor, M. Jones, D. Tiwari, and V . Gadepally, “Sustainable supercomputing for AI: GPU power capping at HPC scale,” inProceedings of the 2023 ACM Symposium on Cloud Computing, 2023, pp. 588–596
2023
-
[140]
RouteLLM: Learning to route LLMs with preference data,
I. Ong, A. Almahairi, V . Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs with preference data,” arXiv:2406.18665 [cs.LG], 2025
2025 arXiv
-
[141]
LoCoML: A framework for real-world ML inference pipelines,
K. Maddireddy, S. K. Methukula, C. Sridhar, and K. Vaidhyanathan, “LoCoML: A framework for real-world ML inference pipelines,” arXiv:2501.14165 [cs.SE], 2025
2025 arXiv
-
[142]
How much does GPT-4 cost?
OpenAI, “How much does GPT-4 cost?” 2024, accessed: 2025-04-25. [Online]. Available: https://help.openai.com/en/articles/ 7127956-how-much-does-gpt-4-cost
2024
-
[143]
About Claude Pro usage,
Anthropic, “About Claude Pro usage,” 2025, accessed: 2025-04-25. [Online]. Available: https://support.anthropic.com/en/articles/ 8324991-about-claude-pro-usage
2025
-
[144]
TypeFly: Flying drones with large language model,
G. Chen, X. Yu, N. Ling, and L. Zhong, “TypeFly: Flying drones with large language model,” arXiv:2312.14950 [cs.RO], 2024
2024 arXiv
-
[145]
A performance analysis of you only look once models for deployment on constrained computational edge devices in drone applications,
L. Rey, A. M. Bernardos, A. D. Dobrzycki, D. Carrami ˜nana, L. Bergesio, J. A. Besada, and J. R. Casar, “A performance analysis of you only look once models for deployment on constrained computational edge devices in drone applications,”Electronics, vol. 14, no. 3, p. 638, Feb. 2025
2025
-
[146]
Efficiently scaling transformer inference,
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems, vol. 5, pp. 606–624, 2023
2023
-
[147]
Reducing activation recomputation in large transformer models,
V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,”Proceedings of Machine Learning and Systems, vol. 5, pp. 341–353, 2023
2023
-
[148]
OPTQ: Accurate post-training quantization for generative pre-trained transformers,
E. Frantar, S. Ashkboos, T. Hoefler, and D.-A. Alistarh, “OPTQ: Accurate post-training quantization for generative pre-trained transformers,” in Proceedings of the International Conference on Learning Representations (ICLR), 2023
2023
-
[149]
AI and memory wall,
A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “AI and memory wall,” arXiv:2403.14123 [cs.LG], 2024
2024 arXiv
-
[150]
Outrageously large neural networks: The sparsely-gated mixture-of- experts layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of- experts layer,” arXiv:1701.06538 [cs.LG], 2017
2017 arXiv
-
[151]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” arXiv:2101.03961 [cs.LG], 2022
2022 arXiv
-
[152]
DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model,
DeepSeek-AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Yang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J...
2024 arXiv
-
[153]
Mixture of parrots: Experts improve memorization more than reasoning,
S. Jelassi, C. Mohri, D. Brandfonbrener, A. Gu, N. Vyas, N. Anand, D. Alvarez-Melis, Y . Li, S. M. Kakade, and E. Malach, “Mixture of parrots: Experts improve memorization more than reasoning,” inProceedings of the International Conference on Learning Representations (ICLR), 2025
2025
-
[154]
Two experts are all you need for steering thinking: Reinforcing cognitive effort in moe reasoning models without additional training,
M. Wang, X. Chen, Y . Wang, Z. He, J. Xu, T. Liang, Q. Liu, Y . Yao, W. Wang, R. Ma, H. Mi, N. Zhang, Z. Tu, X. Li, and D. Yu, “Two experts are all you need for steering thinking: Reinforcing cognitive effort in moe reasoning models without additional training,” arXiv:2505.146...
2025 arXiv
-
[155]
Fast transformer decoding: One write-head is all you need,
N. Shazeer, “Fast transformer decoding: One write-head is all you need,” arXiv:1911.02150 [cs.NE], 2019
1911 arXiv
-
[156]
WaferLLM: A wafer-scale LLM inference system,
C. He, Y . Huang, P. Mu, Z. Miao, J. Xue, L. Ma, F. Yang, and L. Mai, “WaferLLM: A wafer-scale LLM inference system,” arXiv:2502.04563 [cs.LG], 2025
2025 arXiv
-
[157]
Ban the H20: Competing in the Inference Age,
V . Somala, “Ban the H20: Competing in the Inference Age,” https://www.chinatalk.media/p/ban-the-h20-competing-in-the-inference, 2025, accessed: 2025-04-26
2025
-
[158]
SpikeLLM: Scaling up spiking neural network to large language models via saliency-based spiking,
X. Xing, B. Gao, Z. Zhang, D. A. Clifton, S. Xiao, L. Du, G. Li, and J. Zhang, “SpikeLLM: Scaling up spiking neural network to large language models via saliency-based spiking,” arXiv:2407.04752 [cs.LG], 2025
2025 arXiv
-
[159]
A million spiking-neuron integrated circuit with a scalable communication network and interface,
P. A. Merolla, J. V . Arthur, R. Alvarez-Icaza, A. S. Cassidy, J. Sawada, F. Akopyan, B. L. Jackson, N. Imam, C. Guo, Y . Nakamura, B. Brezzo, I. V o, S. K. Esser, R. Appuswamy, B. Taba, A. Amir, M. D. Flickner, W. P. Risk, R. Manohar, and D. S. Modha, “A million spiking-neuro...
2014
-
[160]
Loihi: A neuromorphic manycore processor with on-chip learning,
M. Davies, N. Srinivasa, T.-H. Lin, G. Chinya, Y . Cao, S. H. Choday, G. Dimou, P. Joshi, N. Imam, S. Jain, Y . Liao, C.-K. Lin, A. Lines, R. Liu, D. Mathaikutty, S. McCoy, A. Paul, J. Tse, G. Venkataramanan, Y .-H. Weng, A. Wild, Y . Yang, and H. Wang, “Loihi: A neuromorphic ...
2018
-
[161]
Biological neurons vs deep reinforcement learning: Sample efficiency in a simulated game-world,
F. Habibollahi, M. Khajehnejad, A. Gaurav, and B. J. Kagan, “Biological neurons vs deep reinforcement learning: Sample efficiency in a simulated game-world,” inDeep Reinforcement Learning Workshop NeurIPS 2022, 2022
2022
-
[162]
Brain organoid reservoir computing for artificial intelligence,
H. Cai, Z. Ao, C. Tian, Z. Wu, H. Liu, J. Tchieu, M. Gu, K. Mackie, and F. Guo, “Brain organoid reservoir computing for artificial intelligence,” Nature Electronics, vol. 6, no. 12, pp. 1032–1039, 2023
2023
-
[163]
Directed information flow in computing systems with living neurons,
A. R. Ellis-Mohr and L. R. Varshney, “Directed information flow in computing systems with living neurons,” inProceedings of the 2024 IEEE International Symposium on Information Theory Workshops (ISIT-W), 2024
2024
-
[164]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” arXiv:2201.11903 [cs.CL], 2023
2023 arXiv
-
[165]
Compressed chain of thought: Efficient reasoning through dense representations,
J. Cheng and B. Van Durme, “Compressed chain of thought: Efficient reasoning through dense representations,” arXiv:2412.13171 [cs.CL], 2024
2024 arXiv
-
[166]
Efficient reasoning with hidden thinking,
X. Shen, Y . Wang, X. Shi, Y . Wang, P. Zhao, and J. Gu, “Efficient reasoning with hidden thinking,” arXiv:2501.19201 [cs.CL], 2025
2025 arXiv
-
[167]
TokenSkip: Controllable chain-of-thought compression in LLMs,
H. Xia, Y . Li, C. T. Leong, W. Wang, and W. Li, “TokenSkip: Controllable chain-of-thought compression in LLMs,” arXiv:2502.12067 [cs.CL], 2025
2025
-
[168]
C3oT: Generating shorter chain-of-thought without compromising effectiveness,
Y . Kang, X. Sun, L. Chen, and W. Zou, “C3oT: Generating shorter chain-of-thought without compromising effectiveness,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 23, 2025, pp. 24 312–24 320
2025
-
[169]
Tree of thoughts: Deliberate problem solving with large language models,
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y . Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” inAdvances in Neural Information Processing Systems, 2023, vol. 36, pp. 11 809–11 822
2023
-
[170]
Dynamic cheatsheet: Test-time learning with adaptive memory,
M. Suzgun, M. Yuksekgonul, F. Bianchi, D. Jurafsky, and J. Zou, “Dynamic cheatsheet: Test-time learning with adaptive memory,” arXiv:2504.07952 [cs.LG], 2025. 35
2025 arXiv
-
[171]
Inference scaling for long-context retrieval augmented generation,
Z. Yue, H. Zhuang, A. Bai, K. Hui, R. Jagerman, H. Zeng, Z. Qin, D. Wang, X. Wang, and M. Bendersky, “Inference scaling for long-context retrieval augmented generation,” arXiv:2410.04343 [cs.CL], 2024
2024 arXiv
-
[172]
Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context,
Gemini Team, Google, “Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context,” Google DeepMind, Tech. Rep., Feb. 2024
2024
-
[173]
The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation,
Meta AI, “The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation,” 2025, accessed: 2025-04-25. [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/
2025
-
[174]
Scaling inference-efficient language models,
S. Bian, M. Yan, and S. Venkataraman, “Scaling inference-efficient language models,” arXiv:2501.18107 [cs.LG], 2025
2025 arXiv
-
[175]
Language is primarily a tool for communication rather than thought,
E. Fedorenko, S. T. Piantadosi, and E. A. F. Gibson, “Language is primarily a tool for communication rather than thought,”Nature, vol. 630, pp. 575–586, June 2024
2024
-
[176]
Functional neuroanatomy of deductive inference: a language-independent distributed network,
M. M. Monti, D. N. Osherson, M. J. Martinez, and L. M. Parsons, “Functional neuroanatomy of deductive inference: a language-independent distributed network,”NeuroImage, vol. 37, no. 3, pp. 1005–1016, 2007
2007
-
[177]
The boundaries of language and thought in deductive inference,
M. M. Monti, L. M. Parsons, and D. N. Osherson, “The boundaries of language and thought in deductive inference,”Proceedings of the National Academy of Sciences, vol. 106, no. 30, pp. 12 554–12 559, 2009
2009
-
[178]
Functional specificity for high-level linguistic processing in the human brain,
E. Fedorenko, M. K. Behr, and N. Kanwisher, “Functional specificity for high-level linguistic processing in the human brain,”Proceedings of the National Academy of Sciences, vol. 108, no. 39, pp. 16 428–16 433, 2011
2011
-
[179]
Thought beyond language: Neural dissociation of algebra and natural language,
M. M. Monti, L. M. Parsons, and D. N. Osherson, “Thought beyond language: Neural dissociation of algebra and natural language,”Psychological Science, vol. 23, no. 8, pp. 914–922, 2012
2012
-
[180]
A distinct cortical network for mathematical knowledge in the human brain,
M. Amalric and S. Dehaene, “A distinct cortical network for mathematical knowledge in the human brain,”NeuroImage, vol. 189, pp. 19–31, April 2019
2019
-
[181]
Training large language models to reason in a continuous latent space,
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y . Tian, “Training large language models to reason in a continuous latent space,” arXiv:2412.06769 [cs.CL], 2024
2024 arXiv
-
[182]
Scaling up test-time compute with latent reasoning: A recurrent depth approach,
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein, “Scaling up test-time compute with latent reasoning: A recurrent depth approach,” arXiv:2502.05171 [cs.LG], 2025
2025 arXiv
-
[183]
Imagine while reasoning in space: Multimodal visualization-of-thought,
C. Li, W. Wu, H. Zhang, Y . Xia, S. Mao, L. Dong, I. Vuli ´c, and F. Wei, “Imagine while reasoning in space: Multimodal visualization-of-thought,” arXiv:2501.07542 [cs.CL], 2025
2025 arXiv
-
[184]
Thinking with Images,
OpenAI, “Thinking with Images,” 2025, accessed: 2025-05-05. [Online]. Available: https://openai.com/index/thinking-with-images/
2025
-
[185]
DiffusionCLIP: Text-guided diffusion models for robust image manipulation,
G. Kim, T. Kwon, and J. C. Ye, “DiffusionCLIP: Text-guided diffusion models for robust image manipulation,” arXiv:2110.02711 [cs.CV], 2022
2022 arXiv
-
[186]
Inference-time scaling for diffusion models beyond scaling denoising steps,
N. Ma, S. Tong, H. Jia, H. Hu, Y .-C. Su, M. Zhang, X. Yang, Y . Li, T. Jaakkola, X. Jia, and S. Xie, “Inference-time scaling for diffusion models beyond scaling denoising steps,” arXiv:2501.09732 [cs.CV], 2025
2025 arXiv
-
[187]
Diffusion-LM improves controllable text generation,
X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto, “Diffusion-LM improves controllable text generation,” arXiv:2205.14217 [cs.CL], 2022
2022 arXiv
-
[188]
SSD-LM: Semi-autoregressive simplex-based diffusion language model for text generation and modular control,
X. Han, S. Kumar, and Y . Tsvetkov, “SSD-LM: Semi-autoregressive simplex-based diffusion language model for text generation and modular control,” arXiv:2210.17432 [cs.CL], 2023
2023 arXiv
-
[189]
Discrete diffusion modeling by estimating the ratios of the data distribution,
A. Lou, C. Meng, and S. Ermon, “Discrete diffusion modeling by estimating the ratios of the data distribution,” arXiv:2310.16834 [stat.ML], 2024
2024 arXiv
-
[190]
Scaling diffusion language models via adaptation from autoregressive models,
S. Gong, S. Agarwal, Y . Zhang, J. Ye, L. Zheng, M. Li, C. An, P. Zhao, W. Bi, J. Han, H. Peng, and L. Kong, “Scaling diffusion language models via adaptation from autoregressive models,” arXiv:2410.17891 [cs.CL], 2024
2024 arXiv
-
[191]
Large language diffusion models,
S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y . Lin, J.-R. Wen, and C. Li, “Large language diffusion models,” arXiv:2502.09992 [cs.CL], 2025
2025 arXiv
-
[192]
Fast inference from transformers via speculative decoding,
Y . Leviathan, M. Kalman, and Y . Matias, “Fast inference from transformers via speculative decoding,” inProceedings of the International Conference on Machine Learning, 2023, pp. 19 274–19 286
2023
-
[193]
Multi-agent sampling: Scaling inference compute for data synthesis with tree search-based agentic collaboration,
H. Ye, M. Lin, H. T. Ng, and S. Yan, “Multi-agent sampling: Scaling inference compute for data synthesis with tree search-based agentic collaboration,” arXiv:2412.17061 [cs.CL], 2024
2024 arXiv
-
[194]
Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning,
C. Jin, H. Peng, Q. Zhang, Y . Tang, D. N. Metaxas, and T. Che, “Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning,” arXiv:2504.09772 [cs.AI], 2025
2025 arXiv
-
[195]
Scaling large language model-based multi-agent collaboration,
C. Qian, Z. Xie, Y . Wang, W. Liu, K. Zhu, H. Xia, Y . Dang, Z. Du, W. Chen, C. Yang, Z. Liu, and M. Sun, “Scaling large language model-based multi-agent collaboration,” arXiv:2406.07155 [cs.AI], 2025
2025 arXiv
-
[196]
Computing power and the governance of artificial intelligence,
G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O’Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y . Bengio, and D. Coyle, “Computing power and the governance of artificial intelligen...
2024 arXiv
-
[197]
What does it take to catch a Chinchilla? verifying rules on large-scale neural network training via compute monitoring,
Y . Shavit, “What does it take to catch a Chinchilla? verifying rules on large-scale neural network training via compute monitoring,” arXiv:2303.11341 [cs.LG], 2023
2023 arXiv
-
[198]
Hardware-enabled governance mechanisms: Developing technical solutions to exempt items otherwise classified under export control classification numbers 3A090 and 4A090,
G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, and Z. Winkelman, “Hardware-enabled governance mechanisms: Developing technical solutions to exempt items otherwise classified under export control classification numbers 3A090 and 4A090,” RAND Corporation, Sa...
2024
-
[199]
Open problems in technical AI governance,
A. Reuel, B. Bucknall, S. Casper, T. Fist, L. Soder, O. Aarne, L. Hammond, L. Ibrahim, A. Chan, P. Wills, M. Anderljung, B. Garfinkel, L. Heim, A. Trask, G. Mukobi, R. Schaeffer, M. Baker, S. Hooker, I. Solaiman, A. S. Luccioni, N. Rajkumar, N. Mo ¨es, J. Ladish, N. Guha, J. N...
2024 arXiv
-
[200]
INTELLECT-2: A reasoning model trained through globally decentralized reinforcement learning,
Prime Intellect Team, S. Jaghouar, J. Mattern, J. M. Ong, J. Straube, M. Basra, A. Pazdera, K. Thaman, M. D. Ferrante, F. Gabriel, F. Obeid, K. Erdem, M. Keiblinger, and J. Hagemann, “INTELLECT-2: A reasoning model trained through globally decentralized reinforcement learning,...
2025 arXiv
-
[201]
Pretrained AI models: Performativity, mobility, and change,
L. R. Varshney, N. S. Keskar, and R. Socher, “Pretrained AI models: Performativity, mobility, and change,” arXiv:1909.03290 [cs.CY], 2019
1909 arXiv
-
[202]
A careful examination of large language model performance on grade school arithmetic,
H. Zhang, J. Da, D. Lee, V . Robinson, C. Wu, W. Song, T. Zhao, P. Raja, C. Zhuang, D. Slacket al., “A careful examination of large language model performance on grade school arithmetic,” inAdvances in Neural Information Processing Systems, 2024, vol. 37, pp. 46 819–46 836
2024
-
[203]
Forget what you know about LLMs evaluations – LLMs are like a chameleon,
N. Cohen-Inger, Y . Elisha, B. Shapira, L. Rokach, and S. Cohen, “Forget what you know about LLMs evaluations – LLMs are like a chameleon,” arXiv:2502.07445 [cs.CL], 2025
2025
-
[204]
BIG-Bench extra hard,
M. Kazemi, B. Fatemi, H. Bansal, J. Palowitch, C. Anastasiou, S. V . Mehta, L. K. Jain, V . Aglietti, D. Jindal, P. Chen, N. Dikkala, G. Tyen, X. Liu, U. Shalit, S. Chiappa, K. Olszewska, Y . Tay, V . Q. Tran, Q. V . Le, and O. Firat, “BIG-Bench extra hard,” arXiv:2502.19187 [...
2025 arXiv
-
[205]
Provisions Pertaining to U.S. Investments in Certain National Security Technologies and Products in Countries of Concern,
U.S. Department of the Treasury, “Provisions Pertaining to U.S. Investments in Certain National Security Technologies and Products in Countries of Concern,” Nov. 2024. [Online]. Available: https://www.federalregister.gov/documents/2024/11/15/2024-25422/ provisions-pertaining-t...
2024
-
[206]
Article 51: Classification of General-Purpose AI Models as General-Purpose AI Models with Systemic Risk,
European Union, “Article 51: Classification of General-Purpose AI Models as General-Purpose AI Models with Systemic Risk,” Jul. 2024. [Online]. Available: https://artificialintelligenceact.eu/article/51/
2024
-
[207]
Training compute thresholds: Features and functions in AI regulation,
L. Heim and L. Koessler, “Training compute thresholds: Features and functions in AI regulation,” arXiv:2405.10799 [cs.CY], 2024. 36
2024 arXiv
-
[208]
o3 and o4-mini system card,
OpenAI, “o3 and o4-mini system card,” 2025. [Online]. Available: https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/ o3-and-o4-mini-system-card.pdf
2025
-
[209]
Claude 4 System Card,
Anthropic, “Claude 4 System Card,” 2024. [Online]. Available: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf
2024
-
[210]
AutoML: A survey of the state-of-the-art,
X. He, K. Zhao, and X. Chu, “AutoML: A survey of the state-of-the-art,”Knowledge-based Systems, vol. 212, p. 106622, 2021
2021
-
[211]
Adaptive self-improvement LLM agentic system for ML library development,
G. Zhang, W. Liang, O. Hsu, and K. Olukotun, “Adaptive self-improvement LLM agentic system for ML library development,” arXiv:2502.02534 [cs.CL], 2025
2025
-
[212]
AlphaEvolve: A coding agent for scientific and algorithmic discovery,
A. Novikov, N. Vu, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, A. Mehrabian, M. P. Kumar, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog, “AlphaEvolve: A coding agent for scientific and alg...
2025
-
[213]
ReGenesis: LLMs can grow into reasoning generalists via self-improvement,
X. Peng, C. Xia, X. Yang, C. Xiong, C.-S. Wu, and C. Xing, “ReGenesis: LLMs can grow into reasoning generalists via self-improvement,” arXiv:2410.02108 [cs.CL], 2024
2024 arXiv
-
[214]
Large language models are human-level prompt engineers,
Y . Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba, “Large language models are human-level prompt engineers,” inProceedings of the International Conference on Learning Representations (ICLR), 2022
2022
-
[215]
Frontier Models are Capable of In-context Scheming,
A. Meinke, B. Schoen, J. Scheurer, M. Balesni, R. Shah, and M. Hobbhahn, “Frontier Models are Capable of In-context Scheming,” Apollo Research, Tech. Rep., January 2025. APPENDIXA ADDITIONALANALYSIS FORCOTVS. TOT(1) INFERENCEPOLICIES This appendix provides a supplemental deriv...
2025
-
[217]
Gray line (right axis) shows the percentage-point accuracy gap demonstrating significant differences
Accuracy curves versus compute. Gray line (right axis) shows the percentage-point accuracy gap demonstrating significant differences. 8. Extra compute required by the Chinchilla model to match the accuracy of the global optimum. A stark penalty is incurred over the entire accu...
-
[2025]
Available: https://transformer-circuits.pub/2025/attribution-graphs/methods.html 31
[Online]. Available: https://transformer-circuits.pub/2025/attribution-graphs/methods.html 31
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.