{"id":"8869537a-b93c-40d0-8a33-dbdfd625a2dc","arxiv_id":"2503.05705","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"If frontier AI progress shifts from pre-training compute to inference-time compute, AI governance must be rebuilt around deployment-time capabilities and transparency, with different implications depending on whether that inference happens during deployment or inside the training process.","lead":"This essay argues that if AI labs keep spending more compute at 'thinking time' rather than on ever-larger training runs, the usual levers of AI governance, such as training-compute thresholds, may stop working. It sketches two very different futures, inference used when a model answers questions versus inference used as part of training, and their policy consequences.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The essay's conclusions depend on the under-evidenced premise that pre-training scaling has plateaued near GPT-4; a Reuters anecdote and a 'seems to have fallen' trend line support it, and if pre-training resumes historical growth the governance implications weaken.","rationale":"The reader's weakest_assumption ('pre-training scaling has plateaued near GPT-4') is precisely the load-bearing premise I identify. The paper's own language is hedged ('may,' 'if'), and it explicitly acknowledges the uncertainty about whether pre-training growth has gone to zero or just slowed. Because the essay promises a set of governance implications conditional on a shift to inference scaling, the argument is internally coherent and not circular; the risk is purely empirical. My concrete test would settle the premise's truth using public compute and benchmark data. Since the reader already judged the essay CONDITIONAL with moderate confidence and identified this same assumption, no verdict adjustment is needed; the stress-test confirms rather than overturns the reader's assessment. I also note the unsourced appendix claim about deployment compute exceeding training compute, but that is peripheral to the central claim and does not require a further verdict change.","tokens_in":8916,"tokens_out":5383,"duration_ms":50277,"concrete_test":"Compile Epoch AI's data on the training compute of the largest released models from GPT-4 (2023) through GPT-5/Claude 4/Gemini 2.5 (2025-2026), and compute the year-over-year growth rate of pre-training FLOP. Then compare benchmark deltas (e.g., MMLU, GPQA, AIME) between successive frontier models at fixed inference compute. If pre-training FLOP growth exceeds 3x/year and benchmark deltas match historical scaling-law expectations, the plateau premise is falsified and the governance conclusions lack their stated basis. Alternatively, if growth has indeed fallen below ~1.5x/year with performance gains coming from post-training/inference scaling, the premise holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 asserts 'The end of an era' based on 'reports from unnamed employees' (Hu & Tong, 2024) that scaling pre-training beyond GPT-4 yielded modest gains, and Section 3 concedes it is 'not yet clear' whether the pre-training growth rate has fallen 'to zero or to some fraction of its previous rate.' Yet every governance consequence in the paper—threshold breakage, open-weight decline, business-model change, reduced value of stolen weights, and the entire inference-during-training scenario—requires that inference compute becomes the dominant scaling lever. If pre-training continues at anything close to the historical 5x/year (Epoch AI), a below-threshold model amplified by inference would still be overtaken by new pre-trained models; compute thresholds remain meaningful; open weights remain the main proliferation vector; and the 'first human-level AGI' arrives via pre-training, not inference. The paper offers no systematic evidence for the plateau beyond anecdotal reporting and a single OpenAI benchmark chart; the rate of pre-training growth is treated as an uncertainty, but the central claim is actually conditional on it. Thus the load-bearing assumption is empirical and currently under-supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the apparent plateau in pre-training compute scaling and the rise of inference-time compute may transform AI governance. It distinguishes two scenarios: inference-at-deployment, where extra compute is spent on reasoning during use, and inference-during-training, where extra compute is spent inside the training process (for synthetic data or iterated distillation and amplification). The paper derives a list of governance consequences for each scenario, including reduced value of closed-model weights, reduced strategic importance of open-weight models, higher cost of first human-level systems, breakage of compute-threshold regulation, changes to the frontier-lab business model, and less transparency into state-of-the-art capabilities in the training-internal scenario. The essay is explicitly exploratory, with many hedged claims and caveats.","tokens_in":9112,"tokens_out":4102,"duration_ms":42149,"significance":"The paper is a timely and clearly written mapping of an important potential shift in AI development. Its main value is in laying out a structured set of governance implications that differ sharply between deployment-side and training-side inference scaling, and in flagging the threshold problem in compute-based regulation. The treatment of iterated distillation and amplification, while speculative, is a useful contribution because it connects an existing AI-alignment idea to current frontier-lab practices. The essay's strengths are its conceptual clarity, honest acknowledgment of uncertainty, and the explicit scenario structure. Its main weakness is that the central empirical premise that pre-training scaling has plateaued near GPT-4 level rests on thin evidence, so the governance conclusions are conditional on an assumption that could fail.","major_comments":[{"comment":"The claim that pre-training scaling has plateaued near GPT-4 level is supported only by anonymous Reuters reporting (Hu & Tong 2024) and a single OpenAI chart, and the paper itself later concedes in Section 3 that it is 'not yet clear' whether the pre-training growth rate has fallen to zero or to some fraction of its previous rate. This premise is load-bearing for every governance consequence in Sections 4–6: threshold breakage, the diminished value of model weights, the reduced strategic importance of open-weight models, the business-model shift, and the inference-during-training scenarios. If pre-training continues at something close to the historical 5x/year (Epoch AI), most of the described effects would be substantially delayed or would not occur. The paper would be much stronger if it either presented this as one explicit scenario among two and analyzed the continuation case in parallel, or provided additional systematic evidence (e.g., recent training-run compute records, algorithmic-efficiency trends) for the plateau. As written, the central argument rests on an under-evidenced empirical premise that is acknowledged as uncertain but is then treated as the basis for the rest of the essay.","section":"Section 2 (The end of an era) and Section 3"},{"comment":"The statement that 'deployment compute exceeding total training compute on commercial frontier systems' (footnote ‡‡‡) is asserted without supporting data or a citation, and the footnote text appears to be missing from the manuscript. This claim is load-bearing for the appendix's conclusion that when deployment compute dominates, scaling inference by 10x increases total costs by nearly 10x while scaling pre-training by 10x increases costs by only about 3x. That cost comparison in turn motivates the essay's claim that inference scaling changes the industry's business model. Please either add a verifiable citation or clearly label this as an assumption rather than an established fact, and restore the missing footnote.","section":"Appendix (cost comparison)"},{"comment":"The threshold-breakage example uses a linear effective-OOM conversion (0.7 × OOMs of inference) to claim that a 10^24 FLOP model with 4 OOM of inference scaling would perform at the level of a 10^27 FLOP model. This extrapolates to a 10,000x inference multiplier even though the paper later acknowledges that current inference-scaling techniques hit performance plateaus that cannot be exceeded by any level of compute. The argument would be more convincing if the threshold-breakage conclusion were framed as a growing risk that depends on breakthrough research in inference scaling, with the plateau limitation discussed in the same section rather than in the later paragraph. As it stands, the example overstates the certainty of the threshold problem.","section":"Section 4 (Breaking the strategy of AI governance via compute thresholds)"}],"minor_comments":[{"comment":"Several footnote markers (‡, ‡‡‡) appear in the text but the corresponding footnote text is missing or not rendered in the manuscript; this should be fixed before publication.","section":"Global/typography"},{"comment":"The word 'rapdily' should be 'rapidly'.","section":"Section 6"},{"comment":"The phrase 'overton window' should be capitalized as 'Overton window'.","section":"Section 6"},{"comment":"The 'Effective orders of magnitude' equation would benefit from a citation to the source of the 0.7 coefficient and a note that the original source gives a range (0.5 to 1.0) rather than a fixed value.","section":"Section 4"},{"comment":"The reference to 'Epoch AI (2024)' for the 5x/year pre-training growth rate is generic; please specify the exact page or dataset.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an essay of the kind that can be valuable for the governance community, and the scenario analysis is well structured. My major concern is that the essay's central pivot depends on an empirical premise—pre-training plateau—that is currently supported only by anecdotal reporting. If the author reframes the essay as fully conditional on that premise, or adds real data on recent training runs, the piece would be publishable. The missing footnote content for the deployment-compute claim is a sign the manuscript needs a careful technical check before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a good, honest scenario essay from Toby Ord about what AI governance would look like if inference compute replaces pre-training compute as the main scaling lever. The two-scenario structure — inference at deployment vs. inference during training — is genuinely useful, and I don't know another governance piece that lays out the consequences this clearly. The compute-threshold breakage argument and the open-weights analysis are the strongest parts; the AlphaGo Zero / iterated distillation and amplification discussion is a fair and well-sourced synthesis.\n\nThe paper's main soft spot is the empirical premise that pre-training scaling has plateaued near GPT-4. That rests on a Reuters story quoting unnamed employees and a 'seems to have fallen' trend line. The stress-test note says this is load-bearing, and I think that's only half right. The paper repeatedly says it's not yet clear whether the growth rate went to zero or just slowed, and the governance analysis is presented as conditional on the shift happening. So it isn't circular or deceptive. But the title overstates things, and if pre-training continues to scale at anything like 5x per year, much of the interesting conclusions (threshold breakage, open-weights decline) lose their force. That's worth telling readers up front. Also, the appendix's 'deployment compute already exceeds training compute' claim is asserted without a citation, and the 'effective OOMs' formula is a rule of thumb with no derivation. Both are minor for an essay, but a referee could ask for evidence.\n\nThe citation pattern is fine. It credits Epoch, Christiano, Silver et al., and the relevant scaling work. No sign of self-citation inflation.\n\nWho's this for? Policy people and AI-governance researchers who want a fast, readable mapping of the governance stakes under inference scaling. It's not a technical contribution and doesn't pretend to be. I'd give it a serious referee — it's important enough and well-argued enough to warrant careful engagement, even though the empirical base is shaky. The referee should push the author to either label it explicitly as a conditional scenario analysis in the title and abstract, or bring more evidence for the plateau claim.","headline":"A clear, honest scenario analysis of governance under inference scaling, built on an empirical premise the paper itself flags as uncertain.","tokens_in":9641,"tokens_out":2268,"would_cite":true,"duration_ms":22667,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI governance would need an overhaul if compute scaling shifts from training to inference.","keywords":["inference scaling","AI governance","pre-training compute","training compute thresholds","iterated distillation and amplification","open-weight models","recursive self-improvement","frontier AI"],"falsifier":"Track the compute and capability of successive frontier pre-training runs over the next few years. If leading labs show that orders of magnitude more pre-training compute still yield the capability gains they did from 2020 to 2024, the premise that the pre-training era is over collapses, and with it the urgency of the paper's governance analysis. In the other direction, if inference-scaling techniques reliably plateau well below a 100,000x multiplier for general tasks, the claim that thresholds such as the EU AI Act's $10^{25}$ FLOP can be breached by inference amplification weakens.","tokens_in":8647,"feed_emoji":"⚖️","tokens_out":11503,"duration_ms":86906,"temperature":0.7,"pith_summary":"The paper argues that the era of scaling pre-training compute — the basis of most current AI governance — may be ending, and that a shift to scaling inference compute would upend standard governance assumptions. Whether the new compute is spent at deployment (extra reasoning per user query, as in o1 and R1) or during training (using reasoning models to build better models) leads to sharply different futures. Inference-at-deployment would lower the importance of open-weight models, blunt the impact of the first human-level systems, change the frontier business model, and break regulation keyed to training-compute thresholds. Inference-during-training could either revive pre-training by supplying high-quality synthetic data or, through iterated distillation and amplification, produce rapid recursive self-improvement with far less outside visibility. In either case, the author concludes, the future becomes less predictable, and governance should track which form of inference scaling is actually unfolding.","feed_headline":"Inference compute shift would break AI governance thresholds","feed_subtitle":"Whether new compute goes to deployment or training decides which governance rules survive.","key_machinery":"The argument is carried by a distinction between two destinations for scaled inference compute. Inference-at-deployment spends extra compute per user query, and the paper's rule of thumb is that each order of magnitude of inference compute adds roughly 0.7 orders of magnitude of effective pre-training equivalence. Inference-during-training spends the compute inside the lab, using an inference-amplified model to generate synthetic training data or to guide a search whose outputs are distilled into a new model. The second mechanism, iterated distillation and amplification, is the paper's most consequential piece of machinery: start from a system-1 model, amplify it with inference-time search, distill the amplified behaviour into a new model, and repeat — the loop that carried AlphaGo Zero past world-champion play in forty days, and which the paper argues is a plausible pathway for general LLMs. Against these, the paper sets the existing governance machinery: compute thresholds such as the EU AI Act's $10^{25}$ FLOP and the US executive order's $10^{26}$ FLOP, which draw a bright regulatory line by training compute alone, and which inference scaling threatens to make unworkable.","core_discovery":"The central claim is that the marginal compute driving frontier AI is shifting from pre-training to inference, and that this is not an implementation detail but a change of regime. If the scaled compute is spent at deployment, then the number of simultaneous copies of a new frontier model falls by roughly a factor of 100 per two orders of magnitude of inference scaling; the first human-level systems may cost more than equivalent human labour; model weights become less worth stealing; open-weight models become less attractive and less dangerous; the software-like business model of frontier AI erodes; monolithic data centres lose strategic centrality; and governance via training-compute thresholds (the EU AI Act's $10^{25}$ FLOP, the US executive order's $10^{26}$ FLOP) breaks, because a model trained below the threshold can be amplified by inference to perform at the level of a far larger training run. If the scaled compute is spent during training, the effects are more ambiguous: it could feed high-quality synthetic data back into pre-training, or — in the more consequential variant — it could power iterated distillation and amplification in the manner of AlphaGo Zero, a ladder in which each rung distills an inference-amplified model into a stronger base model. The paper treats this loop as a plausible form of recursive self-improvement for general LLMs that could shorten timelines to transformative AI while remaining invisible to outside observers.","pith_inferences":["My inference: the threshold problem the paper identifies may partially self-correct, because inference costs pass to users; if only well-resourced actors can afford a dangerous level of amplification, the governance target shifts from model weights to who can pay for compute.","My inference: the two scenarios are asymmetric — the deployment scenario is the near-term default, while the distillation–amplification scenario is the higher-consequence tail; a governance regime that tracks the ratio of training to deployment compute at frontier labs would be the natural early-warning metric.","My inference: because inference scaling is tunable per task, a dangerous capability could be concentrated on a few high-value targets at enormous compute multipliers even when average deployment compute stays low, so safety cases should weigh worst-case concentration of inference compute rather than averages.","My inference: the argument's empirical engine is data scarcity, so the thesis is testable by watching whether frontier labs resume rapid pre-training scaling after absorbing synthetic-data techniques; a resumption would weaken the urgency of the governance overhaul without refuting the mechanics of the deployment scenario."],"forward_implications":["Inference-at-deployment could let a model trained below a $10^{25}$ FLOP threshold perform at the level of a $10^{27}$ FLOP training run, breaking threshold-based regulation and pushing governance toward use-based rules.","Inference-at-deployment would make open-weight releases less attractive to users, since the heavy inference costs fall on the deployer rather than the trainer.","Inference-at-deployment could make the first human-level systems more expensive to run than equivalent human labour, creating a window to study or demonstrate them before transformative deployment.","Inference-during-training, if it works through iterated distillation and amplification, could shorten timelines to transformative AGI while keeping the best models hidden from outside observers.","Both scenarios argue for policies that require disclosure of current capabilities and immediate plans, and for monitoring which kind of inference scaling is actually happening."],"supporting_citations":[{"why":"Reuters report, cited in the opening, that leading labs see only modest gains from scaling pre-training beyond GPT-4; it supplies the premise that the pre-training era is ending.","marker":"Hu & Tong, 2024"},{"why":"The o1 release announcement whose chart shows performance rising with post-training and inference compute; it is the template for the inference-scaling scenarios.","marker":"OpenAI 2024"},{"why":"Defines the training-compute-threshold approach to AI regulation (EU AI Act, US executive order) that the paper argues inference scaling breaks.","marker":"Heim & Koessler, 2024"},{"why":"The AlphaGo Zero result, used as the proof of concept that iterated amplification and distillation can climb far beyond human performance.","marker":"Silver et al. 2017"},{"why":"Supplies the amplification–distillation idea in an AI-safety context, which the paper extends to LLM training as a path to recursive self-improvement.","marker":"Christiano 2017"},{"why":"Provides the 5x-per-year pre-training compute growth figure and the algorithmic-efficiency trend used in the cost and weights-value analysis.","marker":"Epoch AI 2024"},{"why":"The Chinchilla scaling relation, used in the appendix to compare how pre-training versus inference scaling changes total compute cost.","marker":"Hoffmann et al., 2022"},{"why":"Quantifies the trade-off between pre-training and inference compute, underpinning the effective-orders-of-magnitude rule of thumb.","marker":"Villalobos & Atkinson 2023"}],"fun_headline_variants":["Inference scaling evades AI compute caps","Shift to inference compute defeats AI law","Inference scaling undermines training thresholds","Deployment inference bypasses AI governance","Compute shift to inference breaks AI thresholds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that pre-training scaling has plateaued near GPT-4 level, so that future frontier progress must come substantially from inference compute; if pre-training resumes rapid scaling, the described governance disruptions lose their force.","fun_headline_variants_meta":{"raw":{"variants":["Inference scaling evades AI compute caps","Shift to inference compute defeats AI law","Inference scaling undermines training thresholds","Deployment inference bypasses AI governance","Compute shift to inference breaks AI thresholds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000349,"raw_usage":{"total_tokens":1907,"prompt_tokens":942,"completion_tokens":965,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":903}},"tokens_in":558,"tokens_out":965,"duration_ms":9428,"temperature":1.0,"reasoning_tokens":903,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:35:04.500226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Track the compute and capability of successive frontier pre-training runs over the next few years. If leading labs show that orders of magnitude more pre-training compute still yield the capability gains they did from 2020 to 2024, the premise that the pre-training era is over collapses, and with it the urgency of the paper's governance analysis. In the other direction, if inference-scaling techniques reliably plateau well below a 100,000x multiplier for general tasks, the claim that thresholds such as the EU AI Act's $10^{25}$ FLOP can be breached by inference amplification weakens.","supporting_citations":[{"cited_title":"This has led to intense speculation that the previous era of scaling pre-training compute could be followed by an era of scaling up inference-compute","cited_arxiv_id":null,"evidence_quote":"The o1 release announcement whose chart shows performance rising with post-training and inference compute; it is the template for the inference-scaling scenarios."}],"review_version":1}