{"id":"26c14d7b-99ea-4bd7-aa3e-d22f50cca5b8","arxiv_id":"2502.15719","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"If pretraining scaling plateaus while capabilities keep rising through reasoning models, compute-based triggers in the EU AI Act and US export controls will miss the main sources of risk.","lead":"A policy essay argues that frontier AI laws built on pretraining compute thresholds will lose their grip if AI progress shifts to inference-time reasoning and other approaches. It introduces the idea of a 'pretraining frontier' and proposes transparency, data and compute monitoring, and regulator capacity building as alternatives.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The regulatory-critique argument rests on the empirical bet that pretraining scaling is ending; the paper's own Section 4 admits that if scaling continues or stalls differently, current compute-based tools remain effective, so the abstract's assertion of an actual paradigm shift outruns the…","rationale":"The reader identified the same load-bearing concern: the argument's practical force depends on the contested empirical claim that the pretraining paradigm is ending. My reading agrees with that assessment. The paper is internally consistent and explicitly hedged in several places: Section 2 says the paradigm 'may be ending,' Section 4 says 'The pretraining paradigm may be coming to an end,' and the conclusion says 'The end of the pretraining paradigm, if it comes.' Those hedges are honest and are a genuine strength. However, the abstract and the strongest claim as formulated by the reader assert a paradigm shift as though it is already undermining the regulatory order, and the policy recommendations in Section 5 are presented as timely correctives rather than as one branch of a scenario tree. Because Section 4 itself concedes that the current compute-centric frameworks remain effective in the scaling-continues and AI-winter scenarios, the only thing separating 'valuable scenario analysis' from 'urgent regulatory diagnosis' is an implicit probability estimate for the paradigm-shift branch. The paper does not supply that estimate, and the evidence it does supply is disputed by sources it cites. This is not an internal contradiction, and it does not warrant rejection; a conditional acceptance is the right posture. The concrete test—tracking post-publication flagship training runs against scaling-law extrapolations—would resolve whether the premise is currently true or merely hypothetical, and would clarify how much weight the essay's conclusions should carry.","tokens_in":22713,"tokens_out":5175,"duration_ms":52736,"concrete_test":"Compile a post-publication record of flagship pretraining runs released after 1 February 2025 (e.g., GPT-5.x, Gemini 3, Claude 4, Llama 4, DeepSeek successors), recording total training FLOPs and token counts where disclosed or credibly estimated, plus benchmark gains on a common suite. Fit the observed compute-performance pairs against the Hoffmann et al. 2022 and Kaplan et al. 2020 scaling-law extrapolations. If 2025-2026 models continue to deliver benchmark improvements in line with or above the extrapolated scaling curve, the premise that the pretraining paradigm is ending is falsified and the paper's urgency claim fails; if gains flatten relative to compute, or substantive capability increases shift to inference-time compute, the premise is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract as 'This paradigm shift presents fundamental challenges... threatening to undermine this new legal order as it emerges,' treats the end of pretraining scaling as a real, near-term development. The evidence offered in Section 2 is a set of press reports that OpenAI, Google, and Anthropic were disappointed with recent runs (refs 52, 65, 78), the finite supply of human text (ref 85), and Sutskever's prediction (ref 82); the paper itself acknowledges refs 17 and 67 disagree, with SemiAnalysis arguing that scaling-law-style progress is continuing through reasoning training infrastructure and post-training. Section 4 then lists three futures: progress stalls into an AI winter, technical breakthroughs overcome the data wall, or the paradigm shifts. In the first two cases the paper concedes existing regulatory frameworks 'would likely be preserved' because compute governance continues to work. Thus the actionable conclusion that regulators must reorient now is not entailed by the paper's own conditional analysis; it depends on a probability judgment about which future is coming. The paper supplies no likelihood estimate or decision-theoretic treatment, and a reader who assigns high probability to continued scaling has no reason to adopt Section 5's proposed transition away from compute-centric triggers. The concern is about an empirical premise, not internal logic: the essay is transparent about its 'if' framing in places, but the abstract and policy recommendations are written as though the shift has already happened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a policy analysis of frontier AI regulation in light of a possible end to the \"pretraining paradigm\"—the assumption that scaling up pretraining compute and data is the primary driver of AI capability gains. It argues that current regulatory instruments (the EU AI Act's 10^25 FLOPs threshold, US export controls on chips, the forthcoming UK approach, and Chinese draft AI law) are premised on pretraining scale as a legible, trackable bottleneck. If that paradigm ends, the paper contends, these regulations become misaligned with a more diffuse and capability-diverse frontier. The paper introduces the concept of the \"pretraining frontier,\" analyzes industry-structure implications such as deconcentration, and proposes regulatory alternatives: increased transparency, data-oriented triggers, wider or inference-focused compute governance, selective information regulation, and capacity-building. The central argument is conditional: much of the text says \"if the pretraining paradigm ends,\" and Section 4 explicitly lists alternative futures in which existing frameworks would remain effective.","tokens_in":22978,"tokens_out":3469,"duration_ms":35669,"significance":"The paper addresses a timely and important policy question with concrete, actionable proposals, and it is unusually honest about the uncertainty surrounding its central premise. It engages with the actual mechanisms of existing laws (compute thresholds, export controls, institutional capacity) and offers a useful taxonomy of possible regulatory responses. Its discussion of evaluating risks, grouping models by capabilities, and using data or inference-compute as alternative bottlenecks is thoughtful and goes beyond generic calls for 'more AI safety research.' The manuscript is not an empirical study; it does not claim to prove that pretraining scaling is ending, and much of the analysis is self-consciously conditional. That being said, the abstract states the paradigm shift as an unfolding fact, while the body concedes that the evidence is contested and that other futures are possible. If reframed to align its claims with its own conditional structure, the paper would be a valuable contribution to the frontier AI governance literature.","major_comments":[{"comment":"The abstract asserts an actual paradigm shift ('This paradigm shift presents fundamental challenges... threatening to undermine this new legal order as it emerges'), but Section 4 (paragraph 2) states that if progress stalls into an AI winter or if technical breakthroughs overcome the data wall, 'the effectiveness of existing regulatory frameworks would likely be preserved' because compute-based governance continues to work. The actionable conclusion that regulators should reorient now is therefore not entailed by the paper's own conditional analysis; it depends on an unstated probability judgment about which future is more likely. The manuscript should either consistently frame its thesis as conditional ('if the pretraining paradigm ends, then...') throughout the abstract and recommendations, or provide an explicit decision-theoretic justification for adapting regulations now despite the uncertainty. Without this, the abstract overstates what the analysis establishes.","section":"Abstract and Section 4"},{"comment":"The evidentiary basis for the load-bearing premise that the pretraining paradigm is ending is thin and the paper itself acknowledges it is contested. The supporting cites are press reports of disappointment at OpenAI, Google, and Anthropic (refs 52, 65, 78), an estimate of the finite supply of human text (ref 85), and a prediction by Ilya Sutskever (ref 82); the paper explicitly notes that refs 17 and 67 disagree, with SemiAnalysis arguing that scaling-law-style progress continues through reasoning and post-training infrastructure. If the manuscript aims to persuade the reader that the paradigm shift is actually occurring, this evidence is insufficient and should be strengthened with quantitative trends (e.g., scaling-law fits, benchmark improvements over compute, model release trajectories). If it is merely a conditional analysis, the premise should be presented as a stipulated scenario rather than as the basis for the abstract's factual-sounding claim.","section":"Section 2"}],"minor_comments":[{"comment":"The manuscript is printed with 'Unpublished working draft. Not for distribution.' at the top of every page; for a journal submission this header should be removed and replaced with a conventional title page.","section":"Title page"},{"comment":"The sentence 'And regulations premised on the idea that scaling with remain the driver of capabilities will have to change.' contains a typo: 'scaling with remain' should read 'scaling will remain'.","section":"Section 2, final paragraph"},{"comment":"References are formatted inconsistently: [31] lacks the article title, [36] is given as a White House fact sheet but is cited for export controls that are also described in [33] and [35], and several URLs are broken across line breaks (e.g., [28], [66]). The list should be cleaned up before publication.","section":"References"},{"comment":"The analysis of the 'Forthcoming UK Frontier AI Bill' is necessarily provisional because the bill had not been released at the time of writing. The paper should state its 'as of' date more prominently in that section, since the analysis may quickly become dated if the bill is published with different triggers.","section":"Section 3.3"},{"comment":"The proposal that cloud providers report inference-compute expenditures above a threshold is interesting but underspecified regarding legal authority, technical feasibility, and privacy implications; a footnote or caveat acknowledging these open questions would strengthen the presentation.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a genuine policy contribution, but it is currently framed as a working draft, and the abstract's language overstates the strength of the underlying evidence. The main fix—reconciling the conditional reasoning in Section 4 with the abstract's declarative tone—is well within the manuscript's scope and should be achievable in revision. I would also note that the paper's fit with a technical journal depends on whether the venue welcomes policy essays; if it does, after revision it could be a solid publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is a policy essay, not a technical paper, and the genuinely new piece is the 'pretraining frontier' construct—the idea that if pretraining scaling tops out, the regulatory choke points built on compute thresholds (EU AI Act Article 51, US export controls, the draft UK bill) lose their target. The paper also maps out a concrete menu of alternatives: inference-compute reporting by cloud providers, data stewardship, algorithmic-transparency regimes, and capacity-building. That synthesis goes beyond prior work (Sastry et al., Narayanan et al.) that flagged the general worry about compute thresholds becoming outdated.\n\nThe paper does well on structure and honesty. It acknowledges the empirical disagreement (refs 17 and 67 argue scaling-style progress continues via reasoning training infrastructure), and Section 4 explicitly lays out three futures—AI winter, technical breakthrough keeping pretraining dominant, and paradigm shift. In the first two, the paper concedes existing compute-based governance 'would likely be preserved.' That is the right kind of conditional honesty.\n\nThe soft spot is real and load-bearing. The abstract and the policy recommendations are written as though the paradigm shift is already happening, while the body only supports it as a scenario. The evidence for the end of pretraining is a handful of press reports that companies were disappointed with recent runs, plus the finite supply of human text. The paper itself notes the disagreement. No likelihood estimate or decision-theoretic framing is provided, so the claim that regulators must reorient now is not entailed by the analysis. A reader who assigns high probability to continued scaling—through reasoning models, synthetic data, or better training infrastructure—has little reason to adopt the proposed shift away from compute-centric triggers. This is not a flaw in internal logic; it is an empirical premise that the paper treats as more settled than its own footnotes allow. Framing the paper explicitly as scenario analysis would fix most of this.\n\nMinor point: the section on the EU AI Act is strong on the Article 51 triggers but slightly overstates how much the 10^25 threshold will 'implicitly' drive coverage; the evaluations trigger and Commission decisions are genuine alternatives, and the paper acknowledges this but then downplays it.\n\nBottom line: for AI governance readers and policymakers, this is a valuable scenario-planning piece and deserves a serious referee; it should be revised to align the abstract with the conditional body. I'd bring it to a reading group and would cite it as an example of forward-looking regulatory analysis.\n\nRecommendation: send to peer review, with the expectation of a major revision focused on reframing the empirical claim.","headline":"A serious, clearly written policy essay whose actionable conclusion is conditional on an empirical bet that pretraining scaling is ending; the abstract overstates the evidence, but the scenario analysis and concrete governance proposals are genuinely useful.","tokens_in":23491,"tokens_out":1927,"would_cite":true,"duration_ms":18053,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frontier AI governance, built on pretraining compute thresholds, will become misaligned if the pretraining paradigm ends; the paper proposes transparency, new bottlenecks, and regulatory capacity as the path forward.","keywords":["frontier AI regulation","pretraining paradigm","compute thresholds","scaling laws","AI safety","transparency","inference-time compute","data wall"],"falsifier":"A single observation would settle much of the debate: if a next-generation model trained with substantially more compute (at or above the EU's FLOP threshold) delivers a clear capability jump over GPT-4-class models and continues to improve according to scaling laws, then the pretraining frontier is not binding and compute-based thresholds remain effective. The paper itself cites sources predicting this outcome, so the distinction is observable in upcoming model releases.","tokens_in":22493,"feed_emoji":"⚖️","tokens_out":5024,"duration_ms":44708,"temperature":0.7,"pith_summary":"Frontier AI laws now being enacted in the EU, US, UK, and China treat pretraining scale as the choke point that makes the technology governable: compute thresholds (like the EU AI Act's $10^{25}$ FLOPs presumption) and chip export controls assume that more pretraining compute is the main route to more capability. The paper argues that this \"pretraining paradigm\" is hitting a wall because good training data is finite and recent large runs disappointed their developers, while capabilities keep advancing through inference-time reasoning and other routes. If the paradigm ends, the paper claims, the regulatory field will become more diffuse: more companies can reach the \"pretraining frontier,\" new bottlenecks replace compute, and current triggers stop tracking risk. The paper introduces the \"pretraining frontier\" as the capability ceiling of pretraining alone and proposes a path forward based on transparency, new natural bottlenecks (data, inference compute, algorithmic innovation), and stronger regulatory capacity. The stakes are immediate because several major legal frameworks come into force this year and may be misaligned with the technology as it actually develops.","feed_headline":"If pretraining stalls, today's AI laws lose their grip","feed_subtitle":"Compute thresholds in the EU AI Act and US chip export controls rest on scaling assumptions that may be ending.","key_machinery":"The object that carries the argument is the \"pretraining frontier\": the capabilities threshold made possible by scaling up pretraining alone, given finite high-quality data and compute. It does the work of defining what changes after the pretraining paradigm: it marks the point at which scaling no longer buys capability, opening the field to more actors and forcing regulation to target new bottlenecks. The other load-bearing mechanism is the compute-threshold trigger (for example, the EU AI Act's $10^{25}$ FLOPs presumption and US Executive Order 14110's $10^{26}$ FLOPs threshold), which the paper analyzes as the legal expression of the pretraining assumption.","core_discovery":"The central claim is that current frontier AI governance is built on the pretraining paradigm—the regularity that scaling up pretraining data and compute produces predictable capability gains—and that this paradigm is ending, so the regulations built on it are becoming misaligned with the technology they are meant to govern. The load-bearing regulatory tools are compute thresholds: the EU AI Act's $10^{25}$ FLOPs presumption, US Executive Order 14110's $10^{26}$ FLOPs trigger, and US export controls on advanced microchips, all of which assume that large pretraining runs are the key input to frontier capability and the natural bottleneck to monitor, control, or exclude. The paper introduces the \"pretraining frontier,\" the capabilities ceiling reachable by scaling pretraining alone, and argues that if it binds, the field deconcentrates as chip scarcity eases, more players reach the frontier, and new sources of progress—inference-time compute, synthetic data, and algorithmic innovation—take over. It then argues that regulators should stop betting on pretraining compute as the universal handle and instead invest in transparency, find new natural bottlenecks (data, inference compute, information), and build regulatory capacity, while protecting rights. If the argument is right, the legal order for frontier AI now being built is aimed at the wrong layer of the technology.","pith_inferences":["A testable corollary of the paper's argument is that the distribution of frontier-capable models should widen over the next few years, with open-weight models reaching GPT-4-class performance even without massive new pretraining runs.","If the pretraining frontier is real, export controls on advanced chips lose strategic value while retaining diplomatic costs; the paper gestures at this but does not develop the geopolitical consequence.","The paper's logic implies that regulators should start building capability-evaluation triggers now, before a crisis makes them necessary, rather than waiting to see which bottleneck stabilizes.","Inference-time compute governance would depend on cloud providers monitoring customer usage, which raises privacy and civil-liberties tradeoffs that the paper acknowledges but does not resolve."],"forward_implications":["If the pretraining frontier binds, compute thresholds such as the EU's FLOP-based presumption and US chip export controls stop tracking capability, because compute no longer predicts gains.","More companies can reach the frontier as chip scarcity eases and Moore's law raises effective compute supply, making the frontier field more diffuse and harder to monitor.","Incumbents may preserve their lead by shifting to inference-time compute scaling and synthetic data, creating new bottlenecks that regulators can target.","Regulation should move toward transparency, monitoring inputs like data and inference compute, and grouping models by risk class rather than by training scale.","If capability gains continue through non-pretraining paths, regulators should count all compute used in development and deployment, not just pretraining compute.","\n"],"supporting_citations":[{"why":"Establishes the neural scaling laws that underpin the pretraining-paradigm assumption regulators rely on.","marker":"[7]"},{"why":"Shows compute-optimal training ratios, the basis for the claim that data and compute constraints bind as scaling continues.","marker":"[34]"},{"why":"Provides the original scaling-law result predicting loss from compute and data, which the paper says regulators use to forecast capabilities.","marker":"[40]"},{"why":"Documents inference-time reasoning as an alternative path beyond pretraining scaling, driving the paradigm-shift claim.","marker":"[44]"},{"why":"Press reporting that OpenAI, Google, and Anthropic were disappointed by recent pretraining runs, the main evidence that the paradigm is ending.","marker":"[52]"},{"why":"Reporting that OpenAI is shifting strategy as GPT improvements slow, supporting the decoupling of pretraining scale and capability gains.","marker":"[65]"},{"why":"Reporting that the next major models are behind schedule and expensive, evidence for the pretraining wall.","marker":"[78]"},{"why":"Estimates the finite stock of human-generated training data, the core mechanism behind the data wall.","marker":"[85]"},{"why":"Argues that computing power is the governance bottleneck; the paper claims this window is closing.","marker":"[76]"},{"why":"The EU AI Act's Article 51 compute threshold that the paper identifies as the key legal trigger tied to pretraining.","marker":"[1]"}],"fun_headline_variants":["Compute-based AI laws may target a dying paradigm","As pretraining hits a wall, AI governance needs a new lever","EU and US AI rules rest on a scaling assumption that's faltering","The pretraining frontier: why today's AI laws could soon miss the mark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands or falls on the premise that the pretraining paradigm is actually ending—that scaling pretraining with more data and compute no longer yields significant capability gains—rather than merely slowing temporarily.","fun_headline_variants_meta":{"raw":{"variants":["Compute-based AI laws may target a dying paradigm","As pretraining hits a wall, AI governance needs a new lever","EU and US AI rules rest on a scaling assumption that's faltering","The pretraining frontier: why today's AI laws could soon miss the mark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000585,"raw_usage":{"total_tokens":2808,"prompt_tokens":1058,"completion_tokens":1750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":1676}},"tokens_in":674,"tokens_out":1750,"duration_ms":12255,"temperature":1.0,"reasoning_tokens":1676,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:37:44.641762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single observation would settle much of the debate: if a next-generation model trained with substantially more compute (at or above the EU's FLOP threshold) delivers a clear capability jump over GPT-4-class models and continues to improve according to scaling laws, then the pretraining frontier is not binding and compute-based thresholds remain effective. The paper itself cites sources predicting this outcome, so the distinction is observable in upcoming model releases.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents inference-time reasoning as an alternative path beyond pretraining scaling, driving the paradigm-shift claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Press reporting that OpenAI, Google, and Anthropic were disappointed by recent pretraining runs, the main evidence that the paradigm is ending."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reporting that OpenAI is shifting strategy as GPT improvements slow, supporting the decoupling of pretraining scale and capability gains."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reporting that the next major models are behind schedule and expensive, evidence for the pretraining wall."}],"review_version":1}