{"id":"7cdb308a-3bd7-4040-ac4f-96d668ad4427","arxiv_id":"2512.13491","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Zipf's law, via differential Heaps and Hilberg laws, forces a power-law lower bound on the excess cross entropy of any entropy-bounded foundation model.","lead":"This paper proves a deductive chain: if text tokens satisfy Zipf's law, then vocabulary growth, block-entropy scaling, and a lower-bound version of neural scaling follow under explicit conditions. It isolates which assumptions each step needs and argues that underparameterization, not overparameterization, is the theoretically optimal regime.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 15(1) is false for non-stationary v=0 processes: Eq. (82) assumes sup_{k≥t} H(K_{k+1}^{k+s}|K_1^t)/s ≥ h, which fails when conditioning on the past drops entropy below the rate; a two-state counterexample violates (77).","rationale":"The reader's weakest assumption (differential Hilberg not derived from plain Hilberg) is a valid empirical gap, but the false Prop15(1) is a more immediate internal correctness risk: it is a stated theorem that fails on a simple two-state process. The central chain for stationary Santa Fe processes may still be intact, because Prop15(2) supplies the needed stationary inequality; however the abstract and overview advertise arbitrary non-stationarity and a systematic derivation, which this counterexample refutes. A revision should remove or repair Prop15(1), make the stationarity assumption explicit in the chain, and then re-evaluate whether Eq. (86) can be justified for natural language corpora. This does not change the conditional verdict but strengthens the conditions attached to it.","tokens_in":20770,"tokens_out":18924,"duration_ms":160376,"concrete_test":"Evaluate Prop. 15(1) on the process K_t~Bernoulli(1/2) for t≤0, K_t=0 for t≥1, Z_k~Bernoulli(1/2), at t=1. Compute LHS of (77): sup_{k≥1} H(X_{k+1}^{k+s}|X_1^1)/s − h = 0−1 = −1, while RHS = C7·inf_{k≥1} V(K_{k+1}^{k+s}\\setminus K_1^1)/s = 0. If this is verified, Prop. 15(1) is false; the theorem must be restricted to stationary (or otherwise corrected) before the non-stationary generality is claimed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Prop. 15(1) claims that for any Santa Fe process with narration K satisfying v=0, Eq. (77) holds. The proof's key step is Eq. (82): sup_{k≥t} H(K_{k+1}^{k+s}|K_1^t)/s ≥ h. But h is the unconditional rate from (33); conditioning on K_1^t can lower conditional block entropy below h s. No stationarity or mixing is assumed in part (1). Counterexample: let K_t be iid Bernoulli(1/2) for t≤0 and K_t=0 for t≥1, with Z_k iid fair bits independent of K. Then V(K_{k+1}^{k+s})≤2, so v=0, while h_X = h_K = 1. For t=1, X_1=(0,Z_0), so every future X_i=(0,Z_0) is known; hence sup_{k≥1} H(X_{k+1}^{k+s}|X_1^1)/s = 0. Eq. (77) gives 0−1 ≥ C7·0, contradiction. Thus the advertised implication from differential Heaps to differential Hilberg for arbitrary non-stationary processes is unsound. The stationary version (Prop. 15(2)) is not affected, but the paper's claim of non-stationary generality and the overview chain (B) as stated are invalid. The differential Hilberg law (86) is therefore only derived for stationary Santa Fe processes, not for arbitrary v=0 processes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper attempts a systematic deductive chain from Zipf's law to Heaps' law, from Heaps' law to Hilberg's hypothesis, and from Hilberg's hypothesis to the neural scaling law, with Santa Fe processes as a running example. The main mathematical contributions are: (i) Propositions 13–14, deriving differential Heaps laws from Zipf-type tail conditions; (ii) Proposition 15, transferring differential Heaps laws to differential Hilberg laws for Santa Fe processes; and (iii) Proposition 16, deriving a lower bound on excess cross-entropy from a differential Hilberg law and entropy-budget constraints on the model, leading to the exponent inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper is written as a formal proof-based consolidation rather than an empirical study, and it candidly lists open problems about tightness and the meaning of compute.","tokens_in":21234,"tokens_out":15284,"duration_ms":136451,"significance":"If the stated assumptions are granted, the proofs provide a useful formal baseline connecting four well-known empirical regularities. The paper is honest about the main limitations: the derived neural-scaling statement is one-sided, and the differential Hilberg law is stronger than the empirically studied plain Hilberg law. Strengths include the explicit assumption-by-assumption organization, the absence of curve-fitting in the derivation, the Santa Fe process as a concrete non-vacuous example, and the clear identification of what would be needed to close the gaps. The work is best viewed as a formal lower-bound theory for scaling exponents rather than a derivation of the equality-form neural scaling law. Its significance is therefore conditional, but it is a useful contribution to the theoretical literature on scaling laws.","major_comments":[{"comment":"Proposition 16 proves only a lower bound on the excess expected cross-entropy. It does not prove the equality-form neural scaling law in Eqs. (8)–(10), and the resulting inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1 are one-sided. The abstract and conclusion nevertheless state that 'the neural scaling law is a consequence' and that the constraints 'produce the neural scaling law.' The open problems in §4 concede that tightness is unresolved. The paper should be recast as deriving testable lower bounds and upper bounds on exponents under explicit assumptions, not as a derivation of the empirical scaling law itself.","section":"§3.3, Eq. (90); Abstract"},{"comment":"Assumption (86) is a differential Hilberg law, which is substantially stronger than the empirically studied Hilberg hypothesis (6): it requires a uniform bound on conditional block entropies for all future starting points, not just the unconditional block entropy H(X_1^t). It is not derived from (6), and Proposition 15 derives it only for stationary Santa Fe processes. Thus the chain advertised in the abstract—Hilberg's hypothesis ⇒ neural scaling—does not follow for the plain Hilberg law. The paper needs either a derivation of (86) from (6) for a relevant class of processes or an explicit statement that the neural-scaling result depends on an unverified strengthening.","section":"§3.3, Eq. (86); §1, Eq. (15)"},{"comment":"The expression in (90) contains the term ((1−y)/(1+y))^{1−β} with y = (c t^{−β}/(1−β))^{1/2}. If y ≥ 1, the base 1−y is non-positive and the non-integer power is not real; moreover, the definition of s_max in (89) can become negative. The proposition as stated is therefore not a valid real inequality for all c, t, n. The statement should explicitly restrict to y < 1 (e.g., c < (1−β)t^β), and state that the bound is vacuous otherwise. This does not affect the asymptotic regime of fixed c and t→∞, but the theorem statement needs a domain condition.","section":"§3.3, Eqs. (90)–(95)"}],"minor_comments":[{"comment":"In the proof of Proposition 15(1), the step sup_{k≥t} H(K_{k+1}^{k+s}|K_1^t)/s ≥ h might appear to assume stationarity. It is actually a consequence of Proposition 6's inf-characterization; adding a pointer would remove ambiguity.","section":"§3.2, Eq. (82)"},{"comment":"Equation (14) is displayed before the variables f, y, and the entropy-budget assumptions are introduced. Consider moving the display after Proposition 16 or defining the terms inline.","section":"§3.3, Eq. (14)"},{"comment":"There are occasional typos, e.g., 'equvalent' in the Conclusion. More importantly, the abstract and introduction should consistently say 'lower bound on excess cross-entropy' rather than 'neural scaling law' in places where only the lower-bound result is meant.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of cs.IT and the core proofs appear sound under their stated assumptions. The main problem is that the advertised implications are broader than the theorems: the central result is a one-sided bound, and the key differential Hilberg assumption is not derived from the plain Hilberg law. With a revision that recalibrates the claims and fixes the domain issue in Eq. (90), the paper would be a solid contribution. The non-stationary counterexample raised in one stress test does not appear to land, because Proposition 6's inf-characterization provides the inequality used in Eq. (82)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful formal chain showing that Zipf's law implies a lower-bound form of the neural scaling law, through differential versions of Heaps' and Hilberg's laws. The genuinely new pieces are the differential laws and the entropy-budget argument for compute and parameters. The abstract oversells the result: Prop 16 proves a lower bound, not the scaling law itself, and the paper's own open problems concede tightness is unresolved.\n\nWhat's good: Props 13 and 14 give clean bounds on the hapax rate under tail conditions, with an explicit IID/Zipf example. Prop 16 is a neat use of subadditivity and the source-coding inequality: under the differential Hilberg law (86) and entropy bounds on the model, it forces a power-law lower bound on expected cross entropy as a function of tokens t and parameters n (and compute c). The Santa Fe process is worked out properly as a toy model satisfying all four laws. No fitting is done; the chain is a genuine derivation from stated assumptions.\n\nSoft spots, in order: (1) The abstract/conclusion claim 'the neural scaling law is a consequence of Zipf's law' is too strong. The theorem gives one-sided bounds, and the exponents gamma_T <= 1-beta, gamma_N <= 1/beta-1 are loose relative to observed gamma_T ~ 0.095, gamma_N ~ 0.076. (2) The load-bearing differential Hilberg law (86) is stronger than the empirically studied plain Hilberg law (6) and is not derived from it except in the stationary Santa Fe case; for arbitrary processes it is an assumption. The paper is upfront about this, but it limits the practical reach. (3) The identification of parameter count and compute with entropy bounds (87)-(88) is a coarse abstraction, acknowledged. Minor: Prop 15(1) states an equality (80) that should carry constants C7/C8; not a real problem.\n\nI disagree with the stress-test note. Its counterexample to Prop 15(1) miscomputes the entropy rate: in that example the block entropy of length s is 1 bit, so H/s = 1/s and h = 0, not 1. Thus (77) gives 0-0 >= C7*0, which holds. Prop 6 does guarantee sup_{k>=t} H(X_{k+1}^{k+s}|X_1^t)/s >= h for each s, so the step in (82) is sound. The non-stationary version of Prop 15(1) survives.\n\nWho it's for: theorists working on scaling law explanations, and quantitative linguists who want the Zipf/Heaps/Hilberg implications made precise. It won't resolve the empirical exponent puzzle, but it gives a clean baseline. I'd cite it for the differential-law formulation.\n\nRecommendation: send to peer review. The core theorems are novel and mostly rigorous, and the author is honest about the gaps. The referee should ask for a rewritten abstract/conclusion that matches the one-sided nature of the result and an explicit discussion of the status of the differential Hilberg law as an assumption for non-stationary processes.","headline":"A rigorous lower-bound chain from Zipf to neural scaling, with the abstract overselling a one-sided result; the non-stationary step survives the stress-test.","tokens_in":21645,"tokens_out":13855,"would_cite":true,"duration_ms":104786,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","60G10","91F20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the neural scaling law of foundation models can be derived deductively from Zipf's law via Heaps' law and Hilberg's hypothesis, under explicit information-theoretic assumptions.","keywords":["neural scaling law","Zipf's law","Heaps' law","Hilberg's hypothesis","Santa Fe processes","cross entropy","entropy rate","power laws"],"falsifier":"On a large corpus, compute sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h for a range of t and s; if this conditional excess entropy decays faster than C (t+s)^{β−1} for any β close to the compression-based estimate 0.8, or vanishes over long horizons, Proposition 16's premise fails and the chain from Hilberg to neural scaling is not applicable to that corpus.","tokens_in":20675,"feed_emoji":"📉","tokens_out":5419,"duration_ms":45269,"temperature":0.7,"pith_summary":"This paper tries to prove that the empirical power-law improvement of large language models is not a mystery of architecture or optimization but a consequence of ordinary token statistics. It builds a formal chain: Zipf's law implies Heaps' law on vocabulary growth; Heaps' law implies Hilberg's hypothesis on entropy growth; Hilberg's hypothesis implies the neural scaling law that ties cross entropy to training tokens, parameters, and compute. The key step is a differential form of Hilberg's law, stronger than the plain law studied empirically, plus entropy budgets that model limited data, parameters, and compute. A toy source called the Santa Fe process is shown to satisfy all four laws, illustrating that large memory and power-law complexity need not mean intuitively complex structure. If the chain holds, the scaling exponents of modern models are upper-bounded by the Hilberg exponent β, with γ_T ≤ 1−β and γ_N ≤ 1/β−1.","feed_headline":"Neural scaling laws trace back to Zipf's law","feed_subtitle":"A deductive chain through Heaps' law and Hilberg's hypothesis explains power-law LLM loss from token statistics.","key_machinery":"The load-bearing object is the differential Hilberg law, Eq. (86): the conditional block entropy per symbol, sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h, is bounded below by C (t+s)^{β−1}. It is a strengthened, horizon-dependent version of the plain Hilberg law, and it is what makes the entropy-budget argument go through. Alongside it sit the entropy budgets (87)–(88), which model compute and parameter counts as caps on conditional and total Shannon entropy, and the Santa Fe process, a toy source in which each token is a pair (K_t, Z_{K_t}) of a Zipf-distributed index and a copied knowledge bit, which lets Heaps-type vocabulary growth be converted into Hilberg-type entropy growth.","core_discovery":"The central claim is Proposition 16: if a stochastic text satisfies the differential Hilberg law — conditional excess entropy per symbol bounded below by a power law in the horizon — and if a trained model's entropy is bounded by compute and parameter budgets, then the worst-case expected cross entropy exceeds the entropy rate by at least the explicit expression in Eq. (90). This yields exponent bounds γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper further derives the differential Hilberg law from a differential Heaps law for stationary Santa Fe processes, and derives that Heaps law from approximate Zipf distributions under mixing conditions. The conclusion is that the observed neural scaling law can","pith_inferences":["The paper leaves open whether real corpora satisfy the differential Hilberg law; measuring that conditional excess entropy directly on large datasets would test the crucial premise without relying on the plain-law proxy.","If the entropy-budget identification of parameter and compute counts is replaced by resource-bounded Kolmogorov complexity, the exponent bounds might tighten and possibly explain why overparameterized models appear better.","A testable extension: across languages or domains with different measured β, the framework predicts correspondingly different neural scaling exponents — a cross-linguistic prediction that current scaling-law studies do not yet examine.","The two-regime Zipf distributions and log-log convex vocabulary growth documented in quantitative linguistics would break the simple single-β picture; the framework suggests they should show up as piecewise or varying scaling exponents in model loss."],"forward_implications":["If the differential Hilberg law holds with exponent β, the neural scaling exponents satisfy γ_T ≤ 1−β and γ_N ≤ 1/β−1, so token-scaling and parameter-scaling exponents are determined by the same language-level exponent.","Under the empirical compression-based estimate β≈0.8, the bounds give γ_T ≤ 0.2 and γ_N ≤ 0.25, which are loose compared with the commonly cited values γ_T≈0.095 and γ_N≈0.076; the paper suggests internet-scale corpora may have a larger β.","The derivation predicts underparameterization (γ_T < γ_N) as the optimal regime when parameters have bounded entropy, in contrast to the overparameterization reported in practice.","The Santa Fe process shows a Zipf-distributed IID narration over random knowledge bits satisfies Heaps' law and Hilberg's law, so Hilberg's hypothesis does not require intuitively complex structure.","Because the final implication (C) allows arbitrary non-stationary processes, the neural-scaling step is the most robust link in the chain once its differential premise is granted."],"fun_headline_variants":["Neural scaling law emerges from Zipf's law","Zipf's law spawns neural scaling","From Zipf's law to neural scaling via Heaps and Hilberg","Zipf's law explains neural scaling","Neural scaling law is a consequence of Zipf's law"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is the differential Hilberg law, Eq. (86) — a horizon-wise power-law lower bound on conditional excess entropy that strengthens the empirically studied plain Hilberg law and is assumed rather than derived from it; if real text satisfies only the plain law, the neural-scaling conclusion does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Neural scaling law emerges from Zipf's law","Zipf's law spawns neural scaling","From Zipf's law to neural scaling via Heaps and Hilberg","Zipf's law explains neural scaling","Neural scaling law is a consequence of Zipf's law"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000805,"raw_usage":{"total_tokens":3346,"prompt_tokens":691,"completion_tokens":2655,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":2578}},"tokens_in":435,"tokens_out":2655,"duration_ms":18790,"temperature":1.0,"reasoning_tokens":2578,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:25:36.462667+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a large corpus, compute sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h for a range of t and s; if this conditional excess entropy decays faster than C (t+s)^{β−1} for any β close to the compression-based estimate 0.8, or vanishes over long horizons, Proposition 16's premise fails and the chain from Hilberg to neural scaling is not applicable to that corpus.","supporting_citations":[],"review_version":1}