{"id":"3cc4c822-8d56-4bee-976f-557bc7b4b381","arxiv_id":"2504.13219","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":16,"one_line_summary":"The paper fits power-law scaling laws for downstream vision tasks and claims a pretraining-data threshold where distilled models stop outperforming non-distilled ones, but the theory's assumptions come from the fitted exponents and the model-size scaling is contradicted by the paper's own figures.","lead":"The paper proposes scaling laws for visual transfer learning when downstream data are scarce, and a distillation boundary where small-to-large knowledge distillation helps at low pretraining data volumes but hurts beyond a threshold. The authors fit power-law exponents on four image classification datasets and derive the crossover condition from those same fitted values.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1's proof asserts F(Dp)>0 for small Dp, but the paper's fitted exponents (alpha'>alpha) make F→−∞ as Dp→0; the guaranteed distillation crossover is not established.","rationale":"The reader's weakest assumption focused on the monotone model-size term M^{−β}, citing Appendix D and Figures 4 and 13–15. That is a real and serious contradiction: the text explicitly says error can increase with model size, so Eq. (1)'s monotone model-size term cannot hold in those regimes. My pass identifies an additional, more direct internal flaw in the proof of Theorem 1: the claim that F(Dp) is positive for small Dp is false under the paper's own fitted exponents, because α'>α forces F→−∞ as Dp→0. This is not an external challenge to the scaling-law convention but a failure of the paper's mathematical derivation. Since the central advertised contribution includes a guaranteed distillation boundary, and the proof of that boundary is unsupported, the rejection stands. I would not manufacture a non-finding here; the internal inconsistency is concrete and testable. The concrete test is deliberately analytic: it requires only the reported exponents and the definition of F, so it settles the proof's validity without new experiments. If the authors can show F(Dp*)>0 and justify an intermediate interval, a revised version might be salvageable, but as written the central claim is not established.","tokens_in":26043,"tokens_out":9886,"duration_ms":105507,"concrete_test":"Analytical check: substitute Table 1 α=0.620 and Table 2 α'=0.702 into Eq. (8) and take Dp→0 with any positive λp, λp'; the term −Dp^{−α'}/λp' dominates, so F→−∞, contradicting the proof's 'positive for small Dp' claim. Then, using the released fits or a refit from the paper's data, evaluate F at the local maximum Dp* from Eq. (14). If F(Dp*)≤0, no crossover occurs under the fitted laws; if F(Dp*)>0, the theorem must be restated with an intermediate interval of distillation superiority rather than a small-Dp superiority claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Concern with Theorem 1's derivation. In Lemma 1, F(Dp)=E1−E2 reduces to Δ + Dp^{−α}/λp − Dp^{−α'}/λp'. The proof states that F is 'positive for small Dp' and then uses the Intermediate Value Theorem together with lim_{Dp→∞}F=Δ<0 to conclude a threshold beyond which standard training beats distillation. This positivity premise is not proven, and it is incompatible with the paper's own fitted exponents. Table 1 and Table 2 give α'−α>0 for both distilled datasets (ImageNet100: 0.702 vs 0.620; TinyImageNet: 0.475 vs 0.412), and Eq. (6) imposes α−α'∈(−1,0), i.e. α'>α. For fixed positive λp and λp', the term −Dp^{−α'}/λp' dominates as Dp→0, so F(Dp)→−∞, not +∞. Thus, under the fitted power laws, F must start negative, possibly cross to positive at some intermediate Dp, and only later cross back; the claimed 'distillation superiority in data-scarce conditions' is not a consequence of the scaling laws. The proof also asserts Δ<0 while treating the M-dependent term in Eq. (10) as negligible even though it is positive: with β<β', M^{−β}/λm − M^{−β'}/λm' is positive before the exponential-decay approximation is applied. Therefore Theorem 1's guaranteed crossover is not derived. Separately, Appendix D concedes that model-size scaling is non-monotonic, which further contradicts the monotone M^{−β} term in Eq. (1) used to define E1. The central claim rests on an unproven and internally inconsistent analytic step.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies scaling behavior of visual transfer learning in data-constrained settings. It proposes an additive power-law model E = E_inf + Dp^{-alpha}/lambda_p + M^{-beta}/lambda_m + Df^{-gamma}/lambda_f for downstream error (Eq. 1) and an analogous law for distilled models with teacher and student model sizes (Eq. 4). It fits these laws on ImageNet100, TinyImageNet, CIFAR100, and CIFAR10 over pretraining data sizes from roughly 64K to 1.3M, finetuning data from 2K to 130K, and model sizes from 2.5M to 38M. The central theoretical contribution is the 'distillation boundary theory' (Theorem 1), which asserts that a critical pretraining-data threshold D_p^* exists: distilled models are better for small Dp and worse for large Dp, under parametric constraints (6). The proof compares E1 and E2 through a differential function F(Dp) and applies the intermediate value theorem. The paper reports empirical error-difference plots (Figures 5 and 6) that it interprets as confirming the boundary.","tokens_in":26604,"tokens_out":7290,"duration_ms":71987,"significance":"If the claimed scaling laws and distillation boundary theorem were valid, the framework would offer practical guidance for choosing between standard fine-tuning and small-to-large distillation under data constraints, and it would extend the scaling-law literature to low-data transfer settings. The paper has genuine strengths: broad experiments across four downstream datasets, multiple model scales and data fractions, and fitted parameter tables. It is also transparent about limitations in Appendix F and about non-monotonic model-size behavior in Appendix D. However, the central theoretical result is not established: the sign analysis in Lemma 1 is internally inconsistent with the fitted exponents, the model-size term is not monotone as assumed, and the theorem's conditions are drawn from the same fits used to confirm the prediction. These are load-bearing concerns that affect the paper's main claim.","major_comments":[{"comment":"The proof asserts that F(Dp) is 'positive for small Dp' and then uses the intermediate value theorem to conclude a crossover. This assertion is not proven and is incompatible with the paper's own fitted exponents. Under Table 1 and Table 2, alpha' > alpha for both ImageNet100 (0.702 vs 0.620) and TinyImageNet (0.475 vs 0.412), and Eq. (6) imposes alpha - alpha' in (-1,0), i.e., alpha' > alpha. For fixed positive lambda_p and lambda_p', the term -Dp^{-alpha'}/lambda_p' dominates as Dp -> 0, so F(Dp) -> -infinity, not +infinity. Thus the claimed guarantee of 'distillation superiority in data-scarce conditions' does not follow from the fitted scaling laws; under those laws F could be negative for all Dp, or cross from negative to positive and back, and the simple one-crossing picture in Lemma 1 is not established.","section":"Section 4.1, Lemma 1 proof, Eq. (12), Tables 1 and 2"},{"comment":"The proof of Delta < 0 is not sound. In Eq. (10), the first model-size term M^{-beta}/lambda_m - M^{-beta'}/lambda_m' is positive under the constraint beta < beta' and lambda_m ~ lambda_m', not negative. The argument that M^{-beta'} is 'vanishingly small' ignores the paper's own setup in which M is reported in millions with values 3, 10, 22, 38 (e.g., Figures 4 and 13-15); for M = 3 and beta' = 5.84, the term is not negligibly small relative to the other terms in Delta. Since the positive model-size contribution could offset the negative terms, the conclusion Delta < 0, which is needed to show that traditional training eventually wins for large Dp, is unsupported.","section":"Section 4.1, Eqs. (9)-(11)"},{"comment":"The paper's own Appendix D states that 'as the downstream model size increases, the error rate and loss do not simply decrease; in some cases, they even increase with model size,' and Figures 13-15 show visibly non-monotonic curves. This directly contradicts the monotone decreasing M^{-beta} term in Eq. (1) and the claim in Obs. 3 and the Figure 4 caption that error and loss decrease as M increases. Because the model-size term is one of the three axes of the proposed scaling law and also enters Theorem 1 through E1 in Eq. (5), the empirical support for the model-size component of the central claim is missing.","section":"Appendix D, Figures 4 and 13-15, Eq. (1), Obs. 3"},{"comment":"The parametric constraints in Eq. (6) are described in Remark 1 as 'derived from extensive experimental evidence,' and the same fitted exponents and coefficients are then used to prove the existence of the crossover in Theorem 1. Since the inequalities E < E', gamma > gamma', beta < beta', alpha - alpha' in (-1,0), and the coefficient near-equalities are satisfied by construction of the fits (Tables 1, 2, and 4), the theorem states a consequence of the fitted parametric family rather than an independent, falsifiable prediction. Confirming the crossover on the same datasets used for fitting is therefore partly circular; a held-out dataset or a prediction made before fitting would be needed to validate the claimed boundary.","section":"Theorem 1, Remark 1, Eq. (6), Tables 1, 2, and 4"}],"minor_comments":[{"comment":"The sentence listing the fitted exponents says 'alpha, beta, gamma, alpha', beta', and gamma'' but Eq. (4) also contains the exponent eta' for teacher model size; please correct the list.","section":"Section 3.4, Eq. (4)"},{"comment":"ImageNet-100 is described as 'containing approximately 1.2 million training images distributed across 1000 categories'; since ImageNet100 is later defined as a 100-class subset of ImageNet1K, this appears to be a typo.","section":"Section 5.1, Pre-training dataset"},{"comment":"The text says 'beta of 4.882 and -5.086 respectively' for ImageNet100 and TinyImageNet, but Table 1 lists 5.086 for TinyImageNet without a minus sign; the discrepancy should be fixed.","section":"Section 5.3, text after Table 1"},{"comment":"Reporting goodness-of-fit measures (e.g., R^2, RMSE, or confidence intervals) for the fitted exponents would materially help readers assess whether the proposed power-law forms are distinguishable from simpler alternatives in the data range considered.","section":"Section 5.3 and Tables 1, 2, 4"}],"recommendation":"reject","confidential_remarks":"The paper's positioning as the 'first practical framework' for data-efficient scaling laws is stronger than its own related-work section supports: Hernandez et al. (2021) and Wang et al. (2023) already propose scaling laws for transfer and model reusing, respectively, and the paper does not clearly delineate the novel prediction beyond a new experimental range. Additionally, the central theorem is conditional on constraints extracted from the same fits used for validation, so it should not be presented as a standalone theoretical result. If the authors were to revise, the most useful path would be to repair the lemma, report fit quality, and test the predicted crossover on a dataset or model configuration not used for fitting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: the experimental study is real, but the central theoretical claim doesn't survive contact with the paper's own numbers. The observed crossover between distillation and direct fine-tuning might be genuine, but this paper doesn't establish it.\n\nWhat's actually new and useful: the experimental sweep is substantial. Four downstream datasets, model sizes from 2.5M to 38M, pretraining data from 64K to 1.3M, fine-tuning data in the 1K-130K range, under a consistent small-to-large distillation protocol. Extending the additive power-law to vision with a model-size term is a reasonable move, and the observation that distillation's edge shrinks as pretraining data grows is practically interesting if it holds up.\n\nThe soft spots are load-bearing. Lemma 1's proof asserts F(D_p) is positive for small D_p and uses that plus lim F = Δ < 0 to get a guaranteed zero crossing. But the paper's own constraint (6) imposes α − α' in (−1, 0), i.e. α' > α, and its own fits give α' > α on both datasets. With α' > α, the −D_p^{−α'}/λ'_p term dominates as D_p → 0, so F(D_p) → −∞, not +∞. The claimed \"distillation superiority in data-scarce conditions\" is the opposite of what the fitted equations imply. The sign is backwards.\n\nSecond, the threshold is circular in the way that matters. Remark 1 says conditions (6) are \"derived from extensive experimental evidence,\" and the proof of Theorem 1 then uses those same fitted values to derive the crossover. The theorem is algebra on the fit, not an independent prediction. The empirical crossover in Figures 5-6 may be real, but the theory section adds nothing beyond the fitted curves.\n\nThird, the model-size term in Eq. (1) is contradicted by the paper's own Appendix D, which says error \"do[es] not simply decrease; in some cases, they even increase with model size,\" matching Figures 4 and 13-15. Obs 3 claims the opposite. A monotone M^{−β} law with β > 0 is not supported by the data it was fitted to.\n\nMinor: the Δ < 0 step waves a positive term out of the sign-determining position, and the CIFAR10 exponents (α = 10.129, β'' = 25.435) suggest fit instability.\n\nWho this is for: a practitioner might get a rule of thumb from the empirical curves, and a referee could usefully sort out whether the crossover is robust. But the theory section as written shouldn't be trusted. Recommendation: send to peer review rather than desk-reject—the empirical question deserves referee time. Expect heavy revision: fix or drop the theorem, reconcile Appendix D with the main text, and recast the boundary as an empirical observation.","headline":"Real experimental effort, but the distillation-boundary proof contradicts the paper's own fitted exponents and the model-size law contradicts its own appendix.","tokens_in":27096,"tokens_out":6620,"would_cite":false,"duration_ms":62178,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Visual transfer error follows an additive power law with a crossover where distillation stops paying off.","keywords":["scaling laws","transfer learning","knowledge distillation","data-efficient learning","vision transformers","power law","downstream fine-tuning","distillation boundary"],"falsifier":"Fit Eq. (1) on a random subset of the data-size/model-size grid and evaluate on held-out sizes for any one of the four datasets; if held-out error swings upward as $M$ increases (as Appendix D shows for some settings) or the residuals do not decay as a power law in $D_p$, $D_f$, and $M$, then the additive law is not the right description and the derived threshold $D^*_p$ lacks predictive content.","tokens_in":25843,"feed_emoji":"📉","tokens_out":9878,"duration_ms":93554,"temperature":0.7,"pith_summary":"Visual-transfer error, not just pretraining error, is claimed to follow an additive power law in three independent scales: pretraining data size $D_p$, model size $M$, and fine-tuning data size $D_f$. On top of that law, the paper argues that knowledge distillation from a small teacher to a larger student obeys its own power law with an extra teacher-size term, and that comparing the two laws forces a sharp crossover: distilled models win in data-scarce regimes, while ordinary pretrain-and-finetune wins once pretraining data exceeds a critical threshold $D^*_p$. If the claim holds, model-reuse strategy becomes a quantitative decision that can be computed from a handful of fitted exponents rather than a rule of thumb.","feed_headline":"Distillation pays off only below a pretraining-data threshold","feed_subtitle":"A fitted power law maps exactly when small-to-large distillation beats standard fine-tuning.","key_machinery":"The central object is the additive power-law error function (Eq. 1) and its distilled analogue (Eq. 4). The mechanism that produces the boundary is the difference $F(D_p)=E_1(D_p)-E_2(D_p)$: with the fitted parameter constraints $E<E'$, $\\gamma>\\gamma'$, $\\beta<\\beta'$, $\\alpha-\\alpha'\\in(-1,0)$, $\\lambda_m\\approx\\lambda'_m$, and $\\lambda_f\\approx\\lambda'_f$, $F$ is positive for small $D_p$ and tends to a negative constant as $D_p\\to\\infty$, so continuity and the intermediate value theorem guarantee a root $D^*_p$ where the two strategies tie. That root is the claimed crossover point.","core_discovery":"The paper's central claim is a two-part scaling description of transfer learning in vision. First, downstream error and cross-entropy loss follow $E(D_p,M,D_f) = E_\\infty + D_p^{-\\alpha}/\\lambda_p + M^{-\\beta}/\\lambda_m + D_f^{-\\gamma}/\\lambda_f$ (and the analogous loss law), with fitted exponents showing pretraining data as the dominant factor. Second, distillation has a boundary: comparing this law with the distilled-model law $E = E_\\infty + D_p^{-\\alpha'}/\\lambda'_p + M_s^{-\\beta'}/\\lambda'_s + D_f^{-\\gamma'}/\\lambda'_f + M_t^{-\\eta'}/\\delta'$ yields a critical pretraining-data threshold $D^*_p$ at which the error difference between the two strategies changes sign, so distillation is superior below the threshold and inferior above it. The paper reports the predicted sign flip, with error-difference curves crossing from positive to negative at a critical data size, across model sizes from 2.5M to 38M parameters.","pith_inferences":["If the additive law holds beyond the four benchmarks, the fitted equations could be read as a budget rule: spend on distillation only when the cost of more pretraining data is higher than the cost of a larger teacher, since the boundary shifts with both model sizes.","The non-monotonic model-size curves the paper itself reports in Appendix D suggest a natural extension: replacing the model-size term with a shape-aware term (depth versus width) would make both the fitted law and the crossover prediction safer outside the 2.5M to 38M parameter range.","The theory implies that 'data-efficient transfer' and 'knowledge distillation' are two ends of the same curve rather than separate techniques: inherited knowledge substitutes for data up to a saturation point, and that point is measurable in advance from a few cheap runs."],"forward_implications":["In data-scarce downstream settings, a small teacher distilling into a larger model should beat the same larger model trained without distillation, so distillation is the recommended strategy below the threshold.","Beyond the critical pretraining-data threshold, the distilled model's inherited bias costs more than it saves, and standard pretrain-and-finetune should be preferred.","Because the distilled law contains explicit teacher-size and student-size terms, the threshold can be re-estimated for any small-to-large pair before running the full experiment.","The fitted exponents rank pretraining data as the strongest lever on downstream performance, suggesting that adding pretraining data is a more efficient use of resources than enlarging the model or adding fine-tuning examples in the tested range."],"supporting_citations":[{"why":"supplies the prior scaling-law-for-transfer baseline (effective data transferred) that this paper extends to downstream vision tasks.","marker":"[16]"},{"why":"provides the small-to-large model-reusing idea on which the distillation setup is built.","marker":"[42]"},{"why":"justifies varying model size by head count rather than overall shape, and supplies the foundational power-law form.","marker":"[20]"},{"why":"establishes the power-law scaling principle for neural models that the paper adapts to downstream error.","marker":"[19]"},{"why":"previous work on distillation scaling laws that the paper's distilled-model equation builds on.","marker":"[6]"},{"why":"supplies the DeiT architecture used for all pretraining, fine-tuning, and distillation experiments.","marker":"[40]"},{"why":"establishes scaling behavior for vision transformers that the downstream law extends.","marker":"[47]"}],"fun_headline_variants":["Distillation wins only when pretraining data is scarce","Scaling laws reveal distillation's critical data threshold","Limited pretraining data? Distillation beats fine-tuning","The turning point where distillation loses to fine-tuning","Data-efficient scaling: when to distill vs fine-tune"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that downstream error decreases monotonically as model size grows, so the term $M^{-\\beta}/\\lambda_m$ in the scaling law is a valid decreasing function; the paper's own Appendix D reports cases where error increases with model size, which would break both the fitted law and the crossover proof.","fun_headline_variants_meta":{"raw":{"variants":["Distillation wins only when pretraining data is scarce","Scaling laws reveal distillation's critical data threshold","Limited pretraining data? Distillation beats fine-tuning","The turning point where distillation loses to fine-tuning","Data-efficient scaling: when to distill vs fine-tune"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1351,"prompt_tokens":1004,"completion_tokens":347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":272}},"tokens_in":620,"tokens_out":347,"duration_ms":3301,"temperature":1.0,"reasoning_tokens":272,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:24:57.185507+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit Eq. (1) on a random subset of the data-size/model-size grid and evaluate on held-out sizes for any one of the four datasets; if held-out error swings upward as $M$ increases (as Appendix D shows for some settings) or the residuals do not decay as a power law in $D_p$, $D_f$, and $M$, then the additive law is not the right description and the derived threshold $D^*_p$ lacks predictive content.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"establishes scaling behavior for vision transformers that the downstream law extends."}],"review_version":1}