{"id":"d6ea3ec4-162d-4a61-8f70-de791d3c4089","arxiv_id":"2608.04448","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Downscaled training gradients are close to native only inside specific noise windows and on mild routes; the mismatch splits into a ratio-governed part and an absolute-size floor that persists at every noise level.","lead":"Training diffusion models on downscaled images usually changes the learning signal, and this paper measures exactly when it does not. It finds a noise-dependent part of the gap that follows the downscale ratio and a fixed size-dependent floor, then uses the safe windows to cut LoRA training time by about 15%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Route-uniformity assumption underpinning Eq. 8 is contradicted by the paper's own Table 11; held-out predictions need re-check with route-specific mismatch curves.","rationale":"The reader's weakest_assumption identifies assumption (iv) and notes the contradiction with Table 11; I agree that this is the most load-bearing place in the argument. The two-term reduction is the paper's central theoretical contribution, and its route-transferability is what separates it from the spectral account. The paper's own Table 11 gives route spreads larger than the claimed ±0.02, and the claim is used verbatim in §4.3 to support the factorization and the statement that 'all route identity lives in J.' A failure of route-uniformity would not collapse the empirical map or the debiasing contribution, which are real, but it would degrade the headline predictive claim from a parameter-light transfer to a per-route fitted curve. The concrete test directly asks whether the held-out predictions survive using route-specific mismatch curves, which settles whether the concern actually lands. I also considered the acknowledged failure of the two-term reduction on the 768 window and the underdetermined floor law; both are real but are explicitly disclosed and instrumented by the authors, whereas the route-uniformity claim is asserted as established fact and contradicted by the table cited in support. The verdict should remain CONDITIONAL: the empirical instrumentation and the map are valuable, but the predictive account needs re-stating and re-testing on this assumption.","tokens_in":32153,"tokens_out":15183,"duration_ms":131691,"concrete_test":"Recompute the Table 4 held-out predictions for 1024→768 and 1280→1024 using each held-out route's own measured ||Δr(σ)|| curve (from Table 11, or a fresh matched-D measurement) instead of the shared curve, while keeping the ratio-governor amplitudes and floors fixed. If the held-out RMSE stays below the pre-registered gates (≈0.09 for 1024→768 and ≈0.05 for 1280→1024), route-uniformity is not the bottleneck. If RMSE degrades by more than about 0.03 or a route falls below the oracle quadratic, the shared-curve assumption is load-bearing and the predictive claim must be restated with a route-dependent data term.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central predictive step of Eq. 8 is assumption (iv): ||B⊥_e|| ≈ a_e ||Δr(σ)|| with a single route-independent mismatch curve. Section 4.3 asserts this curve is 'the same for every route within ±0.02' and cites Appendix Table 11. The table does not support that claim: across the four tabulated routes the relative-L2 excess spreads by 0.037 at σ=0.125 (0.860–0.897) and by 0.075 at σ=1.0 (0.360–0.435), with spreads near 0.05–0.06 at several intermediate points. The held-out predictions of Table 4 for 1024→768 and 1280→1024 are produced by feeding this one shared curve into the two-term reduction, scaled only by a per-route amplitude and a floor. If the mismatch curve is actually route-dependent at the level Table 11 shows, then the route identity is not confined to a_e and the floor; the data term's shape is partly route-specific, and the held-out RMSE of 0.07–0.09 could be an artifact of an effective amplitude absorbing the route offset rather than a confirmation of the decomposition. This is load-bearing because the paper's claimed superiority over the spectral account rests on the route-transferability of the two-term form, and the training-time window is read from the same shared-curve account.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether a downscaled, re-encoded latent used during training preserves the native adapter gradient in a flow-matching DiT. It defines a population demotion distance, derives an exact angular identity (Eq. 6) and a four-term perturbation expansion (Eq. 7), reduces the excess gap to a two-term form (Eq. 8) under four assumptions, and proposes a debiased estimator for finite-draw attenuation. The measured per-route/per-sigma gap map is used to score a spectral account against a gradient-perturbation account, to validate the two-term reduction on two held-out routes, and to construct a sigma-gated LoRA trainer that reports 14.6% training-time savings at a fixed step budget.","tokens_in":32403,"tokens_out":8478,"duration_ms":68442,"significance":"The paper is unusually careful experimentally: it uses pre-registered gates, split-half reliability checks, deterministic kernels, paired draw sets, and a controlled two-adapter replication. The exact angular link (Eq. 6) and the branch decomposition (Eq. 5) are clean and self-contained, and the attempt to separate ratio-governed data effects from token-count-governed graph effects is valuable. If the route-uniformity assumption were supported, the held-out predictive RMSEs of 0.07–0.09 would be a strong result. However, the current evidence for route-uniformity conflicts with the paper's own Table 11, and the 1024→768 window is an acknowledged structural counterexample to Eq. 8, so the central claims need substantive revision before the paper can be accepted.","major_comments":[{"comment":"The claim that the residual mismatch curve is 'the same for every route within ±0.02' is contradicted by the table cited in its support. Table 11 reports relative-L2 cross-grid excess values with spreads of 0.037 at σ=0.125 (0.860–0.897), 0.064 at σ=0.375 (0.808–0.872), and 0.075 at σ=1.0 (0.360–0.435); at most sigma bins the spread exceeds 0.05. Because Eq. 8 consumes one shared ∥Δr(σ)∥ curve and scales it by a per-route amplitude a_e, a route-dependent mismatch shape can be absorbed by a_e and by the fitted floor, so the held-out RMSE values in Table 4 do not by themselves confirm the factorization. The held-out protocol should be re-run with route-specific mismatch curves, or the paper should provide a statistical test that the amplitude-times-shared-curve model is consistent with the residual-probe data within its stated precision.","section":"§4.3, Appendix Table 11, Assumption (iv), Eq. (8)"},{"comment":"The 1024→768 mid-σ window is a structural counterexample to the two-term reduction, not merely a numerical miss. Eq. 8 is the sum of a nonnegative data term and a nonnegative floor, so it cannot produce a gap below the route's own floor; Table 6 shows paired excess values around +0.02 to +0.04 at σ=0.700–0.938 while the 768 floor is 0.06–0.09, and Appendix C confirms that 'a positive two-term reduction cannot produce a gap under its own floor.' Since the abstract and introduction present the reduction as the paper's central result, and §5 stacks this exact 768 window as a second training route, the central claim must be qualified to the domain in which assumptions (i)–(iv) hold, or the expansion must be extended to include the interference term. A limitation note is not sufficient when the flagship route is inside the reduction's failure region.","section":"Eq. (8), §4.2, Appendix C"},{"comment":"The 'safe' verdicts for the 1024→896 and 1024→768 windows clear only the enlarged margin 0.09, not the strict margin ε=0.02 fixed at the outset of §4.1. The abstract says the downscaled gradient 'stays within a small margin of the native one' and §5 calls 1024→896 the 'measured-safe corpus route' without stating the margin; given the title's 'same gradients' framing, this is a material overstatement. The abstract and the safety claims should state the margin explicitly, and the practical training claims in §5 should be tied to the margin at which the route is actually certified.","section":"§4.1–§4.2, abstract"}],"minor_comments":[{"comment":"The text refers to route-uniformity across 'the three main routes' but the table contains four routes; clarify which routes enter the held-out protocol and which are simply descriptive, since 1280→1024 appears both in Table 11 and as a held-out route in Table 4.","section":"§4.3, Table 11"},{"comment":"The pseudocode gates per-sample via 'if σ>σ* and y on-route', but the prose says 'When every sample in the batch is above the gate' the latent is replaced; specify whether mixed batches substitute per sample or are left entirely native, since this affects the reported wall-clock saving.","section":"§5, Table 2"},{"comment":"The endpoint bin (σ=1) is plotted with open markers and later excluded from the safe interval σ∈(0.5,0.94], but this exclusion is stated only in the text; add the endpoint status and its margin explicitly to the figure caption.","section":"§4.2, Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious experimental contribution with a strong measurement apparatus, but the route-uniformity overclaim relative to Table 11 and the 768-window structural failure of Eq. 8 are load-bearing. The authors can likely fix both with additional analysis and honest re-scoping, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know before you read this paper: the central route-uniformity claim in §4.3 is contradicted by the paper's own Table 11, and the held-out predictions in Table 4 rest on that claim. The paper is still worth reading — the debiased estimator and the measured safety map are genuinely new — but the headline two-term reduction is not as well supported as the abstract implies.\n\nWhat is actually new: this is the first gradient-level treatment of training on downscaled images for diffusion transformers. The estimand is carefully defined, and the paper builds a debiased cosine estimator using Spearman's attenuation correction, which addresses a real finite-draw artifact. The interventional data/graph split via a repromote arm is a nice piece of experiment design. The paper is also unusually explicit about margins, verdict resolution, and pre-registered criteria, and it ships code.\n\nSoft spots, in order. First, §4.3 asserts the residual mismatch curve is 'the same for every route within ±0.02' and cites Table 11. The table actually shows spreads of about 0.037 at σ=0.125 and 0.075 at σ=1.0 across the four routes. That is not within ±0.02. Since the held-out predictions are produced by feeding one shared curve into Eq. 8, scaled only by per-route amplitude and floor, the route-transfer claim is stronger than the evidence. The held-out RMSE of 0.07–0.09 could in part reflect the fitted amplitude absorbing route-specific offsets rather than a confirmation of the decomposition. This is the main fix: re-run the predictive protocol with route-specific mismatch curves, or re-state the claim with error bars reflecting the observed spread.\n\nSecond, the two-term reduction does not predict the 768 mid-σ window, where the measured gap dips below the floor. The paper acknowledges this and traces it to negative interference, but it means the reduction's domain is narrower than the abstract suggests. Third, the floor 'law' is a two-anchor exponential fit, and the paper's own identifiability analysis shows other functional forms fit equally well. The limitations section says this, but the conclusion leans on it. Fourth, the training evidence is 480-step LoRA runs on two small corpora, so the 14.6% saving is suggestive, not generalizable. The single-model limitation is acknowledged.\n\nOverall: the paper deserves a serious referee. It is well-written, honest about many limits, and the empirical map and estimator are real contributions. But the route-uniformity claim is load-bearing and contradicted by the paper's own numbers. I would accept it for peer review with major revision, requiring the route-specific re-analysis and re-scaled claims.","headline":"Useful empirical study, but the route-uniformity claim at the heart of the two-term reduction conflicts with the paper's own Table 11.","tokens_in":32961,"tokens_out":4624,"would_cite":false,"duration_ms":38083,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that when training a diffusion transformer on downscaled latents, the gradient gap decomposes into a noise-dependent ratio term and a token-count floor, and that a measured map licenses selective low-resolution LoRA…","keywords":["diffusion transformers","flow matching","downscaled training","gradient similarity","LoRA","token-count floor","noise-conditional training","cosine attenuation correction"],"falsifier":"Measure the cross-grid mean prediction-residual mismatch $\\|\\Delta \\bar{r}(\\sigma)\\|$ at $\\sigma = 1$ on the 1280 to 1120 iso-severity route with a draw count high enough to resolve 0.01, and compare it with the four-route family of Appendix Table 11; a deviation larger than the claimed $\\pm 0.02$ band would jointly falsify assumption (iv) and the exponential-in-target-tokens floor law.","tokens_in":31891,"feed_emoji":"🖼️","tokens_out":9026,"duration_ms":77224,"temperature":0.7,"pith_summary":"Training a diffusion transformer on downscaled latents is cheap only if the downscaled step still produces the same gradient direction as the native step. The paper claims this question reduces to two separable effects: a noise-dependent term governed by the downscale ratio, which decays as the noise level rises, and a $\\sigma$-independent floor set by the absolute token count of the target grid, which no amount of noise removes. With a debiased cosine estimator that removes finite-sample attenuation, the measured $(\\text{route}, \\sigma)$ map confirms this and finds a narrow window on the 1024 to 768 route that no spectral criterion predicts. Restricting LoRA training to the validated windows cuts training time by 14.6% at fixed steps while keeping the endpoint near the native retrain in weight space.","feed_headline":"Downscaled training matches native gradients only inside noise windows","feed_subtitle":"A ratio term and a token-count floor explain the safe 1024 to 896 strip; gated LoRA training saves 14.6% at fixed steps.","key_machinery":"The load-bearing object is the demotion distance $d_{e_0 \\to e}(\\sigma) = 1 - \\cos(\\bar{g}_{\\mathrm{src}}(\\sigma), \\bar{g}_{\\mathrm{dem}}(\\sigma))$ between population mean adapter gradients, and the re-encoding excess $\\mathrm{gap}_e(\\sigma) = d_e(\\sigma) - d_{\\mathrm{reenc}}(\\sigma)$ that subtracts the VAE round trip. The argument decomposes the perturbation $\\delta g$ of the gradient $g = J^\\top r$ into a data branch $\\bar{J}^\\top \\Delta \\bar{r}(\\sigma)$, a graph branch $\\Delta J^\\top \\bar{r}$, and a remainder, then reduces the cosine distance to a rotation-to-scale ratio $\\kappa_{\\mathrm{eff}}$ whose small-perturbation expansion yields a four-term form; assumptions (i) through (iv), including the route-uniformity of the residual mismatch, collapse it to the two-term reduction. The second key piece is the estimator: finite-draw cosine attenuation inflates distances more on coarser grids and manufactures the exact signature of a floor, so the paper applies Spearman's correction for attenuation via per-arm self-cosines and draw-limit extrapolation before any verdict is read.","core_discovery":"The paper's central claim is that a \"demoted\" training step, in which the native latent is replaced by a downscaled, re-encoded one, perturbs the adapter gradient through two structurally different channels. The first is a data branch set by the downscale ratio: its contribution to the gap is proportional to $\\|\\Delta \\bar{r}(\\sigma)\\|^2 / \\|\\bar{g}_{\\mathrm{src}}(\\sigma)\\|^2$ and decays as the noise level rises, just as the spectral premise predicts. The second is a graph branch that is independent of $\\sigma$: it is set by the absolute token count of the target grid, carried partly by the rotary position embedding and partly by the coarse graph's residual capacity, and it cannot be removed by any noise level. The measured map shows 1024 to 896 is safe for $\\sigma \\in (0.5, 0.94]$, 1024 to 768 has a narrow low-excess window $0.65 < \\sigma < 0.95$ that arises from negative data-graph interference, and any route to a 512-token grid is unsafe at every noise level. On this basis the paper claims a selective LoRA trainer saves 14.6% wall-clock at fixed steps while staying near-native in weight space.","pith_inferences":["If the token-count floor is as universal as the paper's law suggests, even training schemes that restage the objective per resolution level carry an irreducible gradient-mismatch floor at coarse grids; output-level success in those schemes may come from the data term alone.","The debiasing result generalizes beyond this model: any comparison of gradients computed on different grid sizes or different draw budgets can manufacture a spurious floor, so cosine-attenuation correction should be standard in gradient-similarity audits.","A natural extension is to predict the map from a closed-form posterior covariance model rather than measuring the empirical mismatch curve; the paper shows second-order data-only closures recover the shape but not the amplitude, so the missing ingredient is likely caption-conditioned, non-Gaussian posterior structure.","The late-scheduled result suggests that scheduling demotion to late training is the most weight-preserving mode; testing whether the safe windows shift during the training trajectory is a direct next experiment."],"forward_implications":["On the paper's model, the spectral criterion \"safe once noise masks the frequencies the coarse grid loses\" is not a sufficient test for training; any low-resolution training proposal should be validated at the gradient level with a debiased estimator before being trusted.","The floor law implies that routes sharing the same target token count share the same $\\sigma$-independent gap regardless of downscale ratio, so the cost of demotion is set by where you land, not by how far you shrink.","Gated demotion on the validated 1024 to 896 route at $\\sigma > 0.5$ yields a 14.6% wall-clock saving at fixed steps with a weight-space endpoint within kernel-noise reach of native; late-scheduled demotion improves the endpoint cosine to about 0.75 with proportionally smaller savings.","The two-term reduction is not universal: the 1024 to 768 low-excess window is a signed failure of the reduction, caused by negative interference between data and graph branches, which the paper measures rather than explains away.","Held-out routes are predicted at about 0.07 to 0.09 RMSE from the ratio governor and an exponential-in-target-tokens floor law, compared with 0.147 to 0.355 RMSE for the spectral family."],"supporting_citations":[{"why":"Supplies the tolerance-parameterized spectral progressive-diffusion account that the paper restates as the spectral baseline and then falsifies at the gradient level.","marker":"Xiao et al., 2026"},{"why":"Provides the correction for attenuation used to remove finite-draw cosine bias that would otherwise mimic the token-count floor.","marker":"Spearman, 1904"},{"why":"Defines the DiT backbone whose quadratic token count and attention cost make resolution the expensive axis.","marker":"Peebles & Xie, 2023"},{"why":"Defines the LoRA adapter class whose gradients are the estimand in all probes and in the selective trainer.","marker":"Hu et al., 2022"},{"why":"Supplies the rotary position embedding whose grid-dependent phase geometry is one component of the sigma-independent floor.","marker":"Su et al., 2024"},{"why":"Anchors the progressive-growing lineage of resolution schedules whose stage-specific objectives the paper contrasts with its unchanged-objective demotion.","marker":"Karras et al., 2018"},{"why":"Exemplifies training-time low-resolution usage where the objective is restaged per pyramid level rather than left native.","marker":"Jin et al., 2024"},{"why":"Supplies the sigma-gated banded rotary schedule used as the optional training-time refinement.","marker":"Zhao et al., 2026"}],"fun_headline_variants":["Downscaled gradient parity: ratio decay plus token floor","Noise windows and token floor decide downscale gradient parity","Gated downscale LoRA saves 14.6% at fixed steps","Safe downscale training only inside measured noise windows","Gradient gap decomposes into ratio decay and token floor"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole two-term reduction leans on assumption (iv): the cross-grid prediction-residual mismatch has one universal $\\sigma$-shape for every route, with route identity entering only as a scalar amplitude; the paper's own Table 11 shows spread up to about 0.075 at $\\sigma = 1$, wider than the claimed $\\pm 0.02$.","fun_headline_variants_meta":{"raw":{"variants":["Downscaled gradient parity: ratio decay plus token floor","Noise windows and token floor decide downscale gradient parity","Gated downscale LoRA saves 14.6% at fixed steps","Safe downscale training only inside measured noise windows","Gradient gap decomposes into ratio decay and token floor"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000953,"raw_usage":{"total_tokens":4122,"prompt_tokens":1058,"completion_tokens":3064,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":2981}},"tokens_in":674,"tokens_out":3064,"duration_ms":19927,"temperature":1.0,"reasoning_tokens":2981,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:40:04.364509+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the cross-grid mean prediction-residual mismatch $\\|\\Delta \\bar{r}(\\sigma)\\|$ at $\\sigma = 1$ on the 1280 to 1120 iso-severity route with a draw count high enough to resolve 0.01, and compare it with the four-route family of Appendix Table 11; a deviation larger than the claimed $\\pm 0.02$ band would jointly falsify assumption (iv) and the exponential-in-target-tokens floor law.","supporting_citations":[],"review_version":2}