{"id":"4d516efd-82a1-42b2-9ecd-0cbfc737f653","arxiv_id":"2412.07972","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A time-dilation schedule makes the mode-probability learning phase survive in high dimension, and the learned flow autoencoder recovers the mixture's p and σ² in two separate phases.","lead":"This paper introduces a time-dilated noise schedule for flow-based generative models and analyzes a two-layer autoencoder learning to sample from a two-mode Gaussian mixture. It shows training separates into two phases, one learning mode probabilities and one learning mode variances, and tests a feature-aware time-sampling method on MNIST.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim—that Θ_d(1) samples suffice and that X̂_2 recovers p and σ²—rests on an unproven sample-symmetry ansatz in the replica calculation; without a stability check, Results 1–3 and Corollary 6 are not established.","rationale":"The reader's weakest_assumption is exactly the load-bearing point I find: the paper's learning theorems are conditional on an unproven sample-symmetry ansatz and a heuristic replica saddle-point computation. The exact-flow analysis (Proposition 1) is carefully argued and the time-dilation construction is plausible; the phase separation itself is not the problem. The empirical MNIST section is suggestive but uses a U-Net and SDE training, not the two-layer denoiser and probability-flow ODE of the theory, so it cannot confirm the ansatz. No machine-checked proof or reproducible theory code is provided. I do not see an additional objection that would change the verdict: the proof of Proposition 1 has a likely typo in the ν-equation (Eq. 24 appears to miss a √d factor), but the argument is repaired by working with M_t=μ·X_t/d in Eq. 27, so it does not alter the conditional acceptance. Thus the appropriate verdict remains CONDITIONAL, i.e., UNCHANGED from the reader.","tokens_in":21648,"tokens_out":17410,"duration_ms":178219,"concrete_test":"Compute the local stability of the sample-symmetric saddle point from Appendix B.1/B.2 against sample-asymmetric perturbations (allow qη^μ, qξ^μ, pη^μ to differ by mode or by sample) by evaluating the Hessian of the replicated free energy at the Result 1/2 fixed point for representative parameters (e.g., p=0.8, σ∈{0.5,1,2}, κ∈{1,4,10}, n∈{1,2,4,8}). If any Hessian eigenvalue is positive—or a 1-step replica-symmetry-breaking ansatz yields higher free energy—the symmetric solution is not the relevant minimizer and the central claim fails. If the symmetric point is stable and the 1-RSB free energy is lower, this check would resolve the main objection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"At Appendix B.1, after defining per-sample overlaps qξ^μ, qη^μ, pη^μ, the text states 'we now assume a sample-symmetry ansatz' and sets them equal across μ before evaluating the saddle point. This is not derived, and no stability or symmetry-breaking check is provided. Since n=Θ_d(1) while d→∞, fluctuations between samples do not average out, and for p≠1/2 the minority-mode samples are not exchangeable with majority-mode samples in the loss; the ansatz is exactly the point where the empirical-risk problem is replaced by a tractable symmetric problem. Results 1 and 2, Corollaries 1, 3, 4, 5, and Result 3 all inherit this assumption, and the authors themselves label the derivation 'at the level of rigor of theoretical physics' (Section 4). Result 3's O(1/n) generation bound additionally assumes |θ̂_t−θ_t|=O(1/n) from a saddle-point expansion that is asserted, not proved. If the true minimizer of (9) breaks sample symmetry, the overlap equations do not describe the learned denoiser and Corollary 6 is unsupported. The MNIST experiment uses a different architecture and training procedure and cannot validate the ansatz.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the training of a two-layer denoising autoencoder (eq. 7) used to parameterize a flow-based generative model for a two-mode Gaussian mixture (eq. 1). It proposes a time-dilation schedule τ_t (eq. 12) that stretches the small-time interval in which the mode probability p is learned, so that this phase survives the d→∞ limit. Using a replica/saddle-point calculation (Appendix B), it derives asymptotic equations for the learned network parameters (Results 1 and 2), claims a phase separation into probability learning and variance learning, and states that with Θ_d(1) samples the learned velocity field generates samples with the correct p and σ² (Corollary 6). The paper also proposes a practical method for identifying important training-time intervals and reports preliminary MNIST experiments (Section 6).","tokens_in":21892,"tokens_out":13294,"duration_ms":115986,"significance":"If the central claim (Corollary 6) were rigorously established, the paper would make a valuable contribution: it gives a concrete noise schedule that fixes the vanishing speciation phase identified in prior work and claims a rare Θ_d(1) sample-complexity guarantee for a flow-based generative model. The authors are transparent that the replica/saddle-point derivations are at the level of theoretical physics, and the time-dilation idea is elegant. The MNIST experiment is a useful sanity check of the practical heuristic. However, as detailed below, the central theoretical result rests on an unproven sample-symmetry ansatz and on asserted O(1/n) estimates, so the main claim is currently conditional rather than established.","major_comments":[{"comment":"The saddle-point derivation of Result 1 (and similarly Result 2) rests on the 'sample-symmetry ansatz' introduced after the partition-function integral, which sets the per-sample overlaps q_ξ^μ, q_η^μ, p_η^μ equal across μ. Since n=Θ_d(1) and there is no averaging over μ, and since samples with s^μ=+1 and s^μ=−1 are not exchangeable when p≠1/2, this ansatz is not a harmless mean-field limit; it replaces the empirical-risk problem by a symmetric problem that may have a different minimizer. No stability analysis or symmetry-breaking check is provided, and the MNIST experiment uses a different architecture and training procedure and therefore does not validate the ansatz. Because Results 1 and 2 and all their corollaries, including Corollary 6, inherit this assumption, the central claim of the paper is not established as stated.","section":"Appendix B.1"},{"comment":"Corollary 1 is derived by verifying that ω=κt and b=tanh^{-1}(2p−1) satisfy the saddle-point equations in the n→∞ limit. This verifies that these values form a stationary point, but it does not prove that the global minimizer of the loss in (9) converges to this point. The effective free energy may have multiple saddle points, and the phase-separation statements in Corollary 5 and the conclusions of Result 3 assume that this particular solution is selected. The text should either prove uniqueness or global optimality of this solution within the sample-symmetric class, or explicitly state that the result is conditional on this selection.","section":"Appendix B.1.1"},{"comment":"Result 3 is not proved. The proof begins with the assertion 'From Results 1 and 2 and their Corollaries 1 and 3, we have that |θ̂_t−θ_t|=O_n(1/n) for all overlaps,' which is not established anywhere in the paper; the saddle-point analysis gives asymptotic equations but does not quantify the convergence in n. The subsequent ODE estimates for ϵ_m, ϵ_η, ζ^m, ζ^η, and ζ^ξ are stated as holding 'with high probability' without a Gronwall argument, without a uniform-in-t bound, and without specifying the probability space or the constants involved. Since the O(1/n) generation error is the entire content of the Θ_d(1) sample-complexity claim, this is a load-bearing missing proof.","section":"Appendix C"}],"minor_comments":[{"comment":"The prefactor of f(x,t) in the velocity-field definition appears to be a typo: it should read ˙β_t − ˙α_t β_t/α_t rather than ˙β_t − ˙α_t α_t/β_t, based on the definition b_t(x)=E[ẋ_t|x_t=x] with x_t=α_t x_0 + β_t x_1. Please correct this and check that the subsequent derivations use the corrected factor.","section":"Equations (4), (10), (33)"},{"comment":"The notation for averages of s^μϕ^μ is introduced only implicitly; in the formulas for msetrain and msetest, the quantities written as 'sϕ' and 'ϕ_s' should be defined explicitly to avoid confusion with the overlap variables.","section":"Corollaries 2 and 4"},{"comment":"Lemma 5 is referenced in the proof of Proposition 2 before it is stated; reorder the presentation or add an explicit pointer so that the reader can find the lemma.","section":"Appendix E"},{"comment":"The y-axis label reads 'proportion of0s' with a missing space; please fix the typo.","section":"Figure 2"},{"comment":"The claim that the proposed training schedule 'validates our hypothesis' is stronger than what the experiment shows, since the experiment uses a different architecture (U-Net with VP SDE) and does not test the theoretical assumptions. I suggest softening this phrasing to 'is consistent with' or 'provides preliminary support for'.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"The main theoretical claim is conditional on the sample-symmetry ansatz, which the paper transparently labels as heuristic. If the authors can add a numerical validation of the saddle-point equations (for example, solving the high-dimensional loss directly for moderate d and comparing the measured overlaps to the predicted ones) or a local stability analysis of the sample-symmetric saddle point, the paper would be substantially strengthened. Otherwise, the abstract's claim that Θ_d(1) samples suffice should be presented as a conjecture. The MNIST application is interesting but preliminary and does not bear on the ansatz. The paper is likely within the scope of cs.LG, though the technical audience may be more in the statistical-physics community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe paper is worth a read: the time-dilation schedule in eq. (12) is a genuinely new idea that makes the mode-probability phase survive the d→∞ limit for unbalanced Gaussian mixtures, and the two-phase description of the learned autoencoder is clean and plausible. The authors are also honest about what is rigorous and what is not. Proposition 1 is proved; the learning results are explicitly labeled 'theoretical physics' level. The rigorous parts—the exact-flow analysis and the infinite-data verification—hold up.\n\nThe soft spot, as your stress-test note says, is the sample-symmetry ansatz in Appendix B. The replica calculation assumes all per-sample overlaps are equal before evaluating the saddle point, and no stability check is given. Since n = Θ_d(1), fluctuations between samples do not average out, and for p ≠ 1/2 the minority-mode samples are structurally different in the loss. If the true minimizer breaks symmetry, the phase-separation picture and the O(1/n) generation bound do not follow. The paper flags its own heuristic level, but that does not make the central claim solid. Result 3 is also a sketch: the Gronwall-style estimates are asserted rather than proved. The MNIST experiment is preliminary and uses a U-Net, so it cannot validate the replica ansatz. That said, the empirical heuristic—sampling training times according to feature-relevant intervals—is reasonable, and the numbers move in the right direction (88% to 81%, closer to the true 80%).\n\nOn balance, the conditional verdict is right. The mode-probability phase disappearing is a real known issue; the time dilation is a simple, clever fix; and the two-phase simplification of the network parameters will be cited. But the learning guarantee is only as strong as an unproven ansatz. A serious referee should ask for a stability analysis of the saddle-point solution, or at least a clear statement in the abstract of what is conjectural.\n\nI would send it to peer review rather than desk-reject it. With a referee who can check the replica step, it could be a solid paper after revision.","headline":"A clever time-dilation schedule with an honest but unproven replica core; the learning guarantee rests on a sample-symmetry ansatz, yet the paper deserves referee time.","tokens_in":22478,"tokens_out":3354,"would_cite":true,"duration_ms":31361,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A time-dilated interpolant makes the mode-probability phase survive the high-dimensional limit, so a single two-layer denoiser learns both p and σ² with Θ_d(1) samples.","keywords":["flow-based generative models","Gaussian mixture","phase transitions","time schedule","sample complexity","denoising autoencoder","diffusion models","feature emergence"],"falsifier":"Train the same two-layer denoiser on an unbalanced two-mode Gaussian mixture (say p = 0.8, d = 5000, n = 128) with the dilated schedule, initialize the network asymmetrically, and measure the empirical distribution of per-sample overlaps and of $\\mu\\cdot\\hat{X}_2/d$; if the overlaps vary across samples in a way the sample-symmetry ansatz forbids, or if the generated mode proportion stays systematically away from p unless $\\kappa$ is sent to infinity, the central claim collapses.","tokens_in":21305,"feed_emoji":"⏱️","tokens_out":8750,"duration_ms":82962,"temperature":0.7,"pith_summary":"Flow-based generative models trained on high-dimensional Gaussian mixtures can fail to learn the relative probability of the modes: the time window in which that probability is decided shrinks like 1/√d and vanishes as d grows. This paper introduces a time dilation that stretches that window to constant duration, and proves that a simple two-layer denoiser then learns the velocity field in two separate phases. In the first phase it estimates only the mode probability p; in the second it estimates only the within-mode variance σ². The authors show that Θ_d(1) samples are enough and that the generated distribution matches p and σ² in the appropriate limits. They also turn the phase structure into a training-schedule heuristic for real data, with preliminary MNIST evidence.","feed_headline":"Time dilation rescues mode-probability learning in diffusion models","feed_subtitle":"A two-layer denoiser with a dilated schedule recovers p and σ² from Θ_d(1) samples.","key_machinery":"The load-bearing object is the time-dilated interpolant together with the overlap variables that track the learned network inside the high-dimensional geometry. The interpolant $x_t = \\alpha_t x_0 + \\beta_t x_1$ with $\\alpha_t = 1-\\tau_t$, $\\beta_t = \\tau_t$, and $\\tau_t = \\kappa t/\\sqrt{d}$ on $[0,1]$ slows time near the speciation window $t \\approx 1/\\sqrt{d}$; this is what keeps the p-learning phase at $O(1)$ duration as $d\\to\\infty$. The analysis then follows the scalar projections of the learned readout and weight vectors onto the mean direction $\\mu$, the per-sample noise directions $z^\\mu$, and the initial noise $x_0^\\mu$: $m = \\mu\\cdot u/d$, $\\omega = \\mu\\cdot w/d$, $q_\\eta^\\mu$, $p_\\eta^\\mu$, $q_\\xi^\\mu$, and so on. Saddle-point equations for these overlaps give closed-form asymptotics in $d\\to\\infty$ and then $n\\to\\infty$, showing the first phase is governed by $\\tanh(b) = 2(p-1/2)$, $m=1$, $\\omega=\\kappa t$, and the second by $c = \\tau\\sigma^2/(1+(\\sigma^2-1)\\tau^2)$, $m=1-c\\tau$. These overlap equations are what connect the learned network parameters to the generated distribution.","core_discovery":"The central claim is that the right noise schedule fixes the failure mode identified in prior asymptotic analyses: with $\\alpha_t = 1 - \\tau_t$ and $\\beta_t = \\tau_t$, where $\\tau_t$ equals $\\kappa t/\\sqrt{d}$ for $t \\in [0,1]$ and then climbs linearly to 1 on $[1,2]$, the learned flow satisfies $\\lim_{\\kappa\\to\\infty} \\lim_{n\\to\\infty} \\lim_{d\\to\\infty} \\mu\\cdot\\hat{X}_2/d \\sim p\\delta_1 + (1-p)\\delta_{-1}$ and $\\lim_{n\\to\\infty} \\lim_{d\\to\\infty} w\\cdot\\hat{X}_2/\\sqrt{d} \\sim N(0,\\sigma^2)$ for $w \\perp \\mu$. Thus the distribution generated by the learned denoiser captures both parameters of the two-mode Gaussian mixture. Along the way the paper characterizes the minimizer of the denoising loss in the $d\\to\\infty$ limit: for $t \\in [0,1]$ the overlaps depend only on p, and for $t \\in [1,2]$ only on $\\sigma^2$, so the network simplifies by ignoring the parameter not relevant to the current phase. This is the sense in which the time-dilated schedule turns diffusion into a staged estimator rather than a single monolithic denoiser.","pith_inferences":["The paper does not claim that the MSE-jump heuristic is guaranteed for arbitrary data; a direct test would be to compute the validation MSE across training times on real datasets and compare the time of its largest jump with the U-Turn class-decision interval.","Because the theory rests on the sample-symmetry ansatz, an immediate stress test is to check numerically whether per-sample overlaps spread out when the network is trained from random initialization on finite n; if they do, the clean phase separation may not survive.","For multi-modal data, the general dilation formula in the appendix suggests a concrete recipe—allocate training time in proportion to the reciprocals of the mode-centre lengths—which could be tested on mixtures with known mode distances.","The staged-learning picture suggests that curriculum over noise levels, not just reweighting of training times, could be beneficial; this is a natural extension the paper does not test."],"forward_implications":["For the two-mode Gaussian mixture, $\\Theta_d(1)$ data samples suffice to learn the velocity field, and the generated samples recover both the mode probability p and the variance $\\sigma^2$ in the stated limits.","The learned denoiser decomposes the estimation problem across time: p first, $\\sigma^2$ second, so the network only needs to represent the parameter relevant to the current phase.","The test MSE has a jump at the phase transition; dilating time near the jump removes the discontinuity, giving a data-driven signature for locating phase transitions.","For a given feature, training more often in the time interval where that feature's class is decided improves accuracy on that feature; MNIST experiments with 20% ones and 80% zeros move generated proportions from 88.2% toward the true 80%.","The time-dilation formula extends to Gaussian mixtures with m modes, producing m+1 phases: one per mode probability and a final variance phase."],"supporting_citations":[{"why":"Supplies the stochastic-interpolant framework: the velocity field as a conditional expectation, the denoising loss, and the exact denoiser formula for Gaussian mixtures.","marker":"Albergo et al. (2023)"},{"why":"Establishes the previous learning analysis for balanced two-mode Gaussian mixtures and the partition-function and overlap method that this paper extends to unbalanced p.","marker":"Cui et al. (2024)"},{"why":"Shows the speciation time is $O(1/\\sqrt{d})$ and that the probability-learning phase disappears without time dilation, which motivates the dilation.","marker":"Biroli et al. (2024)"},{"why":"Identifies the same phase-transition problem when learning unbalanced Gaussian mixtures; the paper contrasts its single-network solution with that work's separate-networks approach.","marker":"Montanari (2023)"},{"why":"Provides the U-Turn method used to find the time interval where a feature's class is decided in the MNIST experiment.","marker":"Sclocchi et al. (2024)"},{"why":"Defines the Variance Preserving SDE used in the real-data experiments.","marker":"Song et al. (2021)"},{"why":"Supplies the U-Net architecture used to parameterize the MNIST denoiser.","marker":"Ronneberger et al. (2015)"}],"fun_headline_variants":["Time-dilated schedule turns diffusion into staged learning","Diffusion schedule fix: network ignores irrelevant parameter per phase","Phase-aware time schedule rescues mode learning in flow models","Two-phase schedule simplifies diffusion: first p, then sigma","New training schedule makes flow models estimate modes in phases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The high-dimensional analysis assumes all training samples contribute identical overlaps in the saddle-point calculation and treats that saddle-point computation as exact; if the true minimizer breaks this symmetry, the phase separation and the $\\Theta_d(1)$ sample guarantee do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Time-dilated schedule turns diffusion into staged learning","Diffusion schedule fix: network ignores irrelevant parameter per phase","Phase-aware time schedule rescues mode learning in flow models","Two-phase schedule simplifies diffusion: first p, then sigma","New training schedule makes flow models estimate modes in phases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1745,"prompt_tokens":980,"completion_tokens":765,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":687}},"tokens_in":596,"tokens_out":765,"duration_ms":8191,"temperature":1.0,"reasoning_tokens":687,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:21:27.793701+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same two-layer denoiser on an unbalanced two-mode Gaussian mixture (say p = 0.8, d = 5000, n = 128) with the dilated schedule, initialize the network asymmetrically, and measure the empirical distribution of per-sample overlaps and of $\\mu\\cdot\\hat{X}_2/d$; if the overlaps vary across samples in a way the sample-symmetry ansatz forbids, or if the generated mode proportion stays systematically away from p unless $\\kappa$ is sent to infinity, the central claim collapses.","supporting_citations":[],"review_version":1}