{"id":"2732ff7c-efa1-4919-a34e-822ed856129c","arxiv_id":"2602.07928","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A per-sample kinetic-energy score along flow-matching trajectories tracks semantic quality and data rarity, with an extreme-energy regime that predicts memorization, and a two-phase inference-time shaping that improves generation.","lead":"This paper proposes KPE, a per-sample 'kinetic energy' score equal to the time integral of squared velocity along a generative ODE trajectory, and shows higher KPE correlates with sharper semantics and rarer samples while extreme KPE predicts memorization. It uses this Goldilocks picture to build a training-free inference-time velocity shaping method that improves FID and reduces memorization.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (11) does not follow from Theorem 4.2: the additive constants in Theorem B.1 diverge non-integrable near t=1, so the claimed Θ(∫−log p̂_t dt)+O(1) integrated energy–density relation is unproven.","rationale":"The reader's verdict is CONDITIONAL, and my analysis does not move it to a different category: the paper has a coherent empirical story and a substantial appendix, but the strongest theoretical guarantee advertised in the abstract and §4.3 is not currently proven. The reader's weakest_assumption focused on posterior dominance near t=0 and in overlapping regions; my concern is distinct and more fundamental: even when posterior dominance holds at every time (as in the d=1, N=1 test), the additive constants in Theorem B.1 are non-integrable near t=1, so the integrated bound (11) does not follow. This is a mathematical gap in the paper, not a disagreement with the empirical correlations. The proposed check is a one-line analytic computation that settles whether the gap is real. If the authors can provide a revised integrated bound with finite additive remainder, or explicitly restrict the theorem to a trajectory class where the constants integrate, the contribution would be materially stronger. Until then, CONDITIONAL remains appropriate: the empirical findings and KTS may stand on independent evidence, but the headline theoretical claim needs repair.","tokens_in":33182,"tokens_out":7829,"duration_ms":87536,"concrete_test":"Analytic check: instantiate the appendix with d=1, N=1, x^(1)=1, γ(t)=t, and the trajectory x(t)=t (the exact ODE solution of the closed-form empirical field). Compute actual KPE = 1/2 and ∫_0^1 −log p̂_t(x(t))dt finite. Then compute C′_t as in Theorem B.1/Lemma B.5: m(t)=−1/(1−t), b_t=1/(1−t), F_t=1/(1−t), so C′_t≥2.5/(1−t)². Verify ∫_0^1 C′_t dt diverges. If it does, Eq. (11) cannot be derived from Eq. (54); the authors would need a tighter trajectory-dependent bound or an additional condition (e.g., uniform integrability of A_t(z(t))) to salvage the integrated Θ relation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theoretical claim is Eq. (11): E(z0→1) = Θ(∫_0^1 −log p̂_t(z(t))dt) + O(1), asserted to follow by integrating the pointwise bound (10) of Theorem 4.2/Theorem B.1. It does not follow as written. In Appendix B.5, for the linear bridge γ(t)=t, c1(t)=1/2, c2(t)=12, but the additive constant C′_t defined in Theorem B.1 and Lemma B.5 has C′_t ≥ const·∥x^(i*)∥²/(1−t)², coming from terms such as m(t)²∥μ_{i*}(t)∥² and ∥b_t∥² with m(t)=−1/(1−t), μ_{i*}(t)=t x^(i*), b_t=x^(i*)/(1−t). Hence ∫_0^1 C′_t dt = ∞. Substituting into the integrated inequalities (54) gives only E ≤ +∞ and E ≥ −∞; the Θ relation and 'O(1)' remainder are not consequences. This is not a technicality about loose constants: take d=1, N=1, x^(1)=1, γ(t)=t, and the exact trajectory x(t)=t. Then posterior dominance holds with λ=1 for all t, KPE = 1/2, and −log p̂_t(z(t)) = −log(1−t)+O(1), whose integral is finite, but the theorem's C′_t ~ 2.5/(1−t)². Thus Eq. (11) fails in the ideal posterior-dominated case where all assumptions of the lemma are met. The paper's theoretical support for the KPE–density connection is therefore not established; posterior dominance was already a limitation, but this integrability gap is independent and more decisive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Kinetic Path Energy (KPE), defined as the time integral of squared velocity along a flow-matching ODE trajectory, and proposes it as a per-sample diagnostic. Empirically, the authors report that high-KPE trajectories correlate with higher CLIP-based semantic fidelity on ImageNet and with lower estimated training-data density on synthetic datasets, CIFAR-10, and ImageNet. The theoretical section analyzes empirical flow matching (EFM), gives a pointwise energy–density bound under a posterior-dominance assumption (Theorem 4.2 / Theorem B.1), and claims an integrated form E = Θ(∫ −log p̂_t dt) + O(1) (Eq. 11). The paper further identifies a terminal 1/(1−t) singularity in closed-form EFM as a memorization mechanism and proposes Kinetic Trajectory Shaping (KTS), a training-free inference-time velocity rescaling that boosts early motion and damps late motion. Experiments on CelebA and ImageNet show tunable FID/memorization trade-offs.","tokens_in":33692,"tokens_out":7199,"duration_ms":66090,"significance":"KPE is a simple, zero-overhead path-level diagnostic, and the Goldilocks picture — moderate, well-timed kinetic effort improves quality while excessive terminal energy causes memorization — is intuitively appealing and supported by the closed-form EFM analysis. The pointwise energy–density theorem has explicit constants and is internally coherent under its posterior-dominance assumption, and the KTS experiments are a useful training-free contribution. However, the integrated energy–density relation stated as Eq. (11) is not established by the supplied proof, and the real-data density evidence is admittedly representation-dependent. If Eq. (11) can be proved with integrable constants, the theoretical claim would be significant; as written, the central theoretical guarantee is overclaimed.","major_comments":[{"comment":"Eq. (11) does not follow from integrating Eq. (10). For the linear bridge γ(t)=t, Lemma B.5 gives C′_t ≥ C−(t), with C−(t) = m(t)²∥μ_{i*}(t)∥²/2 + 2F_t² and F_t ≥ ∥b_t∥ = ∥x^(i*)/(1−t)∥. Hence C′_t ≥ const/(1−t)², so ∫₀¹ C′_t dt = ∞ and the integrated bound (54) is vacuous: E ≤ +∞ and E ≥ −∞. The claimed O(1) remainder is not a consequence. A concrete check: d=1, N=1, x^(1)=1, trajectory x(t)=t satisfies posterior dominance with λ=1 for every t, yet KPE = 1/2 and ∫₀¹ −log p̂_t(z(t)) dt is finite, while the lemma's C′_t is Θ((1−t)⁻²). Thus the current proof cannot establish Eq. (11) even in an ideal posterior-dominated case. This requires either a genuinely integrable additive-constant argument or a reformulation of the density claim as pointwise and conditional only.","section":"§4.3, Remark 4.3; Appendix B.5, Eq. (54)"},{"comment":"Even the pointwise theorem is conditional on λ_{i*}(z,t) ≥ 1−ε at the evaluated (z,t). No argument is provided that sampled trajectories satisfy this condition on the whole interval [0,1]; near t=0 and in overlap regions the responsibilities are comparable, so the bound (10) may fail exactly where the integrated density integral is path-dependent. The empirical density validation does not fill this gap, because it uses k-NN/KDE in a 2D PCA descriptor space, which the authors themselves call a representation-dependent proxy (§4.2 Limitation). The theoretical claim should be explicitly restricted to the posterior-dominance regime, and the real-data claim should be phrased at the same level of generality as the proxy.","section":"§4.3, Theorem 4.2 (posterior dominance assumption)"},{"comment":"The real-data validation of the KPE–density correspondence is carried out in a 2D PCA descriptor space (RGB statistics, Gabor responses, edge density) rather than in pixel space. The authors acknowledge this limitation, but the abstract and Finding 2 state the result as 'low-density manifold frontiers.' This is stronger than the evidence supports. The revised manuscript should carefully separate the descriptor-space proxy result from the geometric statement about the data manifold, or provide validation that the proxy tracks manifold density.","section":"§4.2, Figs. 4–5"}],"minor_comments":[{"comment":"The text says '22D descriptors'; this should read '2D descriptors.'","section":"§4.2"},{"comment":"The x-axis labels show CFG = 1.0, 1.5, 2.0, while the caption and Table 1 refer to CFG 1.0/1.5/4.0. Please reconcile.","section":"Figure 2"},{"comment":"The heading 'Lemma 2' should refer to Lemma B.3 to match the numbering scheme.","section":"Appendix B.4"},{"comment":"Line 10 uses x_{t+Δt} on both sides; this is not a standard time-indexed update. Use an explicit Euler step notation, e.g., x_{k+1} = x_k + η(t_k) v_θ(x_k,t_k) Δt.","section":"Algorithm 1"},{"comment":"Lemma 5.1 and Proposition 5.2 are named 'informal' but are presented as if they are formal results. Either state them with full assumptions in the main text or label them as claims/conjectures.","section":"§5.1 and Appendix D"},{"comment":"The hyperparameters α₀, β₀, τ_split, k are varied in a small grid, but no selection rule, confidence intervals, or multiple-seed statistics are reported. Since KTS is proposed as a practical inference-time method, this should be clarified.","section":"Tables 3–4"}],"recommendation":"major_revision","confidential_remarks":"The integrated energy–density theorem advertised in the abstract and §4.3 is not proven as written. I would not reject the paper, because the pointwise theorem, the closed-form EFM memorization analysis, and the KTS experiments are valuable, but the authors must either repair Eq. (11) with a rigorous integrable-constants argument or explicitly downgrade the claim. The real-data density finding should also be presented as a descriptor-space proxy result unless stronger evidence is provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing you should know: this paper has a useful empirical core and a broken proof in its main theoretical claim. The KPE per-sample diagnostic is simple and the correlations it reports are plausible; the KTS inference-time modulation is a reasonable idea. But Eq. (11) — the integrated energy–density equivalence — does not follow from the stated Theorem B.1, and the fault is not a minor constant.\n\nWhat is genuinely new: framing the action integral along a flow-matching ODE as a per-sample scalar is a clean, zero-overhead diagnostic. The CLIP-score and CLIP-margin increases with KPE are well executed with proper multiple-testing correction. The energy-versus-density negative correlation is convincing on the synthetic datasets, and the real-data analysis is honestly flagged as representation-dependent. The EFM terminal singularity, which they attribute to Bertrand et al., is correctly linked to memorization, and the CelebA checkpoints showing Fmem and KPE rising together while FID plateaus is a nice empirical observation. KTS is a trivial rescaling of the velocity field; the experiments show a believable trade-off between FID and memorization, though the hyperparameters are per-dataset and some settings hurt FID badly.\n\nThe soft spot is the theoretical centerpiece. Theorem B.1 gives pointwise bounds with additive constants C'_t that, for the linear bridge, include terms like ∥x^(i*)∥²/(1-t)². Integrating these along any trajectory produces infinities, so inequality (54) is vacuous: it gives E ≤ ∞ and E ≥ −∞. The claim that Eq. (11) has an O(1) remainder is therefore unsupported. A concrete counterexample to the proof: one data point at x=1, linear bridge, trajectory x(t)=t. Posterior dominance holds with λ=1 at every t. KPE = 1/2, ∫−log p̂_t dt is finite (it is O(1) plus −log(1-t) integrated), and the theorem's constants blow up like (1-t)⁻². So the theorem as stated cannot imply the integrated Θ relation. This is not a loose-constants quibble; the argument as written is invalid. The authors would need a trajectory-aware argument that controls or cancels those divergent terms.\n\nThe empirical sections stand on their own and the memorization mechanism is well grounded. With the theory fixed or downgraded to a pointwise heuristic, the paper is publishable. As is, the main theorem overclaims.\n\nMy recommendation: send it to a serious referee, but insist on major revision. The gap is fixable or the claim can be softened, and the diagnostic is worth having in the literature. No code is released either, which limits reproducibility but does not affect the math.\n\nI would bring it to a reading group to debate the fix.","headline":"KPE diagnostic and KTS are worth a look, but the paper's main energy–density theorem is unproven: the additive constants in the pointwise bound blow up non-integrably at t=1.","tokens_in":34178,"tokens_out":3386,"would_cite":true,"duration_ms":33127,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The kinetic energy accumulated along a flow-matching trajectory — computed for free during sampling — predicts both semantic fidelity and rarity, until extreme energy tips into memorization.","keywords":["flow matching","kinetic path energy","per-sample diagnostic","memorization","energy-density relation","terminal singularity","inference-time control","Goldilocks principle"],"falsifier":"Check the pointwise ratio ∥u*(z,t)∥² / (−log p_t(z)) along trajectories that cross the region where two training components are equally responsible; if the ratio leaves the Θ(1) band implied by Theorem 4.2 in that region, the energy–density proportionality is not universal and holds only in the posterior-dominance regime.","tokens_in":33054,"feed_emoji":"⚡","tokens_out":8926,"duration_ms":92492,"temperature":0.7,"pith_summary":"Flow-based generative models move a particle from noise to data by integrating a learned velocity field, and this paper argues that the total kinetic effort of that journey — Kinetic Path Energy, E = ½∫∥v∥²dt — is a meaningful per-sample quality signal. Empirically, higher E goes with stronger semantic alignment and with trajectories that land in low-density parts of the data manifold, matching a theoretical bound that relates instantaneous energy to the negative log-density of the interpolating mixture. The paper then shows the relation is not monotone: the closed-form, regression-optimal empirical flow-matching velocity contains a 1/(1−t) terminal singularity, forcing extreme-energy trajectories into near-copies of training points, so 98% of CelebA outputs memorize. This Goldilocks picture motivates a training-free, two-phase inference scheme (KTS) that boosts early velocity and damps late velocity, reducing memorization on CelebA by 16% while improving FID, and offering a tunable quality–coverage knob on ImageNet-256.","feed_headline":"Kinetic energy of a sampling path predicts quality, then memorization","feed_subtitle":"A free per-sample diagnostic for flow models: higher energy means clearer images — until it means copying.","key_machinery":"The central object is Kinetic Path Energy, E = ½∫₀¹∥v_θ(x(t),t)∥² dt, an action-like scalar accumulated along the sampling ODE. It does two kinds of work: as a diagnostic, it converts a high-dimensional trajectory into a single per-sample number that empirically tracks CLIP score/margin and estimated density; as a dynamical quantity, it inherits the score-decomposition identity u*(z,t) = α(t)∇ log p_t + β(t)z, which yields the energy–density proportionality under posterior dominance (Theorem 4.2) and reveals the 1/(1−t) softmax singularity in the closed-form empirical velocity (Proposition 5.2), the mechanism that drives memorization. KTS is a time-dependent gain η(t) that boosts velocity be","core_discovery":"The central claim is that the kinetic path energy of a sampling ODE trajectory is a dual indicator: it tracks how semantically precise the generated sample is and how rare its location is on the data manifold. The mechanism is the identity u*(z,t) = α(t)∇z log p_t(z) + β(t)z, expressing the optimal empirical velocity as a mixture-score plus drift; from this the authors prove the affine proportionality ∥u*(z,t)∥² ≍ −log p_t(z) whenever a single training component dominates the posterior. The counterpoint is the terminal singularity of the same closed-form field: because the optimal velocity is 1/(1−t) times a softmax-weighted sum of training atoms, any trajectory that keeps a fixed gap from t","pith_inferences":["If the energy–density proportionality is taken at face value, per-sample KPE becomes a cheap memorization/novelty risk meter during training runs, long before retrieving nearest neighbors; the paper shows KPE and memorization rising together but does not propose this monitoring use explicitly.","The posterior-dominance caveat suggests the KPE–density correlation should be weakest at small t (where all mixture weights are comparable) and near mode boundaries; a practitioner wanting reliable rarity ranking might restrict the diagnostic to the late-time portion of the trajectory, t ≳ 0.5.","The terminal-blow-up argument is generic to any deterministic transport that must land exactly on one of finitely many atoms at t = 1; the same 1/(1−t) cost structure would apply to rectified flows or stochastic interpolants targeting empirical distributions, so the memorization paradox likely extends beyond the specific flow-matching setting studied here."],"forward_implications":["KPE can be read off any ODE-based flow sampler at zero extra cost, giving per-sample quality and rarity estimates without image-level metrics or density estimation.","Memorization in the regression-optimal empirical flow field is a structural property (the 1/(1−t) terminal singularity), not simply an optimization artifact, so matching the regression objective exactly is explicitly the wrong target if memorization avoidance matters.","An inference-time velocity rescaling (boost early, damp late) yields a tunable quality–memorization trade-off on both CelebA and ImageNet without retraining, consistent with the Goldilocks principle.","Under posterior dominance, KPE is a finite-time observable proxy for the integrated negative log-density of the intermediate Gaussian mixture, so the diagnostic is theoretically grounded, not merely heuristic."],"fun_headline_variants":["Sampling energy: more is better until it's copying","Path energy in flow matching: quality up to the limit","KPE: per-sample energy predicts fidelity, then memorization","Goldilocks rule for flow sampling: avoid too-hot paths"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Load-bearing premise: at every point along the trajectory some single training example must dominate the posterior (responsibility ≥ 1−ε) for the theorem ∥u*∥² ≍ −log p_t to hold; near t=0 or between equally weighted examples it fails, and the real-data rarity finding additionally rests on a descriptor-space density proxy, which the authors call representation-dependent.","fun_headline_variants_meta":{"raw":{"variants":["Sampling energy: more is better until it's copying","Path energy in flow matching: quality up to the limit","KPE: per-sample energy predicts fidelity, then memorization","Goldilocks rule for flow sampling: avoid too-hot paths"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1315,"prompt_tokens":735,"completion_tokens":580,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":511}},"tokens_in":479,"tokens_out":580,"duration_ms":6181,"temperature":1.0,"reasoning_tokens":511,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:25:43.347468+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the pointwise ratio ∥u*(z,t)∥² / (−log p_t(z)) along trajectories that cross the region where two training components are equally responsible; if the ratio leaves the Θ(1) band implied by Theorem 4.2 in that region, the energy–density proportionality is not universal and holds only in the posterior-dominance regime.","supporting_citations":[],"review_version":1}