{"id":"2d09b0f0-cfe5-4f12-af42-032bd18e40ba","arxiv_id":"2506.01144","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.","lead":"A new inference-time guidance method, FlowMo, improves motion coherence in text-to-video diffusion models by minimizing patch-wise temporal variance in the model's own latent predictions. It runs on existing models like Wan2.1 and CogVideoX without retraining or extra conditioning, and human raters preferred its output in a study.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Sec. 3.2 correlation does not control for motion magnitude, so FlowMo's variance loss may suppress fast motion rather than target incoherence, undercutting the central claim.","rationale":"The reader's weakest assumption and my assessment converge: the validity of the variance-based proxy is the load-bearing premise. The paper's motivational study in Sec. 3.2 is selective (extreme coherence ratings, motion-filtered) and does not control for the obvious confound of motion magnitude. This is particularly serious because the guidance loss is exactly the variance measure shown in Fig. 2; if that measure is dominated by motion quantity, then FlowMo's effect is equivalent to a motion-reducing regularizer, and the user-study wins could reflect a preference for less dynamic videos rather than genuinely more coherent motion. The paper's counter-evidence, a small aggregate Dynamic Degree drop, is suggestive but not sufficient because it averages over prompts; fast-motion prompts might be disproportionately affected. A stratified re-analysis of the original correlation data, or a per-prompt dynamic-degree measurement on deliberately fast prompts, would settle the question. The reader already recommends conditional acceptance with re-analysis; my stress test identifies the precise condition and a concrete way to achieve it. Thus the verdict remains CONDITIONAL, unchanged, with the condition made explicit.","tokens_in":17640,"tokens_out":4870,"duration_ms":59703,"concrete_test":"Re-analyze the 120-video dataset from Sec. 3.2 by stratifying on the rated motion score (3, 4, 5) or on a computed optical-flow magnitude, and plot the coherent-vs-incoherent mean patch-wise variance (Fig. 2) separately within each stratum. If the variance gap between coherent and incoherent videos collapses or becomes statistically non-significant once motion magnitude is matched, the proxy is confounded by motion magnitude and the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FlowMo's central claim is that minimizing patch-wise temporal variance of appearance-debiased latent predictions (Eqs. 2–3, loss Eq. 8) specifically targets motion incoherence. The supporting correlation study in Sec. 3.2 compares only extreme coherence ratings (1 vs 5), filters for motion score ≥ 3, but does not match or regress out motion magnitude between the coherent and incoherent groups. If the incoherent videos happen to have higher motion magnitude, the variance gap in Fig. 2 could reflect how much the scene changes (motion quantity) rather than whether the change is incoherent. In that case, FlowMo would lower variance by dampening legitimate fast motion, not by repairing artifacts. Table 1's aggregate Dynamic Degree drop is small (<1.5%), but this does not rule out a larger suppression on high-motion prompts, and the user-study preference for 'more coherent motion' could be confounded by reduced motion. The paper's assertion that motion magnitude is preserved (Sec. 4.2) is therefore not established for the cases where the variance signal is strongest. This is the weakest link because the method's mechanism is coherence-specific only if the proxy separates incoherence from motion magnitude.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FlowMo, a training-free, inference-time guidance method for text-to-video diffusion models. At selected early denoising steps, FlowMo computes an appearance-debiased temporal representation from the model's predicted velocity (the L1 distance between consecutive latent frames, Eq. 2), forms a patch-wise temporal variance map (Eq. 3), averages over channels, and uses the maximum value as a loss to take a gradient step on the input latent. Experiments on Wan2.1-1.3B and CogVideoX-5B report VBench Final Score gains of +6.20% and +5.26%, user-study preferences for motion coherence, aesthetic quality, and text alignment, and ablations of the max-variance objective, the debiasing operator, and the timestep selection.","tokens_in":17880,"tokens_out":5096,"duration_ms":57809,"significance":"If the central mechanism holds, FlowMo is a valuable contribution: it is architecture-agnostic, requires no retraining or external conditioning, and is supported by a relatively large human preference study (640 responses per baseline), per-dimension VBench results, and a comparison to FreeInit. The conceptual claim that patch-wise temporal variance in an appearance-debiased latent prediction tracks perceived incoherence is interesting and actionable. However, the motivating evidence is weakened by a missing control for motion magnitude, and the claim that FlowMo improves quality 'without sacrificing' visual quality or prompt alignment is not fully supported by the reported per-dimension VBench degradations.","major_comments":[{"comment":"The correlation study that motivates the variance objective does not control for motion magnitude. The text states that videos were filtered to have a motion score of at least 3, but it does not match or regress out the motion-magnitude rating between the completely incoherent (1) and completely coherent (5) groups. Since Eq. (3) is a temporal variance, it will increase with the amount of motion regardless of coherence, so the gap in Fig. 2 could reflect motion quantity rather than incoherence. This is load-bearing because the whole method minimizes this variance. Please provide the variance conditioned on the motion score, or a matched-pair comparison with equal motion scores, and report the coherence-vs-variance relationship after controlling for motion. In addition, the aggregate VBench Dynamic Degree drop of less than 1.5% does not rule out larger motion suppression on high-motion prompts; report the per-prompt or distributional effect on Dynamic Degree or on human motion-magnitude ratings.","section":"Sec. 3.2, Eqs. (2)-(3), Fig. 2"},{"comment":"The abstract claims FlowMo improves motion coherence 'without sacrificing visual quality or prompt alignment,' but Table 3 shows several per-dimension degradations: CogVideoX Temporal Flickering drops from 99.23% to 96.21%, Human Action drops on both models (Wan2.1: 98.27% to 97.23%; CogVideoX: 97.81% to 95.69%), and Wan2.1 Background Consistency and Color also decline. Since Temporal Flickering and Human Action are directly relevant to temporal coherence and overall quality, the claim of no sacrifice is too strong as stated. Either temper the claim to acknowledge these dimensions, or provide an analysis showing that these drops are not statistically significant or are outweighed in a principled way.","section":"Sec. 4.2, Table 3, App. E"},{"comment":"The ablation study is presented only as qualitative still frames. The central design choices—max versus mean variance, presence of the debiasing operator, and the set of optimized timesteps—are load-bearing for FlowMo's effectiveness, so they should be evaluated with quantitative metrics (e.g., VBench Motion Smoothness, Final Score, or a user study) in addition to the qualitative examples. Without quantitative ablation results, it is difficult to judge whether the differences shown in Fig. 6 are robust or cherry-picked.","section":"Sec. 4.3, Fig. 6"}],"minor_comments":[{"comment":"The word 'denoinsing' appears in the first paragraph of Sec. 3.2 and should be corrected to 'denoising'.","section":"Sec. 3.2"},{"comment":"In the paragraph following Eq. (8), the text says 'the loss in Eq. 7' when referring to the FlowMo loss; this should be Eq. (8).","section":"Sec. 3.3"},{"comment":"The notation 'max_{w∼[W],h∼[H]}' reads as sampling from a distribution; use 'max_{w∈[W],h∈[H]}' for clarity.","section":"Algorithm 1, line 8"},{"comment":"The phrase 'first 12 timesteps of the generation' is not reproducible without knowing the exact scheduler and timestep indices. Please specify the scheduler and the concrete list of timesteps used.","section":"Sec. 4, Implementation details"},{"comment":"The text reports '640 unique responses per baseline' but does not state the number of prompts or participants. Adding these numbers would help readers assess the study design.","section":"Sec. 4.2, User study"},{"comment":"There is a typo in the final sentence: 'produced viseos' should be 'produced videos'.","section":"App. E"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written paper with a substantial empirical component, and the central idea is promising. The main risk is the missing control for motion magnitude in the motivating correlation study, which is load-bearing for the method's mechanism. If the authors can add that analysis, moderate the 'without sacrificing' claim in light of Table 3, and add quantitative ablations, the paper would be suitable for acceptance. I do not see a fundamental circularity or an unfixable flaw; the requested changes are within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FlowMo is a clean and genuinely training-free intervention: take the model's own velocity prediction, compute frame-wise L1 differences to strip shared appearance, take patch-wise temporal variance, maximize over patches, and do a gradient step on the latent during early denoising. That combination is new as far as I can tell, and the paper gives it a fair shake. The user study is real: 640 responses per baseline on VideoJAM prompts, consistent preference for FlowMo on motion quality, aesthetics, and text alignment across two models. VBench Final Score goes up by about 5-6% on both, and the full per-dimension table is in the appendix rather than hidden. The ablation defends the max-patch choice, the debiasing operator, and the early-step schedule. I believe the method does what it claims: it reduces severe temporal artifacts on Wan2.1 and CogVideoX.\n\nThe soft spots are worth naming, in order of severity.\n\nFirst, the motivational correlation in Sec. 3.2 does not control for motion magnitude. You take videos rated at least 3 on motion and compare only the extremes of coherence (1 vs 5). If the incoherent group happens to have more motion, the variance gap in Fig. 2 could reflect how much the scene changes, not whether the change is incoherent. The paper's own explanation for the Dynamic Degree drop (artifacts add spurious motion) is plausible, but it does not settle the question. Since FlowMo directly minimizes patch variance, it could be dampening legitimate fast motion on some prompts. The aggregate Dynamic Degree decrease is under 1.5%, which is reassuring, but per-prompt behavior is not reported. This is the weakest link, and the authors should either match motion across groups or regress it out.\n\nSecond, the closest prior method, 3dv-ton, is cited but never compared. The authors compare to FreeInit adapted to DiTs, which is reasonable, but 3dv-ton minimizes a global 3D variance loss and is the natural baseline for 'variance-based guidance.' A direct comparison would make the contribution crisper.\n\nThird, the abstract says 'without sacrificing visual quality or prompt alignment,' but the VBench breakdown shows CogVideoX Temporal Flickering dropping from 99.23 to 96.21, and Human Action dropping on both models. These are not catastrophic, but the claim is too strong. The text in Sec. 4.2 partially acknowledges decreases, but the abstract should be more careful.\n\nThe central claim - training-free variance guidance improves temporal coherence - survives these issues. The mechanism question is open, but the evidence for the practical effect is solid. This deserves a serious referee; the missing baseline and the motion-magnitude confound are fixable in revision.","headline":"Training-free variance guidance for motion coherence that mostly works, with solid human-preference evidence; the main open question is whether the variance signal targets incoherence or just motion magnitude.","tokens_in":18398,"tokens_out":3159,"would_cite":true,"duration_ms":34817,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FlowMo claims that temporal coherence in text-to-video diffusion models can be improved at inference time by reducing patch-wise temporal variance in appearance-debiased latent predictions, without retraining or external conditioning.","keywords":["text-to-video generation","training-free guidance","temporal coherence","motion artifacts","latent-space variance","flow matching","inference-time optimization","diffusion transformers"],"falsifier":"Compute the same variance statistic on a larger set of generated videos spanning all coherence ratings and matched for motion magnitude; if high variance also appears in coherent videos or fails to separate moderate incoherence from coherence, the guidance signal would be targeting motion magnitude rather than incoherence. A second check would apply FlowMo to a video whose incoherence is confined to patches with low variance; if artifacts persist, the max-patch variance criterion is not sufficient.","tokens_in":17455,"feed_emoji":"🎬","tokens_out":5728,"duration_ms":55745,"temperature":0.7,"pith_summary":"FlowMo asks whether a pre-trained text-to-video diffusion model already contains enough temporal information inside its own latent predictions to fix its motion errors. The paper answers yes: it defines an appearance-debiased temporal representation by taking per-patch distances between consecutive frames, measures patch-wise temporal variance, and uses the maximum-variance patch as a loss to refine the latent during sampling. The method requires no retraining, no external optical flow, and no architectural changes. On Wan2.1-1.3B and CogVideoX-5B, FlowMo improves motion smoothness and overall VBench final score by 6.20% and 5.26% respectively, while human raters prefer it for motion coherence, aesthetic quality, and text alignment. A sympathetic reading is that temporal coherence can be improved by looking inward at the model's own predictions rather than adding external motion priors.","feed_headline":"Training-free fix steadies motion in AI video","feed_subtitle":"Lowering patch-wise latent variance during sampling boosts coherence in Wan2.1 and CogVideoX without extra inputs.","key_machinery":"At the core is the appearance-debiased temporal variance signal. Given the model's predicted velocity $u_{\\theta,t}$ for $F$ latent frames, the operator $\\Delta$ computes the $\\ell^1$-distance between consecutive frames channel-wise, removing shared appearance; then a patch-wise variance tensor $\\sigma^2_{w,h,c}$ is computed across the $F-1$ difference frames and averaged over channels to form a per-patch coherence map $s_{w,h}$. The FlowMo loss is the maximum over patches, $L=\\max_{w,h} s_{w,h}$, which targets the most dynamically incoherent region rather than averaging over mostly static patches. Gradient descent on this loss updates the input latent $z_{t_i}$, after which the diffusion step is recomputed; it is applied only at early-to-mid timesteps where qualitative visualization shows coarse motion emerging. This mechanism carries the argument by converting the correlation in Section 3.2 into an actionable, training-free guidance step.","core_discovery":"On the paper's own terms, the central discovery is that coherent motion corresponds to low temporal variance in an appearance-debiased latent representation, and that this variance can be used as a guidance signal. The paper observes that predictions of flow-matching video models are appearance-biased, so it computes the $\\ell^1$-distance between latent predictions of consecutive frames to cancel shared appearance, then computes variance across frames for each spatial patch and channel. Incoherent videos show consistently higher patch-wise variance than coherent ones, with separation emerging around the fifth denoising step, and coarse motion is established in intermediate steps rather than at the start. FlowMo selects the patch with maximal temporal variance, back-propagates through the prediction to adjust the input latent, recomputes the model prediction, and repeats this refinement at the first twelve timesteps. The paper claims this makes latent transitions smoother and that this maps to smoother pixel-space behavior, significantly boosting motion coherence while preserving or improving visual quality and prompt alignment.","pith_inferences":["Editorial inference: the per-patch variance map that FlowMo uses as a loss could also serve as a cheap automatic localizer of temporal artifacts in generated videos, since it already highlights the regions the optimization targets.","Editorial inference: because FlowMo consumes no external inputs, it could plausibly be stacked on top of trajectory-based or optical-flow conditioning methods to clean residual incoherence they leave behind, though the paper does not test this combination.","Editorial inference: the same variance objective could be folded into training to give video models a richer temporal prior; the paper itself notes that inference-time optimization is bounded by what the pretrained model already knows how to represent."],"forward_implications":["Any pre-trained flow-matching text-to-video model with accessible latent predictions can be guided the same way; the method is architecture-agnostic and needs no fine-tuning.","Generated videos gain motion smoothness and fewer artifacts such as extra limbs and objects that appear or disappear, while dynamic degree drops only slightly because spurious motion is removed.","The method offers a plug-and-play alternative to retraining or external motion signals for improving temporal fidelity.","Applied during the first twelve timesteps, FlowMo also reduces and stabilizes maximal patch-wise variance in later, non-optimized steps, suggesting the coarse motion structure set early controls later coherence."],"supporting_citations":[{"why":"Supplies Wan2.1-1.3B, the primary model used for the motivation experiments, ablations, and comparisons against FreeInit.","marker":"[1]"},{"why":"Supplies CogVideoX-5B, the second model on which FlowMo's generalization is tested.","marker":"[2]"},{"why":"Provides the VideoJAM benchmark prompts used in the user studies and the observation that text-to-video model predictions are biased toward appearance.","marker":"[3]"},{"why":"Supplies VBench, the automatic evaluation suite whose Final Score, Motion Smoothness, and Dynamic Degree report the main quantitative gains.","marker":"[16]"},{"why":"Defines flow matching, the training objective that makes the model's output a velocity estimate in latent space.","marker":"[44]"},{"why":"Inspires the mechanism of optimizing the input latent with gradient descent on an auxiliary loss during inference.","marker":"[47]"},{"why":"Is FreeInit, the closest baseline, which the paper adapts to DiT models and compares against in both user studies and VBench metrics.","marker":"[59]"},{"why":"Supplies classifier-free guidance, applied before and after each FlowMo refinement step.","marker":"[64]"}],"fun_headline_variants":["Training-free guidance uses latent variance to stabilize AI video","Reducing patch variance in latents improves AI motion coherence","FlowMo: training-free variance guidance for steadier AI video","AI video motion made coherent without extra training","Latent variance guidance steadies AI video without retraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the size of frame-to-frame changes in the model's internal prediction, after stripping out shared appearance, marks where motion looks incoherent, and that this was verified only on a small user study comparing extreme coherence ratings.","fun_headline_variants_meta":{"raw":{"variants":["Training-free guidance uses latent variance to stabilize AI video","Reducing patch variance in latents improves AI motion coherence","FlowMo: training-free variance guidance for steadier AI video","AI video motion made coherent without extra training","Latent variance guidance steadies AI video without retraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000498,"raw_usage":{"total_tokens":2425,"prompt_tokens":919,"completion_tokens":1506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":1428}},"tokens_in":535,"tokens_out":1506,"duration_ms":11547,"temperature":1.0,"reasoning_tokens":1428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:49:57.331753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the same variance statistic on a larger set of generated videos spanning all coherence ratings and matched for motion magnitude; if high variance also appears in coherent videos or fails to separate moderate incoherence from coherence, the guidance signal would be targeting motion magnitude rather than incoherence. A second check would apply FlowMo to a video whose incoherence is confined to patches with low variance; if artifacts persist, the max-patch variance criterion is not sufficient.","supporting_citations":[],"review_version":1}