{"id":"0a1ae4bb-ae5a-4f75-94be-d86953595e87","arxiv_id":"2412.00148","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Guiding a pre-trained image-to-video model with hand-designed energy functions discovers multiple distinct, object-focused motions from a single image, without training.","lead":"Motion Modes is a training-free method that finds several different, plausible ways an object in a single photo could move, by guiding a pre-trained video generator during generation. It could give animators and editors quick motion options without retraining or motion-capture data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest comparative claim rests on circular metrics: Table 1 scores diversity/focus with the same guidance energies the method optimizes, and Study 2 measures coverage of stated human expectations, not a head-to-head against human predictions.","rationale":"The reader's weakest_assumption is that the pre-trained Motion-I2V prior must be rich enough to express diverse motions. That is a genuine scope limitation and is candidly acknowledged in the Limitations section; it bounds generality but does not by itself undercut the proposed mechanism. The more load-bearing problem for the paper's strongest claim is the circular evaluation: the quantities in Table 1 are pieces of the very objective the method minimizes, so 'significantly more diverse and focused' is partly an identity. The user studies are a useful partial check, but Study 2 is a coverage test of self-stated expectations, not a comparison against human-generated predictions, and Study 1 is a set-level preference judgment that can be driven by overall quality rather than the claimed disentanglement. This does not mean the method is wrong; the qualitative results are suggestive and the energy design is plausible. It means the central comparative assertion, as stated in the abstract, is currently under-supported. Independent video-level metrics would settle the matter and would likely strengthen the paper. I therefore keep the reader's CONDITIONAL recommendation (UNCHANGED), while shifting the emphasis from prior richness to evaluation circularity.","tokens_in":12670,"tokens_out":6565,"duration_ms":60690,"concrete_test":"Recompute the comparison on the same 28 scenes using metrics from the final rendered videos rather than from the optimized motion fields: (i) focus = magnitude of off-the-shelf RAFT optical flow in background pixels outside a dilated object mask; (ii) diversity = mean pairwise LPIPS (or FVD) across videos; (iii) plausibility = single-video two-alternative forced choice with independent raters, not set-level comparison. If Motion Modes is not best under these metrics at the same sample budget, the quantitative superiority claim fails; if the project releases motions/code, this check is straightforward.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4's quantitative comparison is the main support for the claim that Motion Modes 'surpasses previous methods.' In Table 1, the diverse metric is the average diversity guidance energy E_d(m,X), and the focused metric is 0.5(E_o+E_c), using the object and static-camera guidance energies. These are exactly the terms minimized during guided denoising via Eq. 4: lambda_d E_d + lambda_c E_c + lambda_o E_o + lambda_s E_s. Thus the result that Motion Modes scores best on these metrics is largely a restatement that the optimizer minimized the objective. Tables 2 and 3 inherit the same issue for the ablation conclusions. The user studies provide some non-circular evidence, but Study 2's 'expected' metric is the fraction of a participant's own stated expectations that our motions cover; it is not a comparison with human-generated motion predictions, so the abstract's 'surpassing ... human predictions' is not supported by the experimental design. The remaining non-circular evidence is qualitative examples and preference judgments that do not separate focus/diversity from overall video appeal. Since code is not yet released and no independent quantitative metric is reported, the strongest comparative claim is currently close to tautological.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Motion Modes proposes a training-free, inference-time guidance method for image-to-video generation. Given a static image and an object mask, the method samples the motion prior of a pre-trained flow-based generator (Motion-I2V) while minimizing energy functions that encourage static camera, focused object motion, diversity across sampled motions, and temporal smoothness. The sampled motions are then rendered into videos. The paper evaluates the method on 28 images against several baselines (prompt-based, ControlNet, random arrows, random noise, FPS noise) using energy-based metrics (Tables 1 and 2), two user studies, and qualitative examples, and also demonstrates an arrow-completion application.","tokens_in":12959,"tokens_out":3323,"duration_ms":29906,"significance":"If the central claim holds, the paper offers a practical, training-free way to discover diverse and plausible object motions from a single image, with clear applications in animation and video editing. The method is well-motivated and the user studies provide non-circular evidence that human judges prefer our motions to baselines on plausibility, diversity, and expectedness. The main weakness is that the headline quantitative metrics (Tables 1 and 2) are exactly the energy functions being optimized during guided denoising, making those numbers partly tautological; the abstract's claim of 'surpassing ... human predictions' is not supported by the experimental design. The limitations are honestly stated, including the reliance on the generator's prior and the discrete sampling of continuous motion spaces.","major_comments":[{"comment":"The 'diverse' metric is defined as the average diversity guidance energy E_d, and the 'focused' metric is 0.5(E_o + E_c), which are precisely the terms minimized by the guided denoising step in Eq. (4): x'_t = x_t - ∇_{x_t} E(x0_θ, m, X) with E = λ_d E_d + λ_c E_c + λ_o E_o + λ_s E_s. Consequently, the result that Motion Modes outperforms all baselines on Table 1 is largely a restatement that the optimization achieved its objective; the metric does not independently validate the method. Please add non-circular quantitative metrics (e.g., optical-flow-based plausibility or external motion-quality scores), or present the user studies as the primary quantitative evidence and relegate the energy-based numbers to a sanity check.","section":"Section 4, Table 1, Eq. (4)"},{"comment":"The abstract's claim of 'surpassing ... human predictions regarding plausibility and diversity' is not supported by the experimental design. Study 2 asks participants to list expected motions and then measures the fraction of those expectations that our motions cover (92%); it does not compare our motions against human-generated motion predictions in a head-to-head evaluation. Please either remove this claim or add an experiment where humans produce motion predictions for the same images and are compared against our outputs on plausibility and diversity.","section":"Abstract and Section 4 (User Study II)"},{"comment":"The ablation study uses the same energy-based metrics as Table 1, so the conclusions are partly confounded. For example, removing E_d and then measuring E_d will trivially show a worse diversity score, since the metric is the very term removed; removing E_c and E_o similarly directly increases the focused metric E_f. The qualitative statement that each component improves the diversity-focus tradeoff is thus not supported by independent evidence. Please evaluate the ablations with the user-study protocol or another external metric, or clearly state that the energy-based ablation numbers are expected by construction.","section":"Section 4, Table 2"}],"minor_comments":[{"comment":"Typo: 'Similiar' should be 'Similar' in the sentence beginning 'Similiar to classifier-free guidance'.","section":"Section 3.2"},{"comment":"The tables report no error bars or significance tests. Since the numbers are dominated by the optimization objective, adding variance or confidence intervals would help readers assess stability across the 28 images.","section":"Section 4, Table 1 and Table 2"},{"comment":"The number of images differs between the quantitative evaluation (28), User Study I (27), and User Study II (10). Please clarify the reason for these differences and whether they affect comparability of the reported percentages.","section":"Section 4, User Studies"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of the journal and the inference-time guidance idea is reasonably novel, but the current quantitative evaluation is too self-referential to support the strongest comparative claims. The user studies are the most convincing evidence and should be foregrounded, while the energy-based metrics should be repositioned as a sanity check. I am not concerned about the training-free claim, since backpropagating guidance gradients at inference time is an accepted form of training-free control. The promise of code release is noted and would strengthen reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nBottom line: this is a practical, training-free method for pulling diverse object motions out of a frozen image-to-video generator, and the core idea is sound. The novelty is in the energy combination and the iterative sampling strategy—static-camera, object-motion, diversity, and smoothness terms all applied during denoising, plus a stopping rule that avoids the memory blow-up of batch particle guidance. I don't know of prior work doing exactly this, and the applications (drag-based editing, motion completion) are useful.\n\nThe paper does several things well. Baselines are fair—same Motion-I2V backbone for everyone. The qualitative results are convincing, and the first user study gives independent evidence that people prefer the outputs on plausibility, diversity, and expectation. The ablation is clean, and the limitations section honestly admits the prior's bias, discretization, and speed.\n\nBut the stress-test concern hits a real soft spot. Tables 1 and 2 report diversity and focus as averages of the exact guidance energies being minimized in Eq. 4. Scoring best on those numbers is largely a statement that the optimizer optimized. That circularity extends to the ablation. It doesn't kill the paper, because the user studies and examples carry the weight, but those tables shouldn't be presented as independent quantitative validation. The abstract's claim about 'surpassing human predictions' is also not supported: Study 2 measures how much of a participant's own stated expectations are covered by our motions; it never compares against actual human-generated motion predictions. That phrasing needs to change.\n\nMinor points: no code released yet, which matters for a method whose main value is plug-and-play inference-time control; and the smoothness-energy-explains-focus story in the appendix is plausible but post hoc.\n\nAll told: I'd send this to a serious referee. The method is likely to work as advertised, the flaws are fixable (report independent metrics or re-frame the claims, tone down the human comparison, release code), and the user-study evidence already gives a non-circular foundation. If I were doing work on inference-time video control I'd cite it; I'd also bring it to reading group to talk about evaluation practice.\n\nRecommendation: engage with it, conditional acceptance after the above edits.","headline":"Useful training-free method for diverse object motion discovery; the evaluation leans too heavily on self-confirming metrics and one overclaimed human comparison, but the core idea is sound and worth refereeing.","tokens_in":13453,"tokens_out":2211,"would_cite":true,"duration_ms":20033,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Motion Modes shows that a training-free scheme can probe a pre-trained flow-based image-to-video generator to discover several distinct, plausible motions for a selected object in a static image, while suppressing camera and background…","keywords":["motion prediction","image-to-video generation","diffusion guidance","training-free","diversity sampling","motion fields","object animation","drag-based editing"],"falsifier":"Run the same backbone with 1000 unguided random noise draws on a fixed set of masked-object scenes, and check whether Motion Modes ever returns a motion that is essentially absent among those random samples; if not, the diversity guidance only selects from already-sampleable motions rather than discovering new modes, and if the returned motions are judged implausible by human raters, the plausibility claim is contradicted.","tokens_in":12485,"feed_emoji":"🎬","tokens_out":11296,"duration_ms":96369,"temperature":0.7,"pith_summary":"The paper tries to establish that a pre-trained image-to-video generator can be probed at inference time, without any training or fine-tuning, to discover several distinct, plausible motions for a selected object in a static image. It proposes Motion Modes, which steers the denoising process of a flow-based generator with energy terms that keep the camera static, move the masked object, push each new motion away from previously found ones, and smooth the motion in time. The authors report that on 28 scenes with articulated objects, animals, vehicles, waves, and flags, this guided sampling beats random-noise, prompt-based, arrow-based, and ControlNet baselines on diversity and focus, and that human raters frequently prefer its motions for plausibility, diversity, and expectation. If correct, the result matters because it turns a black-box video prior into an exploration tool for cinematic ideation and drag-based image editing without needing task-specific training data.","feed_headline":"No training needed: probe a video model for object motions","feed_subtitle":"Four guidance energies keep the camera still, move only the mask, and make each suggested motion distinct.","key_machinery":"The load-bearing objects are the flow generator and four guidance energies. The flow generator models motion as a per-pixel 2D offset field over frames and predicts it separately from appearance, so object and camera motion are partly separated from lighting and shadow changes before guidance begins. The energies are: static-camera $E_c$, which penalizes the average motion magnitude outside the object mask; object-motion $E_o$, a soft-inverse activation applied to the difference between average motion inside and outside the mask; diversity $E_d$, a repulsion from previously found motions measured by a masked angular-and-magnitude distance with weights $w_{mag}=0.25$, $w_{angle}=0.75$; and smoothness $E_s$, which penalizes large frame-to-frame changes inside the mask with weights $w_{mag}=0.75$, $w_{angle}=0.25$. At inference, the gradient of the combined energy is applied to the predicted noise-free motion in each denoising step; iterative sampling discards motions with final energy above $\\rho=5.0$ and stops after two consecutive discards.","core_discovery":"The central discovery is that a flow-based image-to-video generator's latent distribution already contains enough diverse, plausible object motions that a composition of simple energy functions can extract them, disentangled from camera motion and other scene changes. Motion Modes represents a motion as a time-dependent 2D vector field, uses a pre-trained flow generator that predicts this field separately from appearance, and during each denoising step perturbs the predicted clean motion in the direction that lowers the weighted sum of four energies. Iterating this guided sampling builds a set of motions; a threshold on the final guidance energy discards implausible samples and stops when the scene seems to have no more new modes. The paper reports that the resulting motions are more focused and more diverse than four baselines at equal sample budget, and that in a user study 96% of generated motions were judged plausible, 92% of participants' expected motions were produced, and 19% were plausible but outside expectation.","pith_inferences":["If the energy recipe transfers across flow-based generators, the practical contribution is an inference-time control layer rather than a new model, meaning a generator upgrade would immediately boost motion discovery without re-deriving guidance.","Sequential repulsion samples modes one at a time, so continuous motion spaces (a laptop sliding anywhere on a desk) are represented by discrete exemplars; clustering guided samples or interpolating between them in latent space could give continuous motion control.","The dependence on the prior's coverage implies a measurable bias gap: for a fixed scene category, comparing Motion Modes' discovered distributions with human-annotated plausible motions would quantify how much of the missing motion diversity is the generator's data bias rather than the guidance scheme.","The arrow-completion application suggests a direct causal test: if replacing a user's single drag arrow with the closest retrieved motion consistently reduces editing artifacts across many users and scenes, that would characterize when detailed motion completion helps versus when it over-constrains."],"forward_implications":["Because the method is training-free, it can be applied to any pre-trained flow-based image-to-video generator that generates motion separately from appearance, so future improvements to the backbone likewise improve motion discovery without retraining the guidance.","Up to six distinct motions are typically sampled per object, and the stopping criterion automatically gives fewer motions for scenes that admit fewer plausible futures.","The discovered motions can be converted into detailed drag-arrow sets for drag-based editors and motion-to-video generators, replacing a single ambiguous arrow with a complete, plausible motion and avoiding artifacts such as a floating train or a squashed drawer.","On the paper's evaluations, the guided motions outperform random noise, random arrows, ControlNet-restricted motion, farthest-point-sampled noise, and LLM-generated prompts on both diversity and focus, and are judged more plausible, diverse, and expected by users.","The same motions can condition multiple videos: the paper shows that different random noises with the same motion follow the motion accurately while differing in small details."],"supporting_citations":[{"why":"Supplies the pre-trained flow-based image-to-video backbone that predicts motion separately from appearance and is used as the base generator for all guided samples.","marker":"[23]"},{"why":"Contributes the repulsive-energy idea for diverse sampling from diffusion models, which the diversity guidance adapts.","marker":"[6]"},{"why":"Provides a part-level motion prior trained on synthetic data; the paper uses it as a baseline and later as one of the drag-based editors in the motion-completion application.","marker":"[15]"},{"why":"A drag-based image editor used to show that a completed Motion Modes motion avoids implausible results compared with a raw arrow.","marker":"[18]"},{"why":"A drag-based video generation approach that does not produce diverse results, marking the gap that Motion Modes addresses.","marker":"[27]"}],"fun_headline_variants":["IrTraining-free motion modes from a single image","AI invents multiple object motions from one photo, no training","Energy-guided flow discovers diverse object motions from stills","Four energies unlock hidden motions in a video model","Zero training, multiple plausible motions: Motion Modes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the pre-trained generator's latent distribution already contains diverse, plausible motions for the given object and scene; if the prior cannot express a motion, Motion Modes cannot discover it.","fun_headline_variants_meta":{"raw":{"variants":["IrTraining-free motion modes from a single image","AI invents multiple object motions from one photo, no training","Energy-guided flow discovers diverse object motions from stills","Four energies unlock hidden motions in a video model","Zero training, multiple plausible motions: Motion Modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1387,"prompt_tokens":880,"completion_tokens":507,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":431}},"tokens_in":496,"tokens_out":507,"duration_ms":4593,"temperature":1.0,"reasoning_tokens":431,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:11:14.122943+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same backbone with 1000 unguided random noise draws on a fixed set of masked-object scenes, and check whether Motion Modes ever returns a motion that is essentially absent among those random samples; if not, the diversity guidance only selects from already-sampleable motions rather than discovering new modes, and if the returned motions are judged implausible by human raters, the plausibility claim is contradicted.","supporting_citations":[{"cited_title":"Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling","cited_arxiv_id":null,"evidence_quote":"Supplies the pre-trained flow-based image-to-video backbone that predicts motion separately from appearance and is used as the base generator for all guided samples."},{"cited_title":"Jaakkola","cited_arxiv_id":null,"evidence_quote":"Contributes the repulsive-energy idea for diverse sampling from diffusion models, which the diversity guidance adapts."},{"cited_title":"Dragapart: Learning a part-level motion prior for articulated objects","cited_arxiv_id":null,"evidence_quote":"Provides a part-level motion prior trained on synthetic data; the paper uses it as a baseline and later as one of the drag-based editors in the motion-completion application."},{"cited_title":"Dragondiffusion: Enabling drag-style manipula- tion on diffusion models","cited_arxiv_id":null,"evidence_quote":"A drag-based image editor used to show that a completed Motion Modes motion avoids implausible results compared with a raw arrow."},{"cited_title":"Draganything: Motion control for any- thing using entity representation, 2024","cited_arxiv_id":null,"evidence_quote":"A drag-based video generation approach that does not produce diverse results, marking the gap that Motion Modes addresses."}],"review_version":1}