{"id":"24dd2ffa-197b-43dd-9ae4-3b114b2db920","arxiv_id":"2505.19901","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An adapter that injects Qwen2VL multimodal features into CogVideoX-I2V improves dynamic range on the authors' new DIVE benchmark, but the SOTA claims rest mainly on that self-designed metric.","lead":"Dynamic-I2V adds a multimodal language model to an existing image-to-video diffusion model, using an adapter to fuse image, text, and answer features, and reports better motion and control in generated videos. The paper also introduces DIVE, a GPT-4o-based benchmark that penalizes static videos, and shows VBench-I2V gives near-top scores to frozen frames.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed 42.5%/7.9%/11.8% DIVE improvements are not reproducible from Table 3; recomputing over the stated baseline gives 54.6%/9.2%/31.1%, so the SOTA claim lacks internal numerical support.","rationale":"The reader's weakest assumption—that DIVE's GPT-4o scoring with an unreleased prompt is a valid gold standard—is a legitimate external-validity concern. My analysis identifies a more immediate, internal problem: the paper's central quantitative claim is not reproducible from its own Table 3. This is load-bearing because the abstract and Section 4.2 explicitly quantify the SOTA improvement, and those numbers are inconsistent with the reported data. If the percentages are erroneous, the SOTA claim is unsupported; if they are computed differently, the paper fails to explain how, making the result non-verifiable. This does not necessarily mean the method is ineffective—the DIVE scores do still favor Dynamic-I2V on DR and DBQ—but the headline numbers must be corrected or clarified before acceptance. I also note the paper's static-video experiment on VBench is a useful and genuinely concerning finding, and the MCA architecture is a plausible engineering contribution. However, the arithmetic mismatch takes precedence because it undermines the main claim as stated. The reader's verdict of CONDITIONAL remains appropriate: the authors should fix the numerical discrepancy, release the DIVE prompts and code, and provide significance tests. I therefore recommend no change to the verdict, while adding this specific missing check.","tokens_in":12885,"tokens_out":6046,"duration_ms":61291,"concrete_test":"Recompute the three percentage improvements from Table 3 using every plausible baseline: (a) CogVideoX-I2V-5B only, (b) DynamiCrafter only, (c) best score per metric among all non-Dynamic-I2V models, (d) average of all non-Dynamic-I2V models. Verify whether any combination yields 42.5%, 7.9%, and 11.8%. If none does, request the raw per-prompt DIVE scores and the exact definitions of DR, DC, and DBQ to trace the discrepancy. Additionally, ask the authors to report 95% confidence intervals or bootstrapped standard errors for the DIVE metrics across the 355 prompts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 4.2 claim that Dynamic-I2V improves over the SOTA baseline CogVideoX-I2V-5B by 42.5% in dynamic range, 7.9% in dynamic controllability, and 11.8% in dynamics-based quality as measured by DIVE. However, these numbers do not match the data in Table 3. Computing relative improvements over CogVideoX-I2V-5B yields (45.77-29.61)/29.61 = 54.6% for DR, (74.44-68.17)/68.17 = 9.2% for DC, and (62.25-47.48)/47.48 = 31.1% for DBQ. Even if one instead compares against the best non-Dynamic-I2V model per metric (DynamiCrafter for DC), the gains are 55.9%, 3.2%, and 33.5% — none equal to the claimed triple. The paper does not report the exact formula or baseline used to arrive at 42.5%, 7.9%, and 11.8%, nor does it provide standard deviations or significance tests for the DIVE scores. This internal inconsistency means the headline SOTA claim is not numerically grounded even if DIVE were accepted as a valid metric. A secondary issue is that Table 2 shows Dynamic-I2V's Total Score (88.84) is slightly below DynamiCrafter (88.85) when camera motion is excluded, further weakening the broad SOTA statement on VBench-I2V.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Dynamic-I2V, an image-to-video generation model that augments CogVideoX-I2V-5B with Qwen2-VL features via a multimodal conditional adapter (MCA), and proposes DIVE, a GPT-4o-based evaluation benchmark for dynamic quality. The authors report state-of-the-art results on DIVE (improvements of 42.5% in DR, 7.9% in DC, 11.8% in DBQ over existing methods) and on VBench-I2V, and additionally present evidence that VBench-I2V's scores are biased toward static videos. The paper includes ablations of the MCA components and a small human study comparing four models.","tokens_in":13255,"tokens_out":6463,"duration_ms":60397,"significance":"The idea of fusing MLLM tokens into a DiT backbone for I2V is timely and potentially impactful, and the DIVE benchmark addresses a real limitation of existing I2V metrics. The static-video sanity check is a falsifiable demonstration that VBench-I2V rewards non-dynamic outputs. However, the central empirical claims are weakened by internal inconsistencies and the reliance on an unreleased evaluation prompt. If the reported gains survive recomputation and the benchmark is released, this could become a useful contribution. The current manuscript does not yet substantiate the SOTA claim.","major_comments":[{"comment":"The abstract and Section 4.2 claim improvements of 42.5%, 7.9%, and 11.8% over existing methods on DIVE. With CogVideoX-I2V-5B as the stated baseline, Table 3 yields relative gains of 54.6% for DR, 9.2% for DC, and 31.1% for DBQ. No baseline or formula reproduces the reported numbers. Please correct the numbers or explicitly specify the baseline and computation used.","section":"Section 4.2 and Abstract"},{"comment":"DIVE's core evaluation prompt and rubric are not provided (the supplementary is referenced but missing), and the human study used only 20 volunteers on 50 videos with no inter-annotator agreement or confidence intervals. Since DIVE is the primary benchmark for the SOTA claim and is introduced by the authors, the lack of a released prompt and of statistical validation prevents independent verification. Please include the prompt, release the evaluation code, and report statistical significance.","section":"Section 4.2 and 4.3"},{"comment":"The VBench-I2V results do not support the broad claim of state-of-the-art performance: Dynamic-I2V's Dynamic Degree (27.15) is substantially below DynamiCrafter (47.40) and SVD (43.17), and its Total Score advantage over CogVideoX-I2V-5B is only 0.24 points. If the authors' thesis is that dynamics matter, the model's low Dynamic Degree on VBench-I2V is concerning. Please clarify how the model can be SOTA on VBench-I2V while being worse on the dynamics dimension that the paper emphasizes.","section":"Table 1 and Section 4.1"},{"comment":"The ablation shows a large jump from +MLLM+MLPs (DR 27.90) to +MCA (DR 45.77), but the paper does not explain what causes the improvement. The only architectural difference between these two conditions appears to be the zero-initialized convolutions in Eq. (3). Please provide an analysis of the contribution of each MCA component (e.g., zero-conv initialization, residual connection) to the final performance.","section":"Table 6"}],"minor_comments":[{"comment":"There is a duplicated phrase: \"Dynamic-I2V consists of Dynamic-I2V consists of a denosing module\"; please fix the typo.","section":"Section 3, first sentence"},{"comment":"The sentence \"These features, fi, fa and ft are then into CA\" is missing a verb (likely \"are then fed into\"); please correct.","section":"Section 3.2, text near Eq. (2)"},{"comment":"The optimizer setting is written as \"betas of [0.9, 0.95]\" with an unclosed bracket; please check the notation.","section":"Section 5.2, Implementation Details"},{"comment":"The protocol for excluding the camera-motion metric is not fully specified: the paper says the exclusion is due to an additional prompt for camera motion, but does not describe that prompt or how the exclusion affects the reported scores. Please clarify.","section":"Table 2 and Section 4.1"},{"comment":"There is a typo: \"Text Conpliance\" should be \"Text Compliance\".","section":"Table 5, header"},{"comment":"The figure is crowded and the text is small; please improve readability or provide a larger version.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising architecture and a useful critique of VBench, but the internal inconsistency in the headline improvement numbers and the unpublished DIVE prompt are serious. I recommend requesting the evaluation prompt and code, and a careful revision of the claims based on recomputed numbers. If the authors cannot reproduce the improvements, the paper should not be accepted in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the MCA adapter and the DIVE benchmark are both reasonable contributions, and the static-video VBench demonstration is genuinely worth seeing. But the paper's three headline percentages are internally inconsistent with Table 3, and the SOTA claim rests on a self-designed metric with an undisclosed GPT-4o prompt. It deserves peer review, not because the claims are airtight, but because the engineering and the benchmark critique are useful enough to be sorted out carefully.\n\nWhat's new: the multimodal conditional adapter (MCA) — zero-conv MLPs fusing Qwen2VL's answer, vision, and query features with T5 features into CogVideoX-I2V — is a clean, lightweight integration that hasn't appeared in exactly this form. The ablation shows it beats simple addition and MLP-only baselines, though the jump from +MLP to +MCA is suspiciously large (DR 27.90 to 45.77) and gets no explanation. DIVE, extending DEVIL to I2V, addresses a real problem: VBench-I2V rewards static videos, which the paper demonstrates convincingly by feeding it duplicated frames.\n\nSoft spots: the abstract claims 42.5%, 7.9%, and 11.8% improvements over the SOTA baseline, but Table 3 implies 54.6%, 9.2%, and 31.1%. The paper doesn't say how the headline numbers were computed. That's a load-bearing inconsistency for the central SOTA claim. Also, DIVE's GPT-4o prompt and rubric are in a missing supplementary, and the only calibration is a 20-person study on 50 videos with no significance testing. On VBench-I2V itself, Dynamic-I2V's Dynamic Degree (27.15) is well below DynamiCrafter (47.40), so the broad SOTA statement only holds on the self-designed metric.\n\nVerdict: this is a plausible engineering paper with a meaningful benchmark critique, but the evaluation needs real work: release code and prompts, correct the arithmetic, run significance tests, and reconcile the VBench dynamic-degree gap. A serious referee could push it into shape. I wouldn't cite it yet, but I'd bring it to a reading group to discuss the static-bias result.","headline":"A useful MLLM-driven adapter and a valid critique of static-biased benchmarks, undercut by headline numbers that don't match the paper's own table.","tokens_in":13809,"tokens_out":2707,"would_cite":false,"duration_ms":27362,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adding a multimodal LLM's visual and textual features to an image-to-video diffusion transformer yields dramatically more dynamic videos, and that a new evaluation metric (DIVE) reveals gains that static-biased…","keywords":["image-to-video generation","multimodal large language model","diffusion transformer","dynamic video evaluation","motion dynamics","video quality benchmark","text-image conditioning","conditional adapter"],"falsifier":"Re-run the DIVE evaluation with a new, larger set of image-text pairs and have fresh human raters rank videos for dynamics without knowing which model produced them; if the rank correlation between DIVE scores and human rankings is not significantly positive (say Spearman's rho below 0.5), the claimed 42.5% dynamic-range advantage would not be a reliable measure of human-perceived motion quality.","tokens_in":12697,"feed_emoji":"🎬","tokens_out":7408,"duration_ms":66883,"temperature":0.7,"pith_summary":"The paper argues that image-to-video generation models fail at complex scenes because they cannot jointly understand what is in the image and what the text asks them to do. To fix this, the authors integrate a multimodal large language model (Qwen2VL) into the CogVideoX-I2V diffusion transformer and fuse its visual and textual features with the T5 text encoder through a lightweight conditional adapter. They also claim that existing benchmarks like VBench-I2V are biased toward static videos, since even a purely static video scores as well as state-of-the-art models, and they propose a new benchmark, DIVE, that directly measures dynamic range, controllability, and dynamics-based quality. Under DIVE, their model reports large improvements over prior methods, including a 42.5% gain in dynamic range. The broader point is that motion dynamics should be a first-class axis of evaluation and that MLLM-based understanding is a promising route to achieving it.","feed_headline":"MLLM-powered video model boosts dynamic range by 42.5%","feed_subtitle":"A new DIVE benchmark scores motion itself, and the model leads on all three dynamic metrics.","key_machinery":"The load-bearing mechanism is the Multimodal Conditional Adapter (MCA), a learnable fusion module that combines three feature streams: the MLLM's vision-token features $f_i$, its answer-token features $f_a$, and the T5 text-encoder features $f_t$. The adapter computes $f_c = Z_m(M_i(f_i) + M_a(f_a)) + (f_t + Z_t(f_t))$, where $M_i$ and $M_a$ are MLPs and $Z_m$ and $Z_t$ are zero-initialized convolutional layers; the zero initialization ensures that early training preserves the base model's behavior and then gradually injects the learned multimodal conditions. This $f_c$ feeds into the diffusion transformer's 3D full-attention module, guiding denoising. The other central object is the DIVE evaluation metric, which uses GPT-4o to assign a dynamic degree (1-5) to each prompt and a dynamic score (0-1) to each generated video, then aggregates them into Dynamic Range, Dynamics Controllability, and Dynamics-Based Quality.","core_discovery":"The central claim is that adding a multimodal LLM to an image-to-video diffusion transformer, through a carefully designed adapter, unlocks substantially more dynamic and controllable video generation, and that existing evaluation protocols hide this improvement because they reward static outputs. Concretely, the paper integrates Qwen2VL into CogVideoX-I2V, feeding the MLLM's vision-token features, answer-token features, and T5 text features into a zero-initialized convolutional adapter (the Multimodal Conditional Adapter, MCA). The authors report state-of-the-art results on the VBench-I2V leaderboard and, more pointedly, on their proposed DIVE benchmark, where they claim improvements of 42.5% in dynamic range, 7.9% in dynamics controllability, and 11.8% in dynamics-based quality over the base CogVideoX-I2V-5B model. The paper also constructs the striking demonstration that a video made by duplicating a single frame scores on par with leading models on VBench-I2V, which motivates DIVE as a dynamic-aware alternative aligned with human rankings in a 20-person study.","pith_inferences":["A natural extension the paper leaves implicit is that the MCA fusion recipe should transfer to other conditional video tasks, such as text-to-video editing, where the same joint image-text understanding is needed.","The paper's static-video attack on VBench suggests that any video benchmark without an explicit motion axis is vulnerable to reward hacking; future benchmark builders could add adversarial static inputs as a standard sanity check.","Because the model was fine-tuned on a filtered subset of OpenVid (123k videos), part of the dynamic-range gain could come from data curation rather than the MLLM alone; a controlled dataset ablation would disentangle these sources.","A testable extension: DIVE's reliance on GPT-4o could be validated by comparing against human dynamics ratings on a larger, independent pool of videos; if the correlation holds, DIVE becomes a cheap standard for dynamic I2V evaluation."],"forward_implications":["If the paper is right, image-to-video models that understand images and text jointly should be preferred for complex prompts, and simple text-encoder-only conditioning is a bottleneck.","VBench-I2V and similar frame-quality-focused benchmarks can be gamed by static videos, so future I2V evaluation should include a motion-dynamics axis like DIVE.","The reported 42.5% improvement in dynamic range means the model can produce both highly dynamic and near-static videos depending on the prompt, which is useful for controllable generation.","The MCA design shows that zero-initialized adapters allow adding MLLM features to a frozen, pretrained DiT without destroying its existing capabilities, suggesting a reusable recipe for other conditioning modalities.","The model's support for diverse image conditions (e.g., providing an image of an explosion type) suggests MLLM conditioning can carry abstract semantic controls beyond the literal first frame."],"supporting_citations":[{"why":"Supplies the Qwen2VL multimodal LLM whose vision-token and answer-token features carry the joint image-text conditioning.","marker":"[38]"},{"why":"Defines the CogVideoX-I2V-5B base model and VBench-I2V comparison that Dynamic-I2V fine-tunes and claims to outperform.","marker":"[43]"},{"why":"Provides the VBench++ test set used by DIVE and the benchmark whose static-video bias the paper demonstrates.","marker":"[15]"},{"why":"Establishes the dynamic-degree/dynamic-score evaluation framework that DIVE adapts to the image-to-video setting.","marker":"[21]"},{"why":"Supplies the OpenVid-1M dataset from which 123k filtered videos are used to fine-tune Dynamic-I2V.","marker":"[26]"}],"fun_headline_variants":["MLLM adapter makes I2V videos 42.5% more dynamic","Old benchmarks reward static videos; new DIVE scores motion","Qwen2VL-powered I2V gains 42.5% motion range","DIVE benchmark exposes I2V static-frame trap","MLLM + DiT: 42.5% more dynamic, 7.9% better control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The DIVE benchmark's GPT-4o-based scoring is assumed to track what humans mean by good, dynamic video, even though only 20 volunteers watched 50 videos to check this.","fun_headline_variants_meta":{"raw":{"variants":["MLLM adapter makes I2V videos 42.5% more dynamic","Old benchmarks reward static videos; new DIVE scores motion","Qwen2VL-powered I2V gains 42.5% motion range","DIVE benchmark exposes I2V static-frame trap","MLLM + DiT: 42.5% more dynamic, 7.9% better control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001112,"raw_usage":{"total_tokens":4681,"prompt_tokens":1040,"completion_tokens":3641,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":3540}},"tokens_in":656,"tokens_out":3641,"duration_ms":25653,"temperature":1.0,"reasoning_tokens":3540,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:03:45.041785+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the DIVE evaluation with a new, larger set of image-text pairs and have fresh human raters rank videos for dynamics without knowing which model produced them; if the rank correlation between DIVE scores and human rankings is not significantly positive (say Spearman's rho below 0.5), the claimed 42.5% dynamic-range advantage would not be a reliable measure of human-perceived motion quality.","supporting_citations":[{"cited_title":"Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024","cited_arxiv_id":null,"evidence_quote":"Provides the VBench++ test set used by DIVE and the benchmark whose static-video bias the paper demonstrates."},{"cited_title":"Openvid-1m: A large-scale high-quality dataset for text-to- video generation, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the OpenVid-1M dataset from which 123k filtered videos are used to fine-tune Dynamic-I2V."}],"review_version":1}