{"id":"7a9b2cc8-830a-4fb2-82df-113b8d1a507b","arxiv_id":"2504.12048","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Modular-Cam combines LLM-based prompt decomposition, per-motion LoRA modules, and ControlNet-based scene conditioning to generate multi-scene videos with camera-view control.","lead":"This paper introduces Modular-Cam, a text-to-video system that uses a large language model to split a complex prompt into scenes and camera actions, then generates each scene with separate camera-motion modules. Because it targets multi-scene videos with smooth transitions and camera control, it addresses a known weakness of current text-to-video models, though the evaluation is mostly qualitative and self-generated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cannot accept the fine-grained camera-control claim: no experiment tests whether synthetic-trained CamOperator LoRA modules transfer to natural prompts or compose additively, and no metric measures camera-action alignment.","rationale":"The central claim has two parts: multi-scene quality/consistency and modular fine-grained camera control. The first is partially supported by qualitative results, CLIP scores, and ablations, though the evaluation is under-powered. The second, and more distinctive, part depends on CamOperator. The paper's quantitative metrics are incapable of validating camera direction; Dynamic Degree measures motion quantity only, and CLIP score measures overall text compliance, not whether a PanRight was actually executed. No metric compares the estimated camera trajectory to the requested action, and no composition experiment appears in the paper or appendix. The reader's weakest_assumption pinpoints exactly this, and I agree. I do not see a reason to move beyond CONDITIONAL, and the existing verdict is already CONDITIONAL, so the recommended verdict is UNCHANGED. A trajectory-classification experiment plus a composition check would settle the concern. Since no code or weights are released, a third party cannot currently run that check, which reinforces the conditional status rather than supporting rejection, because the qualitative examples are suggestive and the claim could survive if the test passes.","tokens_in":14201,"tokens_out":7125,"duration_ms":71983,"concrete_test":"Generate a matched set of natural-scene prompts for each of the six actions and for at least one composed action (e.g., 'ZoomIn + PanLeft'), using fixed random seeds. Run Modular-Cam with the corresponding single or summed LoRA modules, and with no CamOperator as a control. Estimate per-frame camera motion robustly (e.g., RAFT optical flow plus homography or OpenCV ECC) and classify the trajectory direction. Accept the control claim only if: (a) single-action classification accuracy versus the target action is high (>=80%); (b) the composed trajectory's zoom and translation components approximate the sum of the two single-module trajectories; and (c) these results hold on natural prompts, with no synthetic augmentation at inference. This directly tests transfer, additive composition, and whether the generated video follows the intended camera action.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 'CamOperator with Modular Network' makes the central control claim depend on two unverified empirical premises: (i) LoRA modules fine-tuned on ~50 synthetically augmented videos (Experiment Setup) transfer to natural content, and (ii) separately trained modules compose additively as W = W_TT + A_CO * B_CO^T (Eq. 4), so combinations such as ZoomIn+PanLeft can 'theoretically simulate any motion pattern.' No experiment tests either premise. The quantitative evaluation uses Motion Smoothness, Dynamic Degree, Imaging Quality, CLIP score and User Rank; none measures whether the camera moved in the requested direction, so a high Dynamic Degree is compatible with a wrong PanRight or an unrelated dolly. The synthetic ZoomIn data are made by 'gradually reduc[ing] the video screen size,' a digital crop rather than physical camera motion, and 50 training videos is a small basis for transfer. Additive LoRA composition is non-obvious for temporal attention weights and is asserted, not demonstrated. The Appendix Limitations section concedes exact pan/rotation control is future work, narrowing the scope but not repairing the missing verification of the coarse six-action control and composition. If transfer or composition fails, the 'fine-grained control of camera movements' contribution is unsupported.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-16T12:38:39.911028+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}