{"id":"fba7283d-fc35-476b-88a8-e6a3c6935af6","arxiv_id":"2608.01113","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CoT-Edit achieves state-of-the-art instruction-based video editing by generating bounding boxes and enriched instructions with a CoT-enhanced multimodal planner, which guide mask-based diffusion editing.","lead":"CoT-Edit is a video editing system that first uses a multimodal language model to plan where and how to edit, drawing target boxes frame by frame. Because it grounds edits in explicit locations before generating, it is more reliable than text-only models in scenes with similar objects and for physically plausible additions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gemini-judge bias is the load-bearing risk for the physical-plausibility and spatial-localization SOTA claims; blinded human or independent-judge ratings should settle it.","rationale":"The paper is a coherent and well-motivated plan-guide-edit framework. The modular design is clear, the two-stage training is plausible, and the non-Gemini metrics (FVD, CLIPScore, BC, MS, AES) do show CoT-Edit at or near the top, which is independent evidence that the system is competitive. The weakest point is the evaluation of the two claims that make the paper distinctive: physical plausibility and spatial relations are measured only by Gemini, and the planner's CoT output is exactly the kind of structured MLLM reasoning such a judge may favor. That is a load-bearing risk, not a stylistic objection, because the abstract and Section 4.2 build the SOTA claim on those dimensions. I agree with the reader's weakest_assumption. The proposed blinded human study would settle whether the Gemini margins are genuine. Until then, CONDITIONAL is the right verdict; no change from the reader is needed. Additional minor issues—the lower temporal consistency in Table 1, the unrelated reference [23], and the absence of released code/data—are real but secondary and should be fixed without changing the verdict category.","tokens_in":12044,"tokens_out":4828,"duration_ms":41888,"concrete_test":"Run a blinded expert human preference study on the same 100 Koala-36M editing videos used in Section 4.2, comparing CoT-Edit against OmniVideo and Lucy-1.1 on physical plausibility, spatial relations, and instruction following, with each video independently rated by at least three annotators. Report per-item scores, means with confidence intervals, inter-annotator agreement, and the correlation between Gemini ratings and human ratings. If human ratings do not reproduce CoT-Edit's large margins on Physical Rule and Spatial Relation, the SOTA claim on those dimensions collapses to an artifact of the judge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline differentiators—physically consistent object additions and precise spatial localization—are supported in Table 1 entirely by Gemini ratings (Physical Rule, Spatial Relation, Instruction Following, Editing Quality). The CoT planner is itself an MLLM that produces enriched textual instructions and box sequences, so a Gemini judge may systematically prefer outputs that resemble MLLM-style structured reasoning rather than genuinely better physics or grounding. The reported margins over OmniVideo and Lucy-1.1 are large (e.g., Physical Rule 0.741 vs 0.590; Spatial Relation 0.841 vs 0.641), but no prompt template, per-item scores, variance, or significance is reported, and no correlation with human judgment is given in the main text. A secondary internal inconsistency: Table 1 shows CoT-Edit's Temporal Consistency (0.945) is below InsV2V (0.958), InsViE (0.957), OmniVideo (0.954), and Lucy-1.1 (0.961), so the paper should not claim blanket temporal SOTA. Because the central qualitative claims depend on the validity of Gemini as an unbiased judge, this is the most load-bearing unverified premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CoT-Edit, a Plan-Guide-Edit framework for instruction-based video editing. A CoT-enhanced MLLM planner takes input keyframes and a user instruction to produce a temporal bounding-box sequence and an enriched instruction; a box-conditioned mask branch converts these spatial priors into spatiotemporal masks; and a diffusion editor built on Wan2.2 5B fuses masks, enriched text, and video features to render the edited video. The authors report state-of-the-art quantitative results on a 100-video sample from Koala-36M, together with ablations of the CoT planner and the mask branch, and claim reduced data requirements through modular followed by joint training.","tokens_in":12403,"tokens_out":4922,"duration_ms":42858,"significance":"If the empirical claims hold, the paper makes a useful contribution by decoupling semantic planning from spatial execution: the explicit bounding-box sequence is a clean way to inject spatial priors, and the modular two-stage training is a plausible route to reducing aligned-data requirements. The paper also includes a rare and valuable decomposition of the planner, mask branch, and editor, with clear ablations. However, the current evaluation does not yet support the central SOTA claims: the main evidence for physical plausibility and spatial relations comes from an unvalidated Gemini judge, the temporal-consistency numbers contradict the broad 'outperforms' wording, and the mask-branch ablation shows a non-monotonic physical-plausibility pattern. These issues are fixable with additional experiments and corrected reporting, but they are load-bearing for the paper's headline conclusions.","major_comments":[{"comment":"The headline claims of physically consistent object additions and precise spatial localization rest entirely on Gemini ratings (Physical Rule, Spatial Relation, Instruction Following, Editing Quality), but the manuscript reports neither the Gemini prompt template nor per-item scores, variance, significance tests, or correlation with human judgments. Since the CoT planner is an MLLM and Gemini is also used as a judge, the large margins over OmniVideo and Lucy-1.1 (e.g., Physical Rule 0.741 vs 0.590; Spatial Relation 0.841 vs 0.641) could reflect a systematic preference for MLLM-style structured outputs rather than genuine physical or spatial superiority. An independent human evaluation, or at least a second non-MLLM judge with the full rating protocol, is needed to support the SOTA claim.","section":"Section 4.2, Table 1"},{"comment":"The statement that CoT-Edit consistently outperforms all baselines is contradicted by the reported Temporal Consistency: Ours TC=0.945 is below InsV2V (0.958), InsViE (0.957), OmniVideo (0.954), and Lucy-1.1 (0.961). The abstract and conclusion also emphasize temporal coherence, so the SOTA claim must be qualified to exclude temporal consistency, or the discrepancy must be explained.","section":"Section 4.2, Table 1; Abstract; Conclusion"},{"comment":"The mask-branch ablation is not monotonic in physical plausibility: adding the Mask-Connector lowers Physical Rule from 0.674 (E w/ MLLM) to 0.643 (E+M w/ Mc), and the full Reverse-Connector restores it only to 0.681. The paper explains spatial gains but does not discuss this physical-plausibility regression; since physical consistency is one of the two central differentiators, this pattern needs analysis and not just a summary statement that the mask branch enhances spatial understanding.","section":"Section 4.3, Table 2"},{"comment":"The main quantitative comparison is based on 100 randomly sampled Koala-36M videos, with no confidence intervals, no significance tests, and no description of the instruction distribution. The user study is summarized in one sentence with no participant count, protocol, or statistics, and the full details are deferred to a supplementary document that was not provided for review. This level of reporting is insufficient for state-of-the-art claims.","section":"Section 4.1 and Section 4.3"},{"comment":"The internal training dataset of about 100k video editing pairs with precise mask annotations is neither released nor described in detail, so the reader cannot disentangle the contribution of the proposed architecture from the contribution of this unshared data. At minimum, the dataset composition, filtering criteria, and instruction types should be reported so that the reduced-data-dependency claim can be evaluated.","section":"Section 4.1"}],"minor_comments":[{"comment":"There is a spacing typo 'V AEs' and reference [23] appears unrelated to video generation; it should be removed or replaced.","section":"Section 2.1"},{"comment":"The figure caption uses 'VLM planner' while the text uses 'MLLM planner'; the terminology should be unified.","section":"Figure 2 and Section 3.1"},{"comment":"The relationship between the keyframe-aligned bounding box sequence and the full-frame mask sequence is not specified; clarify how sparse keyframe boxes are propagated to all frames and how the planner's empty-box output for non-spatial tasks is handled by the mask branch.","section":"Section 3.1"},{"comment":"The notation 'QVLcrossattn(MLP(V), C^M_l)' is ambiguous: the text says cross-attention is 'modulated by mask features', but the equation does not make the modulation operation explicit; please define the exact mechanism and dimensional transformations.","section":"Section 3.3, Equation (4)"},{"comment":"The Lucy-1.1 row lacks separators between FVD, TC, and MS values ('1488.120.9610.98'), which makes the table hard to read; fix the formatting.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is within scope and the architecture is interesting, but the evaluation is not yet at the level required for a SOTA claim. I would not reject because the issues are addressable: the authors should provide the Gemini evaluation prompt and per-case breakdown, add an independent human or non-MLLM judge, correct the temporal-consistency claim, and analyze the non-monotonic mask-branch ablation. If those are supplied, the paper could become publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is real: instead of asking a diffusion editor to resolve ambiguous instructions from text alone, the paper inserts a CoT-based MLLM planner that outputs explicit bounding boxes and enriched instructions, feeds those into a box-conditioned mask branch, and connects mask and editor features bidirectionally. That decomposition is a sensible way to attack the where-to-edit problem, especially for scenes with multiple similar objects and for object additions that need physical placement. The modular training stages are also reasonable, and the qualitative examples in Figure 4 and the CoT ablation actually show the mechanism working: the planner's boxes clearly help the ping-pong ball trajectory and the dog-swap localization. This is not a restatement of earlier work; the architecture is a new combination and the paper explains it clearly.\n\nThe soft spots are concentrated in the evaluation. First, the paper says it “consistently outperforms” baselines, but Table 1 shows its Temporal Consistency (0.945) is the worst among all compared methods, including InsV2V, InsViE, OmniVideo, and Lucy-1.1. That is a direct overclaim. Second, and more load-bearing, the big advantages in Physical Rule, Spatial Relation, Instruction Following, and Editing Quality all come from Gemini ratings. The planner itself is an MLLM, so a Gemini judge may systematically prefer outputs that look like MLLM-style structured reasoning rather than genuinely better physics. No prompt template, per-item scores, variance, significance, or correlation with human judgments is reported. The user study is mentioned but no protocols or statistics are given, so it does not fill that gap. This is the main thing I would want fixed before trusting the numbers.\n\nMinor issues: no released code or data despite the GitHub link, and the reference list contains an obvious error — entry [23] is a health-labor-market paper sitting in a paragraph about video generation models. That kind of mistake makes me a bit wary, though it does not affect the technical content.\n\nOn balance, the architecture holds up; the problem is the empirical support, not the design logic. The paper deserves a serious referee, but I would not accept it without external validation of the Gemini scores — blinded human or independent-judge comparison, or at minimum a human correlation study — and a corrected TC claim.\n\nThe reader's take is about right. The paper is worth a place in a reading group and worth citing once the evaluation is tightened, but the current SOTA language is stronger than the evidence.","headline":"A genuinely new plan-guide-edit architecture for instruction video editing, with strong qualitative promise, but the headline SOTA claims lean on an unvalidated Gemini judge and one overclaim on temporal consistency.","tokens_in":12758,"tokens_out":961,"would_cite":true,"duration_ms":10488,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CoT-Edit claims that inserting a chain-of-thought multimodal planner that emits bounding boxes before diffusion editing fixes target ambiguity and physical implausibility in instruction-based video editing.","keywords":["instruction-based video editing","Chain-of-Thought reasoning","multimodal large language models","bounding box grounding","diffusion models","video editing","spatial localization","physical plausibility"],"falsifier":"A human perceptual study on the same 100 videos sampled from Koala-36M, where raters choose which edited video selects the intended object and obeys the stated physical motion, would settle the central claim. If CoT-Edit does not beat the best baseline under human judgment on these two dimensions, the reported advantage is an artifact of the judge model rather than a true gain.","tokens_in":11874,"feed_emoji":"🎯","tokens_out":8822,"duration_ms":72503,"temperature":0.7,"pith_summary":"Text-only instruction-based video editing struggles when a scene has several similar objects or when an instruction implies motion, because the model has to guess both what to edit and where. This paper argues that the \"where\" should be decided explicitly before generation begins. A chain-of-thought multimodal large language model examines the video keyframes and the instruction, and outputs a temporal sequence of bounding boxes plus an enriched instruction carrying attributes, spatial relations, and physical constraints. A box-guided branch converts those anchors into masks, and a diffusion editor applies the edit inside the masks while preserving the rest of the video. If the reported results hold, this ordering—plan, guide, edit—is what makes instruction-based video editing localize correctly among similar objects and produce physically plausible additions.","feed_headline":"CoT-Edit: video editing that finds the right object and obeys physics","feed_subtitle":"A chain-of-thought planner adds spatial anchors, so edits hit the intended object and added objects move plausibly.","key_machinery":"The central mechanism is the chain-of-thought enhanced multimodal large language model planner, which acts as a translator from a text instruction to spatial anchors. It decomposes the task into parsing the editing intent, identifying the target across frames, separating camera motion from object motion, and checking physical and cinematic consistency, then emits a keyframe-aligned bounding box sequence and an enriched instruction. The bounding boxes turn mask prediction from an open-ended global search into a local refinement inside known regions, which is what allows precise selection among similar objects. The enriched instruction carries attributes, contact, and motion priors into both the mask branch and the diffusion editor. Two bidirectional connectors—one feeding editor semantics back into the mask branch and one injecting mask features into the editor—couple the two stages so that the mask stays accurate at thin or occluded structures while the editor receives spatial guidance at multiple depths.","core_discovery":"The central claim is that instruction-based video editing becomes controllable when an explicit planning step converts a text instruction into spatial anchors before any pixel is generated. In this design, a chain-of-thought multimodal large language model examines the video keyframes and the user instruction, and produces two things: a temporal sequence of bounding boxes and an enriched instruction that carries attributes, spatial relations, and physical constraints. A mask branch then turns those boxes into spatiotemporal masks, and a diffusion editor fuses masks, enriched text, and video features to produce the final edit. The paper reports that this ordering achieves the best numbers on FVD, CLIPScore, and judge-model-rated physical plausibility, spatial relations, instruction following, and editing quality against six open-source baselines, with the largest margins in spatial relation and physical-rule scores. It also reports that ablations removing the chain-of-thought step lose physically plausible motion, such as a ball following a parabolic bounce.","pith_inferences":["A direct transfer test would apply the same plan-guide-edit decomposition to instruction-based image editing, where a bounding-box plan should resolve ambiguity among similar objects with minimal changes to the mask and editor branches.","The keyframe-aligned box sequence suggests a natural extension to long videos: sample more keyframes or interpolate box trajectories between them; the paper's experiments do not report how performance scales with sequence length.","Because the judge model used for the main scores may systematically prefer outputs that resemble multimodal-LLM planning, re-scoring the same outputs with a judge that sees no planning text, or with human raters, would separate genuine spatial and physical gains from judge bias.","The data-efficiency claim implies an experiment: train the full model at several fractions of the 100k internal pairs; if the planner is the load-bearing component, spatial scores should degrade more slowly than appearance scores as data shrink."],"forward_implications":["Instruction-based editing no longer needs to solve implicit global localization; the planner's bounding boxes restrict the edit to the intended region, which is why scenes with multiple similar objects become tractable.","Physically constrained additions, such as an object following a parabolic trajectory, can be specified through the text and realized through the box sequence, because the boxes encode a spatiotemporal path rather than a single location.","The modular training schedule—separate mask and editor training followed by joint fine-tuning on 100k pairs—reduces the need for large aligned video-instruction datasets.","The framework covers non-spatial tasks such as stylization by letting the planner emit an empty box sequence, so the enriched instruction alone drives the edit.","Because edited content is composited inside the masks and original video is preserved outside, background and temporal consistency are maintained even for localized additions."],"supporting_citations":[{"why":"Defines the instruction-conditioned image editing setup that this work extends to video.","marker":"[2]"},{"why":"Introduces instruction-based video editing and illustrates the data bottleneck that staged training is meant to avoid.","marker":"[29]"},{"why":"Provides a scaled high-quality synthetic dataset used in training the editor branch.","marker":"[1]"},{"why":"Baseline that enforces geometric consistency via explicit alignment; contrast for instruction-driven spatial grounding.","marker":"[22]"},{"why":"MLLM-guided editing baseline relying on implicit attention; shows why explicit boxes are proposed.","marker":"[27]"},{"why":"Closest baseline that also combines a multimodal LLM with diffusion; comparison isolates the value of explicit planning.","marker":"[33]"},{"why":"End-to-end diffusion baseline that lacks explicit spatial grounding; serves as the main contrast for localization claims.","marker":"[34]"},{"why":"Diffusion-only baseline with a large dataset; a strong instruction-injection contrast for the reported gains.","marker":"[45]"},{"why":"Judge model that produces the physical-plausibility, spatial-relation, instruction-following, and editing-quality scores in Table 1.","marker":"[35]"},{"why":"Source of the 100 evaluation videos and instructions used for the main quantitative comparison.","marker":"[39]"}],"fun_headline_variants":["CoT-Edit: CoT plans before it edits, so edits obey physics","Chain-of-thought anchors make video edits physically plausible","Plan-then-edit: CoT spatial anchors for instruction video editing","CoT-Edit: Spatial reasoning before pixels, better video edits","CoT-guided editing: bounding boxes first, then physics-faithful edits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported advantage over baselines depends on the judge model's scores being a fair measure of physical plausibility and spatial relations; if the judge simply favors outputs that look like multimodal-LLM planning, the large gains on those dimensions could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["CoT-Edit: CoT plans before it edits, so edits obey physics","Chain-of-thought anchors make video edits physically plausible","Plan-then-edit: CoT spatial anchors for instruction video editing","CoT-Edit: Spatial reasoning before pixels, better video edits","CoT-guided editing: bounding boxes first, then physics-faithful edits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000627,"raw_usage":{"total_tokens":2911,"prompt_tokens":967,"completion_tokens":1944,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":1850}},"tokens_in":583,"tokens_out":1944,"duration_ms":10988,"temperature":1.0,"reasoning_tokens":1850,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:11:20.107258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human perceptual study on the same 100 videos sampled from Koala-36M, where raters choose which edited video selects the intended object and obeys the stated physical motion, would settle the central claim. If CoT-Edit does not beat the best baseline under human judgment on these two dimensions, the reported advantage is an artifact of the judge model rather than a true gain.","supporting_citations":[{"cited_title":"In- structpix2pix: Learning to follow image editing instructions","cited_arxiv_id":null,"evidence_quote":"Defines the instruction-conditioned image editing setup that this work extends to video."},{"cited_title":"Instructvid2vid: Controllable video editing with natural language instructions","cited_arxiv_id":null,"evidence_quote":"Introduces instruction-based video editing and illustrates the data bottleneck that staged training is meant to avoid."},{"cited_title":"Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025","cited_arxiv_id":null,"evidence_quote":"Baseline that enforces geometric consistency via explicit alignment; contrast for instruction-driven spatial grounding."},{"cited_title":"Lucy edit: Open-weight text-guided video editing, 2025","cited_arxiv_id":null,"evidence_quote":"End-to-end diffusion baseline that lacks explicit spatial grounding; serves as the main contrast for localization claims."},{"cited_title":"Insvie-1m: Effective instruction-based video editing with elaborate dataset construction","cited_arxiv_id":null,"evidence_quote":"Diffusion-only baseline with a large dataset; a strong instruction-injection contrast for the reported gains."},{"cited_title":"Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content","cited_arxiv_id":null,"evidence_quote":"Source of the 100 evaluation videos and instructions used for the main quantitative comparison."}],"review_version":1}