{"id":"d76f1615-b0ea-4fcb-ab43-f97dcdf382d3","arxiv_id":"2608.05049","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new benchmark and evaluation framework for instruction-based video editing that covers spatial, temporal, audio, reference, and reasoning edits, with an accuracy-aware penalty to prevent inflated scores for incorrect edits.","lead":"OmniEdit-Bench is a new benchmark that tests video editing models on five task types, including temporal, audio, and reasoning edits, not just simple object changes. It also introduces a scoring method that reduces scores for visually nice videos that ignore the user's instruction.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The only human-alignment evidence is pooled across tracks; because the headline conclusions are about temporal, audio, and reasoning, the VLM's reliability on exactly those tracks is unverified.","rationale":"Reader's weakest assumption points to VLM-human agreement; I agree that is the load-bearing point, but sharpen it: the missing per-track stratification is what makes the validation uninformative for the paper's own headline. The paper is otherwise careful about taxonomy and prompt design, and the accuracy-aware penalty is a reasonable design choice even if weights are arbitrary. The empirical finding that all models score much lower on temporal/audio/reasoning than spatial may well survive human inspection; the issue is evidentiary. The conditional verdict is right. If the authors release track-level human data and it confirms VLM calibration, the benchmark claim would be substantially stronger; if not, the 'reliable' descriptor should stay conditional. I therefore leave the reader's verdict unchanged.","tokens_in":13247,"tokens_out":5277,"duration_ms":55162,"concrete_test":"Re-compute Table 3 stratified by the five tracks: report N per track, per-score-level human mean/std, MAE, and annotation protocol (number of raters, instructions, and whether human raters saw original and edited video). Accept the reliability claim only if temporal, audio, and reasoning tracks have N >= 30 per track, retain the monotonic human-mean trend across VLM levels, and have track-level MAE not substantially above the pooled value. If these data are unavailable, the headline conclusions about video-specific tracks should be downgraded to provisional.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that OmniEdit-Bench is a reliable testbed rests on VLM scores matching human judgment. Section 4.4 / Table 3 is the only supporting evidence, and it pools all tracks and models, reporting aggregate MAE (0.86 accuracy, 0.77 preservation, 0.55 realism, 0.64 consistency) without sample sizes, per-track breakdowns, annotation protocol, or rater agreement. The paper's main conclusions are not about spatial edits; they are that models are far from satisfactory on the temporal, audio, and reasoning tracks. Those are precisely the tracks where a VLM's judgment is least established, and Appendix D concedes the evaluation 'may not fully capture subtle subjective preferences or nuanced editing quality in borderline cases.' The pooled monotonic trend in Table 3 could be driven by spatial items, while temporal/audio/reasoning items might be mis-scored. Thus the 'reliable and comprehensive testbed' claim is conditional on a track-level validation that the paper does not provide. This does not make the results false; it makes them unsupported as reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OmniEdit-Bench, a benchmark for instruction-based video editing (IVE) that decomposes editing tasks into five tracks—spatial, temporal, audio, reference-based, and reasoning—and evaluates models along four dimensions (accuracy, preservation, realism, and consistency) using a VLM (Gemini-3.1-Pro) supplemented by human judgments. An accuracy-aware penalty mechanism scales the non-accuracy dimensions by normalized accuracy to prevent plausible-but-incorrect edits from being scored highly. The authors evaluate eight open-source and commercial models on the benchmark and report that all models perform far better on spatial edits than on temporal, audio, and reasoning tasks, with commercial models generally ahead on spatial and reference tracks. The paper argues that OmniEdit-Bench is a comprehensive and reliable testbed for IVE and provides a direction for future research.","tokens_in":13478,"tokens_out":6442,"duration_ms":57523,"significance":"If the benchmark and its evaluation protocol are validated, this work would fill a real gap: existing IVE benchmarks mostly inherit image-editing tasks and lack video-specific dimensions such as temporal dynamics, audio, and reasoning. The taxonomy in Fig. 2 is thoughtfully organized, and the accuracy-aware penalty mechanism is a principled response to the known failure of similarity-based metrics to reward instruction fidelity. The evaluation covers both open-source and commercial models, and the inclusion of human-alignment analysis is a strength. However, the paper's central claim of being a 'reliable and comprehensive testbed' rests on the VLM-based scores being trustworthy on the exact tracks where the headline conclusions are negative (temporal, audio, reasoning). The evidence for that reliability is thinner than the text implies, and several reproducibility issues in the reported tables need to be addressed before the claims can be accepted as stated.","major_comments":[{"comment":"The human-alignment evidence is pooled across tracks and lacks the information needed to support the paper's track-specific conclusions. Table 3 reports aggregate MAE values (0.86 accuracy, 0.77 preservation, 0.55 realism, 0.64 consistency) but does not give sample sizes, per-track breakdowns, annotation protocols, or rater agreement. Since the headline finding is that models are far from satisfactory on the temporal, audio, and reasoning tracks, and Appendix D concedes the VLM 'may not fully capture subtle subjective preferences or nuanced editing quality in borderline cases,' the paper needs to show that VLM scores align with human judgments on those specific tracks. Without this, the reliability of the benchmark for the tracks that matter most is unverified.","section":"§4.4 / Table 3"},{"comment":"The table is internally inconsistent: the proportions column is identical for all four dimensions, yet the dimensions have different score distributions, and the overall MAE values do not match the weighted average of the per-level MAEs. For example, weighting realism MAE by the reported proportions gives approximately 0.62, not 0.55, and consistency gives approximately 0.96, not 0.64. The authors should clarify how the proportions and overall MAE are computed, and either provide per-dimension distributions or explain why a single proportion column is appropriate.","section":"Table 3"},{"comment":"The 'Overall Score (Average)' column is not reproducible from the track scores and the stated protocol. For instance, Runway Aleph has available track scores 54.2, 16.5, 15.1, and 20.3, whose arithmetic mean is 26.5, yet the table reports 24.2; Grok Imagine has mean 20.9 but reports 19.0; KlingV3-Omni has mean 37.7 but reports 38.3; and UniVideo has mean 16.1 but reports 14.4. The averaging rule for missing tracks (audio and reference) must be specified exactly, or the overall scores will be misleading.","section":"Table 2"},{"comment":"The final-score formula is not fully specified for the reported 100-point scale. The formula Score = 0.5A + 0.2P' + 0.15R' + 0.15C' yields a maximum of 5 when A=P=R=C=5, yet Table 2 reports scores on a 100-point scale. The paper should state the normalization or rescaling step (e.g., multiplication by 20) explicitly. Additionally, the weights 0.5/0.2/0.15/0.15 are presented without justification or sensitivity analysis; since the weights affect all cross-model comparisons, a brief robustness check would strengthen the claims.","section":"§3.2 / Fig. 5"},{"comment":"The model comparison is not apples-to-apples because different models are evaluated on different subsets of tracks. For example, audio is marked as unsupported for several models, and the reference track is only evaluated for models with spatial reference capabilities, while Seedance 2.0 is excluded from audio and reference entirely. The paper should clarify whether the overall scores are comparable given these missing entries, and should avoid drawing fine-grained ranking conclusions from overall averages over differing track sets.","section":"§4.2 / Table 2"},{"comment":"The paper motivates the accuracy-aware penalty by claiming that existing metrics allow incorrect edits to receive high scores due to strong visual priors, but it never empirically compares its proposed metric against these existing metrics (e.g., CLIP-based similarity, SSIM, or prior benchmark protocols). This is a testable claim: the authors could show that their metric down-weights specific failure cases where existing metrics give high scores. Without such a comparison, the claimed superiority of the evaluation framework over prior metrics remains unsubstantiated.","section":"§2 / §1"}],"minor_comments":[{"comment":"The text contains a typo: 'This category focuses on more general objects inluding animals and vehicles' should read 'including.'","section":"§3.1 / Audio Track"},{"comment":"Figure 1 lists representative tasks but does not clearly indicate which are explicit versus implicit instructions; adding a legend or color coding would help readers map the instructions to the two axes described in the text.","section":"§1 / Fig. 1"},{"comment":"The legend for Fig. 3 includes 'N/A or unsupported,' but it is unclear from the figure which bars correspond to this category; the figure should mark unsupported tracks explicitly or explain the convention in the caption.","section":"Fig. 3"},{"comment":"The paper does not specify the API access date, model version details, or inference settings for the commercial models (e.g., KlingV3-Omni, Runway Aleph, Grok Imagine), which are needed for reproducibility of the benchmark results.","section":"§4.1"},{"comment":"The human evaluation is described only as 'human annotations' without saying how many annotators rated each sample, whether they were expert or crowd annotators, or what instructions they received; these details are necessary to assess the reliability of the human ground truth.","section":"§4.4 / Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a computer vision venue, and the benchmark has clear potential value. However, I have two non-technical concerns for the editor. First, several co-authors are affiliated with the Wan team, and Wan2.7-Edit is one of the evaluated models; while this is not a logical circularity, it should be disclosed prominently to avoid any perception of bias in the model comparisons. Second, the manuscript should be checked for completeness: the project page is mentioned, but no link to code or benchmark data is provided in the paper, which limits reproducibility. These issues do not change my technical recommendation but should be considered during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look if you work on video editing evaluation. Its five-track taxonomy — spatial, temporal, audio, reference, reasoning — is a real step beyond existing benchmarks, which mostly stop at spatial or single-dimension shots. The accuracy-aware penalty, where preservation/realism/consistency are gated by accuracy, is a sensible fix for the known failure mode where visually plausible but wrong edits get high scores. And the model sweep (nine models, both commercial and open) gives a useful first map of where the field stands: spatial is comparatively healthy, temporal/audio/reasoning are weak. That qualitative finding is probably robust.\n\nThe soft spots are real but not fatal. The main one is exactly what your stress-test note flags: Table 3 pools all tracks and models into one alignment table, with no sample sizes, no per-track breakdowns, and no rater protocol. The paper's headline conclusions are about temporal, audio, and reasoning, and those are precisely the tracks where a VLM's judgment is least established. Appendix D even concedes the VLM 'may not fully capture subtle subjective preferences or nuanced editing quality in borderline cases.' So the 'reliable and comprehensive testbed' claim is conditional on track-level validation the paper doesn't report. That doesn't make the results false, but it does mean the central reliability claim is thinner than it should be.\n\nThe other soft spots are minor-to-moderate. The score weights (0.5/0.2/0.15/0.15) and the gating mechanism are stated without sensitivity analysis; changing them could shift rankings. There's no comparison to existing metrics (CLIPScore, etc.), so 'comprehensive' isn't empirically demonstrated against alternatives. And no artifacts are released — no data, prompts, or code — which limits immediate adoption despite the project page. Minor but worth noting: several co-authors are from the Wan team, and Wan2.7-Edit is an evaluated model; that's not a circularity, but it should be disclosed in the paper itself.\n\nBottom line: this is a solid, useful benchmark paper that needs a revision. A serious editor should send it to peer review, with reviewers asked to demand per-track human validation, artifact release, and a comparison to existing metrics. I'd bring it to reading group and would cite it once the validation is airtight.\n\nRecommendation: accept for review, conditional on revision.","headline":"The five-track taxonomy and accuracy-aware penalty are a genuine step forward for IVE evaluation, but the reliability claim rests on a single pooled VLM-human alignment table that is thinner than the paper's headline conclusions need it to be.","tokens_in":13966,"tokens_out":2551,"would_cite":true,"duration_ms":24951,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Instruction-based video editing models, including commercial systems, do well only on spatial tweaks and fall far short on temporal, audio, and reasoning edits, according to a new five-track benchmark that gates all quality scores on…","keywords":["instruction-based video editing","video editing benchmark","temporal editing","audio editing","reasoning-based editing","reference-based editing","accuracy-aware evaluation","vision-language model evaluation"],"falsifier":"Take a stratified random sample of edited videos from the temporal, audio, and reasoning tracks, have independent human raters score them with the same four-dimension rubric, and compare per-track mean absolute error and rank correlation with the VLM scores; if agreement on those tracks is substantially worse than the paper's aggregate MAE of 0.55–0.86, the claimed model weaknesses would not be established.","tokens_in":13093,"feed_emoji":"🎬","tokens_out":11587,"duration_ms":106580,"temperature":0.7,"pith_summary":"Instruction-based video editing lets a user change a video by typing an instruction, but judging whether the edit was actually done has been unreliable. This paper tries to establish a benchmark that measures what editing models can and cannot do across the dimensions that matter for video, not just for static images. It organizes tasks into five tracks — spatial, temporal, audio, reference-based, and reasoning — and scores each edited video on accuracy, preservation, realism, and consistency, with accuracy acting as a gate that shrinks all other scores when the instruction is not followed. Evaluated on this testbed, both open-source and commercial models do reasonably well only on spatial edits and drop sharply on temporal, audio, and reasoning tasks. A sympathetic reader would take the paper's point: the field's progress so far is real but narrow, and future editing systems need to handle motion, sound, and implicit instruction before they are usable in realistic settings.","feed_headline":"Benchmark: video editors fail temporal, audio, reasoning","feed_subtitle":"An accuracy-gated five-track test finds even top models struggle with motion, sound, and reasoning.","key_machinery":"The load-bearing mechanism is the accuracy-aware penalty inside the evaluation pipeline. For accuracy A, preservation P, realism R, and consistency C (each scored 1–5), the final score is 0.5A + 0.2(P × A/5) + 0.15(R × A/5) + 0.15(C × A/5), which makes accuracy a multiplicative gate: if the edit does not follow the instruction, every other dimension shrinks and cannot compensate. The second structural component is the five-track taxonomy — spatial, temporal, audio, reference, reasoning — because it defines what accuracy means in each track and supplies the track-specific evaluation prompts.","core_discovery":"The central claim is that instruction-based video editing is a multidimensional task and that current models are far from satisfactory when measured that way. The paper builds a benchmark of 790 editing tasks distributed over five tracks, and a unified evaluation pipeline in which a vision-language model rates each output on four dimensions with track-specific prompts. Accuracy is treated as the primary signal: the scores for preservation, realism, and consistency are multiplied by accuracy divided by five before aggregation, so a visually plausible but instruction-violating edit cannot receive an inflated grade. The reported results show commercial systems leading on spatial editing, most models scoring below 20 on the temporal track, all models remaining below 30 on reasoning, audio editing largely unsupported or weak, and reference-conditioned editing strong only for models with dedicated reference capabilities. The paper interprets this as evidence that larger datasets and bigger models improve spatial manipulation but do not solve temporal coherence, multimodal alignment, or implicit reasoning.","pith_inferences":["The paper does not test whether the accuracy gate makes preservation, realism, and consistency scores too compressed to distinguish among moderately inaccurate edits; a separate binary 'instruction satisfied' gate plus a graded accuracy score might be more informative.","A natural extension is to use the benchmark's per-track failures as a reward signal for training editing models, with accuracy as a hard constraint and the other dimensions as soft objectives.","The reasoning track suggests a divide-and-conquer architecture the paper does not evaluate: an external reasoner that turns implicit instructions into explicit edit plans, followed by a video editor that executes those plans, may outperform end-to-end systems.","The audio track could double as a testbed for audio-visual synchronization research, since it explicitly requires lip-sync, event-sound alignment, and scene-consistent ambient audio rather than frame-level fidelity."],"forward_implications":["If the findings hold, users should not expect current commercial video editors to handle temporal composition, motion-semantic changes, or counterfactual reasoning; their strong spatial scores do not generalize.","Audio editing is a demonstrated gap, so progress will require models that generate and edit synchronized audio jointly with video rather than treating audio as an afterthought.","Implicit and reasoning-based instructions are largely unsolved, implying that practical editing systems will need explicit planning or reasoning components before generation.","The accuracy-aware penalty offers a simple way to make automatic evaluation respect instruction fidelity and could be reused by other video and image editing benchmarks.","VLM-human agreement on the four dimensions (overall mean absolute error between 0.55 and 0.86) suggests that scalable automatic evaluation can stand in for human annotation in large-scale editing benchmarks."],"supporting_citations":[{"why":"A prior video editing benchmark whose spatial-only task coverage motivates the need for temporal and audio tracks.","marker":"[36]"},{"why":"An earlier benchmark with tasks inherited from image editing, used as the comparison baseline for narrow coverage.","marker":"[6]"},{"why":"A current instruction-guided video editing benchmark whose limitations in video-specific dimensions frame the paper's proposal.","marker":"[5]"},{"why":"Previous work on reference-conditioned editing that the reference track builds on and compares against.","marker":"[26]"},{"why":"Earlier reasoning-involved editing benchmark that motivates the reasoning track.","marker":"[24]"},{"why":"A reasoning-informed visual editing benchmark that frames the implicit-instruction setting.","marker":"[46]"},{"why":"A reference-free metric commonly used for generation quality that the paper argues fails to measure instruction fidelity.","marker":"[12]"},{"why":"A perceptual similarity metric used by editing models that is not aligned with whether the instruction was followed.","marker":"[44]"},{"why":"A commercial model with reference-conditioning capability whose reference-track performance anchors the main results.","marker":"[38]"},{"why":"A representative commercial video editing model evaluated in the benchmark's model comparison.","marker":"[31]"}],"fun_headline_variants":["Video editors flunk audio, temporal tests in new benchmark","New benchmark: AI video editors can't handle time, sound, or logic","OmniEdit-Bench: video editors fail temporal, audio, reasoning","Benchmark: video editors flunk temporal, audio, and reasoning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's conclusions rest on the assumption that the vision-language model's 1-to-5 ratings, produced with the paper's prompts, match human judgments across all five tracks and all four dimensions, including temporal timing, audio-visual sync, and causal reasoning.","fun_headline_variants_meta":{"raw":{"variants":["Video editors flunk audio, temporal tests in new benchmark","New benchmark: AI video editors can't handle time, sound, or logic","OmniEdit-Bench: video editors fail temporal, audio, reasoning","Benchmark: video editors flunk temporal, audio, and reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001089,"raw_usage":{"total_tokens":4551,"prompt_tokens":949,"completion_tokens":3602,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":3526}},"tokens_in":565,"tokens_out":3602,"duration_ms":25188,"temperature":1.0,"reasoning_tokens":3526,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:29:43.024546+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of edited videos from the temporal, audio, and reasoning tracks, have independent human raters score them with the same four-dimension rubric, and compare per-track mean absolute error and rank correlation with the VLM scores; if agreement on those tracks is substantially worse than the paper's aggregate MAE of 0.55–0.86, the claimed model weaknesses would not be established.","supporting_citations":[{"cited_title":"Ve-bench: subjective-aligned benchmark suite for text-driven video editing quality assessment","cited_arxiv_id":null,"evidence_quote":"A prior video editing benchmark whose spatial-only task coverage motivates the need for temporal and audio tracks."},{"cited_title":"Editboard: Towards a comprehensive evaluation benchmark for text-based video editing models","cited_arxiv_id":null,"evidence_quote":"An earlier benchmark with tasks inherited from image editing, used as the comparison baseline for narrow coverage."},{"cited_title":"Revise: Towards reason-informed video editing in unified models with self-reflective learning.arXiv preprint arXiv:2512.09924, 2025","cited_arxiv_id":null,"evidence_quote":"Earlier reasoning-involved editing benchmark that motivates the reasoning track."},{"cited_title":"Clipscore: A reference-free evaluation metric for image captioning","cited_arxiv_id":null,"evidence_quote":"A reference-free metric commonly used for generation quality that the paper argues fails to measure instruction fidelity."},{"cited_title":"Runway aleph.https://runwayml.com, 2024","cited_arxiv_id":null,"evidence_quote":"A representative commercial video editing model evaluated in the benchmark's model comparison."}],"review_version":1}