{"id":"a649ab26-d558-4c51-9e17-673813952a98","arxiv_id":"2506.08003","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"MTV splits audio into speech, effects, and music to separately drive lip sync, event timing, and visual mood in video generation, trained on a new 392K-clip dataset.","lead":"MTV turns audio tracks for speech, sound effects, and music into three separate controls for generating video, synchronizing lip motion, event timing, and visual mood. The paper also introduces a 392K-clip cinematic dataset and claims state-of-the-art results, but the evaluation may be unfair because baseline models were not retrained on the new data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOTA claim rests on an unfair comparison: MTV trains on 392K DEMIX clips with demixed inputs, baselines do not, and 50 videos without error bars cannot support the headline.","rationale":"The reader's weakest assumption identifies the same load-bearing issue I find: the comparison in Sec. 5.1 does not control for training data or conditioning format. MTV is trained on the same distribution from which the test videos are drawn, with additional demixed audio streams that the baselines were not designed to consume. This makes the headline SOTA claim unsupported as reported. I agree with the reader's verdict of REJECT. A fair evaluation could potentially salvage the claim, but no code or data were available to verify the experiments, and the small evaluation set without error bars compounds the problem. My concrete test would settle the fairness concern directly by equalizing training conditions. I do not see an internal inconsistency in the method itself; the weakness is in the evidence for the central claim.","tokens_in":13550,"tokens_out":3565,"duration_ms":47796,"concrete_test":"Fine-tune TempoTokens and Xing et al. on the same DEMIX 392K training clips with the same text captions and demixed audio inputs, using comparable compute and resolution, then re-run Table 2 on the full 1K held-out set with bootstrap confidence intervals. If their FVD, Audio-C, and Sync-C gaps collapse or shrink below significance, the reported SOTA is a training-data/domain advantage rather than a property of the method. If MTV still wins by similar margins, the fairness concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MTV achieves state-of-the-art performance on six metrics. In Sec. 5.1 (Table 2), MTV is trained on 392K DEMIX clips with demixed speech/effects/music conditioning and structured text captions, while TempoTokens and Xing et al. are evaluated in their original configuration and MM-Diffusion is only finetuned, not trained on DEMIX. No baseline receives the same training distribution or the same three-stream demixed conditioning. Because the 50 evaluation videos are sampled from the DEMIX test distribution, MTV's large margins (Audio-C 26.22 vs 7.30, Sync-C 3.17 vs 1.55) may reflect a training-data and conditioning-format advantage rather than the MST-ControlNet design. The absence of error bars or significance tests on 50 videos makes it impossible to separate method quality from noise. The paper itself notes that AV-Align is unsuitable because real videos score lowest (Table 3), further weakening confidence in the alignment metric set, but the primary load-bearing gap is experimental fairness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MTV, an audio-sync video generation framework that demixes input audio into speech, effects, and music tracks and uses a proposed Multi-Stream Temporal ControlNet (MST-ControlNet) to control lip motion, event timing, and visual mood respectively. It also introduces DEMIX, a dataset of 392K cinematic video clips with demixed audio tracks, structured into five overlapped subsets for a multi-stage training strategy. The paper claims state-of-the-art performance across six metrics (FVD, Temp-C, Text-C, Audio-C, Sync-C, Sync-D) in comparison with three recent baselines (MM-Diffusion, TempoTokens, Xing et al.), and reports ablations supporting the value of the proposed modules. The central claim is that MTV achieves state-of-the-art audio-sync video generation.","tokens_in":13781,"tokens_out":4498,"duration_ms":53346,"significance":"If the results are established, the DEMIX dataset is a substantial new resource for audio-sync video generation, and the architectural idea of separately controlling lip motion, event timing, and visual mood via demixed audio tracks is a plausible and potentially useful direction. The paper also ships a clear multi-stage training strategy and quantitative ablations with named components. However, the experimental evidence as presented is not sufficient to support the headline state-of-the-art claim because the baseline comparison is confounded by training data and conditioning, and the evaluation set is too small to support reliable conclusions. The strengths of the dataset and architecture are real, but the core empirical claim requires a controlled comparison and more rigorous statistics.","major_comments":[{"comment":"The comparison with baselines is not a fair test of method quality. MTV is trained on 392K DEMIX video clips with demixed audio tracks and structured text captions, while TempoTokens and Xing et al. are evaluated in their original off-the-shelf configuration, and MM-Diffusion is only finetuned without a description of the data used for finetuning. The reported improvements, especially the very large margins on Audio-C (26.22 vs. 7.30) and Sync-C (3.17 vs. 1.55), may therefore reflect a training-data and conditioning-format advantage rather than the MST-ControlNet design. To support the state-of-the-art claim, the authors should either retrain all baselines on DEMIX with the same demixed conditioning and text structure, or evaluate MTV under each baseline's training setup, and report the results of such a controlled comparison.","section":"Sec. 5.1, Table 2"},{"comment":"Only 50 videos are randomly selected from the testing set for evaluation, and no error bars, confidence intervals, or significance tests are reported. Frechet Video Distance computed on 50 samples is known to have high variance, and all six metrics are affected by the small sample. Without repeated evaluation seeds or a larger evaluation set, it is impossible to tell whether the reported differences between methods are statistically meaningful, even if the protocol were otherwise fair. The authors should increase the number of evaluation videos and report mean and standard deviation or confidence intervals.","section":"Sec. 5.1"},{"comment":"The paper states that AV-Align is unsuitable because real videos score lowest with that metric, yet the abstract and Section 5.1 claim state-of-the-art performance across six metrics, including audio-video alignment metrics. If AV-Align is demonstrably broken for this task, that raises a concern about the validity of the other alignment metrics (Sync-C/Sync-D, Audio-C) as well. The authors should either justify why the remaining alignment metrics are reliable, provide an alternative alignment evaluation (e.g., human study or a metric known to correlate with real videos), or temper the claim about alignment performance. As written, the acknowledged failure of one alignment metric weakens confidence in the alignment results.","section":"Sec. 7.2, Table 3"}],"minor_comments":[{"comment":"The dataset statistics say the five subsets are 'overlapped' but then report 'non-overlapped 392K clips'; please clarify whether the total is the union size after deduplication or the sum of subset sizes.","section":"Sec. 3"},{"comment":"The finetuning procedure for MM-Diffusion is not described: it is unclear what data, number of steps, and hyperparameters were used for the 'over 320K steps' finetuning mentioned in the qualitative comparisons. This information is needed to assess whether that baseline was given a comparable training opportunity.","section":"Sec. 5.1"},{"comment":"The ablation section refers to 'Fig. 6' for the ablation results, but the ablation figure appears to be Fig. 4 in the manuscript; please correct the cross-reference.","section":"Sec. 5.2"},{"comment":"'Frechét Video Distance' should be spelled 'Fréchet Video Distance'.","section":"Sec. 7.2"},{"comment":"The notation for the interval stream is under-specified: please define the interval index i and the division of the latent code into intervals, and clarify how hs_i and he_i are combined when the speech and effects intervals may not align.","section":"Eq. (2)"},{"comment":"The limitation section only mentions the scope of audio demixing categories; it would also be appropriate to note that errors in the demixing stage will propagate to the conditioned generation, since all downstream control depends on the quality of the separated tracks.","section":"Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea and dataset contribution are valuable, but the experimental protocol does not currently support the state-of-the-art claim. The authors should be encouraged to perform a controlled comparison with baselines trained on the same DEMIX distribution and to report statistics over a larger evaluation set. I would also note that the dataset release appears to include only samples rather than the full DEMIX dataset, which would limit reproducibility; this is worth clarifying. The paper is likely to be publishable if these issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hey [Name],\n\nThe headline: the demixing-based multi-stream control is genuinely new and DEMIX is a real resource, but the paper's central SOTA claim rests on an unfair comparison and a 50-video evaluation. I'd send it to review, but only with the expectation that the comparison gets redone.\n\nWhat's actually good: The idea of separating speech, effects, and music as distinct conditioning streams is a real step forward for audio-sync generation. The interval stream for lip/event timing and the holistic stream for mood are sensible, and the ablations show each component earns its keep. The DEMIX dataset, with its demixing filtering and five overlapped subsets, is a substantial contribution on its own. The multi-stage training schedule is also well thought out.\n\nWhere it's soft: The experimental protocol is the load-bearing issue. TempoTokens and Xing et al. are run in their original configuration, while MTV is trained on 392K DEMIX clips. MM-Diffusion is only finetuned. So the large gaps on Audio-C and Sync-C likely reflect training-data and conditioning-format advantages as much as the architecture. On top of that, evaluating on 50 videos without error bars is too thin to support a SOTA claim; FVD on 50 clips is noisy. The AV-Align decision is defensible (real videos score lowest), but they should report it properly in the main text rather than relegate it to the supplementary. Also, no code or data release is mentioned at the time of writing, which makes independent verification hard.\n\nThe stress-test note holds up: the central claim is not established by the experiments as reported. That said, I don't think this is a paper to desk-reject. The core idea is novel, the dataset is valuable, and the ablations are informative. With a fair baseline comparison (fine-tune all baselines on DEMIX, or at least standardize the conditioning format), a larger evaluation set with confidence intervals, and artifact release, this could be a solid paper.\n\nWho's it for: people working on audio-conditioned video generation and controllable video synthesis. I'd bring it to the reading group, and I'd cite the dataset and the architecture even if the SOTA claim is questionable.\n\nRecommendation: send to serious peer review, but flag the comparison as a major-revision issue. If the authors can't fix the comparison, the SOTA claim should be toned down.","headline":"A genuinely novel architecture and a valuable dataset undercut by an unfair baseline comparison and a 50-video evaluation; the SOTA claim needs more support.","tokens_in":14302,"tokens_out":2380,"would_cite":true,"duration_ms":26901,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By separating audio into speech, effects, and music, MTV makes audio-sync video generation precisely controllable.","keywords":["audio-sync video generation","audio demixing","multi-stream temporal control","lip motion synchronization","event timing control","visual mood control","DEMIX dataset"],"falsifier":"Retrain or finetune the baselines on the same DEMIX training data and re-run the six-metric evaluation; if the gap largely disappears, the reported state-of-the-art result reflects the training-data advantage rather than the multi-stream architecture itself.","tokens_in":13357,"feed_emoji":"🎬","tokens_out":5251,"duration_ms":58245,"temperature":0.7,"pith_summary":"The paper introduces MTV, a video-generation framework whose central move is to separate an input audio into speech, effects, and music tracks, routing each track to a different visual control: lip motion, event timing, and visual mood. A sympathetic reader would care because prior audio-to-video models treat audio as a single condition and therefore produce scene-level matches but miss fine synchronization, such as lips matching words or a splash landing exactly when the sound plays. To train MTV, the paper contributes DEMIX, roughly 392K cinematic video clips with demixed audio tracks, organized into five overlapping subsets so the model can learn controls from concrete (lip motion) to global (visual mood) in stages. The reported result is state-of-the-art numbers on six metrics covering video quality, text-video consistency, and audio-video alignment.","feed_headline":"Splitting audio into three tracks sharpens audio-video sync","feed_subtitle":"Speech drives lips, effects drive timing, music drives mood; the model tops six metrics.","key_machinery":"The central object is the Multi-Stream Temporal ControlNet (MST-ControlNet). It consists of an interval stream, which combines speech and effects embeddings through interval interaction blocks and injects them into matching time intervals via cross-attention, and a holistic stream, which encodes music into per-clip style scaling and shifting factors applied uniformly to all frames. This machinery converts an audio waveform into two distinct conditioning routes: one local and synchronous, and one global and atmospheric.","core_discovery":"MTV claims that demixing audio before conditioning turns audio-to-video generation from an under-specified mapping into a set of disentangled, temporally precise controls. Wav2vec features from the speech and effects tracks are processed by an interval stream with cross-attention into per-time-interval video latents, driving lip motion and event timing, while music features are pooled by a holistic stream and injected as a style modulation across all frames, shaping the visual mood. Trained on the DEMIX dataset with a five-stage curriculum, the framework reports the best FVD, temporal consistency, text consistency, audio consistency, and lip-sync scores among the compared methods.","pith_inferences":["A direct test the paper leaves implicit is whether the six-metric lead persists when MTV is trained without its 392K DEMIX advantage and the baselines are given comparable training data; this would separate the value of the architecture from the value of the dataset.","Because the paper states that the approach is limited by the categories provided by upstream audio demixing tools, an obvious extension is to couple MTV with more fine-grained demixers, such as those that separate multiple speakers or individual event sounds, and measure whether sync metrics improve accordingly.","The same speech-effects-music split could be applied to joint video-and-audio generation, where the generated audio could first be demixed and then used to control the generated video, potentially tightening synchronization without requiring pre-recorded audio.","A practical extension is to feed MTV a music track with a clearly annotated emotional arc and test whether the generated visual mood shifts at the annotated boundaries, which would isolate the holistic stream's contribution from the text-to-video prior."],"forward_implications":["If the reported numbers hold, the three audio tracks become independent controls: changing the speech track alters lip motion, changing effects alters event timing, and changing music alters visual mood without retraining.","The DEMIX dataset, with its demixed tracks and five overlapping subsets, provides a training resource and curriculum that other audio-video generation methods can reuse.","Because MTV builds on a pretrained text-to-video generator, it can combine text-specified scenes with audio-specified timing, enabling applications such as turning podcasts and historical recordings into visual narratives.","The interval-versus-holistic stream split gives a template for separating local synchronous controls from global ambient controls in other conditional generation tasks.","The reported six-metric gains imply that audio-visual synchronization can be evaluated and optimized as a multi-stream problem rather than as one holistic audio-conditioning problem."],"supporting_citations":[{"why":"Provides the TempoTokens baseline for audio-to-video generation and introduces the AV-Align metric discussed in the appendix.","marker":"[1]"},{"why":"Provides the Xing et al. baseline for open-domain audio-visual generation with diffusion latent aligners, compared in Table 2.","marker":"[3]"},{"why":"Provides the MM-Diffusion baseline for joint audio-video generation, finetuned and compared as a state-of-the-art method.","marker":"[5]"},{"why":"Supplies the pretrained CogVideoX backbone whose VAE and diffusion transformer weights initialize MTV and give it generative priors.","marker":"[10]"},{"why":"Supplies MVSEP, one of the two demixing tools used in the dataset's dual-demixing filtering to produce speech, effects, and music tracks.","marker":"[53]"},{"why":"Supplies Spleeter, the other demixing tool used to cross-check the speech track and conditionally compare effects or music tracks during filtering.","marker":"[54]"},{"why":"Supplies ImageBind, used to compute the Audio-C metric for audio-video consistency.","marker":"[63]"},{"why":"Supplies the Sync-C and Sync-D metrics used to measure lip-motion synchronization with speech.","marker":"[64]"}],"fun_headline_variants":["Demixed audio tracks steer lips, events, mood in video","Audio demixing improves video sync for speech, effects, music","Three audio streams improve video sync: lips, events, mood","From audio to video: three streams for lip, event, mood sync"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of state-of-the-art performance rests on the assumption that comparing MTV with baselines under their original configurations is a fair test, even though MTV is trained on 392K additional DEMIX clips while some baselines are not.","fun_headline_variants_meta":{"raw":{"variants":["Demixed audio tracks steer lips, events, mood in video","Audio demixing improves video sync for speech, effects, music","Three audio streams improve video sync: lips, events, mood","From audio to video: three streams for lip, event, mood sync"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001544,"raw_usage":{"total_tokens":6138,"prompt_tokens":872,"completion_tokens":5266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":5193}},"tokens_in":488,"tokens_out":5266,"duration_ms":43965,"temperature":1.0,"reasoning_tokens":5193,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:20:07.314760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain or finetune the baselines on the same DEMIX training data and re-run the six-metric evaluation; if the gap largely disappears, the reported state-of-the-art result reflects the training-data advantage rather than the multi-stream architecture itself.","supporting_citations":[{"cited_title":"Diverse and aligned audio-to- video generation via text-to-video model adaptation,","cited_arxiv_id":null,"evidence_quote":"Provides the TempoTokens baseline for audio-to-video generation and introduces the AV-Align metric discussed in the appendix."},{"cited_title":"Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,","cited_arxiv_id":null,"evidence_quote":"Provides the Xing et al. baseline for open-domain audio-visual generation with diffusion latent aligners, compared in Table 2."},{"cited_title":"MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,","cited_arxiv_id":null,"evidence_quote":"Provides the MM-Diffusion baseline for joint audio-video generation, finetuned and compared as a state-of-the-art method."},{"cited_title":"CogVideox: Text-to-video diffusion models with an expert transformer,","cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained CogVideoX backbone whose VAE and diffusion transformer weights initialize MTV and give it generative priors."},{"cited_title":"Cinematic sound demixing","cited_arxiv_id":null,"evidence_quote":"Supplies MVSEP, one of the two demixing tools used in the dataset's dual-demixing filtering to produce speech, effects, and music tracks."},{"cited_title":"Spleeter: a fast and efficient music source separation tool with pre-trained models,","cited_arxiv_id":null,"evidence_quote":"Supplies Spleeter, the other demixing tool used to cross-check the speech track and conditionally compare effects or music tracks during filtering."},{"cited_title":"ImageBind: One embedding space to bind them all,","cited_arxiv_id":null,"evidence_quote":"Supplies ImageBind, used to compute the Audio-C metric for audio-video consistency."},{"cited_title":"Out of time: automated lip sync in the wild,","cited_arxiv_id":null,"evidence_quote":"Supplies the Sync-C and Sync-D metrics used to measure lip-motion synchronization with speech."}],"review_version":1}