{"id":"c0e971ee-4a3f-4c5f-8be9-d1651e887ab5","arxiv_id":"2607.14189","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MultiRef-Compass is a 350-sample benchmark and 14-metric protocol for multi-reference-to-audio-video generation; current models still fail at reference binding and audio-visual consistency.","lead":"MultiRef-Compass is a new benchmark of 350 curated tasks for generating video with sound from multiple reference images plus text. It checks whether current AI systems preserve each reference, bind entities to the right roles, and keep audio and video aligned, and finds clear gaps.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SLS pre-filter retention imbalance (84 vs 47 valid videos; Appendix D.2) makes the Audio-Visual Consistency dimension's cross-model rankings non-comparable, weakening a key diagnostic claim.","rationale":"I read the paper as a serious benchmark/evaluation contribution. The asset-composition pipeline, the four-dimension protocol, and the rejudging mechanism are careful, and the stability analysis in Appendix D.4 gives some evidence that MLLM scores are not pure noise. However, the strongest claim — that teams can use MultiRef-Compass to diagnose MR2AV failures dimension by dimension — relies on all sub-metrics being comparable across models. The SLS sub-metric is the one place where the paper's own data demonstrate a clear comparability failure: retention counts differ by nearly 2x across models after the pre-filter. This is not a matter of taste or external consensus; it is an internally documented measurement artifact that affects a headline ranking. The paper's Limitations section and Appendix D.2 are honest about it, but honesty does not remove the need for a correction in the main results. I therefore identify this as the single most load-bearing concern. The proposed test directly settles it: if the ranking changes on common support, the confound is real; if the ranking survives, the concern is mitigated. Other potential issues — hand-set thresholds, four annotators, MLLM bias — are real but less crisply tied to a specific result. I agree with the reader's weakest_assumption and with the CONDITIONAL verdict; the benchmark is valuable but the SLS results should not be adopted as standard until this is addressed.","tokens_in":27109,"tokens_out":6899,"duration_ms":65607,"concrete_test":"Recompute Table 2's SLS column on the intersection of videos that pass the stable-frontal pre-filter for all evaluated models (or on a per-model balanced random subsample). If the relative ordering of Gemini-Omni, Seedance 2.0, and Kling changes or the margin narrows materially, the reported AVC ranking is an artifact of retention imbalance. As a second check, rerun LatentSync on all dialogue/monologue samples without the pre-filter to quantify the filter's impact on model-level scores.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim — that MultiRef-Compass provides reliable dimension-by-dimension model diagnosis — depends on each metric measuring the same construct across models. Appendix D.2 shows this fails for SLS: an MLLM pre-filter retains 84 valid videos for Gemini-Omni but only 47 for Kling and Seedance 2.0. Because the filter removes videos with profile faces, large head motion, or poor mouth visibility, each model's SLS is averaged over a different, non-random subset. Models that produce more dynamic (possibly natural) faces are evaluated only on their easiest, most stable clips, while stable-frontal models are evaluated on a broader set. The reported SLS gap (Gemini-Omni 4.53 vs Kling 2.86) is thus entangled with the per-model distribution of face stability, not purely with lip-sync quality. This is acknowledged in the Limitations and Appendix D.2, but no correction is applied to the main leaderboard. Since the AVC dimension is one of the four headline dimensions and Gemini-Omni's overall second-place rank depends heavily on its SLS advantage, the confound directly threatens the paper's diagnostic contribution. It is fixable — e.g., by common support or selection-bias correction — but as presented, the AVC rankings are not trustworthy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MultiRef-Compass, a benchmark for multi-reference-to-audio-video (MR2AV) generation. It consists of 350 curated samples constructed through a taxonomy-driven asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. The proposed evaluation protocol spans four dimensions — Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following — with 14 sub-metrics that combine automatic tools, MLLM-based judging, and a rejudging stage. The paper reports experiments on eight MR2AV systems (six closed-source, two open-source), finding that no system is uniformly strong and that open-source systems lag substantially in reference consistency and instruction following. Human-preference alignment is reported via win-rate correlations (Pearson 0.90–0.96), and a repeated-run stability study on 60 samples shows standard deviations below 1%.","tokens_in":27359,"tokens_out":5265,"duration_ms":54766,"significance":"If the protocol is reliable, MultiRef-Compass would be a valuable contribution: it is the first benchmark specifically targeting MR2AV, offers dimension-level diagnosis rather than a single aggregate score, and uses a controlled asset-reuse design that makes sample construction reproducible. The hybrid automatic+MLLM framework with rejudging is a practical template for multimodal evaluation, and the reported human-preference correlations and repeated-run stability are genuine strengths. However, the benchmark's diagnostic value depends on each sub-metric measuring the same construct across models, and the paper's own appendix documents a serious violation of this condition for the Speech-Lip Synchronization metric. Because the headline rankings rely on SLS, the central diagnostic claim is currently not fully supported.","major_comments":[{"comment":"The SLS metric is not comparable across models because of the stable-frontal-face pre-filter. Appendix D.2 states that Gemini-Omni retains 84 valid videos after filtering, while Kling and Seedance 2.0 retain only 47. The pre-filter removes videos with profile faces, large head motion, or poor mouth visibility, so models that produce more dynamic (possibly natural) faces are scored only on their easiest, most stable clips. The Table 2 SLS values (Gemini-Omni 4.5333 vs. Kling 2.8560) therefore confound lip-sync quality with per-model face-stability distributions. This issue is acknowledged in the Limitations and Appendix D.2, but the main leaderboard and the overall capability rankings in Figure 4 still use these unadjusted SLS scores. Since Gemini-Omni's second-place overall rank depends heavily on its AVC advantage, the AVC rankings are not trustworthy as presented. Please compute SLS on","section":"§3.2, Table 2; Appendix D.2"},{"comment":"The paste-naturalness correction is load-bearing: it changes SkyReels-V3 from rank 5 to rank 8 (Figure 6). However, the exact mapping from the MLLM paste-artifact judge's 1–5 score to the coefficient P_paste(r) ∈ [0,1] used in Eq. (1) is never specified. The judge prompt in Appendix E.2 gives a 1–5 scale and 'strict caps' (e.g., clear cutout/halo → max 3), but no normalization, calibration data, or transformation rule is provided. Without this mapping, the EF values in Table 2 and the resulting ranking change are not reproducible or auditable. The same section also introduces the Long-Term Consistency adjustment EFbase = EF × (0.7 + 0.3×LTC) with a hand-set weight that is not justified or sensitivity-tested. Please provide the exact transformation and add sensitivity analyses for the LTC and combination weights.","section":"§E.2, Eq. (1)"},{"comment":"The main results compare models on different evaluation subsets. Seedance 2.0* is evaluated on 282 samples and Gemini-Omni* on 245 samples due to content-safety filtering, while open-source models are evaluated on all 300 samples (without audio metrics). Per-metric applicability further changes the denominator (e.g., SLS only on dialogue samples passing pre-filter). The paper notes that Figure 4 uses the shared subset of successfully generated videos, but Tables 2 and 3 present raw aggregates with bold/underline rankings as if directly comparable across models. This makes the cross-model leaderboard in the main tables unreliable. Please present matched-sample results — i.e., scores computed only on videos generated by all systems — in the main tables, or at least report the sample count for every cell and provide a matched-subset sensitivity analysis.","section":"Tables 2–3 and Figure 4"},{"comment":"Several thresholds and weights are hand-set without sensitivity analysis: the Timbre Similarity score mapping uses thresholds of 0.075, 0.150, 0.225, and 0.300, and the VSLS combination uses weights 0.60/0.40. The calibration for TS is based on 'synthesized auxiliary speech pairs', but no details, sample size, or robustness analysis are given. Appendix D.4 tests judge stochasticity, not parameter sensitivity. Since Board 4 conclusions (e.g., 'neither model reliably preserves input timbre') depend on these thresholds, and the SLS scores depend on the VSLS weights, please provide a sensitivity sweep or empirical justification for these choices.","section":"§E.3, Timbre Similarity and VSLS"}],"minor_comments":[{"comment":"The notation EFbase is used inconsistently: in the main text it is the pre-paste base score, while in Appendix E.2 it is already adjusted by the Long-Term Consistency term. Please use distinct names.","section":"Appendix E.2"},{"comment":"The caption phrase 'using ranking due to different scale of auto and mllm' is grammatically unclear and should be rephrased to 'ranks are used because automatic and MLLM scores are on different scales'.","section":"Figure 4 caption"},{"comment":"The paper lists 'Code|Dataset' in the abstract, but no URL or release instructions are given in the text. Please provide the link or a statement of availability.","section":"Abstract"},{"comment":"The human-preference alignment table reports Pearson correlations but does not report the number of pairwise comparisons, inter-annotator agreement, or confidence intervals. With only four annotators and no agreement statistics, the strength of the validation is hard to judge.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central idea is timely and the dataset construction is thoughtfully documented, but the SLS confound and the unmatched-sample presentations in the main tables are load-bearing for the paper's diagnostic claims. In addition, several author affiliations are with Kling Team while Kling 3.0 is one of the evaluated systems; a conflict-of-interest statement or independent third-party replication would increase confidence in the benchmark's fairness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the benchmark is the contribution, not the rankings. MultiRef-Compass is the first evaluation suite I know of that treats multi-reference-to-audio-video generation as its own task, and the construction shows care: 350 samples assembled from reusable asset packs, boards that separate same-identity multi-view from multi-entity binding, and a 14-metric protocol that keeps diagnostics separate from a single aggregate score. The external human-preference alignment (Pearson 0.90–0.96) and repeated-run stability (std < 1% on 60 samples) are real evidence that the protocol measures something stable. The rejudging stage is a sensible fix for MLLM judge inconsistency, and the examples in D.3 are convincing.\n\nThe soft spot is where your stress test lands. Appendix D.2 reports that the SLS pre-filter retains 84 valid videos for Gemini-Omni but only 47 for Kling and Seedance 2.0. That means the lip-sync scores are averaged over different, non-random subsets. A model that generates more dynamic or profile faces is scored only on its easiest clips. The paper flags this in the Limitations and in D.2, which is honest, but it still reports the raw SLS numbers in the main leaderboard and lets them drive the AVC dimension and Gemini-Omni's second-place rank. That is a genuine comparability problem, and it is fixable: report SLS on a common-support subset, apply a selection-bias correction, or at minimum present per-model retention counts alongside the scores in the main tables.\n\nOther concerns are minor. Human validation uses four annotators with no inter-annotator agreement reported; that is thin but not fatal. Some thresholds (e.g., timbre similarity calibration) are hand-set without sensitivity analysis. And the Code|Dataset link is a placeholder in the text I saw, so reproducibility is currently unverifiable. None of this breaks the core benchmark claim. The SLS issue is the only one that directly undermines a headline dimension.\n\nFor whom: anyone building or evaluating MR2AV systems. The diagnostic breakdown is more useful than the leaderboard. It deserves a serious referee; with the SLS correction and released artifacts, it could become a standard testbed.","headline":"A genuinely useful MR2AV benchmark whose cross-model SLS numbers are not comparable as reported—fixable, but the current leaderboard overstates the AVC dimension.","tokens_in":27936,"tokens_out":2332,"would_cite":true,"duration_ms":24115,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes MultiRef-Compass as a benchmark that evaluates multi-reference-to-audio-video generation along four separate dimensions, and its experiments on eight systems show no model is uniformly strong.","keywords":["multi-reference-to-audio-video generation","benchmark","evaluation protocol","MLLM-as-a-Judge","reference binding","audio-visual consistency","video generation evaluation","multi-view identity"],"falsifier":"Run the evaluation on the same 350 samples but, for the lip-sync metric, restrict all models to the intersection of videos that survive the stable-frontal-face filter; if Gemini-Omni, which retains 84 valid videos, no longer outscores Kling and Seedance (47 each) on the matched subset, the reported audio-visual ranking is an artifact of sample selection rather than true synchronization quality.","tokens_in":26937,"feed_emoji":"🎬","tokens_out":5981,"duration_ms":57514,"temperature":0.7,"pith_summary":"Multi-reference-to-audio-video (MR2AV) generation asks a model to take several reference images, a text instruction, and sometimes video or audio references, then produce a synchronized audio-video clip. The paper argues that existing benchmarks cannot evaluate this setting because they focus on text prompts, single reference images, or audio-visual alignment in isolation. MultiRef-Compass is introduced as a benchmark of 350 curated samples built from a controllable asset-composition pipeline, together with an evaluation protocol that scores outputs on four dimensions—Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following—via 14 sub-metrics that combine automatic tools with an MLLM-as-a-Judge procedure. Results on eight models show substantial room for improvement on all dimensions and, notably, no system is uniformly strong, with the largest weaknesses in multi-reference binding and audio-visual consistency. The significance is that developers can now see exactly where a model fails rather than relying on a single aggregate number.","feed_headline":"No video model masters multi-reference generation yet","feed_subtitle":"A 350-sample test of 8 systems finds weak spots in identity binding and audio-video sync.","key_machinery":"The load-bearing machinery is the evaluation protocol: four dimensions—Basic Quality (visual, audio, anatomical), Reference Consistency (entity fidelity, detail preservation, binding correctness), Audio-Visual Consistency (lip sync, event-sound matching, source correctness, timbre similarity), and Instruction Following (visual, audio, speech content, temporal order)—with 14 sub-metrics. A rule-based router activates only applicable metrics per sample (e.g., lip sync only for dialogue samples). Two distinguishing components carry much of the argument: the paste-naturalness coefficient that discounts entity similarity when a model copy-pastes reference content, and the rejudging stage that com","core_discovery":"The paper's central claim is that MR2AV generation is a distinct task the community does not currently measure, and that MultiRef-Compass fills that gap. The benchmark's construction pipeline recombines reusable asset packs (multi-view subjects, objects, scenes, voices, videos) into three boards plus a challenge board, producing 350 samples that require cross-reference understanding, compositional binding, and natural visual integration. Its evaluation protocol decomposes quality into four dimensions and 14 sub-metrics, including an entity-fidelity score calibrated by an MLLM-estimated paste-naturalness coefficient to stop models that simply copy a reference image into the frame, and a rejud","pith_inferences":["The authors leave it implicit that the four dimensions could be tuned independently; a factor analysis on model outputs would reveal whether the dimensions are truly separable or whether a model that is good at one dimension tends to be good at others, and if they correlate the diagnostic claim weakens.","The SLS metric's dependence on a stable-frontal-face filter suggests the audio-visual ranking is partly a ranking of which models happen to keep faces still; a lip-sync metric tolerant to head motion, which the authors call for, could change the ordering of Gemini-Omni versus Kling and Seedance.","The asset-composition pipeline could generate much larger or harder reference sets (for example, pairs of similar-looking subjects) specifically to stress binding, since the main benchmark standardizes to three reference images.","A natural extension is to measure whether the benchmark's dimension scores actually predict downstream user satisfaction in a creative workflow, something the paper's human-preference correlation tests only indirectly."],"forward_implications":["Scores on Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following can be reported separately, so a model that looks strong on an aggregate score may still fail at binding references or syncing audio.","The 350 samples and 14 metrics give future MR2AV systems a fixed test bed; the reported results for eight models are a baseline that others can compare against.","The paste-naturalness factor means models that paste reference images into output will be downgraded on entity fidelity even when their raw similarity is high, changing rankings relative to older metrics.","The rejudging stage makes MLLM-based evaluation more auditable by correcting inconsistent score interpretations across models on the same checklist item.","The omni-reference design means the same protocol can be extended as models accept video or audio references, with Board 4 already demonstrating voice-timbre evaluation."],"fun_headline_variants":["MultiRef-Compass: 8 models fail at multi-reference sync","New benchmark finds audio-video models weak on multi-ref","Multi-reference video generation: no model passes yet","350 tests expose gaps in audio-video multi-ref binding","MultiRef-Compass: the missing benchmark for MR2AV tasks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reported model rankings assume that the MLLM judge and the automatic tools evaluate every model fairly and consistently, and the paper shows this assumption is fragile for lip-sync, where only videos with stable frontal faces are scored, leaving some models with many more valid samples than others.","fun_headline_variants_meta":{"raw":{"variants":["MultiRef-Compass: 8 models fail at multi-reference sync","New benchmark finds audio-video models weak on multi-ref","Multi-reference video generation: no model passes yet","350 tests expose gaps in audio-video multi-ref binding","MultiRef-Compass: the missing benchmark for MR2AV tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1459,"prompt_tokens":781,"completion_tokens":678,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":595}},"tokens_in":525,"tokens_out":678,"duration_ms":7087,"temperature":1.0,"reasoning_tokens":595,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:08:36.401781+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the evaluation on the same 350 samples but, for the lip-sync metric, restrict all models to the intersection of videos that survive the stable-frontal-face filter; if Gemini-Omni, which retains 84 valid videos, no longer outscores Kling and Seedance (47 each) on the matched subset, the reported audio-visual ranking is an artifact of sample selection rather than true synchronization quality.","supporting_citations":[],"review_version":1}