REVIEW 4 major objections 4 minor 1 cited by
This paper establishes MultiRef-Compass as a benchmark that evaluates multi-reference-to-audio-video generation along four separate dimensions, and its experiments on eight systems show no model is uniformly strong.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 03:08 UTC pith:RZNQADHY
load-bearing objection A genuinely useful MR2AV benchmark whose cross-model SLS numbers are not comparable as reported—fixable, but the current leaderboard overstates the AVC dimension. the 4 major comments →
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that MR2AV generation is a distinct task the community does not currently measure, and that MultiRef-Compass fills that gap. The benchmark's construction pipeline recombines reusable asset packs (multi-view subjects, objects, scenes, voices, videos) into three boards plus a challenge board, producing 350 samples that require cross-reference understanding, compositional binding, and natural visual integration. Its evaluation protocol decomposes quality into four dimensions and 14 sub-metrics, including an entity-fidelity score calibrated by an MLLM-estimated paste-naturalness coefficient to stop models that simply copy a reference image into the frame, and a rejud
What carries the argument
The load-bearing machinery is the evaluation protocol: four dimensions—Basic Quality (visual, audio, anatomical), Reference Consistency (entity fidelity, detail preservation, binding correctness), Audio-Visual Consistency (lip sync, event-sound matching, source correctness, timbre similarity), and Instruction Following (visual, audio, speech content, temporal order)—with 14 sub-metrics. A rule-based router activates only applicable metrics per sample (e.g., lip sync only for dialogue samples). Two distinguishing components carry much of the argument: the paste-naturalness coefficient that discounts entity similarity when a model copy-pastes reference content, and the rejudging stage that com
Load-bearing premise
The reported model rankings assume that the MLLM judge and the automatic tools evaluate every model fairly and consistently, and the paper shows this assumption is fragile for lip-sync, where only videos with stable frontal faces are scored, leaving some models with many more valid samples than others.
What would settle it
Run the evaluation on the same 350 samples but, for the lip-sync metric, restrict all models to the intersection of videos that survive the stable-frontal-face filter; if Gemini-Omni, which retains 84 valid videos, no longer outscores Kling and Seedance (47 each) on the matched subset, the reported audio-visual ranking is an artifact of sample selection rather than true synchronization quality.
If this is right
- Scores on Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following can be reported separately, so a model that looks strong on an aggregate score may still fail at binding references or syncing audio.
- The 350 samples and 14 metrics give future MR2AV systems a fixed test bed; the reported results for eight models are a baseline that others can compare against.
- The paste-naturalness factor means models that paste reference images into output will be downgraded on entity fidelity even when their raw similarity is high, changing rankings relative to older metrics.
- The rejudging stage makes MLLM-based evaluation more auditable by correcting inconsistent score interpretations across models on the same checklist item.
- The omni-reference design means the same protocol can be extended as models accept video or audio references, with Board 4 already demonstrating voice-timbre evaluation.
Where Pith is reading between the lines
- The authors leave it implicit that the four dimensions could be tuned independently; a factor analysis on model outputs would reveal whether the dimensions are truly separable or whether a model that is good at one dimension tends to be good at others, and if they correlate the diagnostic claim weakens.
- The SLS metric's dependence on a stable-frontal-face filter suggests the audio-visual ranking is partly a ranking of which models happen to keep faces still; a lip-sync metric tolerant to head motion, which the authors call for, could change the ordering of Gemini-Omni versus Kling and Seedance.
- The asset-composition pipeline could generate much larger or harder reference sets (for example, pairs of similar-looking subjects) specifically to stress binding, since the main benchmark standardizes to three reference images.
- A natural extension is to measure whether the benchmark's dimension scores actually predict downstream user satisfaction in a creative workflow, something the paper's human-preference correlation tests only indirectly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MultiRef-Compass, a benchmark for multi-reference-to-audio-video (MR2AV) generation. It consists of 350 curated samples constructed through a taxonomy-driven asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. The proposed evaluation protocol spans four dimensions — Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following — with 14 sub-metrics that combine automatic tools, MLLM-based judging, and a rejudging stage. The paper reports experiments on eight MR2AV systems (six closed-source, two open-source), finding that no system is uniformly strong and that open-source systems lag substantially in reference consistency and instruction following. Human-preference alignment is reported via win-rate correlations (Pearson 0.90–0.96), and a repeated-run stability study on 60 samples shows standard deviations below 1%.
Significance. If the protocol is reliable, MultiRef-Compass would be a valuable contribution: it is the first benchmark specifically targeting MR2AV, offers dimension-level diagnosis rather than a single aggregate score, and uses a controlled asset-reuse design that makes sample construction reproducible. The hybrid automatic+MLLM framework with rejudging is a practical template for multimodal evaluation, and the reported human-preference correlations and repeated-run stability are genuine strengths. However, the benchmark's diagnostic value depends on each sub-metric measuring the same construct across models, and the paper's own appendix documents a serious violation of this condition for the Speech-Lip Synchronization metric. Because the headline rankings rely on SLS, the central diagnostic claim is currently not fully supported.
major comments (4)
- [§3.2, Table 2; Appendix D.2] The SLS metric is not comparable across models because of the stable-frontal-face pre-filter. Appendix D.2 states that Gemini-Omni retains 84 valid videos after filtering, while Kling and Seedance 2.0 retain only 47. The pre-filter removes videos with profile faces, large head motion, or poor mouth visibility, so models that produce more dynamic (possibly natural) faces are scored only on their easiest, most stable clips. The Table 2 SLS values (Gemini-Omni 4.5333 vs. Kling 2.8560) therefore confound lip-sync quality with per-model face-stability distributions. This issue is acknowledged in the Limitations and Appendix D.2, but the main leaderboard and the overall capability rankings in Figure 4 still use these unadjusted SLS scores. Since Gemini-Omni's second-place overall rank depends heavily on its AVC advantage, the AVC rankings are not trustworthy as presented. Please compute SLS on
- [§E.2, Eq. (1)] The paste-naturalness correction is load-bearing: it changes SkyReels-V3 from rank 5 to rank 8 (Figure 6). However, the exact mapping from the MLLM paste-artifact judge's 1–5 score to the coefficient P_paste(r) ∈ [0,1] used in Eq. (1) is never specified. The judge prompt in Appendix E.2 gives a 1–5 scale and 'strict caps' (e.g., clear cutout/halo → max 3), but no normalization, calibration data, or transformation rule is provided. Without this mapping, the EF values in Table 2 and the resulting ranking change are not reproducible or auditable. The same section also introduces the Long-Term Consistency adjustment EFbase = EF × (0.7 + 0.3×LTC) with a hand-set weight that is not justified or sensitivity-tested. Please provide the exact transformation and add sensitivity analyses for the LTC and combination weights.
- [Tables 2–3 and Figure 4] The main results compare models on different evaluation subsets. Seedance 2.0* is evaluated on 282 samples and Gemini-Omni* on 245 samples due to content-safety filtering, while open-source models are evaluated on all 300 samples (without audio metrics). Per-metric applicability further changes the denominator (e.g., SLS only on dialogue samples passing pre-filter). The paper notes that Figure 4 uses the shared subset of successfully generated videos, but Tables 2 and 3 present raw aggregates with bold/underline rankings as if directly comparable across models. This makes the cross-model leaderboard in the main tables unreliable. Please present matched-sample results — i.e., scores computed only on videos generated by all systems — in the main tables, or at least report the sample count for every cell and provide a matched-subset sensitivity analysis.
- [§E.3, Timbre Similarity and VSLS] Several thresholds and weights are hand-set without sensitivity analysis: the Timbre Similarity score mapping uses thresholds of 0.075, 0.150, 0.225, and 0.300, and the VSLS combination uses weights 0.60/0.40. The calibration for TS is based on 'synthesized auxiliary speech pairs', but no details, sample size, or robustness analysis are given. Appendix D.4 tests judge stochasticity, not parameter sensitivity. Since Board 4 conclusions (e.g., 'neither model reliably preserves input timbre') depend on these thresholds, and the SLS scores depend on the VSLS weights, please provide a sensitivity sweep or empirical justification for these choices.
minor comments (4)
- [Appendix E.2] The notation EFbase is used inconsistently: in the main text it is the pre-paste base score, while in Appendix E.2 it is already adjusted by the Long-Term Consistency term. Please use distinct names.
- [Figure 4 caption] The caption phrase 'using ranking due to different scale of auto and mllm' is grammatically unclear and should be rephrased to 'ranks are used because automatic and MLLM scores are on different scales'.
- [Abstract] The paper lists 'Code|Dataset' in the abstract, but no URL or release instructions are given in the text. Please provide the link or a statement of availability.
- [Table 4] The human-preference alignment table reports Pearson correlations but does not report the number of pairwise comparisons, inter-annotator agreement, or confidence intervals. With only four annotators and no agreement statistics, the strength of the validation is hard to judge.
Circularity Check
No significant circularity: the benchmark's metrics are externally grounded and its empirical claims are measurements, not derivations from their own inputs.
full rationale
MultiRef-Compass does not contain a derivation chain in which a prediction reduces by construction to its inputs. The paper's central claims are (1) that the benchmark is comprehensive and (2) that evaluated MR2AV systems show uneven performance across four dimensions. These are empirical claims supported by concrete evaluation protocols. The metrics are either standard external tools (DOVER++, Audiobox, LatentSync, ECAPA-TDNN, SigLIP, YOLO-World) or MLLM rubrics with explicit, model-independent checklists. No metric is fitted to the model rankings it then 'predicts.' The EF formula EFfinal = EFbase × Ppaste(r) is a post-hoc correction for copy-and-paste artifacts, with Ppaste defined by an independent 'Paste Artifact Judge' prompt; it calibrates a similarity score but does not encode the ranking. The timbre-similarity threshold (>0.30) is calibrated on auxiliary synthesized same-timbre clips, external to the evaluated models, so it is not a fitted input renamed as a prediction. The rejudging stage uses the same MLLM family as the initial judge, which could raise bias concerns, but the paper validates the resulting scores against human preference win rates (Table 4: Pearson 0.90–0.96). That is an external benchmark, not circular self-confirmation. The SLS pre-filter retention imbalance (84 vs 47 valid videos, Appendix D.2) is a genuine measurement-validity limitation and is explicitly acknowledged; it does not make the SLS score equivalent to the number of retained videos by construction, and the paper cautions that VSLS should be interpreted with retained sample counts. Self-citations to prior Compass benchmarks (T2AV-Compass, LongAV-Compass) are used as related work and evaluation-protocol context, not as the load-bearing justification for the benchmark's validity. There is no uniqueness theorem, ansatz, or definitional equivalence imported from the authors' prior work. The paper's conclusions are therefore not circular; at most they inherit typical benchmark-validity risks, which the limitations section discloses.
Axiom & Free-Parameter Ledger
free parameters (5)
- EF human face/body combination weight =
0.6 face / 0.4 body
- Long-Term Consistency weights =
0.7 + 0.3 * LTC
- VSLS combination weights =
0.60 * Cnorm + 0.40 * Dnorm
- Timbre similarity thresholds =
0.075, 0.150, 0.225, 0.300; saturation at cosine > 0.30
- SLS stable-frontal-face pre-filter criteria =
strict frontal/half-body/upper-body framing conditions
axioms (4)
- domain assumption MLLM-as-a-Judge (Gemini 3.1 Pro) judgments are a valid proxy for human preference across all four evaluation dimensions.
- domain assumption Embedding and detection tools (SigLIP, YOLO-World, InsightFace, DINOv2, ECAPA-TDNN, LatentSync) generalize to cartoon subjects, large pose/expression changes, and occlusions.
- domain assumption GPT-5.5-generated structured prompts plus manual review produce benchmark samples that fairly represent MR2AV tasks.
- domain assumption A 350-sample benchmark and reduced valid subsets are sufficient for stable cross-model rankings.
read the original abstract
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises $350$ carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.
Figures
Forward citations
Cited by 1 Pith paper
-
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Mixed-data SFT plus Hierarchical Coverage-Discounted GRPO yields open-source SOTA multi-reference image-grounded video captions on MRVBench without hurting general captioning.
Reference graph
Works this paper leans on
-
[1]
Videophy:Evaluatingphysicalcommonsense for video generation
Bansal, H.; Lin, Z.; Xie, T.; Zong, Z.; Yarom, M.; Bit- ton, Y.; Jiang, C.; Sun, Y.; Chang, K.-W.; and Grover, A.2025. Videophy:Evaluatingphysicalcommonsense for video generation. InICLR, volume 2025, 102075– 102121
2025
-
[2]
Bao, F.; Xiang, C.; Yue, G.; He, G.; Zhu, H.; Zheng, K.; Zhao, M.; Liu, S.; Wang, Y.; and Zhu, J. 2024. Vidu: a highly consistent, dynamic and skilled text-to- video generator with diffusion models.arXiv preprint arXiv:2405.04233
Pith/arXiv arXiv 2024
-
[3]
Cao, Z.; Wang, T.; Wang, J.; Wang, Y.; Zhang, Y.; Wang, J.; Chen, J.; Deng, M.; Guo, Y.; Liao, C.; et al. 2025. T2av-compass: Towards unified evalua- tion for text-to-audio-video generation.arXiv preprint arXiv:2512.21094
Pith/arXiv arXiv 2025
-
[4]
Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; and Shan, Y. 2024. YOLO-World: Real-Time Open- VocabularyObjectDetection. InCVPR,16901–16911
2024
-
[5]
Desplanques, B.; Thienpondt, J.; and Demuynck, K. 2020.Ecapa-tdnn:Emphasizedchannelattention,prop- agationandaggregationintdnnbasedspeakerverifica- tion.arXiv preprint arXiv:2005.07143
Pith/arXiv arXiv 2020
-
[6]
Google DeepMind. 2026. Gemini 3.1 Pro: A smarter model for your most complex tasks. https://blog.google/innovation-and-ai/models-and- research/gemini-models/gemini-3-1-pro/. Accessed: 2026-07-15
2026
-
[7]
Google DeepMind. 2026. Gemini Omni — Google DeepMind. https://deepmind.google/models/gemini- omni/. Accessed: 2026-07-15
2026
-
[8]
HappyHorse:TheSimple,All-in- OneAIVideoPlatform
HappyHorse.2026. HappyHorse:TheSimple,All-in- OneAIVideoPlatform. https://www.happyhorse.com. Accessed: 2026-07-15
2026
-
[9]
He, X.; Jiang, D.; Zhang, G.; Ku, M.; Soni, A.; Siu, S.; Chen, H.; Chandra, A.; Jiang, Z.; Arulraj, A.; et al
-
[10]
Hu, L. 2024. Animate anyone: Consistent and control- lableimage-to-videosynthesisforcharacteranimation. InCVPR, 8153–8163
2024
-
[11]
Huang, Z.; He, Y.; Yu, J.; Zhang, F.; Si, C.; Jiang, Y.; Zhang, Y.; Wu, T.; Jin, Q.; Chanpaisit, N.; et al
-
[12]
Huang, Z.; Zhang, F.; Xu, X.; He, Y.; Yu, J.; Dong, Z.; Ma, Q.; Chanpaisit, N.; Si, C.; Jiang, Y.; et al. 2025. Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE TPAMI
2025
-
[13]
InCVPR, 21807–21818
Vbench: Comprehensive benchmark suite for video generative models. InCVPR, 21807–21818
-
[14]
Jiang,Z.;Han,Z.;Mao,C.;Zhang,J.;Pan,Y.;andLiu, Y. 2025. Vace: All-in-one video creation and editing. InICCV, 17191–17202
2025
-
[15]
C.; and Liu, Z
Jiang, Y.; Wu, T.; Yang, S.; Si, C.; Lin, D.; Qiao, Y.; Loy, C. C.; and Liu, Z. 2024. Videobooth: Diffusion- based video generation with image prompts. InCVPR, 6689–6700
2024
-
[16]
Li, D.; Fei, Z.; Li, T.; Dou, Y.; Chen, Z.; Yang, J.; Fan, M.; Xu, J.; Wang, J.; Gu, B.; et al. 2026. Skyreels-v3 technique report.arXiv preprint arXiv:2601.17323
arXiv 2026
-
[17]
Li, C.; Zhang, C.; Xu, W.; Lin, J.; Xie, J.; Feng, W.; Peng, B.; Chen, C.; and Xing, W. 2024. Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision.arXiv preprint arXiv:2412.09262
Pith/arXiv arXiv 2024
-
[18]
Liu, Y.; Cun, X.; Liu, X.; Wang, X.; Zhang, Y.; Chen, H.; Liu, Y.; Zeng, T.; Chan, R.; and Shan, Y. 2024. Evalcrafter: Benchmarking and evaluating large video generation models. InCVPR, 22139–22149
2024
-
[19]
Liu, T.; Shi, Y.; Zhu, X.; Tang, J.; Yang, L.; Wang, Q.; Zhang, Z.; Tang, Y.; Wang, F.; Dong, Y.; et al. 2026. LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV.arXiv preprint arXiv:2605.26244
Pith/arXiv arXiv 2026
-
[20]
Ma, Y.; He, Y.; Wang, H.; Wang, A.; Shen, L.; Qi, C.; Ying, J.; Cai, C.; Li, Z.; Shum, H.-Y.; et al. 2025. Follow-your-click: Open-domain regional image ani- mationviamotionprompts.InAAAI,volume39,6018– 6026
2025
-
[21]
Liu, Y.; Li, L.; Ren, S.; Gao, R.; Li, S.; Chen, S.; Sun, X.; and Hou, L. 2023. Fetv: A benchmark for fine- grained evaluation of open-domain text-to-video gen- eration.NeurIPS, 36: 62352–62387
2023
-
[22]
Qi,Z.;Shi,P.;Wang,S.;Zhang,C.;Zhao,F.;Ying,Z.; Pan, D.; Yang, X.; He, Z.; and Dai, T. 2025. T2veval: Benchmarkdatasetandobjectiveevaluationmethodfor t2v-generated videos.Displays, 103178
2025
-
[23]
OpenAI. 2026. GPT-5.5 Model. https://developers. openai.com/api/docs/models/gpt-5.5. Accessed:2026- 07-12
2026
-
[24]
Shi,Y.;Liu,J.;Guan,Y.;Wu,Z.;Zhang,Y.;Wang,Z.; Lin, W.; Hua, J.; Wang, Z.; Chen, X.; et al. 2025. Ma- vors: Multi-granularity video representation for mul- timodal large language model. InProceedings of the 33rd ACM International Conference on Multimedia, 10994–11003
2025
-
[25]
Seedance, T.; Chen, D.; Chen, L.; Chen, X.; Chen, Y.; Chen, Z.; Chen, Z.; Cheng, F.; Cheng, T.; Cheng, Y.; et al. 2026. Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2604.14148
Pith/arXiv arXiv 2026
-
[26]
Tjandra, A.; Wu, Y.-C.; Guo, B.; Hoffman, J.; Ellis, B.; Vyas, A.; Shi, B.; Chen, S.; Le, M.; Zacharov, N.; et al. 2025. Meta audiobox aesthetics: Unified auto- maticqualityassessmentforspeech,music,andsound. arXiv preprint arXiv:2502.05139
Pith/arXiv arXiv 2025
-
[27]
Team, K.; Chen, J.; Ci, Y.; Du, X.; Feng, Z.; Gai, K.; Guo,S.;Han,F.;He,J.;He,K.;etal.2025.Kling-Omni Technical Report.arXiv preprint arXiv:2512.16776
Pith/arXiv arXiv 2025
-
[28]
Wang, X.; Yuan, H.; Zhang, S.; Chen, D.; Wang, J.; Zhang, Y.; Shen, Y.; Zhao, D.; and Zhou, J. 2023. Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 36: 7594–7611
2023
-
[29]
Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.- W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. 2025. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314
Pith/arXiv arXiv 2025
-
[30]
Wu, H.; Zhang, E.; Liao, L.; Chen, C.; Hou, J.; Wang, A.; Sun, W.; Yan, Q.; and Lin, W. 2023. Exploring VideoQualityAssessmentonUserGeneratedContents from Aesthetic and Technical Perspectives. InICCV, 20144–20154
2023
-
[31]
Univbench:To- wards unified evaluation for video foundation models
Wei, J.; Zhang, X.; Li, Y.; Wang, Y.; Zhang, Y.; Chen, Z.;Tang,Z.;Xu,W.;andLiu,Z.2026. Univbench:To- wards unified evaluation for video foundation models. InCVPR, 25654–25666
2026
-
[32]
Yang,J.;Xia,B.;Chu,R.;Wang,D.;Xia,W.;Mou,Z.; Zhong, T.; Zhao, Y.; and Yang, W. 2026. AVBench: Human-Aligned and Automated Evaluation Bench- mark for Audio-Video Generative Models.arXiv preprint arXiv:2605.24652
Pith/arXiv arXiv 2026
-
[33]
H.; Yan, H.; Liu, J.-W.; Zhang, C.; Feng, J.; and Shou, M
Xu, Z.; Zhang, J.; Liew, J. H.; Yan, H.; Liu, J.-W.; Zhang, C.; Feng, J.; and Shou, M. Z. 2024. Magican- imate: Temporally consistent human image animation using diffusion model. InCVPR, 1481–1490
2024
-
[34]
Yuan, S.; He, X.; Deng, Y.; Ye, Y.; Huang, J.; Ma, C.; Luo, J.; Yuan, L.; et al. 2026. Opens2v-nexus: A de- tailedbenchmarkandmillion-scaledatasetforsubject- to-video generation.NeurIPS, 38
2026
-
[35]
Yin, S.; Wu, C.; Liang, J.; Shi, J.; Li, H.; Ming, G.; andDuan,N.2023. Dragnuwa:Fine-grainedcontrolin video generation by integrating text, image, and trajec- tory.arXiv preprint arXiv:2308.08089
Pith/arXiv arXiv 2023
-
[36]
Zhang, A.; Lei, L.; Kong, D.; Wang, Z.; Xu, J.; Song, F.;Guo,C.-L.;Liu,C.;Li,F.;andChen,J.2025. UI2V- Bench: An Understanding-based Image-to-video Gen- eration Benchmark.arXiv preprint arXiv:2509.24427
arXiv 2025
-
[37]
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L
-
[38]
Zhou, Z.; Lai, Z.; Wang, R.; Yang, Y.; Xing, Z.; Yang, Y.; Dai, Q.; Qiu, L.; and Luo, C. 2026. AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evalua- tionofText-to-Audio-VideoGeneration.arXivpreprint arXiv:2604.08540. Appendix A MR2AV Definition Multi-reference-to-audio-video(MR2AV)generationsynthesizesnewaudio-videocontentfrommultiplereferencesa...
Pith/arXiv arXiv 2026
-
[40]
Zhang, Y.; Luo, Z.; Yan, Q.; He, W.; Jiang, B.; Chen, X.; and Han, K. 2025. OmniEval: A Benchmark for EvaluatingOmni-modalModelswithVisual,Auditory, and Textual Inputs.arXiv preprint arXiv:2506.20960
Pith/arXiv arXiv 2025
-
[91]
**Sound consistency: ** Does the heard effect plausibly match the visible action, object, material, force, and environment?
-
[102]
not_applicable
**Temporal alignment: ** Does the effect start, continue, repeat, and stop at the correct visible time? 11 12Include: impacts, contacts, footsteps, clicks, scraping, pouring, rustling, object handling, doors, tools, and other Foley; scene-grounded environmental effects such as visible rain, wind-driven movement, traffic, birds, water, or machinery when th...
-
[141]
What requested subject/object/scene/action/relation/state/framing is visible?
-
[151]
**Content Completeness ** - Are all required words, sentences, or key information present?
-
[152]
What part of the requirement is satisfied?
-
[162]
**Content Correctness ** - Is the spoken content free from substitutions, additions, or distortions?
-
[163]
What part is missing, unclear, incomplete, or contradicted?
-
[173]
**Language/Locale Match ** - If a language is specified, does the speech use that language?
-
[174]
22- **2 = Weak related evidence **: something relevant visible, but core requirement mostly missing
Which evidence level best matches: no evidence / weak related evidence / partial / mostly complete / complete direct? 18 19## Scoring (1-5 per item) 20 21- **1 = No evidence **: requirement completely absent, clearly wrong, or contradicted. 22- **2 = Weak related evidence **: something relevant visible, but core requirement mostly missing. 23- **3 = Parti...
-
[184]
19 20## Cross-language Translation Rule 21 22If Chinese speech is required but the generated speech is a semantically correct English translation, assign score **2**
**Ignore Speaker Assignment ** - Do not penalize wrong speaker, swapped speakers, narrator voice, or off-screen voice. 19 20## Cross-language Translation Rule 21 22If Chinese speech is required but the generated speech is a semantically correct English translation, assign score **2**. Do not assign 2 when the translation changes, omits, or invents key inf...
-
[2023]
InICCV, 11975–11986
SigmoidLossforLanguageImagePre-Training. InICCV, 11975–11986
-
[2024]
InEMNLP, 2105–2123
Videoscore:Buildingautomaticmetricstosimu- latefine-grainedhumanfeedbackforvideogeneration. InEMNLP, 2105–2123
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.