T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Chenxi Liao; Jiahao Wang; Jiaheng Liu; Jialu Chen; Jiaming Wang; Miao Deng; Tao Wang; Yanghai Wang; Yize Zhang; Yuanxing Zhang

arxiv: 2512.21094 · v2 · pith:XGE7J4DDnew · submitted 2025-12-24 · 💻 cs.CV

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang

show 5 more authors

Yubin Guo Chenxi Liao Yize Zhang Zhaoxiang Zhang Jiaheng Liu

This is my paper

classification 💻 cs.CV

keywords evaluationrealismt2av-compassaudiocross-modalfollowinggenerationinstruction

0 comments

read the original abstract

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
cs.CV 2026-05 conditional novelty 7.0

MSAVBench is the first comprehensive benchmark for multi-shot audio-video generation, spanning video, audio, shot, and reference dimensions with an adaptive evaluation framework that reaches 91.5% Spearman correlation...
Do Joint Audio-Video Generation Models Understand Physics?
cs.SD 2026-05 unverdicted novelty 7.0

Current joint audio-video generation models lack robust physical commonsense, especially during transitions and when prompted for impossible behaviors.
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
cs.SD 2026-04 unverdicted novelty 7.0

VidAudio-Bench benchmarks V2A and VT2A models across four audio categories, revealing poor speech/singing performance and a tension between visual alignment and text instruction following.