A new multi-control benchmark for mixed audio generation shows current models trade off acoustic fidelity, speech quality, semantic alignment, and temporal control.
A middle-aged male captain aged 40-50 calmly stated in standard English,' The plane is about to land.'
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MMAG: A Multi-Control Mixed Audio Generation Benchmark
A new multi-control benchmark for mixed audio generation shows current models trade off acoustic fidelity, speech quality, semantic alignment, and temporal control.