BLAB is the first benchmark to evaluate audio language models on hour-scale audio, and state-of-the-art models score below 22% exact match on most of its tasks.
Longbench: A bilingual, multitask benchmark for long context understanding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BLAB: Brutally Long Audio Bench
BLAB is the first benchmark to evaluate audio language models on hour-scale audio, and state-of-the-art models score below 22% exact match on most of its tasks.