REVIEW 12 cited by
AudioBench: A Universal Benchmark for Audio Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments.
Forward citations
Cited by 12 Pith papers
-
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.
-
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
SOVA-Bench is a new evaluation framework for speech LLMs covering knowledge, recognition, linguistic and paralinguistic understanding, and semantic and acoustic generation quality.
-
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Audio-visual segmentation models are shown to rely on visual salience rather than audio; a new robustness benchmark and a balanced-training method largely correct this behavior under negative audio conditions.
-
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
Audio LLMs fine-tuned with token-level distillation against an LLM teacher can predict speech quality scores and generate natural-language descriptions, including A/B comparisons.
-
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
A speech-only QA benchmark built from MMLU shows that current end-to-end spoken language models perform near or below random guessing and are brittle to audio changes.
-
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
The authors release MNSC, the largest standardized multitask spoken Singlish corpus, and SingAudioLLM, a multimodal model that sets strong baselines on ASR, spoken QA, dialogue summarization, and paralinguistic QA.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
ALAS: An Automatic Latent Alignment Score for Audio Language Models
ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.
-
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.
-
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.
Discussion (0). Continue with ORCID to comment.