REVIEW 17 cited by
AudioBench: A Universal Benchmark for Audio Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and voice understanding (paralinguistic). Despite recent advancements, there lacks a comprehensive benchmark for AudioLLMs on instruction following capabilities conditioned on audio signals. AudioBench addresses this gap by setting up datasets as well as desired evaluation metrics. Besides, we also evaluated the capabilities of five popular models and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-sourced evaluation toolkit, data, and leaderboard will offer a robust testbed for future model developments.
Forward citations
Cited by 17 Pith papers
-
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.
-
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
SOVA-Bench is a new evaluation framework for speech LLMs covering knowledge, recognition, linguistic and paralinguistic understanding, and semantic and acoustic generation quality.
-
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
Open-source audio-language models perform far below humans on a new 600-item benchmark of temporal reasoning in sound, and their accuracy does not track a proposed perturbation-based uncertainty measure.
-
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Audio-visual segmentation models are shown to rely on visual salience rather than audio; a new robustness benchmark and a balanced-training method largely correct this behavior under negative audio conditions.
-
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
Audio LLMs fine-tuned with token-level distillation against an LLM teacher can predict speech quality scores and generate natural-language descriptions, including A/B comparisons.
-
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
A speech-only QA benchmark built from MMLU shows that current end-to-end spoken language models perform near or below random guessing and are brittle to audio changes.
-
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
The authors release MNSC, the largest standardized multitask spoken Singlish corpus, and SingAudioLLM, a multimodal model that sets strong baselines on ASR, spoken QA, dialogue summarization, and paralinguistic QA.
-
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
A query-guided token pruning method improves speech-LLM accuracy on a new long-form audio benchmark while cutting computation.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
ALAS: An Automatic Latent Alignment Score for Audio Language Models
ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.
-
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.
-
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.
-
Multimodal Financial Foundation Models (MFFMs): Progress, Prospects, and Challenges
A position and survey paper argues that multimodal financial foundation models are the next step beyond text-only financial AI, and it organizes current data, benchmarks, models, and challenges around that claim.
-
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
A new Singapore-tailored AudioLLM, built by fusing a fine-tuned Whisper encoder with SEA-LION V3, shows strong in-domain speech recognition but mixed task-understanding gains versus a cascaded baseline.
-
WavChat: A Survey of Spoken Dialogue Models
WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.
Discussion (0). Continue with ORCID to comment.