Pith. sign in

REVIEW 11 cited by

AHELM: A Holistic Evaluation of Audio-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.21376 v2 pith:TIBXHO4Y submitted 2025-08-29 cs.AI cs.CL

AHELM: A Holistic Evaluation of Audio-Language Models

classification cs.AI cs.CL
keywords modelsalmsahelmaudioacrossaspectsdatasetsaudio-language
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Evaluations of audio-language models (ALMs) -- multimodal models that take interleaved audio and text as input and output text -- are hindered by the lack of standardized benchmarks; most benchmarks measure only one or two capabilities and omit evaluative aspects such as fairness or safety. Furthermore, comparison across models is difficult as separate evaluations test a limited number of models and use different prompting methods and inference parameters. To address these shortfalls, we introduce AHELM, a benchmark that aggregates various datasets -- including 2 new synthetic audio-text datasets called PARADE, which evaluates the ALMs on avoiding stereotypes, and CoRe-Bench, which measures reasoning over conversational audio through inferential multi-turn question answering -- to holistically measure the performance of ALMs across 10 aspects we have identified as important to the development and usage of ALMs: audio perception, knowledge, reasoning, emotion detection, bias, fairness, multilinguality, robustness, toxicity, and safety. We also standardize the prompts, inference parameters, and evaluation metrics to ensure equitable comparisons across models. We test 14 open-weight and closed-API ALMs from 3 developers and 3 additional simple baseline systems each consisting of an automatic speech recognizer and a language model. Our results show that while Gemini 2.5 Pro ranks top in 5 out of 10 aspects, it exhibits group unfairness ($p=0.01$) on ASR tasks whereas most of the other models do not. We also find that the baseline systems perform reasonably well on AHELM, with one ranking 6th overall despite having only speech-to-text capabilities. For transparency, all raw prompts, model generations, and outputs are available on our website at https://crfm.stanford.edu/helm/audio/v1.0.0. AHELM is intended to be a living benchmark and new datasets and models will be added over time.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VoxSafeBench: Not Just What Is Said, but Who, How, and Where

    cs.SD 2026-04 unverdicted novelty 8.0

    VoxSafeBench reveals that speech language models recognize social norms from text but fail to apply them when acoustic cues like speaker or scene determine the appropriate response.

  2. RedVox: Safety and Fairness Gaps in Speech Models Across Languages

    cs.CL 2026-06 unverdicted novelty 7.0

    RedVox benchmark shows speech model safety and fairness vulnerabilities persist under non-adversarial conditions, worsen in non-English languages, and increase with spoken inputs.

  3. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI

    eess.AS 2026-05 accept novelty 7.0

    The paper delivers a unified framework for fairness in speech technologies by formalizing seven definitions, organizing research into three paradigms, diagnosing pipeline-specific biases, and mapping mitigations to th...

  4. PRiSM: Benchmarking Phone Realization in Speech Models

    cs.CL 2026-01 conditional novelty 7.0

    PRiSM benchmarks phone recognition in speech models with intrinsic transcription and extrinsic downstream probes, finding that multilingual training and encoder-CTC architectures perform most consistently while LALMs ...

  5. VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

    eess.AS 2026-04 conditional novelty 6.5

    Open-ended evaluation with real speech reveals that gender and accent cues cause statistically significant, task-dependent distributional shifts in recommendations from 12 large audio-language models.

  6. RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

    cs.SD 2026-07 conditional novelty 6.0

    No voice AI system dominates all capabilities; naturalness, expressiveness, identity stability, audio sensitivity, and transcription robustness vary independently, so voice AI should be evaluated as a multidimensional...

  7. Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities

    eess.AS 2026-05 unverdicted novelty 6.0

    Introduces the MUSA benchmark and evaluates LALMs showing that strong single-speaker performance fails to ensure robust selective attention under multilingual interference, with errors from source confusion and unreso...

  8. AudioMosaic: Contrastive Masked Audio Representation Learning

    cs.LG 2026-05 unverdicted novelty 6.0

    AudioMosaic learns general-purpose audio representations through contrastive pre-training with structured spectrogram masking, reaching state-of-the-art results on standard benchmarks and improving audio-language tasks.

  9. VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

    eess.AS 2026-04 unverdicted novelty 6.0

    VIBE evaluates generative biases in large audio-language models with real-world speech and open-ended tasks, showing that gender cues produce larger distributional shifts than accent cues across 11 tested models.

  10. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

  11. AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

    cs.SD 2025-09 unverdicted novelty 6.0

    AU-Harness introduces an efficient unified evaluation framework for audio LLMs featuring batch optimizations, multi-turn dialogue support, and standardized protocols for fair comparisons.