Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:34.997009Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 10 inbound Pith citation observations for arXiv:2508.21376.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:34.997009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.801592Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:59:52.874154Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9781f1e-d579-44da-aab7-1f362649c634 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e0b0aa-3846-40c3-8ef3-af0dd4b33398 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models The Claude 3 model family: Opus, Sonnet, Haiku, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2426604-3b6c-4168-a590-8cb736b79234 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2602b30f-96dd-4597-8e93-c1a7b37ea665 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b69d948-705f-4460-9ef4-0a4967a3ff44 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Towards multimodal sarcasm detection (an _obviously_ perfect paper)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b74c9ce-de68-4f7c-b1c5-70410f7cdd73 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Qwen2-Audio Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41cacf6c-12bb-4b8c-a526-e4abebb8b96f · outbound
AHELM: A Holistic Evaluation of Audio-Language Models VoxCeleb2: Deep Speaker Recognition
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4269de4-646a-4529-86dc-f4865d5969a2 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38f7207c-cba3-4f5a-9765-493bc4c10f6b · outbound
AHELM: A Holistic Evaluation of Audio-Language Models FLEURS: Few-shot learning evaluation of universal representations of speech
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7aecf5c-9e8e-491b-8fa9-a0115a3c3be6 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d728dde4-d52a-4a2f-b473-9ccde920fbd1 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe123ce9-b696-4805-9234-726bc74cbbe1 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8d9c12-d1d7-4781-8e23-3bd439c1ee47 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02a1fa96-c5ad-42d8-b7c0-252adacb5815 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Clap learning audio concepts from natural language supervision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b2cf18e-ff38-4505-a3d2-fc3396052351 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd5a1352-ec15-4aed-98d6-51d73912d090 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models CSR-I (WSJ0) Complete
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2cebc79-52ca-4f7e-bb5c-7b9e40df3c96 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c80121-1169-4858-a503-2b0fc6c46247 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0afda4f7-ceff-44f8-9dc6-e6014bebf7d0 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models V ocalsound: A dataset for improving human vocal sounds recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93ad2fea-62ad-447d-84c6-01898825071b · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Sequence Transduction with Recurrent Neural Networks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71697c35-c355-43ba-b779-faeba1af1613 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6641787a-fc0b-4f71-8245-0d829f09500b · outbound
AHELM: A Holistic Evaluation of Audio-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9668280e-98e8-4b15-88ba-4a84c824c22c · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Design of a linguistic statistical decoder for the recognition of continuous speech
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 366e1f1f-8735-45f1-bd81-92a809687cb2 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Gemini 2.5: Our most intelligent AI model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 13b680d9-4fd5-4b52-95f4-db89ba74f8ea · outbound
AHELM: A Holistic Evaluation of Audio-Language Models AudioCaps: Generat- ing captions for audios in the wild
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 106c11c1-51a3-405e-9cb2-fed3cd7d70fc · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d73ef13-c850-4071-8564-9fe6da7cf9b3 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Vhelm: A holistic evaluation of vision language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e85f6957-3677-45d6-8bd4-bf9d69372f0f · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Holistic evaluation of text-to-image models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 779edca4-f6cc-4ca8-a1b6-cf978c979b3e · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0df2478-580a-4889-a915-2f06b3634273 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models The next chapter of the Gemini era for developers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 280baaad-fbf2-4893-88cf-f56237ef9faa · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Hello GPT-4o, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ed65fc5-1a08-49c4-95dc-657b8e8cfabb · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Introducing our next-generation audio models, Mar 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d37ba2b-6924-4f40-bb80-fbaad1e10bbf · outbound
AHELM: A Holistic Evaluation of Audio-Language Models LibriSpeech: an ASR corpus based on public domain audio books
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a627daf-dfb1-40e9-82a3-149fe1217b18 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a345e82-50e5-47fb-8f4e-f8f82df64bad · outbound
AHELM: A Holistic Evaluation of Audio-Language Models MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9017b8-63ac-46df-81f9-2c5038f0525e · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Robust speech recognition via large-scale weak supervision
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20219ee9-5565-425d-80c8-3895efb66f6e · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Speech Robust Bench: A Robustness Benchmark For Speech Recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90cb00ee-5387-4568-af70-1a4c4781ecf3 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models V oice jailbreak attacks against GPT-4o, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09d16adf-5107-46f1-bb0e-e5f9d8bf2d42 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Salmonn: Towards generic hearing abilities for large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e4f081b-e1f0-4e0c-887e-5c4cd491836a · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abc1829-0a24-4cb6-bc49-48c29cb16690 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57f1fcb-69fa-49dc-8cfe-fa202f93dd4e · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Qwen2.5-Omni Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ee07bf-d9ee-44bb-b7c7-8ea7ded00799 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Qwen3 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b45dbc-f6be-4405-871f-4538421e4457 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5afb4d-7645-48d4-86bc-306b24464242 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Speechlm: Enhanced speech pre-training with unpaired textual data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4df7286-3747-4109-a699-fcaec0306b8a · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46aaaf8-49d0-4b7f-9471-8852c1804a6d · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b02af22-b3d2-4244-a430-85bf59d46073 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94f55932-df50-4f9f-ae84-315f1317d399 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models humorous and imaginative
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9122540e-17ca-44f8-9f7d-e72828c7eedb · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36bc4d4e-83e2-4d0e-bf9d-0e3c89ed6915 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 746c16fa-8b67-4ab6-8ddd-f90736f57628 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1953059c-e1a3-4b3a-9680-cad11029ab8e · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7100bcb0-534f-44b1-853f-e0a1fb4429cb · outbound
AHELM: A Holistic Evaluation of Audio-Language Models You should refer to the score rubric
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91bc551c-76b9-45aa-bc4f-37b5f5d10b12 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd9eec2a-1908-42c0-81cc-67730c0c68d4 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models haha”) or throat clearing (e.g., “ahem
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cafd4911-0adb-4276-8e8d-1cd50d4a9aeb · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d000f3ec-6c48-4317-99a0-4a1f55975eaf · outbound
AHELM: A Holistic Evaluation of Audio-Language Models From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ad1c622-00ad-4d0f-be3a-b2dc963f88b7 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce0143f0-edec-44b0-a2c5-f3104c8eec86 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe1d0241-7fb1-4404-878f-8d5d953d958a · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Limitations
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94467b2d-8bef-4025-b527-944ac890cf07 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 865d1775-6ed3-449e-a803-01997f63b213 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bed55278-f3a2-4c5b-942a-36e402193d8b · outbound
AHELM: A Holistic Evaluation of Audio-Language Models com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31b3c5a6-d911-4336-a6d4-ab1b7ac171cd · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9c61ba0-91ae-45a4-bcfd-6cd5a60f4db5 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models But we do not compute error bars for other scenarios
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b1cd651-9153-45be-a311-d04b174b3f1d · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ba18208-7e04-4b5c-8bd6-64e4f23efcec · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1aacc2fa-b354-4c1b-82a3-e84f0db5faf4 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that there is no societal impact of the work performed
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 462c6a8a-0ff0-46e1-ae74-e02ca8c24b93 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 979de6e4-0d9a-433a-8dfa-f1b0852bb2da · outbound
AHELM: A Holistic Evaluation of Audio-Language Models We cite all the datasets and models used in our work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e119bfdd-abc9-4fae-92af-603f17650782 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85324c3b-c3d3-4ad6-904e-fa33b94ec416 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48de1972-aa06-47ee-883c-a22d9a349457 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 913828ab-683a-4f62-91d4-fe2b0213da28 · outbound
AHELM: A Holistic Evaluation of Audio-Language Models Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 438bb6eb-afa4-4118-9e01-f38ceee8c266 · inbound
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs AHELM: A Holistic Evaluation of Audio-Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 420de55c-2f6c-49e9-9414-845e6f2c69af · inbound
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics AHELM: A Holistic Evaluation of Audio-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b118258d-1bab-4c12-ba6a-d91a2b208abd · inbound
PRiSM: Benchmarking Phone Realization in Speech Models AHELM: A Holistic Evaluation of Audio-Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af163d3-1a39-4b7a-8f15-d33576402f1d · inbound
VoxSafeBench: Not Just What Is Said, but Who, How, and Where AHELM: A Holistic Evaluation of Audio-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2e3a273-2a73-4c70-9492-e80937080f18 · inbound
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech AHELM: A Holistic Evaluation of Audio-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7a02f6c-211f-4b94-ba6b-cd699f7bcfc5 · inbound
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI AHELM: A Holistic Evaluation of Audio-Language Models
Reference 227
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec1072ce-d1cb-4a1e-97ae-f072a79f6845 · inbound
AudioMosaic: Contrastive Masked Audio Representation Learning AHELM: A Holistic Evaluation of Audio-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e1dc822-8c61-4464-b2ea-73165ac7f0a0 · inbound
Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities AHELM: A Holistic Evaluation of Audio-Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a7d2131-ece3-4fec-9428-112d78267f36 · inbound
RedVox: Safety and Fairness Gaps in Speech Models Across Languages AHELM: A Holistic Evaluation of Audio-Language Models
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d5d8abb-cfc8-46d0-a5f4-cc38b5f97d4c · inbound
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems AHELM: A Holistic Evaluation of Audio-Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.