Pith. sign in

Paper Citation Record · LEDGER

AHELM: A Holistic Evaluation of Audio-Language Models

As of 6 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 10 inbound Pith citation observations for arXiv:2508.21376.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21376 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:34.997009Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.801592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:52.874154Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9781f1e-d579-44da-aab7-1f362649c634 · outbound

This paper cites GPT-4 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.697071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.697071Z digest=sha256:6a4637a9257fee7b894859ae7133f337cd321c2a5c689a3b7059823f59021fe6

Observation 51e0b0aa-3846-40c3-8ef3-af0dd4b33398 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models The Claude 3 model family: Opus, Sonnet, Haiku, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.739040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:33.778200Z digest=sha256:5bb398fd53eeb1f7e22b85cfbe2d37e0c35806dfeb0fd6b1ee92c30882b5eb59

Observation a2426604-3b6c-4168-a590-8cb736b79234 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

AHELM: A Holistic Evaluation of Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.923565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.923565Z digest=sha256:528f465c21c45bf5e3cf7f74ea9e87ff84a4b144e79ab072f0e2252053d46b1b

Observation 2602b30f-96dd-4597-8e93-c1a7b37ea665 · outbound

This paper cites Qwen Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.102234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.102234Z digest=sha256:b743de03b1f7f33e22be6b39e6c21691a04c2033abafad42f8c63144e3eeeb4d

Observation 6b69d948-705f-4460-9ef4-0a4967a3ff44 · outbound

This paper cites Towards multimodal sarcasm detection (an _obviously_ perfect paper).

AHELM: A Holistic Evaluation of Audio-Language Models Towards multimodal sarcasm detection (an _obviously_ perfect paper)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.731229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.243183Z digest=sha256:9c72543cdb7843a4d0bc4463a745919b3f76e877452f035e2846096c45abecfc

Observation 2b74c9ce-de68-4f7c-b1c5-70410f7cdd73 · outbound

This paper cites Qwen2-Audio Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.352067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.352067Z digest=sha256:5a211dc322d553b4c6287afcd73168387ca53b6d5092bf2b3694b06d1d3f946f

Observation 41cacf6c-12bb-4b8c-a526-e4abebb8b96f · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models VoxCeleb2: Deep Speaker Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.385152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.385152Z digest=sha256:d4a667063b910680983725ef66b1e6ccf2d7a43c4c443dcf07db8bbfd14ae65f

Observation e4269de4-646a-4529-86dc-f4865d5969a2 · outbound

This paper cites Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit.

AHELM: A Holistic Evaluation of Audio-Language Models Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.722887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.540549Z digest=sha256:13b26c1b8d885afbdd163c4e02f5a56b2fd67ddb521ddd6dac6d4badc312b667

Observation 38f7207c-cba3-4f5a-9765-493bc4c10f6b · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech.

AHELM: A Holistic Evaluation of Audio-Language Models FLEURS: Few-shot learning evaluation of universal representations of speech

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.714660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.635896Z digest=sha256:e1b3e3e9bec4c0f636e956e76fd2807e8a2112fb30a35e6b4ba1f7b3faa1f77e

Observation c7aecf5c-9e8e-491b-8fa9-a0115a3c3be6 · outbound

This paper cites MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector.

AHELM: A Holistic Evaluation of Audio-Language Models MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.199235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.753589Z digest=sha256:e5bc63b227e05b3416ff406a55f714e2c8daab62ef37be167d34688d892a9792

Observation d728dde4-d52a-4a2f-b473-9ccde920fbd1 · outbound

This paper cites Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.706024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.775890Z digest=sha256:cdf583950037937cb0d6d03edc08988be9d7f762a12a342d5ac5ff5d8eeb2ee5

Observation fe123ce9-b696-4805-9234-726bc74cbbe1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

AHELM: A Holistic Evaluation of Audio-Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.779733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.779733Z digest=sha256:55c44005e23711bcaaf14f583df5c9a4ae983fa226d125612590357b98f4dd1b

Observation 9a8d9c12-d1d7-4781-8e23-3bd439c1ee47 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

AHELM: A Holistic Evaluation of Audio-Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.697626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.783925Z digest=sha256:1dad77f68b4ab605acf3e890658b57323282fa3ecb3e1a5fb299908010a42b65

Observation 02a1fa96-c5ad-42d8-b7c0-252adacb5815 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Clap learning audio concepts from natural language supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.689202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.787398Z digest=sha256:531537e42a7d0cbf1f737854634a475bab49ad5b6a7ba43095506d6e083d9d02

Observation 6b2cf18e-ff38-4505-a3d2-fc3396052351 · outbound

This paper cites Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images.

AHELM: A Holistic Evaluation of Audio-Language Models Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.170240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.790680Z digest=sha256:94032c93740bb386da0c3f36b63da7f0e0d5ddfd4f0671330e0619c934122591

Observation dd5a1352-ec15-4aed-98d6-51d73912d090 · outbound

This paper cites CSR-I (WSJ0) Complete.

AHELM: A Holistic Evaluation of Audio-Language Models CSR-I (WSJ0) Complete

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.680312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.794039Z digest=sha256:68a6dee390a6cb99655a00aa76df12de32e0593779175c55c1223a77761becec

Observation a2cebc79-52ca-4f7e-bb5c-7b9e40df3c96 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.797552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.797552Z digest=sha256:53a37bab7e400725f5165fd0395637e023d24a304d17d7ae0c944e5a131139d0

Observation 80c80121-1169-4858-a503-2b0fc6c46247 · outbound

This paper cites Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities.

AHELM: A Holistic Evaluation of Audio-Language Models Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.669917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.801817Z digest=sha256:7985349bb05017f456eb972eee99561af3db87737eab5d9f8922f1610fd34812

Observation 0afda4f7-ceff-44f8-9dc6-e6014bebf7d0 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

AHELM: A Holistic Evaluation of Audio-Language Models V ocalsound: A dataset for improving human vocal sounds recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.661754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.805468Z digest=sha256:5ca089fb8c9613dfd536d22e696acc79a812c31e444a6f06dd4e4b6a68d93198

Observation 93ad2fea-62ad-447d-84c6-01898825071b · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

AHELM: A Holistic Evaluation of Audio-Language Models Sequence Transduction with Recurrent Neural Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.808889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.808889Z digest=sha256:bf4c450923936298f1b835da504a7512ca0156bb9e9c123d318d72d29a0783a8

Observation 71697c35-c355-43ba-b779-faeba1af1613 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

AHELM: A Holistic Evaluation of Audio-Language Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.653539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.812080Z digest=sha256:144d8dfd2e5209e4fe10418a0a7fba8d0d5b8b223202bd7aede3d145650843f2

Observation 6641787a-fc0b-4f71-8245-0d829f09500b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AHELM: A Holistic Evaluation of Audio-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.815275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.815275Z digest=sha256:3c566ecc8883382e2bbf76255b8ac9cea417c8aca347540844a3272e50755b9b

Observation 9668280e-98e8-4b15-88ba-4a84c824c22c · outbound

This paper cites Design of a linguistic statistical decoder for the recognition of continuous speech.

AHELM: A Holistic Evaluation of Audio-Language Models Design of a linguistic statistical decoder for the recognition of continuous speech

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.645260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.818670Z digest=sha256:5334db6cd46a569591a9595da0efab389d4d4d6f4e8b15dc0d4d38fffa31b44b

Observation 366e1f1f-8735-45f1-bd81-92a809687cb2 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 2.5: Our most intelligent AI model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.636480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.821970Z digest=sha256:4f86c9226a76d3a92fc38cb526b0eda82572d8520ddee3c1a18d17ee42bb612c

Observation 13b680d9-4fd5-4b52-95f4-db89ba74f8ea · outbound

This paper cites AudioCaps: Generat- ing captions for audios in the wild.

AHELM: A Holistic Evaluation of Audio-Language Models AudioCaps: Generat- ing captions for audios in the wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.628926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.825065Z digest=sha256:88c0aa15a967de8433a8377b152e19db238757078392badacfe5cc271c875fb3

Observation 106c11c1-51a3-405e-9cb2-fed3cd7d70fc · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

AHELM: A Holistic Evaluation of Audio-Language Models Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.620586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.828327Z digest=sha256:d319ee3ff42edf1899166a94c7b0f18abfb5c7b1d454ae6f16a93a375ed2cd99

Observation 7d73ef13-c850-4071-8564-9fe6da7cf9b3 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.

AHELM: A Holistic Evaluation of Audio-Language Models Vhelm: A holistic evaluation of vision language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.611540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.832092Z digest=sha256:6ab2ca156251f3a6034ab4a57b60cb29a5ec74f44d52aaf2be378d136e3fe1ce

Observation e85f6957-3677-45d6-8bd4-bf9d69372f0f · outbound

This paper cites Holistic evaluation of text-to-image models.

AHELM: A Holistic Evaluation of Audio-Language Models Holistic evaluation of text-to-image models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.601221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.835355Z digest=sha256:1a9766512665106d9fd9dce480aeea2b2e05acb98a3cf9555f08e3b74f00fb20

Observation 779edca4-f6cc-4ca8-a1b6-cf978c979b3e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.591876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.838505Z digest=sha256:b383c5ce585d4a11db45b0164da19bd10c3851b1463003407a406ca6639219e4

Observation e0df2478-580a-4889-a915-2f06b3634273 · outbound

This paper cites The next chapter of the Gemini era for developers.

AHELM: A Holistic Evaluation of Audio-Language Models The next chapter of the Gemini era for developers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.582196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.842252Z digest=sha256:d1ca9bb19f600b8e1eaf143c5ba2859e00d6ea786311d72d13ed1f5585542d9e

Observation 280baaad-fbf2-4893-88cf-f56237ef9faa · outbound

This paper cites Hello GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models Hello GPT-4o, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.572830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.845109Z digest=sha256:b771c84b3d92b5aa79e7f4c3703df0f6ab637153cbe98ced76b70a2165d9ee6f

Observation 0ed65fc5-1a08-49c4-95dc-657b8e8cfabb · outbound

This paper cites Introducing our next-generation audio models, Mar 2025.

AHELM: A Holistic Evaluation of Audio-Language Models Introducing our next-generation audio models, Mar 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.563178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.848693Z digest=sha256:93cd834df46585313c7ecf92713cad13ea7fead2ebc11fb0c3d25f716a2e1b16

Observation 1d37ba2b-6924-4f40-bb80-fbaad1e10bbf · outbound

This paper cites LibriSpeech: an ASR corpus based on public domain audio books.

AHELM: A Holistic Evaluation of Audio-Language Models LibriSpeech: an ASR corpus based on public domain audio books

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.553483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.851608Z digest=sha256:0f10cf8f153c6317a104ab398df02ee34b74bbac5c9a0498eee739d87c48fadc

Observation 0a627daf-dfb1-40e9-82a3-149fe1217b18 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions.

AHELM: A Holistic Evaluation of Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.545031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.855344Z digest=sha256:ccc7ff0942dc51bd4c97e66bac0b62493e83c3b4d4691d54ca885a349d2f305c

Observation 4a345e82-50e5-47fb-8f4e-f8f82df64bad · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

AHELM: A Holistic Evaluation of Audio-Language Models MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.858777Z digest=sha256:70e31b8224d8ec724376b4ea40d1174a5fc53c70c4441788dd64437b3ac6d3ac

Observation 0c9017b8-63ac-46df-81f9-2c5038f0525e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.862427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.862427Z digest=sha256:54b32b28f231db05cf0c38688b3feb17fbd8b55cf0a1bc4dbe71972902f13eaf

Observation 20219ee9-5565-425d-80c8-3895efb66f6e · outbound

This paper cites Speech Robust Bench: A Robustness Benchmark For Speech Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech Robust Bench: A Robustness Benchmark For Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.865850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.865850Z digest=sha256:951fe49aeed30d291be079c4c8d91dbc06e86af84d4e1ee0504342dad7bc92e1

Observation 90cb00ee-5387-4568-af70-1a4c4781ecf3 · outbound

This paper cites V oice jailbreak attacks against GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models V oice jailbreak attacks against GPT-4o, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.530681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.869568Z digest=sha256:dd1e1fc51f0a15602db80d6cf463ab1029735f5d3cdc56f34d118545e5ebe647

Observation 09d16adf-5107-46f1-bb0e-e5f9d8bf2d42 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

AHELM: A Holistic Evaluation of Audio-Language Models Salmonn: Towards generic hearing abilities for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.521493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.872720Z digest=sha256:cf2660259df19ef1adb86017a724791a9bb199cc2c2149fec0555f023ab14e55

Observation 4e4f081b-e1f0-4e0c-887e-5c4cd491836a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.875777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.875777Z digest=sha256:b58a605c30aa4cd5e622489f314e326cfe1ebe5ed18053ed6583f4e73a1070a2

Observation 3abc1829-0a24-4cb6-bc49-48c29cb16690 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

AHELM: A Holistic Evaluation of Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.879225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.879225Z digest=sha256:f2bfca55144afabaf4a4df9e20a5ede0d1388d0fbd59b5dc30dc8ba6b690e1aa

Observation b57f1fcb-69fa-49dc-8cfe-fa202f93dd4e · outbound

This paper cites Qwen2.5-Omni Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2.5-Omni Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.882641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.882641Z digest=sha256:3e71a6fa37a93037af641b616501f1114fc86a58d9c44c616a2b103675029783

Observation 81ee07bf-d9ee-44bb-b7c7-8ea7ded00799 · outbound

This paper cites Qwen3 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.886194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.886194Z digest=sha256:28f2f82a9609c1af09e65d677607016983bf1579ceb77946d059abf993f160bb

Observation d1b45dbc-f6be-4405-871f-4538421e4457 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

AHELM: A Holistic Evaluation of Audio-Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.889693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.889693Z digest=sha256:ae01489f309f87596a06972f04110ebb47648408c4e3a020445ccbfa69a86c63

Observation 4d5afb4d-7645-48d4-86bc-306b24464242 · outbound

This paper cites Speechlm: Enhanced speech pre-training with unpaired textual data.

AHELM: A Holistic Evaluation of Audio-Language Models Speechlm: Enhanced speech pre-training with unpaired textual data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.512709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.893240Z digest=sha256:159556840f0653ef5e35536559e5f7fcb9e0dbbea987d86945940905071af5d4

Observation e4df7286-3747-4109-a699-fcaec0306b8a · outbound

This paper cites Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese.

AHELM: A Holistic Evaluation of Audio-Language Models Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.896569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.896569Z digest=sha256:78c33cb25808a904f04c4a49e2d56f9467d03e63eb639f64dac8e88de384b686

Observation f46aaaf8-49d0-4b7f-9471-8852c1804a6d · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.503933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.901477Z digest=sha256:cbefd2d6450218e66c673d15909f3c13f8365cc9c90a770ad91767e8c53dd40f

Observation 6b02af22-b3d2-4244-a430-85bf59d46073 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.495380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.905490Z digest=sha256:7c90c2ee8a6a11e4ffc822a3e4da102c9b69455aab846189ecdd2229c5e9d491

Observation 94f55932-df50-4f9f-ae84-315f1317d399 · outbound

This paper cites humorous and imaginative.

AHELM: A Holistic Evaluation of Audio-Language Models humorous and imaginative

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.487336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.908871Z digest=sha256:9f0a6b370541d79b369bf822a02168e40092dfa1522b71b4ffa44a7d3f18affe

Observation 9122540e-17ca-44f8-9f7d-e72828c7eedb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.478978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.912676Z digest=sha256:b232a9010753a7d4fcb08e05f955c371d4b77b86d92c13c24968c8f483db1b95

Observation 36bc4d4e-83e2-4d0e-bf9d-0e3c89ed6915 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.470623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.916173Z digest=sha256:a73c65e9e72c3c2fec086dfea3fb94a58dfee7d7a3e917aa727131f705ad4bfb

Observation 746c16fa-8b67-4ab6-8ddd-f90736f57628 · outbound

This paper cites E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content.

AHELM: A Holistic Evaluation of Audio-Language Models E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.461880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.919399Z digest=sha256:e885577a4837f6db2b92367e7f4c2ec909bbe5b3ab60e90ff15c0a8a0450188e

Observation 1953059c-e1a3-4b3a-9680-cad11029ab8e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.922630Z digest=sha256:f0156c718c06281fc85119c24a9b5be477a0f78924c558e3778338cf460318f8

Observation 7100bcb0-534f-44b1-853f-e0a1fb4429cb · outbound

This paper cites You should refer to the score rubric.

AHELM: A Holistic Evaluation of Audio-Language Models You should refer to the score rubric

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.445050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.926181Z digest=sha256:3228bbc01976a4a823395ebb8b1d06cdbcb0d0413b04f49a249bf78ca253055c

Observation 91bc551c-76b9-45aa-bc4f-37b5f5d10b12 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.437420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.929628Z digest=sha256:aaceae70f1c2f629f258508e3b9bae22a0be0240148306ab55c1ebf68fca2930

Observation fd9eec2a-1908-42c0-81cc-67730c0c68d4 · outbound

This paper cites haha”) or throat clearing (e.g., “ahem.

AHELM: A Holistic Evaluation of Audio-Language Models haha”) or throat clearing (e.g., “ahem

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.429327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.933794Z digest=sha256:70f91fbb82d069bf029fac7b80a113d0a0cc38aa1bc38563b22c269123da43e0

Observation cafd4911-0adb-4276-8e8d-1cd50d4a9aeb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.421481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.937749Z digest=sha256:a5d118db6d1158385b5665562df33a9d933d5b830b91beca5b1c4b8dd2a96f24

Observation d000f3ec-6c48-4317-99a0-4a1f55975eaf · outbound

This paper cites From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash.

AHELM: A Holistic Evaluation of Audio-Language Models From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.413407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.940870Z digest=sha256:ab3f40b2a55b11243eb2aacd2dfcfe97ed961edaa72496e714e74f36848295ba

Observation 0ad1c622-00ad-4d0f-be3a-b2dc963f88b7 · outbound

This paper cites When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack.

AHELM: A Holistic Evaluation of Audio-Language Models When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.404930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.944391Z digest=sha256:8a973043d10870af452424b6aaa8f2f11a96b8f0e77626a05ce1fc9c6872431d

Observation ce0143f0-edec-44b0-a2c5-f3104c8eec86 · outbound

This paper cites We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5.

AHELM: A Holistic Evaluation of Audio-Language Models We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.396940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.947659Z digest=sha256:87c194651eccfda09e499e83025904eced22e49ad5a5b2e12f39cc7240a83990

Observation fe1d0241-7fb1-4404-878f-8d5d953d958a · outbound

This paper cites Limitations.

AHELM: A Holistic Evaluation of Audio-Language Models Limitations

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.388203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.950813Z digest=sha256:d94d29461189114b2e6845a2bb744bfd33eaead30afdc61000537d27386888e3

Observation 94467b2d-8bef-4025-b527-944ac890cf07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.379320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.954108Z digest=sha256:001805438dc1561a84a648d46e75867862c53826be118036d2a54b45b3593bbf

Observation 865d1775-6ed3-449e-a803-01997f63b213 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.371248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.957811Z digest=sha256:1de34190916779ffe812c5af197e2e136e5aeff33d0ffb549bcd46c47eba0cff

Observation bed55278-f3a2-4c5b-942a-36e402193d8b · outbound

This paper cites com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1.

AHELM: A Holistic Evaluation of Audio-Language Models com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.362856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.961310Z digest=sha256:0a87361748bb509584e7800bb6da25d9d5b88fd4bbcf442bd042139f5e276b2a

Observation 31b3c5a6-d911-4336-a6d4-ab1b7ac171cd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.354147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.964523Z digest=sha256:0197c2017ab12b42425f4478d2cea174b9bc26109342ff12853fc5010c084a7d

Observation c9c61ba0-91ae-45a4-bcfd-6cd5a60f4db5 · outbound

This paper cites But we do not compute error bars for other scenarios.

AHELM: A Holistic Evaluation of Audio-Language Models But we do not compute error bars for other scenarios

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.344723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.967592Z digest=sha256:f8fc3208319ac6af5422d608c60be2d0d770d4153b3937f6c0a1a685dfb367cd

Observation 2b1cd651-9153-45be-a311-d04b174b3f1d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.335836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.970937Z digest=sha256:c5b66c997e4ef6241b18ae56345e662af5172df9adb75b2ffcd19f00a1b310fd

Observation 0ba18208-7e04-4b5c-8bd6-64e4f23efcec · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.327529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.974113Z digest=sha256:1e73af6fc1a65af6d2447067e6309432c006e602f000b8b3d174ee3cb7b12913

Observation 1aacc2fa-b354-4c1b-82a3-e84f0db5faf4 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.318849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.977488Z digest=sha256:b3923b46eb52a1c0496ab8380292de77b400f865b9254104ced7a7f89e201076

Observation 462c6a8a-0ff0-46e1-ae74-e02ca8c24b93 · outbound

This paper cites Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata.

AHELM: A Holistic Evaluation of Audio-Language Models Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.309287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.980801Z digest=sha256:76e7cdb4da35023c012ce924d233fc9b73163f0b182605f6cf7361fcf4f41536

Observation 979de6e4-0d9a-433a-8dfa-f1b0852bb2da · outbound

This paper cites We cite all the datasets and models used in our work.

AHELM: A Holistic Evaluation of Audio-Language Models We cite all the datasets and models used in our work

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.300092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.983943Z digest=sha256:eef32a40f81400ff3bd4a924d91cd2afb1984b962ffd92a5507a7489267cf2b8

Observation e119bfdd-abc9-4fae-92af-603f17650782 · outbound

This paper cites PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio.

AHELM: A Holistic Evaluation of Audio-Language Models PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.289975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.987392Z digest=sha256:8e6f194f40c675bb6dc5653e17294097d0a4a2c71ae3881b53cf7155e738ce17

Observation 85324c3b-c3d3-4ad6-904e-fa33b94ec416 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.279914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.990404Z digest=sha256:5cf039052235d5555a9460b15fa62879f683c8a788d0903371dbb226e51e5a37

Observation 48de1972-aa06-47ee-883c-a22d9a349457 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.269972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.993270Z digest=sha256:e3c92b4e69447b31625232ad09be26e30e06b46e62046fac0d05e1424aee02f9

Observation 913828ab-683a-4f62-91d4-fe2b0213da28 · outbound

This paper cites Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark.

AHELM: A Holistic Evaluation of Audio-Language Models Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.259930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T14:22:34.997009Z digest=sha256:ae79826e8572400cfbf59a87aaa94a4e3b6fd325601c7f8ee3addf21f0676893

Pith citing papers

Observation 438bb6eb-afa4-4118-9e01-f38ceee8c266 · inbound

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs cites this paper.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs AHELM: A Holistic Evaluation of Audio-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.185900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:bf051889914466dff08dec9915d7b20bedf64f56d8d56ed954d33d3004bf60fd

Observation 420de55c-2f6c-49e9-9414-845e6f2c69af · inbound

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics cites this paper.

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics AHELM: A Holistic Evaluation of Audio-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:52:45.801592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:52:45.801592Z digest=sha256:a819acf161e3419843232a7ba05a709a4746e3d18ee3cf58db02ca919053f10c

Observation b118258d-1bab-4c12-ba6a-d91a2b208abd · inbound

PRiSM: Benchmarking Phone Realization in Speech Models cites this paper.

PRiSM: Benchmarking Phone Realization in Speech Models AHELM: A Holistic Evaluation of Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:31.131370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:31.131370Z digest=sha256:0f24d0330686e07a567da8e77c4c5997038ec9c418edfe0dee1b77cc78914e7e

Observation 3af163d3-1a39-4b7a-8f15-d33576402f1d · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where AHELM: A Holistic Evaluation of Audio-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:22.088783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:2a55f436f2254b303c37b0a5410bae8ce2ee543d8253dd3b66ce99a9f078a1d0

Observation a2e3a273-2a73-4c70-9492-e80937080f18 · inbound

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech cites this paper.

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech AHELM: A Holistic Evaluation of Audio-Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.226416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:46:50.923340Z digest=sha256:0e4be75127714faf9f6d15b94ed1856129362f2ba9995029298fe4d4bd3aec2c

Observation c7a02f6c-211f-4b94-ba6b-cd699f7bcfc5 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI AHELM: A Holistic Evaluation of Audio-Language Models

Reference 227

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:55:29.197360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:74d11a7d7a28815f472a98a4142f0c94d13b0bd91528e82798b2a418758a0c53

Observation ec1072ce-d1cb-4a1e-97ae-f072a79f6845 · inbound

AudioMosaic: Contrastive Masked Audio Representation Learning cites this paper.

AudioMosaic: Contrastive Masked Audio Representation Learning AHELM: A Holistic Evaluation of Audio-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.861492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:89454cf5e374a24e510bafd2f344c0527d4fbda0f871d48a61e06b01d9e42350

Observation 6e1dc822-8c61-4464-b2ea-73165ac7f0a0 · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities AHELM: A Holistic Evaluation of Audio-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.430007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:7815ffc1e041c7fb7314e9798d77acaa05ce86f3b055995a6392d4717050c34b

Observation 0a7d2131-ece3-4fec-9428-112d78267f36 · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages AHELM: A Holistic Evaluation of Audio-Language Models

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.875641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:45cde3dadb7117332e6d8feb7863f3f25aa5e75de95b593377bb57dc3a14e44b

Observation 6d5d8abb-cfc8-46d0-a5f4-cc38b5f97d4c · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems AHELM: A Holistic Evaluation of Audio-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.780637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.780637Z digest=sha256:d34aaa9c5056a25ae7d7cad8045ec62a7f88d4f10ee3a641743c644f043d979b