Pith. sign in

Paper Citation Record · LEDGER

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 8 inbound Pith citation observations for arXiv:2501.04962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04962 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:24:18.303318Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:52:46.606480Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:34:53.233804Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fad0d4c6-c0c6-4c99-a06d-7ca0ab98dc5d · outbound

This paper cites Qwen Technical Report.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.125848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.125848Z digest=sha256:b555ef1eef04a0d20168ab3a9889eab9cb591b28439f06aed39094d99c3e3c21

Observation 19db4261-2075-4f44-8b65-a0faa0eeb656 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.131575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.131575Z digest=sha256:e3057b282c4cc67ce66d97a99bc14d7d3405f982d3682f8104f7fb547d9576e0

Observation ae460761-6a92-4fc1-a873-73aaa59c493d · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:19.065504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.136983Z digest=sha256:2605c7687253f71676927bb3f1126c1fdcb4929053f92c710003b893a79e0723

Observation 48959be5-ca3c-47ec-8509-976e29cc2c12 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.142196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.142196Z digest=sha256:69acf1243c60b08d57e639192811d83c9f43ba5e9cdafeec07a4e4464a83a6b1

Observation 6337ac7c-7266-48b3-a938-dccfb69095f8 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.146813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.146813Z digest=sha256:2da219416b4aa47cc8ed43230a9ad986ef4030ed9e54bd641f3afaf5df1c644f

Observation 45c47c5f-d066-4259-ad05-7aed70135de9 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.151012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.151012Z digest=sha256:bc8f347a617f5d0bf09d3c13297804e550ece7c91e4848ee92e1a65e8afc633d

Observation 5729fd66-3f1a-4f94-9881-412aad4d1d11 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.156186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.156186Z digest=sha256:289650040dc26341ff6124e9d71e9f37972380bf37fd9f41f6788402c36d7e7b

Observation 28d4fa19-a834-4a3d-ad53-419245d4cc2e · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Recent Advances in Speech Language Models: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.160313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.160313Z digest=sha256:3eb62477ed55609686cc34d776a19ecd823ef94c9ee3cd9f7c7b9c17ee7d4fcb

Observation 20498099-e377-4725-baf0-c6a9a9693ffc · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:19.027611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.165155Z digest=sha256:70c68286db7c97f49493e270299d850c1c2927485a7d239ea4a03d5af229a029

Observation 87f6de81-a8db-4e41-a0a7-51a20e363dd2 · outbound

This paper cites The Llama 3 Herd of Models.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.170213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.170213Z digest=sha256:1fb05ae35d5658b615eee5de49f06b3a289b79b27d2582fc6008be57ef06f938

Observation 2dab59f3-c066-4cf6-9df0-9844c6856dd4 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:19.012691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.174919Z digest=sha256:d6e57a3f7bf039a4ed8efd58c0aa30eeba1218e5d2b6796a5937411df996f2ba

Observation 9bd79c1e-cadf-4018-b856-f3bcd298affc · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:18.997765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.179641Z digest=sha256:7cc73f0f180758eab2ea9b8a06edefb39e9f48d97c808f233181082f23c8ecb9

Observation 59220eb8-a44f-4616-8dce-4fe6a104445a · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.184897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.184897Z digest=sha256:404e8fd76d36c00b3dd1918a577534a431f61c4042b5021c28586a41883403ce

Observation 11c44623-f80e-45ca-b817-daca864c1833 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.189539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.189539Z digest=sha256:e12c312ed9a99e332a760c0056e6d7314d4f303cceadf286e2481b017b734f0b

Observation 2c23c7e0-2b4d-43ea-b1ad-4b84f35a5472 · outbound

This paper cites GPT-4o System Card.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.194347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.194347Z digest=sha256:3061e9cf5d2928376e429cd34c9ab17c0cc692f9153649bc85e3576ff8d8ba0f

Observation 8ac7f3f8-c0bd-4f39-a57c-07655dc3e435 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.199063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.199063Z digest=sha256:dd301b526be5a707752b4030dba1d46446cb7ef4545e747c70d3e5261c5b591a

Observation a4558be5-065a-4adc-aa60-7e2a46257eb7 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.203736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.203736Z digest=sha256:0e052ba7b8d0d9e81a6fd0b87b894e516ff8217a3b932c8caea09d360fc19251

Observation 35d5d75b-8a3d-4077-befc-d7ede46348de · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.208085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.208085Z digest=sha256:7f622f239906a4fe6a24748cde2707c71c5125314292a4fe148a99620fa0eb90

Observation 7e48b1de-50d6-4fab-8c26-e80b9894cb5c · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.212520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.212520Z digest=sha256:2480f3045c8992ccd3d207346e438babc0487bce7bec68d974342fbc13b54052

Observation c7fa2ae9-7670-4df1-941a-87d2747c17be · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:18.963414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.217092Z digest=sha256:e6dcd071a3d5a1f4dee2ea72a4c6f642d1712d3b90dcd83cd77c9d597d4fc308

Observation 2f05ebf8-432e-4e85-804a-3db526405eb3 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Spirit LM: Interleaved Spoken and Written Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.222343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.222343Z digest=sha256:c5da805b11248d86a539c4dcd2bd46db0366ba474adcd10dfcf95b19daed28f5

Observation 9476f41d-b8e0-4dca-b298-6d78b03cab51 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.226969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.226969Z digest=sha256:586d138f239753ac03641456e3c8e0370dcc6168063801325fe96e8deb0f0649

Observation dd32c903-250c-4d7e-98f9-42f9efa62f44 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.231794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.231794Z digest=sha256:a8bf78e5bbcdf1acc39a46aeb10efa6d3ddc221407d02b6f8d48ad430ee1f948

Observation 6fd8c8c5-92cb-4ce2-b95b-7f854feb565c · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.236015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.236015Z digest=sha256:ee0291f5c65f4e6902db532b5b7d8d0e8f5bde8d96f5da93f88f9c70e4752b38

Observation 610b4d0a-f33f-4c57-846c-ff000d6244f4 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.240361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.240361Z digest=sha256:368b61bb0d25552d205ab07350cbedbd028d3c4954bc14cce7cc5b976a9f4f13

Observation 08c65c45-4b6a-431a-a241-c0b627dcb9b5 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:18.937319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.244358Z digest=sha256:bcefc679f04f1052f2b858f4f41755117f57dd45352d24058e29af0f48ac44e2

Observation 87d70989-59de-4d26-8a40-1e03013178f3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.248265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.248265Z digest=sha256:dc9ee8013377d399283bc0fbd4c8422bfe575a89742d19d9d1ecb536e9cd3820

Observation 4e06397e-35b5-46fa-a8a7-9956ebe86798 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.252562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.252562Z digest=sha256:434d3eb81b83662a755acb51bf1da655d85d76277613765e018e5e3bcd041df3

Observation e87f1a87-d1f4-4327-82aa-c7ad876f3b05 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.256256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.256256Z digest=sha256:5f2defd48c57f007064962d8828bd169f8385e8faaf59750e5fa928036455118

Observation ee6b270a-b4b1-47dc-9918-6963d52c28d9 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.260746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.260746Z digest=sha256:edbe2b1d4a67bc20e9e87c4933085fbd5d8a36c1e31ad30e0dee97a1aa2730df

Observation c2d80cbd-c6ae-4db9-ab2e-d97aa2f75a18 · outbound

This paper cites Chi, Quoc V.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Chi, Quoc V

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.265248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.265248Z digest=sha256:b65c38bba1ca1a56fa75b05dead00bedbd6c02c7ab0cdebd243af8c05d12f5d5

Observation aea2079a-dcd9-43dc-91f0-e893ef2e77d4 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.269589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.269589Z digest=sha256:04f60fb25e3ddd5e3f72de7d3ec8f73217fceb513b771bce373236585ec847dc

Observation 65f4bbd2-44dc-41a5-a091-32cf45b75a47 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.274403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.274403Z digest=sha256:51fc975574facaae0a86597edbc40875267df697f236fdfa91bd53a5f1322e34

Observation 3d5e7723-6d36-414f-a1d0-1646215d8e72 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.278793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.278793Z digest=sha256:fafa72a4031c1bb2cdd9a5c165cd072d2b29a232619d2a615b68965c8b5ecec4

Observation 4d7153d9-eb86-47e6-b56b-c68a5c328885 · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.283701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.283701Z digest=sha256:a251d910b2adf99c1a04c26e0ce5cb84b76992b5b8e2aa821b28a0d008d13443

Observation cadb4aaa-8ca5-4d76-bd9f-41a3c2e5502b · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:24:18.891239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:24:18.288518Z digest=sha256:70fef6ba4b600f08d0365357dea5776a7ccd50aed17bd8ad97a978fe4dbb7c5d

Observation c5b943a3-0338-4ef5-8930-60100bf10b9f · outbound

This paper cites an unresolved cited work.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.293030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.293030Z digest=sha256:061a05e7e584e40af358b8cbf4c51d734400aff6f4db0e0374a24bf2cd8de81f

Observation 4373fd6b-6959-48ad-b7dd-65d7a90e62a7 · outbound

This paper cites online" 'onlinestring :=.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.297626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.297626Z digest=sha256:3be8d840ef798235e98d547f21490db8f8005d21224c39ed13e35eaf6fc22853

Observation 87d9d31f-fc97-445d-a6d7-35aab337e6d7 · outbound

This paper cites write newline.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.303318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.303318Z digest=sha256:9ae9860745685148ffe777bba6a68499cbf3cef60de607b906608783d5fc277c

Pith citing papers

Observation 3664a2fe-b193-4ef5-afd8-b995ec8f5b92 · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.156425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.156425Z digest=sha256:62e81e10847cb37f805cfdcae0d33b8562fae37f72ef8d595114bce3d5d3b211

Observation 808752eb-2d33-4068-9394-84cdbf49e175 · inbound

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey cites this paper.

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.236630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T13:32:57.771753Z digest=sha256:1a1c98f80fb4ce4d012d592208f69ca31f5104c9ccae2515f4b46c192d0a553e

Observation 7658def9-1a65-4a98-950f-4ccdda30307a · inbound

Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility cites this paper.

Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:30.950009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:07:30.950009Z digest=sha256:c0159758aedc9b0a5398c7cbac08bb50cc6c085a9663fe07a31fc01a985774fb

Observation 4352680a-1322-4fa0-b674-cc5878118272 · inbound

Audio-Aware Large Language Models as Judges for Speaking Styles cites this paper.

Audio-Aware Large Language Models as Judges for Speaking Styles VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.186300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.186300Z digest=sha256:214dad282c0233d24404e9d2fe83bdce1fd71648fa03db0f7fec51e0ec0628dd

Observation 8e701956-24d2-46c2-8c4c-b82a9999b347 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.394319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.394319Z digest=sha256:2a764975b7c3cb420303c0514791c046893b82ae50b106d014091cde0d9ad3c9

Observation a829bc1c-343f-4a9a-a9f3-66f18d204e8a · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.318974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.318974Z digest=sha256:17c013facaced8fc595fe177b396bf99fa3043f27c71b880d194ac867609cfbb

Observation 9aac127c-6064-45bb-80dd-44417d324763 · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:17:57.483334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:629042c29126005a865242d9de632ca9b97d964bbda9b41e33ef0b11f58fa99b

Observation a0d3d4a0-7ec4-41f1-858a-e52641e48c35 · inbound

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models cites this paper.

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:52:46.606480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:52:46.606480Z digest=sha256:cf49de0797ae0b708763ba45c52c638816a2d04769bb66946107708b0df45482