Pith. sign in

Paper Citation Record · LEDGER

URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2502.17810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17810 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:36:06.472390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db7cb67f-4973-4c2f-9e03-74dc6668fbff · inbound

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey cites this paper.

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.242696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T13:32:57.771753Z digest=sha256:c235bece7684c480e5979269660f56cad19d134c9fa3f1eb1fe156faf0dd62b7

Observation 106fd36e-5574-4659-a943-269086e2d72c · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:59:51.100939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:12c754792e84f5e301426c36faaf991b7456f2baea9bc58a416a4debf24ba998

Observation d17a3726-3a7b-4486-a695-6b0d4d77566b · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.472390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.472390Z digest=sha256:3ca95cbac575394475753f6913afdde37088e4f93ade3239603f1c434a1a6639

Observation ccd32e67-ce37-47b6-92ba-2e7f39352a61 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:52:35.642915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:79913de7f8a1a4d8335d99c03511210a608e2024da3a8da221046c1c0b1acb4b

Observation b60aa5e7-3069-4b7e-a59c-a4ba0bbc85f4 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:46:03.542900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:54091c46524187c019c6841083b6a0ffe16eba2c16c3f2917b63b62b2aa899a8

Observation aba706f3-94e5-400c-b97d-419f63c5aace · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.937202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.937202Z digest=sha256:4715c4b2a7c6f676cfe9250384649f20133c13db7f7285caffb94b62922ab3ce

Observation c3c10ec2-4a90-4f64-8d47-f4c4f0029954 · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.504670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:aa395fda4998cc2495537ad94ed98441df3874e8cdfaa2a25dd1ac7c0ba33671

Observation 5be8038e-bcc9-4550-ba44-09d2b089ce63 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.186750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:8e44131de79567ed6321f4b25e3e39ed9cfc70c31eda79c97e627b0541fa7236

Observation cab538ae-c1ca-426a-bfce-4a2d31dfbdcf · inbound

Evaluating the Expressive Appropriateness of Speech in Rich Contexts cites this paper.

Evaluating the Expressive Appropriateness of Speech in Rich Contexts URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:56:14.573707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:56:11.473572Z digest=sha256:e9d299cceca0f9b286f1eaf2ee5c98a5f5d2012c583b0e2df3b5d6b973c96748

Observation 47f6c4d8-2d33-41ab-8c32-a3db3e5085b2 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:49.047504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0c48dd65d17fc08c997320d4e52c55c7e6390c1068b98327592736d692639f69

Observation 330ba7bf-eb71-4b78-873f-bddb8994d002 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:55.060576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:2079aa774e29c0defd185e92a5c20feef2a89736d7a12c8c395351b5122b794c

Observation d981c1f0-b2bf-43c5-9c63-56797b1abc01 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.301296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:d19f1d1c66df5a15829dc8f32f4be8644357d7864c9e8cf8cea621a2e11fe44f

Observation bdd1c086-e010-41c2-8fd5-4207109c55dd · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:09:24.437109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:eee4ae5475c34a372f55665b9717ab511c521d2ddfb916a21e35056216c77d6d

Observation 4cad0548-0107-41e4-a95f-f154e1f5ac0e · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.009937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:a2abfc2f29fc45bdb69e3090796d312301715a01dd87ac4a9696afcabe39cb23

Observation b94e9da7-36be-463a-aa74-4ad1a086ea3d · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.235234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:b56f0308ab0f86ace006124d6ec6754ce48c919e2965bb438e31e0be4b9134e3

Observation 30a4e08a-3965-4cca-931a-e968a2deb774 · inbound

Liberating LLM Capabilities in Full-Duplex Speech Models cites this paper.

Liberating LLM Capabilities in Full-Duplex Speech Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:15:09.030108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T00:13:14.701980Z digest=sha256:f2b8835386c119dc1206bb31518b5bbde8f9ac400109d2b2182ea3303a638823

Observation 533cc13a-9506-4390-bb4e-c9a1d0468fef · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:18.079021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:18.079021Z digest=sha256:f6e49b9a9ec039085e72c5dfcf7b58931ebc5712180eac6f45617b5565c4272c