Pith. sign in

REVIEW 17 cited by

URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.17810 v3 pith:EGJGGHV5 submitted 2025-02-25 cs.CL eess.AS

classification cs.CLeess.AS
keywords modelssdmsbenchmarkdialogueevaluationsspokenuro-benchevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios. To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model's abilities in Understanding, Reasoning, and Oral conversation. Evaluations on our proposed benchmark reveal that current open-source SDMs perform rather well in daily QA tasks, but lag behind their backbone LLMs in terms of instruction-following ability and also suffer from catastrophic forgetting. Their performance in advanced evaluations of paralinguistic information and audio understanding remains subpar, highlighting the need for further research in this direction. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    PolySpeech-100 is a new benchmark for native-level speech comprehension across 110 linguistic variants that evaluates 22 models and reports E2E advantages on dialects, robustness gaps on low-resource languages, and de...

  2. DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

    eess.AS 2026-05 unverdicted novelty 7.0 of 10

    DuplexSLA introduces a three-channel full-duplex architecture that synchronizes continuous user audio, discrete assistant audio, and rate-limited textual actions inside a single backbone for native turn-taking and in-...

  3. Evaluating the Expressive Appropriateness of Speech in Rich Contexts

    eess.AS 2026-05 unverdicted novelty 7.0 of 10

    CEAEval is a context-aware evaluation system for speech expressive appropriateness, supported by a new Mandarin dataset with multi-dimensional human annotations and a model that outperforms prior systems.

  4. VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    VITA-QinYu is the first expressive end-to-end spoken language model supporting role-playing and singing alongside conversation, trained on 15.8K hours of data and outperforming prior models on expressiveness and conve...

  5. Liberating LLM Capabilities in Full-Duplex Speech Models

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    LWS is a text-first paradigm for full-duplex speech LLMs that treats visible writing as a primary output channel alongside audio input and spoken response, implemented via token schema and synthetic per-second annotations.

  6. Game-Time: Evaluating Temporal Dynamics in Spoken Language Models

    eess.AS 2025-09 unverdicted novelty 7.0 of 10

    Game-Time Benchmark shows spoken language models handle basic tasks but degrade sharply under temporal constraints like tempo adherence and synchronized responses.

  7. Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Spoken math models that emit a 40%-compressed reasoning trace between question and answer beat full-reasoning baselines by ~3 accuracy points while using roughly one third of the text tokens.

  8. VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A Chinese benchmark built on real human speech evaluates large audio language models across instruction following, knowledge, and robustness, revealing large performance gaps.

  9. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    MPS proposes a dual-brain architecture separating formulation reasoning from articulation to achieve real-time CoT in SLMs with accuracy comparable to full pre-computation but much lower latency.

  10. Step-Audio 2 Technical Report

    cs.CL 2025-07 unverdicted novelty 6.0 of 10

    Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and c...

  11. Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

    eess.AS 2025-05 accept novelty 6.0 of 10

    The survey introduces a four-category taxonomy for LALM evaluations and reviews benchmarks across general auditory processing, knowledge reasoning, dialogue, and fairness-safety.

  12. Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    A dual-agent closed-loop system integrates Theory of Mind reasoning with multimodal video generation to create social avatars that outperform full-information baselines on dialogue quality under information asymmetry.

  13. DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

    eess.AS 2026-05 unverdicted novelty 5.0 of 10

    DuplexSLA is a dual-stream three-channel full-duplex model that synchronizes continuous user audio, discrete assistant audio, and rate-limited action text for native turn-taking and in-conversation tool calling.

  14. A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

    cs.SD 2026-05 unverdicted novelty 5.0 of 10

    A survey of Large Audio Language Models that establishes a taxonomy of trustworthiness vulnerabilities and proposes a Defense-in-Depth roadmap for audio intelligence.

  15. Qwen3.5-Omni Technical Report

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Qwen3.5-Omni scales an omnimodal model to hundreds of billions of parameters with 256k context, introduces ARIA for stable speech synthesis, and reports SOTA performance on 215 audio-visual benchmarks while adding mul...

  16. VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

    cs.CL 2025-09 reject novelty 5.0 of 10

    A 65.6-hour movie-dialogue benchmark for spoken role-playing agents, with an evaluation framework whose main judge is also an evaluated model.

  17. A Survey of Audio Reasoning in Multimodal Foundation Models

    eess.AS 2026-05 unverdicted novelty 2.0 of 10

    A survey that provides a unified formulation of audio reasoning and reviews advances across Audio-to-Text, Audio-to-Speech, Audio-Visual, and Agentic paradigms while discussing challenges and future directions.

Pith tools