RedVox benchmark shows speech model safety and fairness vulnerabilities persist under non-adversarial conditions, worsen in non-English languages, and increase with spoken inputs.
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
6 Pith papers cite this work. Polarity classification is still indexing.
abstract
As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks, bypassing traditional transcription-based pipelines. Whether this integration improves ST quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 6 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable solution overall, but most recent SpeechLLMs can match or even outperform cascades in various settings while SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.
fields
cs.CL 6years
2026 6verdicts
UNVERDICTED 6representative citing papers
Ouvia is a user-centered evaluation framework for speech translation usability in real-world scenarios, showing limited usability rates and the superiority of QA-based metrics.
Proposes STEL task with protocol and dataset; shows XCOMET and Qwen2.5-Omni label errors at roughly half human precision and that speech processing is required.
Meta-evaluation on gender and prosody contrastive datasets finds text and speech quality estimation metrics fall short at assessing speech-specific features, including newly trained SpeechCOMET models.
Pearmut is a platform that makes end-to-end human evaluation of translations as easy as automatic metrics by supporting DA, ESA, MQM and features like document context and attention checks.
A cascaded SimulST system using Parakeet and Qwen 3.5 with adaptive black-box policies and RAG context achieves +5.82 XCOMET-XL improvement on En→De for IWSLT 2026.
citing papers explorer
-
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
RedVox benchmark shows speech model safety and fairness vulnerabilities persist under non-adversarial conditions, worsen in non-English languages, and increase with spoken inputs.
-
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
Ouvia is a user-centered evaluation framework for speech translation usability in real-world scenarios, showing limited usability rates and the superiority of QA-based metrics.
-
Automatic Labelling of Speech Translation Errors
Proposes STEL task with protocol and dataset; shows XCOMET and Qwen2.5-Omni label errors at roughly half human precision and that speech processing is required.
-
Why We Need Speech to Evaluate Speech Translation
Meta-evaluation on gender and prosody contrastive datasets finds text and speech quality estimation metrics fall short at assessing speech-specific features, including newly trained SpeechCOMET models.
-
Pearmut: Human Evaluation of Translation Made Trivial
Pearmut is a platform that makes end-to-end human evaluation of translations as easy as automatic metrics by supporting DA, ESA, MQM and features like document context and attention checks.
-
MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task
A cascaded SimulST system using Parakeet and Qwen 3.5 with adaptive black-box policies and RAG context achieves +5.82 XCOMET-XL improvement on En→De for IWSLT 2026.