OmniVoice introduces a diffusion language model-style non-autoregressive TTS system that directly maps text to multi-codebook acoustic tokens, scaling zero-shot synthesis to over 600 languages with SOTA results on multilingual benchmarks using 581k hours of open data.
Habibi: Laying the open-source foundation of unified-dialectal arabic speech synthesis
2 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
The authors construct and evaluate an end-to-end speech-to-speech pipeline for Algerian Dialect by adapting Whisper for ASR, transformer embeddings for NLU, and a neural TTS on custom dialectal data.
citing papers explorer
-
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
OmniVoice introduces a diffusion language model-style non-autoregressive TTS system that directly maps text to multi-codebook acoustic tokens, scaling zero-shot synthesis to over 600 languages with SOTA results on multilingual benchmarks using 581k hours of open data.
-
Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
The authors construct and evaluate an end-to-end speech-to-speech pipeline for Algerian Dialect by adapting Whisper for ASR, transformer embeddings for NLU, and a neural TTS on custom dialectal data.