A ~50-hour-per-language multi-accent S2ST dataset for four Nigerian languages shows few-shot AudioLLMs beat fine-tuned cascaded/E2E systems on speech-to-text, while speech-to-speech remains comparable and underdeveloped.
InProceedings of the Seventh Conference on Machine Translation (WMT), pages 46–68, Abu Dhabi, United Arab Emirates (Hybrid)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
A ~50-hour-per-language multi-accent S2ST dataset for four Nigerian languages shows few-shot AudioLLMs beat fine-tuned cascaded/E2E systems on speech-to-text, while speech-to-speech remains comparable and underdeveloped.