REVIEW 2 cited by
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.
Forward citations
Cited by 2 Pith papers
-
It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems
End-to-end speech translation systems translate idioms worse than text-based systems, frequently producing literal or incorrect outputs, across German and Russian to English.
-
GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
Fine-tuning SeamlessM4T-v2 directly for end-to-end speech translation is competitive, and ASR-encoder initialization adds about 1 to 5 BLEU for languages unseen by the base model.
Discussion (0). Continue with ORCID to comment.