REVIEW 3 cited by
Recent Advances in Direct Speech-to-text Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges -- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.
Forward citations
Cited by 3 Pith papers
-
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
For Turkish-French and Pashto-French conversational speech translation, the best cascaded pipelines combine Whisper or Microsoft ASR with Google or Microsoft MT, and component rankings are mostly stable across pipelines.
-
Addressing speaker gender bias in large scale speech translation systems
Fine-tuning a large speech translation model on GPT-4-reformulated gender-balanced training data raises MuST-SHE feminine-form accuracy from about 10% to over 84% without BLEU loss.
-
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.
Discussion (0). Continue with ORCID to comment.