Pith. sign in

REVIEW 3 cited by

Recent Advances in Direct Speech-to-text Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11646 v1 pith:XEHF4GWM submitted 2023-06-20 cs.CL eess.AS

classification cs.CLeess.AS
keywords datamodelingtranslationworkapplicationburdendirectdirections
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges -- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages

    cs.CL 2025-06 conditional novelty 6.0 of 10

    For Turkish-French and Pashto-French conversational speech translation, the best cascaded pipelines combine Whisper or Microsoft ASR with Google or Microsoft MT, and component rankings are mostly stable across pipelines.

  2. Addressing speaker gender bias in large scale speech translation systems

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Fine-tuning a large speech translation model on GPT-4-reformulated gender-balanced training data raises MuST-SHE feminine-form accuracy from about 10% to over 84% without BLEU loss.

  3. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

Pith tools