Pith. sign in

REVIEW

Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.07575 v3 pith:VSGJRCUJ submitted 2019-09-17 cs.CL eess.AS

classification cs.CLeess.AS
keywords pre-trainingend-to-endfine-tuningspeechconsistentencodermethodsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model outperforms baselines 2.2 BLEU on a large benchmark dataset.

Discussion (0). Sign in to comment.

Pith tools