Pith. sign in

REVIEW 2 cited by

Word Alignment by Fine-tuning Embeddings on Parallel Corpora

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.08231 v4 pith:H4AL4DY5 submitted 2021-01-20 cs.CL

classification cs.CL
keywords wordalignmentparallellanguagemodelspre-trainedcorporademonstrate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Word alignment over parallel corpora has a wide variety of applications, including learning translation lexicons, cross-lingual transfer of language processing tools, and automatic evaluation or analysis of translation outputs. The great majority of past work on word alignment has worked by performing unsupervised learning on parallel texts. Recently, however, other work has demonstrated that pre-trained contextualized word embeddings derived from multilingually trained language models (LMs) prove an attractive alternative, achieving competitive results on the word alignment task even in the absence of explicit training on parallel data. In this paper, we examine methods to marry the two approaches: leveraging pre-trained LMs but fine-tuning them on parallel text with objectives designed to improve alignment quality, and proposing methods to effectively extract alignments from these fine-tuned models. We perform experiments on five language pairs and demonstrate that our model can consistently outperform previous state-of-the-art models of all varieties. In addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs. Our aligner, AWESOME (Aligning Word Embedding Spaces of Multilingual Encoders), with pre-trained models is available at https://github.com/neulab/awesome-align

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Word Level Timestamp Generation for Automatic Speech Recognition and Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.

  2. Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Machine-translated task data is the best average parallel-data source for cross-lingual transfer of vision-language encoders, but authentic caption-like data beats it in some languages, and multilingual training helps...

Pith tools