Pith. sign in

REVIEW 2 cited by

BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.04588 v1 pith:MVLYFWKT submitted 2021-09-09 cs.CL

classification cs.CL
keywords translationmodelspre-trainedbertcontextualizedembeddingslanguagebibert
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of bidirectional encoders using masked language models, such as BERT, on numerous natural language processing tasks has prompted researchers to attempt to incorporate these pre-trained models into neural machine translation (NMT) systems. However, proposed methods for incorporating pre-trained models are non-trivial and mainly focus on BERT, which lacks a comparison of the impact that other pre-trained models may have on translation performance. In this paper, we demonstrate that simply using the output (contextualized embeddings) of a tailored and suitable bilingual pre-trained language model (dubbed BiBERT) as the input of the NMT encoder achieves state-of-the-art translation performance. Moreover, we also propose a stochastic layer selection approach and a concept of dual-directional translation model to ensure the sufficient utilization of contextualized embeddings. In the case of without using back translation, our best models achieve BLEU scores of 30.45 for En->De and 38.61 for De->En on the IWSLT'14 dataset, and 31.26 for En->De and 34.94 for De->En on the WMT'14 dataset, which exceeds all published numbers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeistBERT: Breathing Life into German NLP

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A 126M-parameter German BERT, pretrained further on 1.3TB of mixed German text, beats other base models on most tested German NLP benchmarks.

  2. Transferable Modeling Strategies for Low-Resource LLM Tasks: A Prompt and Alignment-Based Approach

    cs.CL 2025-07 reject novelty 2.0 of 10

    A prompt-and-alignment fine-tuning recipe is claimed to beat multilingual baselines on MLQA, XQuAD, and PAWS-X under low-resource data, but lacks reproducible experimental detail.

Pith tools