REVIEW 1 cited by
A baseline revisited: Pushing the limits of multi-segment models for context-aware translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper addresses the task of contextual translation using multi-segment models. Specifically we show that increasing model capacity further pushes the limits of this approach and that deeper models are more suited to capture context dependencies. Furthermore, improvements observed with larger models can be transferred to smaller models using knowledge distillation. Our experiments show that this approach achieves competitive performance across several languages and benchmarks, without additional language-specific tuning and task specific architectures.
Forward citations
Cited by 1 Pith paper
-
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
Decoder self-attention heads attending the target-side antecedent are the most influential for pronoun disambiguation, and fine-tuning them yields up to 5 percentage points improvement on contrastive tests.
Discussion (0). Continue with ORCID to comment.