Pith. sign in

REVIEW 8 cited by

Does Neural Machine Translation Benefit from Larger Context?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.05135 v1 pith:KLKSLCJ7 submitted 2017-04-17 stat.ML cs.CLcs.LG

classification stat.MLcs.CLcs.LG
keywords translationmachineneurallargermodelspredictionpronountrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a neural machine translation architecture that models the surrounding text in addition to the source sentence. These models lead to better performance, both in terms of general translation quality and pronoun prediction, when trained on small corpora, although this improvement largely disappears when trained with a larger corpus. We also discover that attention-based neural machine translation is well suited for pronoun prediction and compares favorably with other approaches that were specifically designed for this task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    GRAFT reports improved document-level machine translation by segmenting documents into discourse units, modeling dependencies between them as a DAG, and translating each unit with context from its graph predecessors.

  2. Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Decoder self-attention heads attending the target-side antecedent are the most influential for pronoun disambiguation, and fine-tuning them yields up to 5 percentage points improvement on contrastive tests.

  3. Context-Aware Monolingual Repair for Neural Machine Translation

    cs.CL 2019-09 accept novelty 6.0 of 10

    A monolingual sequence-to-sequence model, trained on round-trip translations, repairs cross-sentence inconsistencies in neural machine translation output.

  4. Enhancing Context Modeling with a Query-Guided Capsule Network for Document-level Translation

    cs.CL 2019-09 conditional novelty 6.0 of 10

    A query-guided capsule network that clusters preceding-sentence context into perspectives and adds a source-target regularization loss gives small BLEU and Meteor gains on TED and Europarl En-De translation, but not on News.

  5. One Model to Learn Both: Zero Pronoun Prediction and Translation

    cs.CL 2019-09 conditional novelty 6.0 of 10

    By jointly predicting and translating zero pronouns in one model, the authors improve Chinese-English BLEU by 5.3 points and Japanese-English BLEU by 2.1 points over a strong baseline.

  6. Reference Network for Neural Machine Translation

    cs.CL 2019-08 conditional novelty 6.0 of 10

    A Reference Network built on local coordinate coding gives NMT decoders a compressed global context from the training corpus and improves BLEU on Zh-En and En-De by roughly 1.3 to 2.7 points over strong baselines.

  7. Improving Context-aware Neural Machine Translation with Target-side Context

    cs.CL 2019-09 conditional novelty 5.0 of 10

    Sharing decoder states across sentences yields consistent, modest BLEU gains, suggesting target-side context is best injected into the decoder.

  8. Bidirectional Context-Aware Hierarchical Attention Network for Document Understanding

    cs.CL 2019-08 conditional novelty 5.0 of 10

    Context-aware sentence encoding and bidirectional document encoding improve HAN accuracy by up to 0.46 percentage points on three document classification benchmarks.

Pith tools