REVIEW 4 major objections 5 minor 66 references
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A few attention heads carry pronoun disambiguation in context-aware machine translation, and tuning them to attend the antecedent raises accuracy by up to 5 percentage points without BLEU loss.
desk verdict A useful map of where pronoun disambiguation lives in single-encoder MT, but the causal claim about underutilized heads is undercut by an off-target intervention and same-benchmark head selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the Modifying Heads intervention. For a chosen head and a relation Y to X, the formula replaces the pre-softmax scores of tokens in X so that their post-softmax attention sum equals a chosen value C, while keeping the pre-softmax scores of all other tokens unchanged. This lets the authors test, head by head, whether forcing more or less attention onto a pronoun-antecedent relation changes disambiguation accuracy. The complementary machinery is head tuning: freeze all parameters except the Q and K projections of a selected head and train it with an MSE loss toward the modified pre-softmax scores, thereby solidifying the intervention into the model weights.
What would settle it
Run the same head-tuning pipeline on randomly selected heads or on heads that were classified as non-attending and non-responsive: if those heads produce the same or similar ContraPro gains, the improvement is not specific to the identified disambiguation heads. Alternatively, hold the attention mass on all other token relations fixed while strengthening only the target relation and check whether the accuracy gain survives; if it disappears, the reported effect is an artifact of global attention redistribution.
Extended reading notes
Core claim
The central claim is that in context-aware Transformer translation, a handful of decoder self-attention heads causally support pronoun disambiguation by attending to the target-side antecedent, and some of those heads are underused. Evidence comes from three complementary measurements: average attention scores, point-biserial correlation between a head's attention on a relation and the model being correct, and a controlled modification of pre-softmax attention scores. The strongest result is that fine-tuning selected heads (for example head d-6-4 on the target-pronoun-to-target-antecedent relation) raises ContraPro accuracy from 81.46% to 86.42% in the sentence-level OPUS-MT model and produces similar gains in context-aware models, with no BLEU drop. The paper also concludes that target-side context is more impactful than source-side context, that the most relevant heads sit in higher decoder layers, and that some heads show the same disambiguation behavior across English-to-German and English-to-French in the multilingual model.
Load-bearing premise
The conclusions depend on the intervention being clean: changing the pre-softmax scores for one head-relation pair is assumed to alter only that relation's attention mass while leaving the rest of the model's behavior intact, yet the paper acknowledges off-target effects at C=0.99.
Editorial extensions
If this is right
- If the identified heads are genuinely responsible, context-aware MT improvements can be targeted at a few decoder heads rather than requiring full-model retraining.
- Target-side context, meaning the previously generated target sentence, is the main carrier of disambiguation information; source-side antecedent relations matter less.
- Underutilized heads mean current models have latent disambiguation capacity that parameter-level tuning can unlock without sacrificing translation quality.
- The overlap between improvements from pairs of heads is below 30%, suggesting the heads play complementary roles, so tuning several together could compound the gain.
- In multilingual models, some heads exhibit the same disambiguation behavior across language pairs, hinting at shared cross-lingual mechanisms.
Reading between the lines
- The paper's causal reading assumes the Modifying Heads intervention changes only the target relation's attention mass; because the paper itself discards accuracy losses at C=0.99 as off-target effects, a cleaner test would verify that attention on other token relations is unchanged when one relation is strengthened.
- If the tuning result transfers from contrastive ranking to generative decoding, the same method could be applied to other context-dependent phenomena such as deixis, ellipsis, and lexical cohesion, where contrastive test sets exist.
- The head-tuning recipe uses gold target context; in a real system the context is the model's own output, so error propagation may reduce the measured gains unless the tuned heads are trained on predicted context.
- A direct falsifier would be to fine-tune the same number of randomly chosen heads: if random heads produce similar ContraPro gains, the improvement is not specific to the identified disambiguation heads.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates which attention heads in context-aware machine translation models (OPUS-MT en-de and NLLB-200, fine-tuned on IWSLT 2017) are responsible for pronoun disambiguation in English-to-German and English-to-French. Using the ContraPro and LCPT contrastive datasets, the authors measure per-head attention to five pronoun-antecedent or pronoun-pronoun relations, correlate those attention scores with disambiguation correctness, and then intervene by modifying pre-softmax attention scores (Eq. 2) to force a head to spend a fraction C of its attention on a relation of interest. They categorize heads as responsive or non-responsive, identify several heads that appear to be underutilized, and fine-tune selected heads to reproduce the modified attention behavior. The paper reports that fine-tuning the most promising heads improves ContraPro accuracy by up to about 5 percentage points without reducing BLEU (Table 6).
Significance. If the causal claims were substantiated, this would be a useful contribution to interpretability for machine translation and to targeted, head-level fine-tuning: the paper provides a clear experimental pipeline, public code, and a transparent mathematical derivation in Appendix A. The finding that decoder self-attention heads matter most for target-side context integration is plausible and consistent with prior work. However, the central claim that the intervention and fine-tuning improve disambiguation by strengthening the antecedent relation is not uniquely supported, because the intervention also imposes a hard cap on all non-antecedent attention, and the evaluation loop selects heads on the same contrastive test set used for the final measurement. The paper also explicitly discards accuracy changes that contradict its expected pattern. These issues prevent the current version from validating the underutilized-head mechanism, though they are addressable with control experiments and held-out evaluation.
major comments (4)
- [Section 3.3, Eq. (2)] The Modifying Heads intervention is not relation-specific. As derived in Appendix A, setting the post-softmax mass on the target subset X to C forces the total post-softmax mass on all tokens outside X to be exactly 1−C. At C=0.99, every non-antecedent token combined receives only 1% of the head's attention, regardless of its original score. The paper itself concedes this in Section 5 by discarding accuracy losses at C=0.99 as 'possibly resulting from the decreased attention scores for other token-to-token relations, not investigated in this work.' Consequently, the observed accuracy gains at C=0.99 may stem from off-target suppression rather than from increased attention to the antecedent relation, and the 'underutilized heads' interpretation is not uniquely supported. A control intervention that redirects the same total mass C to a matched set of non-antecedent tokens is needed to separate the two mechanisms.
- [Section 6, Eq. (10) and Table 6] The head-tuning objective reproduces the same off-target effect as the modifying-head intervention. The target pre-softmax scores are taken from Eq. (2), and although the outside pre-softmax scores are frozen from the unmodified model, the softmax denominator still rescales so that the post-softmax attention to all non-antecedent tokens is capped at 1−C. Thus the fine-tuned heads also learn to suppress non-antecedent attention, and the accuracy improvements in Table 6 are equally consistent with that suppression mechanism. The claim that tuning validates the antecedent-attention role requires a control head trained to allocate the same total mass to a non-antecedent subset; without it, the mechanism remains indistinguishable.
- [Sections 5 and 6] The head-selection and evaluation loop is circular. Heads are selected for fine-tuning based on their modification gains on ContraPro (Section 5), and the final accuracy in Table 6 is reported on the same ContraPro dataset. Using CTXPRO to construct a training set does not break the loop because the head choice itself is still based on ContraPro performance. This makes the reported improvements partly self-fulfilling. An evaluation on a held-out contrastive split or an independent pronoun-disambiguation test set is required to support the claim that the improvement is genuine and generalizes.
- [Section 5] The paper's classification into 'attending and positively responsive,' 'attending and negatively responsive,' and similar categories relies on selectively discarding opposing evidence. Section 5 states that losses at C=0.99 and gains at C=0.01 are ignored. This one-sided treatment makes the categories depend on the authors' prior rather than on a consistent behavioral rule. The full modification curves in Appendix G should be used to define categories symmetrically, and the significance of the reported accuracy differences (e.g., the 1.6 percentage point gain for head e-6-1) should be assessed with appropriate statistical tests.
minor comments (5)
- [Section 5.1] Head e-6-1 is described as 'negatively responsive for all models and positively responsive for the context-aware-3 model,' but the category definitions in Section 5 are mutually exclusive; a head that increases accuracy when modified to 0.99 cannot be negatively responsive by the given definition. Please clarify whether the head should instead be labeled fully responsive for that model.
- [Appendix D] There is a typo: 'OpisMT' should be 'OpusMT'.
- [Section 3.3 and Appendix A] The notation X d and Xd (or X^d and X_d) is visually indistinguishable in the printed equations, making the derivation of Eq. (2) difficult to follow. Please use distinct typography (e.g., calligraphic X for the full key set and italic X for the subset of interest).
- [Table 6] No measure of variance or significance is reported for the tuned-head accuracy improvements. Given that some differences are small (e.g., 80.68 vs 80.78 for the context-aware-1 model with c-5-1), reporting multiple seeds or confidence intervals would strengthen the claims.
- [Section 3.3] The description says the method 'preserves the pre-softmax attention scores H for all other target tokens,' but post-softmax scores for those tokens are not preserved; they are rescaled by the denominator. Please rephrase to avoid misleading the reader into thinking the intervention has no off-target effect.
Circularity Check
No significant circularity: the head-modification and fine-tuning interventions are external manipulations, and the paper's own limitation statements identify confounds rather than definitional equivalences.
full rationale
The paper's derivation chain is self-contained and does not reduce to its inputs by construction. Section 3.3 (Eq. 2) defines an external intervention that rewrites pre-softmax scores for the relation subset X so that post-softmax attention to X equals C; this is a mathematical construction, but the resulting accuracy measurements are empirical outcomes of the modified model, not quantities defined by Eq. 2. The 'attending and positively responsive' categories in Section 5 are post-hoc interpretations of observed accuracy changes, not definitions that make the underutilized-heads conclusion true by construction. The fine-tuning step (Appendix F, Eq. 10) sets target scores equal to the modified scores for j in X and the unmodified model's scores otherwise, and trains on CTXPRO/IWSLT data unrelated to ContraPro; although the same ContraPro set is used both to select heads/relations and to report the +5 pp gain in Table 6, which is a selection-bias/leakage concern, the tuned accuracy is not algebraically forced by the training target and can be lower than the modified-model accuracy (Table 6). The paper itself flags the main confound in Section 5: losses at C=0.99 are 'possibly resulting from the decreased attention scores for other token-to-token relations, not investigated in this work'; this is an off-target-effect validity threat to the causal interpretation, but it is not a case of a prediction being equivalent to its input by construction. The only self-citation (Maka et al., 2024) appears in Related Work as an example of a multi-encoder architecture and is not load-bearing. Therefore no significant circularity is found.
Assumptions & free parameters
free parameters (2)
- Head categorization thresholds
- Selected heads for fine-tuning =
4 per OpusMT model (d-6-4, d-6-6, d-6-7, c-5-1)
assumptions (3)
- ad hoc to paper Attention scores are computed as softmax over keys, and modifying pre-softmax scores as in Eq. 2 only changes the attention mass on the target relation while leaving all other values fixed.
- domain assumption Token identity is preserved across Transformer layers, so input-side context cue tokens TC+1 correspond to predicted tokens TC.
- domain assumption The contrastive datasets ContraPro and LCPT, with provided gold target context, are valid measures of pronoun disambiguation ability.
Cite this review
Pith. "Pith review of Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models." pith.science (2026). https://pith.science/paper/4PUF47Z6
@misc{pith2026241211187,
author = {Pith},
title = {Pith review of: Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4PUF47Z6}},
note = {Machine review of arXiv:2412.11187}
}
read the original abstract
In this paper, we investigate the role of attention heads in Context-aware Machine Translation models for pronoun disambiguation in the English-to-German and English-to-French language directions. We analyze their influence by both observing and modifying the attention scores corresponding to the plausible relations that could impact a pronoun prediction. Our findings reveal that while some heads do attend the relations of interest, not all of them influence the models' ability to disambiguate pronouns. We show that certain heads are underutilized by the models, suggesting that model performance could be improved if only the heads would attend one of the relations more strongly. Furthermore, we fine-tune the most promising heads and observe the increase in pronoun disambiguation accuracy of up to 5 percentage points which demonstrates that the improvements in performance can be solidified into the models' parameters.
Figures
Figures from the paper (20 more)
Reference graph
Works this paper leans on
-
[1]
Samira Abnar and Willem Zuidema. 2020. https://doi.org/10.18653/v1/2020.acl-main.385 Quantifying attention flow in transformers . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190--4197, Online. Association for Computational Linguistics
-
[2]
Ruchit Agrawal, Marco Turchi, and Matteo Negri. 2018. Contextual handling in neural machine translation: Look behind, ahead and on both sides. In Proceedings of the 21st Annual Conference of the European Association for Machine Translation, pages 31--40
work page 2018
-
[3]
Guangsheng Bao, Yue Zhang, Zhiyang Teng, Boxing Chen, and Weihua Luo. 2021. https://doi.org/10.18653/v1/2021.acl-long.267 G -transformer for document-level machine translation . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Lo...
-
[4]
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2018. https://api.semanticscholar.org/CorpusID:53215110 Identifying and controlling important neurons in neural machine translation . ArXiv, abs/1811.01157
arXiv 2018
-
[5]
Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. https://doi.org/10.18653/v1/N18-1118 Evaluating discourse phenomena in neural machine translation . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , pages 1304--1...
-
[6]
Nikolay Bogoychev. 2021. https://doi.org/10.18653/v1/2021.blackboxnlp-1.28 Not all parameters are born equal: Attention is mostly what you need . In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 363--374, Punta Cana, Dominican Republic. Association for Computational Linguistics
-
[7]
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020. On identifiability in transformers. In 8th International Conference on Learning Representations (ICLR 2020)(virtual). International Conference on Learning Representations
work page 2020
-
[8]
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022. Recurrent memory transformer. Advances in Neural Information Processing Systems, 35:11079--11091
2022
Show all 66 references
-
[9]
Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian St \"u ker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. 2017. https://aclanthology.org/2017.iwslt-1.1 Overview of the IWSLT 2017 evaluation campaign . In Proceedings of the 14th Internat...
2017
-
[10]
Linqing Chen, Junhui Li, Zhengxian Gong, Min Zhang, and Guodong Zhou. 2022. https://doi.org/10.1145/3526215 One type context is not enough: Global context-aware neural machine translation . ACM Trans. Asian Low-Resour. Lang. Inf. Process., 21(6)
2022 doi
-
[11]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. https://doi.org/10.18653/v1/W19-4828 What does BERT look at? an analysis of BERT ' s attention . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NL...
2019 doi
-
[12]
Yukun Feng, Feng Li, Ziang Song, Boyuan Zheng, and Philipp Koehn. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.105 Learn to remember: Transformer with recurrent memory for document-level machine translation . In Findings of the Association for Computational Linguistic...
2022 doi
-
[13]
Patrick Fernandes, Kayo Yin, Emmy Liu, Andr \'e Martins, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.acl-long.36 When does translation require context? a data-driven, multilingual exploration . In Proceedings of the 61st Annual Meeting of the Association for Comp...
2023 doi
-
[14]
Patrick Fernandes, Kayo Yin, Graham Neubig, and Andr \'e F. T. Martins. 2021. https://doi.org/10.18653/v1/2021.acl-long.505 Measuring and increasing context usage in context-aware machine translation . In Proceedings of the 59th Annual Meeting of the Association for Computatio...
2021 doi
-
[15]
G \'a llego, Belen Alastruey, Carlos Escolano, and Marta R
Javier Ferrando, Gerard I. G \'a llego, Belen Alastruey, Carlos Escolano, and Marta R. Costa-juss \`a . 2022. https://doi.org/10.18653/v1/2022.emnlp-main.599 Towards opening the black box of neural machine translation: Source and target interpretations of the transformer . In ...
2022 doi
-
[16]
Harritxu Gete, Thierry Etchegoyhen, and Gorka Labaka. 2023. https://aclanthology.org/2023.eamt-1.15 What works when in context-aware neural machine translation? In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 147--156, Ta...
2023
-
[17]
Mozhdeh Gheini, Xiang Ren, and Jonathan May. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.132 Cross-attention is all you need: A dapting pretrained T ransformers for machine translation . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Proce...
2021 doi
-
[18]
Akshay Goindani and Manish Shrivastava. 2021. https://aclanthology.org/2021.ranlp-1.52 A dynamic head importance computation mechanism for neural machine translation . In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021)...
2021
-
[19]
Christian Hardmeier. 2012. https://doi.org/10.4000/discours.8726 Discourse in statistical machine translation: A survey and a case study . Discours-Revue de linguistique, psycholinguistique et informatique, 11
2012 doi
-
[20]
Jingjing Huo, Christian Herold, Yingbo Gao, Leonard Dahlmann, Shahram Khadivi, and Hermann Ney. 2020. https://aclanthology.org/2020.wmt-1.71 Diving deep into context-aware neural machine translation . In Proceedings of the Fifth Conference on Machine Translation, pages 604--61...
2020
-
[21]
Sarthak Jain and Byron C. Wallace. 2019. https://doi.org/10.18653/v1/N19-1357 A ttention is not E xplanation . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and...
2019 doi
-
[22]
Sebastien Jean, Stanislas Lauly, Orhan Firat, and Kyunghyun Cho. 2017. Does neural machine translation benefit from larger context? arXiv preprint arXiv:1704.05135
2017 arXiv
-
[23]
Jae-young Jo and Sung-Hyon Myaeng. 2020. https://doi.org/10.18653/v1/2020.acl-main.311 Roles and utilization of attention heads in transformer-based neural language models . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3404-...
2020 doi
-
[24]
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.574 Attention is not only a weight: Analyzing transformers with vector norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Pro...
2020 doi
-
[25]
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.373 I ncorporating R esidual and N ormalization L ayers into A nalysis of M asked L anguage M odels . In Proceedings of the 2021 Conference on Empirical Methods ...
2021 doi
-
[26]
Anna Langedijk, Hosein Mohebbi, Gabriele Sarti, Willem Zuidema, and Jaap Jumelet. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.296 D ecoder L ens: Layerwise interpretation of encoder-decoder transformers . In Findings of the Association for Computational Linguistics: ...
2024 doi
-
[27]
Pierre Lison, J \"o rg Tiedemann, and Milen Kouylekov. 2018. https://aclanthology.org/L18-1275 O pen S ubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora . In Proceedings of the Eleventh International Conference on Language Resources an...
2018
-
[28]
Amin Farajian, Rachel Bawden, Michael Zhang, and Andr \'e F
Ant \'o nio Lopes, M. Amin Farajian, Rachel Bawden, Michael Zhang, and Andr \'e F. T. Martins. 2020. https://aclanthology.org/2020.eamt-1.24 Document-level neural MT : A systematic comparison . In Proceedings of the 22nd Annual Conference of the European Association for Machin...
2020
-
[29]
Shuming Ma, Dongdong Zhang, and Ming Zhou. 2020. https://doi.org/10.18653/v1/2020.acl-main.321 A simple and effective unified encoder for document-level machine translation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3505...
2020 doi
-
[30]
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022. https://doi.org/10.1145/3546577 Post-hoc interpretability for neural nlp: A survey . ACM Comput. Surv., 55(8)
2022 doi
-
[31]
Suvodeep Majumde, Stanislas Lauly, Maria Nadejde, Marcello Federico, and Georgiana Dinu. 2022. A baseline revisited: Pushing the limits of multi-segment models for context-aware translation. arXiv preprint arXiv:2210.10906
2022 arXiv
-
[32]
Pawe Maka, Yusuf Semerci, Jan Scholtes, and Gerasimos Spanakis. 2024. https://aclanthology.org/2024.findings-eacl.127 Sequence shortening for context-aware machine translation . In Findings of the Association for Computational Linguistics: EACL 2024, pages 1874--1894, St. Juli...
2024
-
[33]
Sameen Maruf, Andr \'e F. T. Martins, and Gholamreza Haffari. 2019. https://doi.org/10.18653/v1/N19-1313 Selective attention for context-aware neural machine translation . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational...
2019 doi
-
[34]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2024. Locating and editing factual associations in gpt. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA. Curran Associates Inc
2024
-
[35]
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018. https://doi.org/10.18653/v1/D18-1325 Document-level neural machine translation with hierarchical attention networks . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Pro...
2018 doi
-
[36]
Wafaa Mohammed and Vlad Niculae. 2024. https://aclanthology.org/2024.findings-eacl.113 On measuring context utilization in document-level MT systems . In Findings of the Association for Computational Linguistics: EACL 2024, pages 1633--1643, St. Julian ' s, Malta. Association ...
2024
-
[37]
Hosein Mohebbi, Willem Zuidema, Grzegorz Chrupa a, and Afra Alishahi. 2023. https://doi.org/10.18653/v1/2023.eacl-main.245 Quantifying context mixing in transformers . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistic...
2023 doi
-
[38]
Mathias M \"u ller, Annette Rios, Elena Voita, and Rico Sennrich. 2018. https://doi.org/10.18653/v1/W18-6307 A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation . In Proceedings of the Third Conference on Machine Transla...
2018 doi
-
[39]
NLLB Team , Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia-Gonzalez, Pran...
2022
-
[40]
Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova Dassarma, Tom Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, K...
2022 arXiv
-
[41]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...
2002
-
[42]
Matt Post and Marcin Junczys-Dowmunt. 2023. Escaping the sentence-level paradigm in machine translation. arXiv preprint arXiv:2304.12959
2023 arXiv
-
[43]
Gabriele Sarti, Grzegorz Chrupa a, Malvina Nissim, and Arianna Bisazza. 2023. Quantifying the plausibility of context reliance in neural machine translation. arXiv preprint arXiv:2310.01188
2023 arXiv
-
[44]
Shazeer and Mitchell Stern
Noam M. Shazeer and Mitchell Stern. 2018. https://api.semanticscholar.org/CorpusID:4786918 Adafactor: Adaptive learning rates with sublinear memory cost . ArXiv, abs/1804.04235
2018 arXiv
-
[45]
Zewei Sun, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Shujian Huang, Jiajun Chen, and Lei Li. 2022. https://doi.org/10.18653/v1/2022.findings-acl.279 Rethinking document-level neural machine translation . In Findings of the Association for Computational Linguistics: ACL 2022, page...
2022 doi
-
[46]
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. https://doi.org/10.18653/v1/P19-1452 BERT rediscovers the classical NLP pipeline . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593--4601, Florence, Italy. Association for ...
2019 doi
-
[47]
o rg Tiedemann, Mikko Aulamo, Daria Bakshandaeva, Michele Boggia, Stig-Arne Gr \
J \"o rg Tiedemann, Mikko Aulamo, Daria Bakshandaeva, Michele Boggia, Stig-Arne Gr \"o nroos, Tommi Nieminen, Alessandro Raganato\, Yves Scherrer, Raul Vazquez, and Sami Virpioja. 2023. https://doi.org/10.1007/s10579-023-09704-w Democratizing neural machine translation with OP...
2023 doi
-
[48]
J \"o rg Tiedemann and Yves Scherrer. 2017. https://doi.org/10.18653/v1/W17-4811 Neural machine translation with extended context . In Proceedings of the Third Workshop on Discourse in Machine Translation, pages 82--92, Copenhagen, Denmark. Association for Computational Linguistics
2017 doi
-
[49]
J \"o rg Tiedemann and Santhosh Thottingal. 2020. OPUS-MT — B uilding open translation services for the W orld. In Proceedings of the 22nd Annual Conferenec of the European Association for Machine Translation (EAMT), Lisbon, Portugal
2020
-
[50]
Mariya Toneva and Leila Wehbe. 2019. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). Curran Associates Inc., Red Hook, NY, USA
2019
-
[51]
Zhaopeng Tu, Yang Liu, Zhengdong Lu, Xiaohua Liu, and Hang Li. 2017. https://doi.org/10.1162/tacl_a_00048 Context gates for neural machine translation . Transactions of the Association for Computational Linguistics, 5:87--99
2017 doi
-
[52]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30
2017
-
[53]
Jesse Vig and Yonatan Belinkov. 2019. https://doi.org/10.18653/v1/W19-4808 Analyzing the structure of attention in a transformer language model . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63--76, Florence, It...
2019 doi
-
[54]
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 a . https://doi.org/10.18653/v1/D19-1081 Context-aware monolingual repair for neural machine translation . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint...
2019 doi
-
[55]
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 b . https://doi.org/10.18653/v1/P19-1116 When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion . In Proceedings of the 57th Annual Meeting of the Associ...
2019 doi
-
[56]
Elena Voita, Rico Sennrich, and Ivan Titov. 2021. https://doi.org/10.18653/v1/2021.acl-long.91 Analyzing the source and target contributions to predictions in neural machine translation . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistic...
2021 doi
-
[57]
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 c . https://doi.org/10.18653/v1/P19-1580 Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned . In Proceedings of the 57th Annual Meeting of the Associa...
2019 doi
-
[58]
Rachel Wicks and Matt Post. 2023. https://doi.org/10.18653/v1/2023.wmt-1.42 Identifying context-dependent translations for evaluation set production . In Proceedings of the Eighth Conference on Machine Translation, pages 452--467, Singapore. Association for Computational Linguistics
2023 doi
-
[59]
Sarah Wiegreffe and Yuval Pinter. 2019. https://doi.org/10.18653/v1/D19-1002 Attention is not not explanation . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (...
2019 doi
-
[60]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020 doi
-
[61]
Billy T. M. Wong and Chunyu Kit. 2012. https://aclanthology.org/D12-1097 Extending machine translation evaluation metrics with lexical cohesion to document level . In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational...
2012
-
[62]
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, Andr \'e F. T. Martins, and Graham Neubig. 2021. https://doi.org/10.18653/v1/2021.acl-long.65 Do context-aware translation models pay the right attention? In Proceedings of the 59th Annual Meeting of the Association ...
2021 doi
-
[63]
Pei Zhang, Boxing Chen, Niyu Ge, and Kai Fan. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.81 Long-short term masking transformer: A simple but effective baseline for document-level neural machine translation . In Proceedings of the 2020 Conference on Empirical Methods in...
2020 doi
-
[64]
Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen, and Alexandra Birch. 2021. Towards making the most of context in neural machine translation. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages ...
2021
-
[65]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[66]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.