REVIEW 2 major objections 2 minor 2 cited by
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
T0 review · 2 major / 2 minor · reviewed 2026-05-14 · grok-4.3
Pith's one-line read Automatic evaluation metrics and LLM judges correlate poorly with professional translators on creativity in literary texts and bias toward machine outputs.
desk verdict The paper documents LLM judges favoring machine translations on creativity scores in literary MT, backed by a new multi-genre dataset, but the annotations lack reported agreement stats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A dataset of literary translations across human, machine, and post-edited modalities, annotated by professional translators for creative shifts and errors, used to measure alignment with automatic metrics and LLM judges.
What would settle it
A new set of annotations on the same dataset by a different group of professional literary translators that produces substantially different creativity scores from the original annotations.
Extended reading notes
Core claim
Automatic evaluation metrics and LLM-as-a-judge evaluations correlate poorly with professional literary translators' assessments of creativity, and LLM judges display a systematic bias that favors machine-translated texts while penalizing creative and culturally appropriate solutions, with performance dropping further on poetry and similar literary genres.
Load-bearing premise
Detailed annotations by experienced professional literary translators constitute an objective and reliable ground truth for measuring creativity and translation quality across genres and modalities.
Editorial extensions
If this is right
- Automatic metrics cannot serve as reliable substitutes for professional judgment when creativity is the focus of evaluation.
- LLM-as-a-judge methods introduce a consistent preference for literal machine outputs over creative human solutions.
- Evaluation accuracy declines markedly for poetry and other highly literary genres.
- New automatic tools are required that treat creative out-of-routine solutions as valid rather than errors.
Reading between the lines
- If automatic scores are used for quality control, translation workflows may systematically undervalue creative post-editing by humans.
- The same bias pattern could appear in AI evaluation of other creative writing tasks such as story generation or script adaptation.
- Explicit cultural and creative criteria would need to be built into future evaluation frameworks to reduce the observed mismatch.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that automatic evaluation metrics (AEMs) and LLM-as-a-judge methods correlate poorly with professional literary translators' assessments of translation quality and creativity (creative shifts and errors) across three language pairs, three genres, and three modalities (human translation, machine translation, post-editing). It further reports that LLM judges exhibit systematic bias favoring machine-translated outputs while penalizing creative and culturally appropriate solutions, with performance degrading for more literary genres such as poetry.
Significance. If the empirical findings hold after addressing reporting gaps, the work provides useful evidence of limitations in current automatic tools for evaluating creativity in literary translation and motivates development of new metrics. The construction of a multi-genre, multi-modality dataset annotated by experienced professionals is a concrete contribution that can support future research.
major comments (2)
- [Annotation methodology] Annotation methodology section: no inter-annotator agreement statistics (Fleiss' kappa, Cohen's kappa, or equivalent) are reported for the creativity labels (creative shifts & errors) assigned by the professional translators. Without these figures, the low correlations and reported LLM bias could reflect label noise rather than genuine misalignment with the ground truth.
- [Results] Results section: the abstract and main results provide no sample sizes (number of translations or segments per genre/modality), exact implementations of the AEMs and LLM prompts, correlation coefficients with confidence intervals, or statistical tests (e.g., p-values for differences between modalities). These omissions prevent assessment of whether the claimed poor correlations and systematic bias are statistically robust.
minor comments (2)
- [Abstract] Abstract: add concrete numbers (e.g., total segments annotated, number of annotators, range of correlation values) to make the headline claims more informative.
- [Dataset construction] Dataset description: clarify selection criteria for the literary texts and any filtering applied after annotation.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback, which has helped us strengthen the reporting in our manuscript. We address each major comment below and have revised the paper to incorporate the requested details on annotation reliability and statistical reporting.
read point-by-point responses
-
Referee: [Annotation methodology] Annotation methodology section: no inter-annotator agreement statistics (Fleiss' kappa, Cohen's kappa, or equivalent) are reported for the creativity labels (creative shifts & errors) assigned by the professional translators. Without these figures, the low correlations and reported LLM bias could reflect label noise rather than genuine misalignment with the ground truth.
Authors: We agree that inter-annotator agreement metrics are essential to demonstrate label reliability. Although the original submission omitted these statistics, the annotations were performed by multiple professional translators with overlapping segments. In the revised manuscript we now report Cohen's kappa values for the creativity shift and error labels (computed on the double-annotated subset), which show substantial agreement. This addition confirms that the observed low correlations with automatic metrics are unlikely to stem from label noise. revision: yes
-
Referee: [Results] Results section: the abstract and main results provide no sample sizes (number of translations or segments per genre/modality), exact implementations of the AEMs and LLM prompts, correlation coefficients with confidence intervals, or statistical tests (e.g., p-values for differences between modalities). These omissions prevent assessment of whether the claimed poor correlations and systematic bias are statistically robust.
Authors: We acknowledge the reporting gaps. The revised version now includes explicit sample sizes (number of segments per genre, language pair, and modality) in both the abstract and Results section. We have added the precise AEM implementations, full LLM prompts, Pearson/Spearman correlations with 95% confidence intervals, and p-values from appropriate statistical tests comparing modalities. These details are also summarized in a new table for clarity and confirm the robustness of the reported poor correlations and LLM bias. revision: yes
Circularity Check
No circularity: purely empirical comparison to external annotations
full rationale
The paper constructs a dataset of literary translations across modalities and genres, obtains detailed annotations for creativity from experienced professional translators, and reports correlations between these annotations and both automatic evaluation metrics and LLM-as-a-judge outputs. No equations, derivations, fitted parameters, or self-referential definitions appear in the provided text. The central claims rest on observed empirical mismatches rather than any reduction of predictions to inputs by construction. Self-citations, if present, are not load-bearing for the core result. This is a standard empirical evaluation study whose validity depends on annotation quality and inter-annotator agreement (a separate reliability concern), not on circular logic.
Assumptions & free parameters
assumptions (1)
- domain assumption Professional literary translators' judgments provide a reliable and objective measure of creativity and translation quality
Cite this review
Pith. "Pith review of Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations." pith.science (2026). https://pith.science/paper/IPZIX56F
@misc{pith2026260513596,
author = {Pith},
title = {Pith review of: Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations},
year = {2026},
howpublished = {\url{https://pith.science/paper/IPZIX56F}},
note = {Machine review of arXiv:2605.13596}
}
read the original abstract
This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess how well these tools align with professionals when evaluating translation, creativity (creative shifts & errors), and see if they can substitute laborious manual annotations. A dataset of literary translations across three modalities (human translation, machine translation, and post-editing), three genres and three language pairs was created and annotated in detail for creativity by experienced professional literary translators. The results show that both AEMs and LLM-as-a-judge evaluations correlate poorly with professional evaluations on creativity, with LLM-as-a-judge showing a systematic bias in favour of machine-translated texts and penalising creative and culturally appropriate solutions. Moreover, performance is consistently worse for more literary genres such as poetry. This highlights fundamental limitations of current automatic evaluation tools for literary translation and the need to create new tools that do not frequently consider out of routine translations as errors.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
AI translation of literary texts is "fine", but readers still prefer human translations
Human readers prefer human literary translations over AI-generated ones for immersion and clarity despite finding MT adequate and struggling to identify the source.
-
LitSeg: Narrative-Aware Document Segmentation for Literary RAG
LitSeg segments literary texts using narrative analysis via multi-stage prompting and offers a distilled lightweight version for efficient use in RAG systems.
Reference graph
Works this paper leans on
-
[1]
Aho, Alfred V. and Ullman, Jeffrey D. , title =. 1972 , volume=1, publisher =
work page 1972
-
[2]
Interspeech 2006 --- Ninth International Conference on Spoken Language Processing , address=
Unsupervised language model adaptation using latent semantic marginals , author=. Interspeech 2006 --- Ninth International Conference on Spoken Language Processing , address=. 2006 , pages=
work page 2006
- [3]
-
[4]
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year=1981, title=. Journal of the Association for Computing Machinery , volume=28, issue=1, pages=
work page 1981
-
[5]
Coling 2008, 22nd International Conference on Computational Linguistics , address=
Anne Gledson and John Keane , year=2008, title=. Coling 2008, 22nd International Conference on Computational Linguistics , address=
work page 2008
-
[6]
Dan Gusfield , title=
-
[7]
Yik-Cheung Tam and Tanja Schultz , year=2007, title=. Proceedings of ICASSP 2007, International Conference on Acoustics, Speech, and Signal Processing , address=
work page 2007
-
[8]
MATEO : MA chine T ranslation E valuation O nline
Vanroy, Bram and Tezcan, Arda and Macken, Lieve. MATEO : MA chine T ranslation E valuation O nline. Proceedings of the 24th Annual Conference of the European Association for Machine Translation. 2023
work page 2023
Show all 130 references
-
[9]
2020 , eprint=
BERTScore: Evaluating Text Generation with BERT , author=. 2020 , eprint=
2020
-
[10]
Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle=
Kishore Papineni and Salim Roukos and Todd Ward and Wei-Jing Zhu , year=. Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle=. doi:10.3115/1073083.1073135 , pages=
-
[11]
A Study of Translation Edit Rate with Targeted Human Annotation
Matthew Snover and Dorr, Bonnie and Schwartz, Rich and Micciulla, Linnea and Makhoul, John. A Study of Translation Edit Rate with Targeted Human Annotation. Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers. 2006
2006
-
[12]
chr F deconstructed: beta parameters and n-gram weights
Popovi \'c , Maja. chr F deconstructed: beta parameters and n-gram weights. Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers. 2016. doi:10.18653/v1/W16-2341
2016 doi
-
[13]
COMET : A Neural Framework for MT Evaluation
Rei, Ricardo and Stewart, Craig and Farinha, Ana C and Lavie, Alon. COMET : A Neural Framework for MT Evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18653/v1/2020.emnlp-main.213
2020 doi
-
[14]
BLEURT : Learning Robust Metrics for Text Generation
Sellam, Thibault and Das, Dipanjan and Parikh, Ankur. BLEURT : Learning Robust Metrics for Text Generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.704
2020 doi
-
[15]
L i T rans P ro QA : An LLM -based Literary Translation Evaluation Metric with Professional Question Answering
Zhang, Ran and Zhao, Wei and Macken, Lieve and Eger, Steffen. L i T rans P ro QA : An LLM -based Literary Translation Evaluation Metric with Professional Question Answering. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18...
2025 doi
-
[16]
Prompting C hat GPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
He, Sui. Prompting C hat GPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts. Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1). 2024
2024
-
[17]
Poetry and Emotions: Investigating the Limitations of AI Translation
Priya Sharma and Tanuja Yadav. Poetry and Emotions: Investigating the Limitations of AI Translation. Proceedings of International Conference on Innovations in Data Science. 2026
2026
-
[18]
AL-ĪMĀN Research Journal , author=
A Comparative Analysis of Machine Translation and Human Translation: Efficacy of Poetry Translation from Urdu to English , volume=. AL-ĪMĀN Research Journal , author=. 2026 , month=. doi:10.63283/IRJ.04.01/02 , abstractNote=
2026 doi
-
[19]
Don ' t Go Far Off: An Empirical Study on Neural Poetry Translation
Chakrabarty, Tuhin and Saakyan, Arkadiy and Muresan, Smaranda. Don ' t Go Far Off: An Empirical Study on Neural Poetry Translation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.577
2021 doi
-
[20]
The Translator ' s Canvas: Using LLM s to Enhance Poetry Translation
Resende, Nat \'a lia and Hadley, James. The Translator ' s Canvas: Using LLM s to Enhance Poetry Translation. Proceedings of the 16th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track). 2024
2024
-
[21]
A Detailed Comparative Analysis of Automatic Neural Metrics for Machine Translation:
Mukherjee, Aniruddha and Hassija, Vikas and Chamola, Vinay and Gupta, Karunesh Kumar , journal=. A Detailed Comparative Analysis of Automatic Neural Metrics for Machine Translation:. 2025 , volume=
2025
-
[22]
A Fine-Grained Analysis of BERTS core
Hanna, Michael and Bojar, Ond r ej. A Fine-Grained Analysis of BERTS core. Proceedings of the Sixth Conference on Machine Translation. 2021
2021
-
[23]
Extrinsic Evaluation of Machine Translation Metrics
Nikita Moghe and Tom Sherborne and Mark Steedman and Alexandra Birch. Extrinsic Evaluation of Machine Translation Metrics. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.730
2023 doi
-
[24]
Results of WMT 23 Metrics Shared Task: Metrics Might Be Guilty but References Are Not Innocent
Freitag, Markus and Mathur, Nitika and Lo, Chi-kiu and Avramidis, Eleftherios and Rei, Ricardo and Thompson, Brian and Kocmi, Tom and Blain, Frederic and Deutsch, Daniel and Stewart, Craig and Zerva, Chrysoula and Castilho, Sheila and Lavie, Alon and Foster, George. Results of...
2023 doi
-
[25]
BLEU , METEOR , BERTS core: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-Oriented Text
Saadany, Hadeel and Orasan, Constantin. BLEU , METEOR , BERTS core: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-Oriented Text. Proceedings of the Translation and Interpreting Technology Online Conference. 2021
2021
-
[26]
Findings of the WMT 25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets
Kocmi, Tom and Artemova, Ekaterina and Avramidis, Eleftherios and Bawden, Rachel and Bojar, Ond r ej and Dranch, Konstantin and Dvorkovich, Anton and Dukanov, Sergey and Fishel, Mark and Freitag, Markus and Gowda, Thamme and Grundkiewicz, Roman and Haddow, Barry and Karpinska,...
2025 doi
-
[27]
Freitag, Markus and Rei, Ricardo and Mathur, Nitika and Lo, Chi-kiu and Stewart, Craig and Avramidis, Eleftherios and Kocmi, Tom and Foster, George and Lavie, Alon and Martins, Andr \'e F. T. Results of WMT 22 Metrics Shared Task: Stop Using BLEU -- Neural Metrics Are Better a...
2022 doi
-
[28]
and Rei, Ricardo and Stigt, Daan van and Coheur, Luisa and Colombo, Pierre and Martins, Andr \'e F
Guerreiro, Nuno M. and Rei, Ricardo and Stigt, Daan van and Coheur, Luisa and Colombo, Pierre and Martins, Andr \'e F. T. x COMET : Transparent Machine Translation Evaluation through Fine-grained Error Detection. Transactions of the Association for Computational Linguistics. 2...
2024 doi
-
[29]
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
Zouhar, Vil \'e m and Ding, Shuoyang and Currey, Anna and Badeka, Tatyana and Wang, Jenyuan and Thompson, Brian. Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2...
2024 doi
-
[30]
ACES : Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics
Amrhein, Chantal and Moghe, Nikita and Guillou, Liane. ACES : Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics. Proceedings of the Seventh Conference on Machine Translation (WMT). 2022. doi:10.18653/v1/2022.wmt-1.44
2022 doi
-
[31]
Evaluating Automatic Metrics with Incremental Machine Translation Systems
Wu, Guojun and Cohen, Shay B and Sennrich, Rico. Evaluating Automatic Metrics with Incremental Machine Translation Systems. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.169
2024 doi
-
[32]
To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation
Kocmi, Tom and Federmann, Christian and Grundkiewicz, Roman and Junczys-Dowmunt, Marcin and Matsushita, Hitokazu and Menezes, Arul. To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation. Proceedings of the Sixth Conference on Machine Tran...
2021
-
[33]
DEMETR : Diagnosing Evaluation Metrics for Translation
Karpinska, Marzena and Raj, Nishant and Thai, Katherine and Song, Yixiao and Gupta, Ankita and Iyyer, Mohit. DEMETR : Diagnosing Evaluation Metrics for Translation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/20...
2022 doi
-
[34]
The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation
Fernandes, Patrick and Deutsch, Daniel and Finkelstein, Mara and Riley, Parker and Martins, Andr \'e and Neubig, Graham and Garg, Ankush and Clark, Jonathan and Freitag, Markus and Firat, Orhan. The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Mach...
2023 doi
-
[35]
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
Lu, Qingyu and Qiu, Baopu and Ding, Liang and Zhang, Kanjian and Kocmi, Tom and Tao, Dacheng. Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1...
2024 doi
-
[36]
Evolutionary Studies in Imaginative Culture , year =
Nehal Ali AbdulGhaffar , title =. Evolutionary Studies in Imaginative Culture , year =
-
[37]
and Zerva, Chrysoula and Farinha, Ana C and Maroti, Christine and C
Rei, Ricardo and Treviso, Marcos and Guerreiro, Nuno M. and Zerva, Chrysoula and Farinha, Ana C and Maroti, Christine and C. de Souza, Jos \'e G. and Glushkova, Taisiya and Alves, Duarte and Coheur, Luisa and Lavie, Alon and Martins, Andr \'e F. T. C omet K iwi: IST -Unbabel 2...
2022 doi
-
[38]
Findings of the WMT 25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Kocmi, Tom and Agrawal, Sweta and Artemova, Ekaterina and Avramidis, Eleftherios and Briakou, Eleftheria and Chen, Pinzhen and Fadaee, Marzieh and Freitag, Markus and Grundkiewicz, Roman and Hou, Yupeng and Koehn, Philipp and Kreutzer, Julia and Mansour, Saab and Perrella, Ste...
2025 doi
-
[39]
M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task
Juraska, Juraj and Deutsch, Daniel and Finkelstein, Mara and Freitag, Markus. M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task. Proceedings of the Ninth Conference on Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.35
2024 doi
-
[40]
Findings of the WMT 24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet
Kocmi, Tom and Avramidis, Eleftherios and Bawden, Rachel and Bojar, Ond r ej and Dvorkovich, Anton and Federmann, Christian and Fishel, Mark and Freitag, Markus and Gowda, Thamme and Grundkiewicz, Roman and Haddow, Barry and Karpinska, Marzena and Koehn, Philipp and Marie, Ben...
2024 doi
-
[41]
Can Automatic Metrics Assess High-Quality Translations?
Agrawal, Sweta and Farinhas, Ant \'o nio and Rei, Ricardo and Martins, Andre. Can Automatic Metrics Assess High-Quality Translations?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.802
2024 doi
-
[42]
and Platek, Ondrej and Sivaprasad, Adarsa
Schmidtova, Patricia and Mahamood, Saad and Balloccu, Simone and Dusek, Ondrej and Gatt, Albert and Gkatzia, Dimitra and Howcroft, David M. and Platek, Ondrej and Sivaprasad, Adarsa. Automatic Metrics in Natural Language Generation: A survey of Current Evaluation Practices. Pr...
2024 doi
-
[43]
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature
Thai, Katherine and Karpinska, Marzena and Krishna, Kalpesh and Ray, Bill and Inghilleri, Moira and Wieting, John and Iyyer, Mohit. Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature. Proceedings of the 2022 Conference on Empir...
2022 doi
-
[44]
2025 , eprint=
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation , author=. 2025 , eprint=
2025
-
[45]
H i MATE : A Hierarchical Multi-Agent Framework for Machine Translation Evaluation
Zhang, Shijie and Li, Renhao and Wang, Songsheng and Koehn, Philipp and Yang, Min and Wong, Derek F. H i MATE : A Hierarchical Multi-Agent Framework for Machine Translation Evaluation. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. doi:10.18653/v1...
2025 doi
-
[46]
M - MAD : Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation
Feng, Zhaopeng and Su, Jiayuan and Zheng, Jiamei and Ren, Jiahan and Zhang, Yan and Wu, Jian and Wang, Hongwei and Liu, Zuozhu. M - MAD : Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation. Proceedings of the 63rd Annual Meeting of the Association ...
2025 doi
-
[47]
The Atlantic , year =
Jeremy Klemin , title =. The Atlantic , year =
-
[48]
The Bookseller , year =
Emily Warner , title =. The Bookseller , year =
-
[49]
The price of debiasing automatic metrics in natural language evaluation
Chaganty, Arun and Mussmann, Stephen and Liang, Percy. The price of debiasing automatic metrics in natural language evaluation. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18653/v1/P18-1060
2018 doi
-
[50]
How Good Are LLM s for Literary Translation, Really? Literary Translation Evaluation with Humans and LLM s
Zhang, Ran and Zhao, Wei and Eger, Steffen. How Good Are LLM s for Literary Translation, Really? Literary Translation Evaluation with Humans and LLM s. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: H...
2025 doi
-
[51]
The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation
Klie, Jan-Christoph and Bugert, Michael and Boullosa, Beto and Eckart de Castilho, Richard and Gurevych, Iryna. The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation. Proceedings of the 27th International Conference on Computational Linguisti...
2018
-
[52]
The Challenges of Using Neural Machine Translation for Literature
Matusov, Evgeny. The Challenges of Using Neural Machine Translation for Literature. Proceedings of the Qualities of Literary Machine Translation. 2019
2019
-
[53]
Machine vs Human Translation of Formal Neologisms in Literature: Exploring
Laura Noriega-Santi. Machine vs Human Translation of Formal Neologisms in Literature: Exploring. Revista Tradum. 2023 , url=
2023
-
[54]
Information , VOLUME =
Corpas Pastor, Gloria and Noriega-Santiáñez, Laura , TITLE =. Information , VOLUME =. 2024 , NUMBER =
2024
-
[55]
Large Language Models Effectively Leverage Document-level Context for Literary Translation, but Critical Errors Persist
Karpinska, Marzena and Iyyer, Mohit. Large Language Models Effectively Leverage Document-level Context for Literary Translation, but Critical Errors Persist. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.41
2023 doi
-
[56]
Translation Spaces , volume =
Dorothy Kenny and Marion Winters , title =. Translation Spaces , volume =. 2020 , doi =
2020
-
[57]
ELOPE: English Language Overseas Perspectives and Enquiries , volume =
Tjaša Mohar and Sara Orthaber and Tomaž Onič , title =. ELOPE: English Language Overseas Perspectives and Enquiries , volume =. 2020 , doi =
2020
-
[58]
Proceedings of the Qualities of Literary Machine Translation
Would MT kill creativity in literary retranslation?. Proceedings of the Qualities of Literary Machine Translation. 2019
2019
-
[59]
Creativity in Translation: Machine Translation as a Constraint for Literary Texts , Volume =
Ana Guerberof-Arenas and Antonio Toral , Journal =. Creativity in Translation: Machine Translation as a Constraint for Literary Texts , Volume =
-
[60]
The Impact of Post-Editing and Machine Translation on Creativity and Reading Experience , Volume =
Ana Guerberof-Arenas and Antonio Toral , Journal =. The Impact of Post-Editing and Machine Translation on Creativity and Reading Experience , Volume =
-
[61]
To be or not to be: A translation reception study of a literary text translated into
Ana Guerberof-Arenas and Antonio Toral , Journal =. To be or not to be: A translation reception study of a literary text translated into
-
[62]
Frontiers in Digital Humanities , VOLUME=
Toral, Antonio and Wieling, Martijn and Way, Andy , TITLE=. Frontiers in Digital Humanities , VOLUME=. 2018 , URL=. doi:10.3389/fdigh.2018.00009 , ISSN=
2018 doi
-
[63]
B y GPT 5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
Belouadi, Jonas and Eger, Steffen. B y GPT 5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.406
2023 doi
-
[64]
Evaluating Diversity in Automatic Poetry Generation
Chen, Yanran and Gr. Evaluating Diversity in Automatic Poetry Generation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1097
2024 doi
-
[65]
Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , articleno =
Chakrabarty, Tuhin and Laban, Philippe and Agarwal, Divyansh and Muresan, Smaranda and Wu, Chien-Sheng , title =. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , articleno =. 2024 , isbn =. doi:10.1145/3613904.3642731 , abstract =
2024 doi
-
[66]
Re-evaluating the Role of B leu in Machine Translation Research
Callison-Burch, Chris and Osborne, Miles and Koehn, Philipp. Re-evaluating the Role of B leu in Machine Translation Research. 11th Conference of the E uropean Chapter of the Association for Computational Linguistics. 2006
2006
-
[67]
Tangled up in BLEU : Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics
Mathur, Nitika and Baldwin, Timothy and Cohn, Trevor. Tangled up in BLEU : Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.448
2020 doi
-
[68]
Why We Need New Evaluation Metrics for NLG
Novikova, Jekaterina and Du s ek, Ond r ej and Cercas Curry, Amanda and Rieser, Verena. Why We Need New Evaluation Metrics for NLG. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. doi:10.18653/v1/D17-1238
2017 doi
-
[69]
AI -Assisted Human Evaluation of Machine Translation
Zouhar, Vil \'e m and Kocmi, Tom and Sachan, Mrinmaya. AI -Assisted Human Evaluation of Machine Translation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long ...
2025 doi
-
[70]
INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback
Xu, Wenda and Wang, Danqing and Pan, Liangming and Song, Zhenqiao and Freitag, Markus and Wang, William and Li, Lei. INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Langu...
2023 doi
-
[71]
G -Eval: NLG Evaluation using Gpt-4 with Better Human Alignment
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang. G -Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.em...
2023 doi
-
[72]
Tagged Span Annotation for Detecting Translation Errors in Reasoning LLM s
Yeom, Taemin and Ryu, Yonghyun and Choi, Yoonjung and Bak, Jinyeong. Tagged Span Annotation for Detecting Translation Errors in Reasoning LLM s. Proceedings of the Tenth Conference on Machine Translation. 2025. doi:10.18653/v1/2025.wmt-1.62
2025 doi
-
[73]
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
Qiu, Ziliang and Hu, Renfen. Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.550
2025 doi
-
[74]
Human Translation of Stylistic Neologisms in English Language Chick Lit into Ukrainian , url=
Machine vs. Human Translation of Stylistic Neologisms in English Language Chick Lit into Ukrainian , url=. Respectus Philologicus , author=. 2025 , month=. doi:10.15388/RESPECTUS.2025.48.9 , abstractNote=
2025 doi
-
[75]
2024 , eprint=
CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints , author=. 2024 , eprint=
2024
-
[76]
Automatic Creativity Measurement in Scratch Programs Across Modalities , year=
Kovalkov, Anastasia and Paaßen, Benjamin and Segal, Avi and Pinkwart, Niels and Gal, Kobi , journal=. Automatic Creativity Measurement in Scratch Programs Across Modalities , year=
-
[77]
Prompting Large Language Models for Idiomatic Translation
Castaldo, Antonio and Monti, Johanna. Prompting Large Language Models for Idiomatic Translation. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024
2024
-
[78]
Machine Translation with Large Language Models: Prompting, Few-shot Learning, and Fine-tuning with QL o RA
Zhang, Xuan and Rajabi, Navid and Duh, Kevin and Koehn, Philipp. Machine Translation with Large Language Models: Prompting, Few-shot Learning, and Fine-tuning with QL o RA. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.43
2023 doi
-
[79]
Proceedings of the 40th International Conference on Machine Learning , pages =
Prompting Large Language Model for Machine Translation: A Case Study , author =. Proceedings of the 40th International Conference on Machine Learning , pages =. 2023 , editor =
2023
-
[80]
Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops , articleno =
Gao, Yuan and Wang, Ruili and Hou, Feng , title =. Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops , articleno =. 2024 , isbn =. doi:10.1145/3700410.3702123 , abstract =
2024 doi
-
[81]
Optimizing Machine Translation through Prompt Engineering: An Investigation into C hat GPT ' s Customizability
Yamada, Masaru. Optimizing Machine Translation through Prompt Engineering: An Investigation into C hat GPT ' s Customizability. Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track. 2023
2023
-
[82]
Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Models: An Initial Multi-layered Tabular Review , volume=
Opaluwah, Adeyola , year=. Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Models: An Initial Multi-layered Tabular Review , volume=. American Journal of Computer Science and Technology , publisher=. doi:10.11648/j.ajcst.20250802...
-
[83]
Efficiently Exploring Large Language Models for Document-Level Machine Translation with In-context Learning
Cui, Menglong and Du, Jiangcun and Zhu, Shaolin and Xiong, Deyi. Efficiently Exploring Large Language Models for Document-Level Machine Translation with In-context Learning. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.finding...
2024 doi
-
[84]
`Can make mistakes'
Egdom, Gys-Walt and Declercq, Christophe and Kosters, Onno. `Can make mistakes'. Prompting C hat GPT to Enhance Literary MT output. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024
2024
-
[85]
Creative shifts as a means of measuring and promoting translation Creativity , Volume =
Gerrit Bayer-Hohenwarter , Journal =. Creative shifts as a means of measuring and promoting translation Creativity , Volume =. 2011 , doi =
2011
-
[86]
Extending CREAMT : Leveraging Large Language Models for Literary Translation Post-Editing
Castaldo, Antonio and Castilho, Sheila and Moorkens, Joss and Monti, Johanna. Extending CREAMT : Leveraging Large Language Models for Literary Translation Post-Editing. Proceedings of Machine Translation Summit XX: Volume 1. 2025
2025
-
[87]
To MT or not to MT : An eye-tracking study on the reception by D utch readers of different translation and creativity levels
Gerrits, Kyo and Guerberof-Arenas, Ana. To MT or not to MT : An eye-tracking study on the reception by D utch readers of different translation and creativity levels. Proceedings of Machine Translation Summit XX: Volume 1. 2025
2025
-
[88]
Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =
Chakrabarty, Tuhin and Laban, Philippe and Wu, Chien-Sheng , title =. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =. 2025 , isbn =. doi:10.1145/3706598.3713559 , abstract =
2025 doi
-
[89]
STYLISTICALLY-AWARE
Pragya Tewari and Anurag Singh Baghel , Journal =. STYLISTICALLY-AWARE. 2026 , doi =
2026
-
[90]
Impact of translation workflows with and without MT on textual characteristics in literary translation
Daems, Joke and Ruffo, Paola and Macken, Lieve. Impact of translation workflows with and without MT on textual characteristics in literary translation. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024
2024
-
[91]
Literary translation as a three-stage process: machine translation, post-editing and revision
Macken, Lieve and Vanroy, Bram and Desmet, Luca and Tezcan, Arda. Literary translation as a three-stage process: machine translation, post-editing and revision. Proceedings of the 23rd Annual Conference of the European Association for Machine Translation. 2022
2022
-
[92]
2025 , eprint=
MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment , author=. 2025 , eprint=
2025
-
[93]
2026 , eprint=
Agent-as-a-Judge , author=. 2026 , eprint=
2026
-
[94]
Translation Studies: The State of the Art
Paul Kussmaul , title =. Translation Studies: The State of the Art. Proceedings of the First James S. Holmes Symposium on Translation Studies , publisher =. 1991 , editor =
1991
-
[95]
1995 , address =
Paul Kussmaul , title =. 1995 , address =
1995
-
[96]
Intercultural Faultlines Research Models in Translation Studies: v
Paul Kussmaul , title =. Intercultural Faultlines Research Models in Translation Studies: v. 1: Textual and Cognitive Aspects , editor =. 2000 , pages =
2000
-
[97]
Behind the Mind: Methods, Models and Results in Translation Process Research , editor =
Gerrit Bayer-Hohenwarter , title =. Behind the Mind: Methods, Models and Results in Translation Process Research , editor =. 2009 , pages =
2009
-
[98]
New Approaches in Translation Process Research , editor =
Gerrit Bayer-Hohenwarter , title =. New Approaches in Translation Process Research , editor =. 2010 , pages =
2010
-
[99]
Tracks and Treks in Translation Studies: Selected Papers from the EST Conference, Leuven 2010 , pages =
Gerrit Bayer-Hohenwarter , title =. Tracks and Treks in Translation Studies: Selected Papers from the EST Conference, Leuven 2010 , pages =. 2013 , doi =
2010
-
[100]
The Routledge Handbook of Translation and Cognition , editor =
Gerrit Bayer-Hohenwarter and Paul Kussmaul , title =. The Routledge Handbook of Translation and Cognition , editor =. 2020 , pages =
2020
-
[101]
Kaufman and John Baer , title =
James C. Kaufman and John Baer , title =. Creativity Research Journal , volume =. 2012 , publisher =. doi:10.1080/10400419.2012.649237 , URL =
2012 doi
-
[102]
Multidimensional quality metrics: a flexible system for assessing translation quality
Lommel, Arle Richard and Burchardt, Aljoscha and Uszkoreit, Hans. Multidimensional quality metrics: a flexible system for assessing translation quality. Proceedings of Translating and the Computer 35. 2013
2013
-
[103]
GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4
Kocmi, Tom and Federmann, Christian. GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.64
2023 doi
-
[104]
The Guardian , author =
Dutch publisher to use. The Guardian , author =. 2024 , keywords =
2024
-
[105]
GPTS core: Evaluate as You Desire
Fu, Jinlan and Ng, See-Kiong and Jiang, Zhengbao and Liu, Pengfei. GPTS core: Evaluate as You Desire. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. d...
2024 doi
-
[106]
Is C hat GPT a Good NLG Evaluator? A Preliminary Study
Wang, Jiaan and Liang, Yunlong and Meng, Fandong and Sun, Zengkui and Shi, Haoxiang and Li, Zhixu and Xu, Jinan and Qu, Jianfeng and Zhou, Jie. Is C hat GPT a Good NLG Evaluator? A Preliminary Study. Proceedings of the 4th New Frontiers in Summarization Workshop. 2023. doi:10....
2023 doi
-
[107]
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM -as-a-Judge
Zhang, Qiyuan and Wang, Yufei and Jiang, Yuxin and Li, Liangyou and Wu, Chuhan and Wang, Yasheng and Jiang, Xin and Shang, Lifeng and Tang, Ruiming and Lyu, Fuyuan and Ma, Chen. Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM -as-a-Judge. Proceedings o...
2025 doi
-
[108]
Findings of the WMT 25 Shared Task on Automated Translation Evaluation Systems: Linguistic Diversity is Challenging and References Still Help
Lavie, Alon and Hanneman, Greg and Agrawal, Sweta and Kanojia, Diptesh and Lo, Chi-Kiu and Zouhar, Vil \'e m and Blain, Frederic and Zerva, Chrysoula and Avramidis, Eleftherios and Deoghare, Sourabh and Sindhujan, Archchana and Wang, Jiayi and Adelani, David Ifeoluwa and Thomp...
2025 doi
-
[109]
A global analysis of metrics used for measuring performance in natural language processing
Blagec, Kathrin and Dorffner, Georg and Moradi, Milad and Ott, Simon and Samwald, Matthias. A global analysis of metrics used for measuring performance in natural language processing. Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP. 2022. doi:10.1...
2022 doi
-
[110]
Quality Expectations of Machine Translation
Way, Andy. Quality Expectations of Machine Translation. Translation Quality Assessment: From Principles to Practice. 2018. doi:10.1007/978-3-319-91241-7_8
2018 doi
-
[111]
Is post-editing really faster than human translation? , Volume =
Silvia Terribile , Journal =. Is post-editing really faster than human translation? , Volume =. 2024 , doi =
2024
-
[112]
The Role of Creativity , booktitle =
Rojo, Ana , publisher =. The Role of Creativity , booktitle =. doi:https://doi.org/10.1002/9781119241485.ch19 , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781119241485.ch19 , year =
-
[113]
Translation and Creativity in the 21st Century
Susan Bassnett and Lawrence Venuti and Jan Pedersen and Ivana Hostová , Journal =. Translation and Creativity in the 21st Century. , Volume =. 2022 , doi =
2022
-
[114]
2024 , eprint=
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control , author=. 2024 , eprint=
2024
-
[115]
Christine Hall. 2025
2025
-
[116]
Nielbo , title =
Mia Jacobsen and Yuri Bizzoni and Pascale Feldkamp Moreira and Kristoffer L. Nielbo , title =. Proceedings of the Computational Humanities Research Conference 2024 , pages =
2024
-
[117]
Quantitative analysis of fanfictions’ popularity , Volume =
Zhivar Sourati Hassan Zadeh and Nazanin Sabri and Houmaan Chamani and Behnam Bahrak , Journal =. Quantitative analysis of fanfictions’ popularity , Volume =. 2021 , doi =
2021
-
[118]
International Journal of Human–Computer Interaction , volume =
Roi Alfassi and Angelora Cooper and Zoe Mitchell and Mary Calabro and Orit Shaer and Osnat Mokryn , title =. International Journal of Human–Computer Interaction , volume =. 2026 , publisher =. doi:10.1080/10447318.2025.2531272 , URL =
2026 doi
-
[119]
Optimising C hat GPT for creativity in literary translation: A case study from E nglish into D utch, C hinese, C atalan and S panish
Du, Shuxiang and Arenas, Ana Guerberof and Toral, Antonio and Gerrits, Kyo and Borillo, Josep Marco. Optimising C hat GPT for creativity in literary translation: A case study from E nglish into D utch, C hinese, C atalan and S panish. Proceedings of Machine Translation Summit ...
2025
-
[120]
2024 , eprint=
Large Language Models are Inconsistent and Biased Evaluators , author=. 2024 , eprint=
2024
-
[121]
2025 , journal =
Detecting and Evaluating Bias in Large Language Models: Concepts, Methods, and Challenges , author=. 2025 , journal =
2025
-
[122]
and Jeffrey D
Aho, Alfred V. and Jeffrey D. Ullman. 1972. The Theory of Parsing, Translation and Compiling , volume 1. Prentice- Hall , Englewood Cliffs, NJ
1972
-
[123]
American Psychological Association . 1983. Publications Manual . American Psychological Association, Washington, DC
1983
-
[124]
Association for Computing Machinery . 1983. Computing Reviews , 24(11):503--512
1983
-
[125]
Kozen, and Larry J
Chandra, Ashok K., Dexter C. Kozen, and Larry J. Stockmeyer. 1981. Alternation. Journal of the Asso\-ciation for Computing Machinery , 28(1):114--133
1981
-
[126]
Gledson, Anne, and John Keane. 2008a. Measuring Topic Homogeneity and its Application to Dictionary-Based Word-Sense Disambiguation. Coling 2008, 22nd International Conference on Computational Linguistics , Manchester, UK. 273--280
2008
-
[127]
Gledson, Anne, and John Keane. 2008b. Using Web-Search Results to Measure Word-group Similarity. Coling 2008, 22nd International Conference on Computational Linguistics , Manchester, UK. 281--288
2008
-
[128]
Gusfield, Dan. 1997. Algorithms on Strings, Trees and Sequences . Cambridge University Press, Cambridge, UK
1997
-
[129]
Tam, Yik-Cheung and Tanja Schultz. 2006. Unsupervised Language Model Adaptation Using Latent Semantic Marginals. Interspeech 2006 -- ICSLP, Ninth International Conference on Spoken Language Processing , Pittsburgh, Pennsylvania, paper 1705-Thu1A2O.2
2006
-
[130]
Tam, Yik-Cheung and Tanja Schultz. 2007. Correlated Latent Semantic Model for Unsupervised Language Model Adaptation. Proceedings of ICASSP 2007, International Conference on Acoustics, Speech, and Signal Processing , Honolulu, Hawaii, Vol. IV, 41--44
2007
Reviewed May 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.