Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

T0 review · 2 major / 2 minor · reviewed 2026-05-14 · grok-4.3

Pith's one-line read Automatic evaluation metrics and LLM judges correlate poorly with professional translators on creativity in literary texts and bias toward machine outputs.

desk verdict The paper documents LLM judges favoring machine translations on creativity scores in literary MT, backed by a new multi-genre dataset, but the annotations lack reported agreement stats. read the letter →

arxiv 2605.13596 v1 pith:IPZIX56F submitted 2026-05-13 cs.CL

classification cs.CL
keywords literarytranslationcreativityevaluationautomaticmetricsLLM-as-a-judgemachinehumanbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper builds a dataset of literary translations in three modalities, three genres, and three language pairs, then has experienced professional translators annotate them in detail for creative shifts and errors. It compares these annotations against scores from automatic evaluation metrics and from LLM-as-a-judge setups. Both automatic approaches show weak alignment with the professionals, and LLM judges systematically rate machine-translated versions higher while marking creative, culturally fitting solutions as mistakes. The gap widens for poetry and other highly literary genres. The work concludes that existing tools cannot yet replace manual expert judgment and that new methods are needed to treat creative deviations as valid rather than errors.

What carries the argument

A dataset of literary translations across human, machine, and post-edited modalities, annotated by professional translators for creative shifts and errors, used to measure alignment with automatic metrics and LLM judges.

What would settle it

A new set of annotations on the same dataset by a different group of professional literary translators that produces substantially different creativity scores from the original annotations.

Watch

Extended reading notes

Core claim

Automatic evaluation metrics and LLM-as-a-judge evaluations correlate poorly with professional literary translators' assessments of creativity, and LLM judges display a systematic bias that favors machine-translated texts while penalizing creative and culturally appropriate solutions, with performance dropping further on poetry and similar literary genres.

Load-bearing premise

Detailed annotations by experienced professional literary translators constitute an objective and reliable ground truth for measuring creativity and translation quality across genres and modalities.

Editorial extensions

If this is right

  • Automatic metrics cannot serve as reliable substitutes for professional judgment when creativity is the focus of evaluation.
  • LLM-as-a-judge methods introduce a consistent preference for literal machine outputs over creative human solutions.
  • Evaluation accuracy declines markedly for poetry and other highly literary genres.
  • New automatic tools are required that treat creative out-of-routine solutions as valid rather than errors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If automatic scores are used for quality control, translation workflows may systematically undervalue creative post-editing by humans.
  • The same bias pattern could appear in AI evaluation of other creative writing tasks such as story generation or script adaptation.
  • Explicit cultural and creative criteria would need to be built into future evaluation frameworks to reduce the observed mismatch.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper claims that automatic evaluation metrics (AEMs) and LLM-as-a-judge methods correlate poorly with professional literary translators' assessments of translation quality and creativity (creative shifts and errors) across three language pairs, three genres, and three modalities (human translation, machine translation, post-editing). It further reports that LLM judges exhibit systematic bias favoring machine-translated outputs while penalizing creative and culturally appropriate solutions, with performance degrading for more literary genres such as poetry.

Significance. If the empirical findings hold after addressing reporting gaps, the work provides useful evidence of limitations in current automatic tools for evaluating creativity in literary translation and motivates development of new metrics. The construction of a multi-genre, multi-modality dataset annotated by experienced professionals is a concrete contribution that can support future research.

major comments (2)
  1. [Annotation methodology] Annotation methodology section: no inter-annotator agreement statistics (Fleiss' kappa, Cohen's kappa, or equivalent) are reported for the creativity labels (creative shifts & errors) assigned by the professional translators. Without these figures, the low correlations and reported LLM bias could reflect label noise rather than genuine misalignment with the ground truth.
  2. [Results] Results section: the abstract and main results provide no sample sizes (number of translations or segments per genre/modality), exact implementations of the AEMs and LLM prompts, correlation coefficients with confidence intervals, or statistical tests (e.g., p-values for differences between modalities). These omissions prevent assessment of whether the claimed poor correlations and systematic bias are statistically robust.
minor comments (2)
  1. [Abstract] Abstract: add concrete numbers (e.g., total segments annotated, number of annotators, range of correlation values) to make the headline claims more informative.
  2. [Dataset construction] Dataset description: clarify selection criteria for the literary texts and any filtering applied after annotation.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback, which has helped us strengthen the reporting in our manuscript. We address each major comment below and have revised the paper to incorporate the requested details on annotation reliability and statistical reporting.

read point-by-point responses
  1. Referee: [Annotation methodology] Annotation methodology section: no inter-annotator agreement statistics (Fleiss' kappa, Cohen's kappa, or equivalent) are reported for the creativity labels (creative shifts & errors) assigned by the professional translators. Without these figures, the low correlations and reported LLM bias could reflect label noise rather than genuine misalignment with the ground truth.

    Authors: We agree that inter-annotator agreement metrics are essential to demonstrate label reliability. Although the original submission omitted these statistics, the annotations were performed by multiple professional translators with overlapping segments. In the revised manuscript we now report Cohen's kappa values for the creativity shift and error labels (computed on the double-annotated subset), which show substantial agreement. This addition confirms that the observed low correlations with automatic metrics are unlikely to stem from label noise. revision: yes

  2. Referee: [Results] Results section: the abstract and main results provide no sample sizes (number of translations or segments per genre/modality), exact implementations of the AEMs and LLM prompts, correlation coefficients with confidence intervals, or statistical tests (e.g., p-values for differences between modalities). These omissions prevent assessment of whether the claimed poor correlations and systematic bias are statistically robust.

    Authors: We acknowledge the reporting gaps. The revised version now includes explicit sample sizes (number of segments per genre, language pair, and modality) in both the abstract and Results section. We have added the precise AEM implementations, full LLM prompts, Pearson/Spearman correlations with 95% confidence intervals, and p-values from appropriate statistical tests comparing modalities. These details are also summarized in a new table for clarity and confirm the robustness of the reported poor correlations and LLM bias. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: purely empirical comparison to external annotations

full rationale

The paper constructs a dataset of literary translations across modalities and genres, obtains detailed annotations for creativity from experienced professional translators, and reports correlations between these annotations and both automatic evaluation metrics and LLM-as-a-judge outputs. No equations, derivations, fitted parameters, or self-referential definitions appear in the provided text. The central claims rest on observed empirical mismatches rather than any reduction of predictions to inputs by construction. Self-citations, if present, are not load-bearing for the core result. This is a standard empirical evaluation study whose validity depends on annotation quality and inter-annotator agreement (a separate reliability concern), not on circular logic.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on treating professional literary translator annotations as the authoritative benchmark for creativity; no free parameters or invented entities are introduced.

assumptions (1)
  • domain assumption Professional literary translators' judgments provide a reliable and objective measure of creativity and translation quality
    Used as the ground-truth reference against which AEMs and LLM judges are evaluated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations." pith.science (2026). https://pith.science/paper/IPZIX56F

@misc{pith2026260513596,
  author       = {Pith},
  title        = {Pith review of: Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IPZIX56F}},
  note         = {Machine review of arXiv:2605.13596}
}
read the original abstract

This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess how well these tools align with professionals when evaluating translation, creativity (creative shifts & errors), and see if they can substitute laborious manual annotations. A dataset of literary translations across three modalities (human translation, machine translation, and post-editing), three genres and three language pairs was created and annotated in detail for creativity by experienced professional literary translators. The results show that both AEMs and LLM-as-a-judge evaluations correlate poorly with professional evaluations on creativity, with LLM-as-a-judge showing a systematic bias in favour of machine-translated texts and penalising creative and culturally appropriate solutions. Moreover, performance is consistently worse for more literary genres such as poetry. This highlights fundamental limitations of current automatic evaluation tools for literary translation and the need to create new tools that do not frequently consider out of routine translations as errors.

Figures

Figures reproduced from arXiv: 2605.13596 by the authors.

Figure 1
Figure 1. Example of a UCP in the ST with a Reproduction in the PE and a CS in the HT with English glosses underneath. The creative shift annotations were done by 3 doctoral students who are proficient with the framework that were also translators and were na￾tive or proficient speakers for the language pair. They received the texts and the UCP annotations (created by 2 of the doctoral students) and could mark a solution as C… view at source ↗
Figure 2
Figure 2. Heat map of the Spearman correlations between AEMs and professional annotations. Poem Short story Thriller ST TT From HT PE MT HT PE MT HT PE MT EN NL Human 3 8 16 4 5 17 6 11 9 LLM 8 12 12 14 8 12 8 9 9 Match 0 3 6 2 0 2 1 4 1 EN CA Human 9 7 16 3 7 25 4 7 18 LLM 13 16 9 14 6 8 9 10 5 Match 4 4 7 1 1 4 0 1 4 RU NL Human 7 6 14 4 14 14 9 15 28 LLM 19 23 24 10 13 10 9 13 20 Match 4 5 8 2 6 4 0 5 12 [PITH_FULL_IMAGE:… view at source ↗
Figure 3
Figure 3. Scatterplot of CI from professional and LLM anno￾tations. Colouring indicates the modality and shapes genre. 4.3 Genre Our last RQ investigates whether the correlations between AEMs and LLM-as-a-judge for the first 16The shapes reflect genre, see Section 4.3. 17MT scores lowest, followed by PE and topped by HT (as shown in the figure by the green dots towards the lower end of the Y-axis, the blue in between and the … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Heat map of Spearman correlations between AEMs and professional evaluations across genres. # Errors Poem Short Story Thriller Human 84 93 107 LLM 137 95 93 Match 42 22 28 [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Box plots for errors, creative shift (CS) and creativ￾ity index (CI) per each modality level. 0 50 100 150 Poem ShortStory Thriller Genre ErrorPoints 0 5 10 15 Poem ShortStory Thriller Genre CS −150 −100 −50 0 50 Poem ShortStory Thriller Genre C.I [PITH_FULL_IMAGE:fig…
Figure 6
Figure 6. Figure 6: Box plots for errors, creative shift (CS) and creativ￾ity index (CI) per each genre level. increase in costs and environmental impact, but we did try out multiple models and prompts us￾ing selected texts. Specifically, we tried out TSA (Yeom et al., 2025), AutoMQM (Fer…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI translation of literary texts is "fine", but readers still prefer human translations

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Human readers prefer human literary translations over AI-generated ones for immersion and clarity despite finding MT adequate and struggling to identify the source.

  2. LitSeg: Narrative-Aware Document Segmentation for Literary RAG

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    LitSeg segments literary texts using narrative analysis via multi-stage prompting and offers a distilled lightweight version for efficient use in RAG systems.

Reference graph

Works this paper leans on

130 extracted references · 130 canonical work pages · cited by 2 Pith papers

  1. [1]

    and Ullman, Jeffrey D

    Aho, Alfred V. and Ullman, Jeffrey D. , title =. 1972 , volume=1, publisher =

  2. [2]

    Interspeech 2006 --- Ninth International Conference on Spoken Language Processing , address=

    Unsupervised language model adaptation using latent semantic marginals , author=. Interspeech 2006 --- Ninth International Conference on Spoken Language Processing , address=. 2006 , pages=

  3. [3]

    1983 , publisher=

    Publications. 1983 , publisher=

  4. [4]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year=1981, title=. Journal of the Association for Computing Machinery , volume=28, issue=1, pages=

  5. [5]

    Coling 2008, 22nd International Conference on Computational Linguistics , address=

    Anne Gledson and John Keane , year=2008, title=. Coling 2008, 22nd International Conference on Computational Linguistics , address=

  6. [6]

    Dan Gusfield , title=

  7. [7]

    Proceedings of ICASSP 2007, International Conference on Acoustics, Speech, and Signal Processing , address=

    Yik-Cheung Tam and Tanja Schultz , year=2007, title=. Proceedings of ICASSP 2007, International Conference on Acoustics, Speech, and Signal Processing , address=

  8. [8]

    MATEO : MA chine T ranslation E valuation O nline

    Vanroy, Bram and Tezcan, Arda and Macken, Lieve. MATEO : MA chine T ranslation E valuation O nline. Proceedings of the 24th Annual Conference of the European Association for Machine Translation. 2023

Show all 130 references
  1. [9]

    2020 , eprint=

    BERTScore: Evaluating Text Generation with BERT , author=. 2020 , eprint=

  2. [10]

    Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle=

    Kishore Papineni and Salim Roukos and Todd Ward and Wei-Jing Zhu , year=. Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle=. doi:10.3115/1073083.1073135 , pages=

  3. [11]

    A Study of Translation Edit Rate with Targeted Human Annotation

    Matthew Snover and Dorr, Bonnie and Schwartz, Rich and Micciulla, Linnea and Makhoul, John. A Study of Translation Edit Rate with Targeted Human Annotation. Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers. 2006

  4. [12]

    chr F deconstructed: beta parameters and n-gram weights

    Popovi \'c , Maja. chr F deconstructed: beta parameters and n-gram weights. Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers. 2016. doi:10.18653/v1/W16-2341

  5. [13]

    COMET : A Neural Framework for MT Evaluation

    Rei, Ricardo and Stewart, Craig and Farinha, Ana C and Lavie, Alon. COMET : A Neural Framework for MT Evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18653/v1/2020.emnlp-main.213

  6. [14]

    BLEURT : Learning Robust Metrics for Text Generation

    Sellam, Thibault and Das, Dipanjan and Parikh, Ankur. BLEURT : Learning Robust Metrics for Text Generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.704

  7. [15]

    L i T rans P ro QA : An LLM -based Literary Translation Evaluation Metric with Professional Question Answering

    Zhang, Ran and Zhao, Wei and Macken, Lieve and Eger, Steffen. L i T rans P ro QA : An LLM -based Literary Translation Evaluation Metric with Professional Question Answering. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18...

  8. [16]

    Prompting C hat GPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts

    He, Sui. Prompting C hat GPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts. Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1). 2024

  9. [17]

    Poetry and Emotions: Investigating the Limitations of AI Translation

    Priya Sharma and Tanuja Yadav. Poetry and Emotions: Investigating the Limitations of AI Translation. Proceedings of International Conference on Innovations in Data Science. 2026

  10. [18]

    AL-ĪMĀN Research Journal , author=

    A Comparative Analysis of Machine Translation and Human Translation: Efficacy of Poetry Translation from Urdu to English , volume=. AL-ĪMĀN Research Journal , author=. 2026 , month=. doi:10.63283/IRJ.04.01/02 , abstractNote=

  11. [19]

    Don ' t Go Far Off: An Empirical Study on Neural Poetry Translation

    Chakrabarty, Tuhin and Saakyan, Arkadiy and Muresan, Smaranda. Don ' t Go Far Off: An Empirical Study on Neural Poetry Translation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.577

  12. [20]

    The Translator ' s Canvas: Using LLM s to Enhance Poetry Translation

    Resende, Nat \'a lia and Hadley, James. The Translator ' s Canvas: Using LLM s to Enhance Poetry Translation. Proceedings of the 16th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track). 2024

  13. [21]

    A Detailed Comparative Analysis of Automatic Neural Metrics for Machine Translation:

    Mukherjee, Aniruddha and Hassija, Vikas and Chamola, Vinay and Gupta, Karunesh Kumar , journal=. A Detailed Comparative Analysis of Automatic Neural Metrics for Machine Translation:. 2025 , volume=

  14. [22]

    A Fine-Grained Analysis of BERTS core

    Hanna, Michael and Bojar, Ond r ej. A Fine-Grained Analysis of BERTS core. Proceedings of the Sixth Conference on Machine Translation. 2021

  15. [23]

    Extrinsic Evaluation of Machine Translation Metrics

    Nikita Moghe and Tom Sherborne and Mark Steedman and Alexandra Birch. Extrinsic Evaluation of Machine Translation Metrics. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.730

  16. [24]

    Results of WMT 23 Metrics Shared Task: Metrics Might Be Guilty but References Are Not Innocent

    Freitag, Markus and Mathur, Nitika and Lo, Chi-kiu and Avramidis, Eleftherios and Rei, Ricardo and Thompson, Brian and Kocmi, Tom and Blain, Frederic and Deutsch, Daniel and Stewart, Craig and Zerva, Chrysoula and Castilho, Sheila and Lavie, Alon and Foster, George. Results of...

  17. [25]

    BLEU , METEOR , BERTS core: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-Oriented Text

    Saadany, Hadeel and Orasan, Constantin. BLEU , METEOR , BERTS core: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-Oriented Text. Proceedings of the Translation and Interpreting Technology Online Conference. 2021

  18. [26]

    Findings of the WMT 25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets

    Kocmi, Tom and Artemova, Ekaterina and Avramidis, Eleftherios and Bawden, Rachel and Bojar, Ond r ej and Dranch, Konstantin and Dvorkovich, Anton and Dukanov, Sergey and Fishel, Mark and Freitag, Markus and Gowda, Thamme and Grundkiewicz, Roman and Haddow, Barry and Karpinska,...

  19. [27]

    Freitag, Markus and Rei, Ricardo and Mathur, Nitika and Lo, Chi-kiu and Stewart, Craig and Avramidis, Eleftherios and Kocmi, Tom and Foster, George and Lavie, Alon and Martins, Andr \'e F. T. Results of WMT 22 Metrics Shared Task: Stop Using BLEU -- Neural Metrics Are Better a...

  20. [28]

    and Rei, Ricardo and Stigt, Daan van and Coheur, Luisa and Colombo, Pierre and Martins, Andr \'e F

    Guerreiro, Nuno M. and Rei, Ricardo and Stigt, Daan van and Coheur, Luisa and Colombo, Pierre and Martins, Andr \'e F. T. x COMET : Transparent Machine Translation Evaluation through Fine-grained Error Detection. Transactions of the Association for Computational Linguistics. 2...

  21. [29]

    Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains

    Zouhar, Vil \'e m and Ding, Shuoyang and Currey, Anna and Badeka, Tatyana and Wang, Jenyuan and Thompson, Brian. Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2...

  22. [30]

    ACES : Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics

    Amrhein, Chantal and Moghe, Nikita and Guillou, Liane. ACES : Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics. Proceedings of the Seventh Conference on Machine Translation (WMT). 2022. doi:10.18653/v1/2022.wmt-1.44

  23. [31]

    Evaluating Automatic Metrics with Incremental Machine Translation Systems

    Wu, Guojun and Cohen, Shay B and Sennrich, Rico. Evaluating Automatic Metrics with Incremental Machine Translation Systems. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.169

  24. [32]

    To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation

    Kocmi, Tom and Federmann, Christian and Grundkiewicz, Roman and Junczys-Dowmunt, Marcin and Matsushita, Hitokazu and Menezes, Arul. To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation. Proceedings of the Sixth Conference on Machine Tran...

  25. [33]

    DEMETR : Diagnosing Evaluation Metrics for Translation

    Karpinska, Marzena and Raj, Nishant and Thai, Katherine and Song, Yixiao and Gupta, Ankita and Iyyer, Mohit. DEMETR : Diagnosing Evaluation Metrics for Translation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/20...

  26. [34]

    The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

    Fernandes, Patrick and Deutsch, Daniel and Finkelstein, Mara and Riley, Parker and Martins, Andr \'e and Neubig, Graham and Garg, Ankush and Clark, Jonathan and Freitag, Markus and Firat, Orhan. The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Mach...

  27. [35]

    Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models

    Lu, Qingyu and Qiu, Baopu and Ding, Liang and Zhang, Kanjian and Kocmi, Tom and Tao, Dacheng. Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1...

  28. [36]

    Evolutionary Studies in Imaginative Culture , year =

    Nehal Ali AbdulGhaffar , title =. Evolutionary Studies in Imaginative Culture , year =

  29. [37]

    and Zerva, Chrysoula and Farinha, Ana C and Maroti, Christine and C

    Rei, Ricardo and Treviso, Marcos and Guerreiro, Nuno M. and Zerva, Chrysoula and Farinha, Ana C and Maroti, Christine and C. de Souza, Jos \'e G. and Glushkova, Taisiya and Alves, Duarte and Coheur, Luisa and Lavie, Alon and Martins, Andr \'e F. T. C omet K iwi: IST -Unbabel 2...

  30. [38]

    Findings of the WMT 25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation

    Kocmi, Tom and Agrawal, Sweta and Artemova, Ekaterina and Avramidis, Eleftherios and Briakou, Eleftheria and Chen, Pinzhen and Fadaee, Marzieh and Freitag, Markus and Grundkiewicz, Roman and Hou, Yupeng and Koehn, Philipp and Kreutzer, Julia and Mansour, Saab and Perrella, Ste...

  31. [39]

    M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task

    Juraska, Juraj and Deutsch, Daniel and Finkelstein, Mara and Freitag, Markus. M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task. Proceedings of the Ninth Conference on Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.35

  32. [40]

    Findings of the WMT 24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet

    Kocmi, Tom and Avramidis, Eleftherios and Bawden, Rachel and Bojar, Ond r ej and Dvorkovich, Anton and Federmann, Christian and Fishel, Mark and Freitag, Markus and Gowda, Thamme and Grundkiewicz, Roman and Haddow, Barry and Karpinska, Marzena and Koehn, Philipp and Marie, Ben...

  33. [41]

    Can Automatic Metrics Assess High-Quality Translations?

    Agrawal, Sweta and Farinhas, Ant \'o nio and Rei, Ricardo and Martins, Andre. Can Automatic Metrics Assess High-Quality Translations?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.802

  34. [42]

    and Platek, Ondrej and Sivaprasad, Adarsa

    Schmidtova, Patricia and Mahamood, Saad and Balloccu, Simone and Dusek, Ondrej and Gatt, Albert and Gkatzia, Dimitra and Howcroft, David M. and Platek, Ondrej and Sivaprasad, Adarsa. Automatic Metrics in Natural Language Generation: A survey of Current Evaluation Practices. Pr...

  35. [43]

    Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature

    Thai, Katherine and Karpinska, Marzena and Krishna, Kalpesh and Ray, Bill and Inghilleri, Moira and Wieting, John and Iyyer, Mohit. Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature. Proceedings of the 2022 Conference on Empir...

  36. [44]

    2025 , eprint=

    MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation , author=. 2025 , eprint=

  37. [45]

    H i MATE : A Hierarchical Multi-Agent Framework for Machine Translation Evaluation

    Zhang, Shijie and Li, Renhao and Wang, Songsheng and Koehn, Philipp and Yang, Min and Wong, Derek F. H i MATE : A Hierarchical Multi-Agent Framework for Machine Translation Evaluation. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. doi:10.18653/v1...

  38. [46]

    M - MAD : Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation

    Feng, Zhaopeng and Su, Jiayuan and Zheng, Jiamei and Ren, Jiahan and Zhang, Yan and Wu, Jian and Wang, Hongwei and Liu, Zuozhu. M - MAD : Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation. Proceedings of the 63rd Annual Meeting of the Association ...

  39. [47]

    The Atlantic , year =

    Jeremy Klemin , title =. The Atlantic , year =

  40. [48]

    The Bookseller , year =

    Emily Warner , title =. The Bookseller , year =

  41. [49]

    The price of debiasing automatic metrics in natural language evaluation

    Chaganty, Arun and Mussmann, Stephen and Liang, Percy. The price of debiasing automatic metrics in natural language evaluation. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18653/v1/P18-1060

  42. [50]

    How Good Are LLM s for Literary Translation, Really? Literary Translation Evaluation with Humans and LLM s

    Zhang, Ran and Zhao, Wei and Eger, Steffen. How Good Are LLM s for Literary Translation, Really? Literary Translation Evaluation with Humans and LLM s. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: H...

  43. [51]

    The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation

    Klie, Jan-Christoph and Bugert, Michael and Boullosa, Beto and Eckart de Castilho, Richard and Gurevych, Iryna. The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation. Proceedings of the 27th International Conference on Computational Linguisti...

  44. [52]

    The Challenges of Using Neural Machine Translation for Literature

    Matusov, Evgeny. The Challenges of Using Neural Machine Translation for Literature. Proceedings of the Qualities of Literary Machine Translation. 2019

  45. [53]

    Machine vs Human Translation of Formal Neologisms in Literature: Exploring

    Laura Noriega-Santi. Machine vs Human Translation of Formal Neologisms in Literature: Exploring. Revista Tradum. 2023 , url=

  46. [54]

    Information , VOLUME =

    Corpas Pastor, Gloria and Noriega-Santiáñez, Laura , TITLE =. Information , VOLUME =. 2024 , NUMBER =

  47. [55]

    Large Language Models Effectively Leverage Document-level Context for Literary Translation, but Critical Errors Persist

    Karpinska, Marzena and Iyyer, Mohit. Large Language Models Effectively Leverage Document-level Context for Literary Translation, but Critical Errors Persist. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.41

  48. [56]

    Translation Spaces , volume =

    Dorothy Kenny and Marion Winters , title =. Translation Spaces , volume =. 2020 , doi =

  49. [57]

    ELOPE: English Language Overseas Perspectives and Enquiries , volume =

    Tjaša Mohar and Sara Orthaber and Tomaž Onič , title =. ELOPE: English Language Overseas Perspectives and Enquiries , volume =. 2020 , doi =

  50. [58]

    Proceedings of the Qualities of Literary Machine Translation

    Would MT kill creativity in literary retranslation?. Proceedings of the Qualities of Literary Machine Translation. 2019

  51. [59]

    Creativity in Translation: Machine Translation as a Constraint for Literary Texts , Volume =

    Ana Guerberof-Arenas and Antonio Toral , Journal =. Creativity in Translation: Machine Translation as a Constraint for Literary Texts , Volume =

  52. [60]

    The Impact of Post-Editing and Machine Translation on Creativity and Reading Experience , Volume =

    Ana Guerberof-Arenas and Antonio Toral , Journal =. The Impact of Post-Editing and Machine Translation on Creativity and Reading Experience , Volume =

  53. [61]

    To be or not to be: A translation reception study of a literary text translated into

    Ana Guerberof-Arenas and Antonio Toral , Journal =. To be or not to be: A translation reception study of a literary text translated into

  54. [62]

    Frontiers in Digital Humanities , VOLUME=

    Toral, Antonio and Wieling, Martijn and Way, Andy , TITLE=. Frontiers in Digital Humanities , VOLUME=. 2018 , URL=. doi:10.3389/fdigh.2018.00009 , ISSN=

  55. [63]

    B y GPT 5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models

    Belouadi, Jonas and Eger, Steffen. B y GPT 5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.406

  56. [64]

    Evaluating Diversity in Automatic Poetry Generation

    Chen, Yanran and Gr. Evaluating Diversity in Automatic Poetry Generation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1097

  57. [65]

    Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , articleno =

    Chakrabarty, Tuhin and Laban, Philippe and Agarwal, Divyansh and Muresan, Smaranda and Wu, Chien-Sheng , title =. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , articleno =. 2024 , isbn =. doi:10.1145/3613904.3642731 , abstract =

  58. [66]

    Re-evaluating the Role of B leu in Machine Translation Research

    Callison-Burch, Chris and Osborne, Miles and Koehn, Philipp. Re-evaluating the Role of B leu in Machine Translation Research. 11th Conference of the E uropean Chapter of the Association for Computational Linguistics. 2006

  59. [67]

    Tangled up in BLEU : Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics

    Mathur, Nitika and Baldwin, Timothy and Cohn, Trevor. Tangled up in BLEU : Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.448

  60. [68]

    Why We Need New Evaluation Metrics for NLG

    Novikova, Jekaterina and Du s ek, Ond r ej and Cercas Curry, Amanda and Rieser, Verena. Why We Need New Evaluation Metrics for NLG. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. doi:10.18653/v1/D17-1238

  61. [69]

    AI -Assisted Human Evaluation of Machine Translation

    Zouhar, Vil \'e m and Kocmi, Tom and Sachan, Mrinmaya. AI -Assisted Human Evaluation of Machine Translation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long ...

  62. [70]

    INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback

    Xu, Wenda and Wang, Danqing and Pan, Liangming and Song, Zhenqiao and Freitag, Markus and Wang, William and Li, Lei. INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Langu...

  63. [71]

    G -Eval: NLG Evaluation using Gpt-4 with Better Human Alignment

    Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang. G -Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.em...

  64. [72]

    Tagged Span Annotation for Detecting Translation Errors in Reasoning LLM s

    Yeom, Taemin and Ryu, Yonghyun and Choi, Yoonjung and Bak, Jinyeong. Tagged Span Annotation for Detecting Translation Errors in Reasoning LLM s. Proceedings of the Tenth Conference on Machine Translation. 2025. doi:10.18653/v1/2025.wmt-1.62

  65. [73]

    Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models

    Qiu, Ziliang and Hu, Renfen. Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.550

  66. [74]

    Human Translation of Stylistic Neologisms in English Language Chick Lit into Ukrainian , url=

    Machine vs. Human Translation of Stylistic Neologisms in English Language Chick Lit into Ukrainian , url=. Respectus Philologicus , author=. 2025 , month=. doi:10.15388/RESPECTUS.2025.48.9 , abstractNote=

  67. [75]

    2024 , eprint=

    CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints , author=. 2024 , eprint=

  68. [76]

    Automatic Creativity Measurement in Scratch Programs Across Modalities , year=

    Kovalkov, Anastasia and Paaßen, Benjamin and Segal, Avi and Pinkwart, Niels and Gal, Kobi , journal=. Automatic Creativity Measurement in Scratch Programs Across Modalities , year=

  69. [77]

    Prompting Large Language Models for Idiomatic Translation

    Castaldo, Antonio and Monti, Johanna. Prompting Large Language Models for Idiomatic Translation. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024

  70. [78]

    Machine Translation with Large Language Models: Prompting, Few-shot Learning, and Fine-tuning with QL o RA

    Zhang, Xuan and Rajabi, Navid and Duh, Kevin and Koehn, Philipp. Machine Translation with Large Language Models: Prompting, Few-shot Learning, and Fine-tuning with QL o RA. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.43

  71. [79]

    Proceedings of the 40th International Conference on Machine Learning , pages =

    Prompting Large Language Model for Machine Translation: A Case Study , author =. Proceedings of the 40th International Conference on Machine Learning , pages =. 2023 , editor =

  72. [80]

    Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops , articleno =

    Gao, Yuan and Wang, Ruili and Hou, Feng , title =. Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops , articleno =. 2024 , isbn =. doi:10.1145/3700410.3702123 , abstract =

  73. [81]

    Optimizing Machine Translation through Prompt Engineering: An Investigation into C hat GPT ' s Customizability

    Yamada, Masaru. Optimizing Machine Translation through Prompt Engineering: An Investigation into C hat GPT ' s Customizability. Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track. 2023

  74. [82]

    Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Models: An Initial Multi-layered Tabular Review , volume=

    Opaluwah, Adeyola , year=. Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Models: An Initial Multi-layered Tabular Review , volume=. American Journal of Computer Science and Technology , publisher=. doi:10.11648/j.ajcst.20250802...

  75. [83]

    Efficiently Exploring Large Language Models for Document-Level Machine Translation with In-context Learning

    Cui, Menglong and Du, Jiangcun and Zhu, Shaolin and Xiong, Deyi. Efficiently Exploring Large Language Models for Document-Level Machine Translation with In-context Learning. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.finding...

  76. [84]

    `Can make mistakes'

    Egdom, Gys-Walt and Declercq, Christophe and Kosters, Onno. `Can make mistakes'. Prompting C hat GPT to Enhance Literary MT output. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024

  77. [85]

    Creative shifts as a means of measuring and promoting translation Creativity , Volume =

    Gerrit Bayer-Hohenwarter , Journal =. Creative shifts as a means of measuring and promoting translation Creativity , Volume =. 2011 , doi =

  78. [86]

    Extending CREAMT : Leveraging Large Language Models for Literary Translation Post-Editing

    Castaldo, Antonio and Castilho, Sheila and Moorkens, Joss and Monti, Johanna. Extending CREAMT : Leveraging Large Language Models for Literary Translation Post-Editing. Proceedings of Machine Translation Summit XX: Volume 1. 2025

  79. [87]

    To MT or not to MT : An eye-tracking study on the reception by D utch readers of different translation and creativity levels

    Gerrits, Kyo and Guerberof-Arenas, Ana. To MT or not to MT : An eye-tracking study on the reception by D utch readers of different translation and creativity levels. Proceedings of Machine Translation Summit XX: Volume 1. 2025

  80. [88]

    Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =

    Chakrabarty, Tuhin and Laban, Philippe and Wu, Chien-Sheng , title =. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =. 2025 , isbn =. doi:10.1145/3706598.3713559 , abstract =

  81. [89]

    STYLISTICALLY-AWARE

    Pragya Tewari and Anurag Singh Baghel , Journal =. STYLISTICALLY-AWARE. 2026 , doi =

  82. [90]

    Impact of translation workflows with and without MT on textual characteristics in literary translation

    Daems, Joke and Ruffo, Paola and Macken, Lieve. Impact of translation workflows with and without MT on textual characteristics in literary translation. Proceedings of the 1st Workshop on Creative-text Translation and Technology. 2024

  83. [91]

    Literary translation as a three-stage process: machine translation, post-editing and revision

    Macken, Lieve and Vanroy, Bram and Desmet, Luca and Tezcan, Arda. Literary translation as a three-stage process: machine translation, post-editing and revision. Proceedings of the 23rd Annual Conference of the European Association for Machine Translation. 2022

  84. [92]

    2025 , eprint=

    MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment , author=. 2025 , eprint=

  85. [93]

    2026 , eprint=

    Agent-as-a-Judge , author=. 2026 , eprint=

  86. [94]

    Translation Studies: The State of the Art

    Paul Kussmaul , title =. Translation Studies: The State of the Art. Proceedings of the First James S. Holmes Symposium on Translation Studies , publisher =. 1991 , editor =

  87. [95]

    1995 , address =

    Paul Kussmaul , title =. 1995 , address =

  88. [96]

    Intercultural Faultlines Research Models in Translation Studies: v

    Paul Kussmaul , title =. Intercultural Faultlines Research Models in Translation Studies: v. 1: Textual and Cognitive Aspects , editor =. 2000 , pages =

  89. [97]

    Behind the Mind: Methods, Models and Results in Translation Process Research , editor =

    Gerrit Bayer-Hohenwarter , title =. Behind the Mind: Methods, Models and Results in Translation Process Research , editor =. 2009 , pages =

  90. [98]

    New Approaches in Translation Process Research , editor =

    Gerrit Bayer-Hohenwarter , title =. New Approaches in Translation Process Research , editor =. 2010 , pages =

  91. [99]

    Tracks and Treks in Translation Studies: Selected Papers from the EST Conference, Leuven 2010 , pages =

    Gerrit Bayer-Hohenwarter , title =. Tracks and Treks in Translation Studies: Selected Papers from the EST Conference, Leuven 2010 , pages =. 2013 , doi =

  92. [100]

    The Routledge Handbook of Translation and Cognition , editor =

    Gerrit Bayer-Hohenwarter and Paul Kussmaul , title =. The Routledge Handbook of Translation and Cognition , editor =. 2020 , pages =

  93. [101]

    Kaufman and John Baer , title =

    James C. Kaufman and John Baer , title =. Creativity Research Journal , volume =. 2012 , publisher =. doi:10.1080/10400419.2012.649237 , URL =

  94. [102]

    Multidimensional quality metrics: a flexible system for assessing translation quality

    Lommel, Arle Richard and Burchardt, Aljoscha and Uszkoreit, Hans. Multidimensional quality metrics: a flexible system for assessing translation quality. Proceedings of Translating and the Computer 35. 2013

  95. [103]

    GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4

    Kocmi, Tom and Federmann, Christian. GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4. Proceedings of the Eighth Conference on Machine Translation. 2023. doi:10.18653/v1/2023.wmt-1.64

  96. [104]

    The Guardian , author =

    Dutch publisher to use. The Guardian , author =. 2024 , keywords =

  97. [105]

    GPTS core: Evaluate as You Desire

    Fu, Jinlan and Ng, See-Kiong and Jiang, Zhengbao and Liu, Pengfei. GPTS core: Evaluate as You Desire. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. d...

  98. [106]

    Is C hat GPT a Good NLG Evaluator? A Preliminary Study

    Wang, Jiaan and Liang, Yunlong and Meng, Fandong and Sun, Zengkui and Shi, Haoxiang and Li, Zhixu and Xu, Jinan and Qu, Jianfeng and Zhou, Jie. Is C hat GPT a Good NLG Evaluator? A Preliminary Study. Proceedings of the 4th New Frontiers in Summarization Workshop. 2023. doi:10....

  99. [107]

    Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM -as-a-Judge

    Zhang, Qiyuan and Wang, Yufei and Jiang, Yuxin and Li, Liangyou and Wu, Chuhan and Wang, Yasheng and Jiang, Xin and Shang, Lifeng and Tang, Ruiming and Lyu, Fuyuan and Ma, Chen. Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM -as-a-Judge. Proceedings o...

  100. [108]

    Findings of the WMT 25 Shared Task on Automated Translation Evaluation Systems: Linguistic Diversity is Challenging and References Still Help

    Lavie, Alon and Hanneman, Greg and Agrawal, Sweta and Kanojia, Diptesh and Lo, Chi-Kiu and Zouhar, Vil \'e m and Blain, Frederic and Zerva, Chrysoula and Avramidis, Eleftherios and Deoghare, Sourabh and Sindhujan, Archchana and Wang, Jiayi and Adelani, David Ifeoluwa and Thomp...

  101. [109]

    A global analysis of metrics used for measuring performance in natural language processing

    Blagec, Kathrin and Dorffner, Georg and Moradi, Milad and Ott, Simon and Samwald, Matthias. A global analysis of metrics used for measuring performance in natural language processing. Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP. 2022. doi:10.1...

  102. [110]

    Quality Expectations of Machine Translation

    Way, Andy. Quality Expectations of Machine Translation. Translation Quality Assessment: From Principles to Practice. 2018. doi:10.1007/978-3-319-91241-7_8

  103. [111]

    Is post-editing really faster than human translation? , Volume =

    Silvia Terribile , Journal =. Is post-editing really faster than human translation? , Volume =. 2024 , doi =

  104. [112]

    The Role of Creativity , booktitle =

    Rojo, Ana , publisher =. The Role of Creativity , booktitle =. doi:https://doi.org/10.1002/9781119241485.ch19 , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781119241485.ch19 , year =

  105. [113]

    Translation and Creativity in the 21st Century

    Susan Bassnett and Lawrence Venuti and Jan Pedersen and Ivana Hostová , Journal =. Translation and Creativity in the 21st Century. , Volume =. 2022 , doi =

  106. [114]

    2024 , eprint=

    The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control , author=. 2024 , eprint=

  107. [115]

    Christine Hall. 2025

  108. [116]

    Nielbo , title =

    Mia Jacobsen and Yuri Bizzoni and Pascale Feldkamp Moreira and Kristoffer L. Nielbo , title =. Proceedings of the Computational Humanities Research Conference 2024 , pages =

  109. [117]

    Quantitative analysis of fanfictions’ popularity , Volume =

    Zhivar Sourati Hassan Zadeh and Nazanin Sabri and Houmaan Chamani and Behnam Bahrak , Journal =. Quantitative analysis of fanfictions’ popularity , Volume =. 2021 , doi =

  110. [118]

    International Journal of Human–Computer Interaction , volume =

    Roi Alfassi and Angelora Cooper and Zoe Mitchell and Mary Calabro and Orit Shaer and Osnat Mokryn , title =. International Journal of Human–Computer Interaction , volume =. 2026 , publisher =. doi:10.1080/10447318.2025.2531272 , URL =

  111. [119]

    Optimising C hat GPT for creativity in literary translation: A case study from E nglish into D utch, C hinese, C atalan and S panish

    Du, Shuxiang and Arenas, Ana Guerberof and Toral, Antonio and Gerrits, Kyo and Borillo, Josep Marco. Optimising C hat GPT for creativity in literary translation: A case study from E nglish into D utch, C hinese, C atalan and S panish. Proceedings of Machine Translation Summit ...

  112. [120]

    2024 , eprint=

    Large Language Models are Inconsistent and Biased Evaluators , author=. 2024 , eprint=

  113. [121]

    2025 , journal =

    Detecting and Evaluating Bias in Large Language Models: Concepts, Methods, and Challenges , author=. 2025 , journal =

  114. [122]

    and Jeffrey D

    Aho, Alfred V. and Jeffrey D. Ullman. 1972. The Theory of Parsing, Translation and Compiling , volume 1. Prentice- Hall , Englewood Cliffs, NJ

  115. [123]

    American Psychological Association . 1983. Publications Manual . American Psychological Association, Washington, DC

  116. [124]

    Association for Computing Machinery . 1983. Computing Reviews , 24(11):503--512

  117. [125]

    Kozen, and Larry J

    Chandra, Ashok K., Dexter C. Kozen, and Larry J. Stockmeyer. 1981. Alternation. Journal of the Asso\-ciation for Computing Machinery , 28(1):114--133

  118. [126]

    Gledson, Anne, and John Keane. 2008a. Measuring Topic Homogeneity and its Application to Dictionary-Based Word-Sense Disambiguation. Coling 2008, 22nd International Conference on Computational Linguistics , Manchester, UK. 273--280

  119. [127]

    Gledson, Anne, and John Keane. 2008b. Using Web-Search Results to Measure Word-group Similarity. Coling 2008, 22nd International Conference on Computational Linguistics , Manchester, UK. 281--288

  120. [128]

    Gusfield, Dan. 1997. Algorithms on Strings, Trees and Sequences . Cambridge University Press, Cambridge, UK

  121. [129]

    Tam, Yik-Cheung and Tanja Schultz. 2006. Unsupervised Language Model Adaptation Using Latent Semantic Marginals. Interspeech 2006 -- ICSLP, Ninth International Conference on Spoken Language Processing , Pittsburgh, Pennsylvania, paper 1705-Thu1A2O.2

  122. [130]

    Tam, Yik-Cheung and Tanja Schultz. 2007. Correlated Latent Semantic Model for Unsupervised Language Model Adaptation. Proceedings of ICASSP 2007, International Conference on Acoustics, Speech, and Signal Processing , Honolulu, Hawaii, Vol. IV, 41--44

Pith tools

Reviewed May 14, 2026 · model on record in the stance chip above.