Pith. sign in

REVIEW 1 minor 300 references

SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

T0 review · 0 major / 1 minor · reviewed 2026-05-09 · grok-4.3

Pith's one-line read A benchmark evaluates language models on everyday knowledge across more than 30 languages and cultures without permitting training on the test data.

desk verdict This is a standard shared-task overview that scales an existing cultural knowledge benchmark to 30+ low-resource languages and draws decent participation, but it introduces no new methods or independent validation. read the letter →

arxiv 2605.02601 v1 submitted 2026-05-04 cs.CL

classification cs.CL
keywords sharedtaskmultilingualNLPculturalknowledgelow-resourcelanguagesLLMevaluationeverydaybenchmarkquestionanswering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper presents a shared task designed to measure how well natural language processing systems can handle everyday knowledge questions in a wide range of languages and cultures. It relies on an extended set of manually created questions covering more than 30 language-culture pairs, with emphasis on low-resource languages from different continents. The task features two question formats, short-answer and multiple-choice, and enforces strict evaluation-only use of the data. Submissions from 62 teams were analyzed to identify effective strategies and persistent difficulties in model behavior for under-represented groups.

What carries the argument

The extended benchmark of everyday knowledge questions in short-answer and multiple-choice formats applied to more than 30 language-culture pairs.

What would settle it

A finding that the benchmark questions systematically miss or misrepresent knowledge held by speakers of the included languages would undermine the task's validity as a measure of cultural adaptability.

Watch

Extended reading notes

Core claim

The paper establishes that by organizing this evaluation-focused task on an extended benchmark of everyday knowledge, it is possible to gather comparable results from many systems and uncover shared insights about challenges in handling linguistic and cultural diversity, particularly for low-resource settings.

Load-bearing premise

The questions in the benchmark accurately represent typical everyday knowledge in each of the covered cultures without introducing bias.

Editorial extensions

If this is right

  • The no-training rule ensures that results reflect genuine generalization rather than memorization.
  • Analysis of top systems reveals common approaches to multilingual question answering.
  • The task highlights open questions around model misalignment with cultural contexts.
  • Performance on low-resource languages indicates areas needing further development in NLP systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Such benchmarks could inform the creation of more inclusive AI models that respect cultural differences.
  • Extending this approach to additional knowledge domains or languages might expose further limitations in current technology.
  • The observed challenges suggest that evaluation methods themselves may need refinement to better capture cultural nuances.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 1 minor

Summary. The manuscript presents SemEval-2026 Task 7, a shared task for evaluating the adaptability of LLMs and NLP systems to everyday knowledge across diverse languages and cultures. The task data are an extended version of the manually constructed BLEnD benchmark covering more than 30 language-culture pairs (predominantly low-resource languages). It defines two tracks—Short-Answer Questions (SAQ) and Multiple-Choice Questions (MCQ)—with strict rules prohibiting any use of the data for training, fine-tuning, or few-shot adaptation. The paper reports 140 registered participants, 62 final submissions, 19 system description papers, and provides analysis of the best-performing systems, common approaches, and open challenges in evaluation, misalignment, and model behavior for low-resource languages and under-represented cultures.

Significance. If the reported results and analysis hold, the work supplies a large-scale, culturally diverse evaluation framework that can serve as a reference benchmark for assessing cultural knowledge and alignment in LLMs, especially in low-resource settings. The high participation rate and the evaluation-only constraint strengthen the reliability of any comparative findings, while the analysis of adopted modeling strategies offers practical insights for future work on cross-cultural NLP.

minor comments (1)
  1. [Abstract] Abstract: the statement that results and analysis are reported would be strengthened by an explicit forward reference to the relevant section or table containing the quantitative performance metrics and error analysis of the top systems.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive review, accurate summary of the task, and recommendation to accept. The feedback correctly identifies the value of the evaluation-only constraint and the insights from high participation rates.

Circularity Check

0 steps flagged · score 1.0 of 10

Descriptive shared-task paper with no derivations or load-bearing circularity

full rationale

The manuscript is a standard SemEval task description whose central statements are factual descriptions of data provenance and participation statistics. It contains no equations, fitted parameters, predictions, or modeling derivations. The single self-citation to Myung et al. 2024 simply identifies the source benchmark being extended; this reference is not used to justify any internal claim that would otherwise be unsupported, nor does any result reduce to the citation by construction. All other content (track definitions, submission rules, result reporting) is observational and does not rely on unverified self-referential logic.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The paper introduces no free parameters, mathematical axioms, or invented entities; it relies on standard assumptions in NLP evaluation such as the validity of benchmark-based testing and participant compliance with rules.

assumptions (1)
  • domain assumption NLP systems can be meaningfully evaluated on knowledge benchmarks without training or fine-tuning on the test data itself.
    The task explicitly prohibits any form of model modification using the benchmark data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures." pith.science (2026). https://pith.science/paper/2605.02601

@misc{pith2026260502601,
  author       = {Pith},
  title        = {Pith review of: SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2605.02601}},
  note         = {Machine review of arXiv:2605.02601}
}
read the original abstract

We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures.

Figures

Figures reproduced from arXiv: 2605.02601 by the authors.

Figure 1
Figure 1. Data creation pipeline: recruitment of local speakers for annotation, native-speaker quality control without the use of LLMs or search engines, and exten￾sion to 17 additional languages. 2024; Pawar et al., 2024; Liu et al., 2025). How￾ever, such models often exhibit substantial limi￾tations in culture-specific knowledge, particularly when handling under-resourced languages or non￾Western regions. They tend to gener… view at source ↗
Figure 2
Figure 2. Language–culture pairs represented in our BLEnD benchmark. Africa: Arabic (Algeria, Egypt, Morocco), Amharic (Ethiopia), Hausa (Northern Nigeria). Asia: Assamese (Assam, India), Azerbaijani (Azerbaijan), Mandarin (China), Indonesian (Indone￾sia), Javanese (West Java, Indonesia), Persian (Iran), Korean (North and South Korea), Arabic (Saudi Ara￾bia), Japanese (Japan), Tagalog (Philippines), Tamil (Sri Lanka, Singapor… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 300 canonical work pages

  1. [1]

    Wangkongqiang at

    Wang, Kongqiang and Zhang, Peng and Tan, Qingli , booktitle=. Wangkongqiang at

  2. [2]

    Tekanlou, Hadi Bayrami Asl and Bakhtiyarzadeh, Mahdi and Razmara, Jafar , booktitle=

  3. [3]

    Bogdanova, Liliia and Sun, Shiran and Han, Lifeng and Amat-Lefort, Natalia and Plaza-del-Arco, Flor Miriam , booktitle=

  4. [4]

    Almanza, Danileth and Serrano, Jairo and Puertas, Edwin and Martinez Santos, Juan Carlos , booktitle=

  5. [5]

    Ning, Jingke , booktitle=

  6. [6]

    king001 at

    Jin, Meizhi and Meng, Zhichao and Yin, Junqi and Jiang, Lianxin and Li, Jianyu , booktitle=. king001 at

  7. [7]

    chengtang at

    Tang, Cheng and Meng, Zhichao and Jin, Meizhi , booktitle=. chengtang at

  8. [8]

    Yao, Xiao and Yang, Liang , booktitle=

Show all 300 references
  1. [9]

    Adam, Faisal Muhammad and Aliyu, Lukman Jibril and Aji, Sani and Abubakar, Abdulhamid and Shuaibu, Aliyu Rabiu , booktitle=

  2. [10]

    Al Ghussin, Yusser and Gurgurov, Daniil and Hamidullah, Yasser and van Genabith, Josef and España-Bonet, Cristina and Ostermann, Simon , booktitle=

  3. [11]

    , booktitle=

    Adjei, Isaac Nyadu and Aryal, Saurav K. , booktitle=

  4. [12]

    Rahman, Mohammad Marufur and Ailneni, Rakshitha Rao and Harabagiu, Sanda , booktitle=

  5. [13]

    Singh, Aditya and Das, Rickarya , booktitle=

  6. [14]

    uir-cis-7 at

    Gao, Jianning and Mao, Xianling and Shi, Shumin and Zhaxi, Duanzhi and Sun, Yingbo and Li, Xiandeng and Li, Binyang , booktitle=. uir-cis-7 at

  7. [15]

    Iranmanesh, Reihaneh and Frieder, Ophir and Goharian, Nazli , booktitle=

  8. [16]

    Yam, Yen Yee and Yam, Hong Meng , booktitle=

  9. [17]

    Song, Jiwoo and Yeom, Sihyeong and Kim, Harksoo , booktitle=

  10. [18]

    Sriram, Swetha Krishna and Sekar, Nirupama , booktitle=

  11. [19]

    2024 , journal=

    Aya 23: Open Weight Releases to Further Multilingual Progress , author=. 2024 , journal=

  12. [20]

    Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce

    Ousidhoum, Nedjma and Beloucif, Meriem and Mohammad, Saif M. Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.1...

  13. [21]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Culture is not trivia: Sociocultural theory for cultural nlp , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  14. [22]

    Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens

    Alkhamissi, Mai and Xiao, Yunze and AlKhamissi, Badr and Diab, Mona T. Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens. Findings of the A ssociation for C omputational L inguistics: EACL 2026. 2026

  15. [23]

    arXiv preprint arXiv:2305.14456 , year=

    Having Beer after Prayer? Measuring Cultural Bias in Large Language Models , author=. arXiv preprint arXiv:2305.14456 , year=

  16. [24]

    arXiv preprint arXiv:2205.01068 , year=

    Opt: Open pre-trained transformer language models , author=. arXiv preprint arXiv:2205.01068 , year=

  17. [25]

    arXiv preprint arXiv:2211.05100 , year=

    Bloom: A 176b-parameter open-access multilingual language model , author=. arXiv preprint arXiv:2211.05100 , year=

  18. [26]

    arXiv preprint arXiv:2307.09288 , year=

    Llama 2: Open foundation and fine-tuned chat models , author=. arXiv preprint arXiv:2307.09288 , year=

  19. [27]

    Language and Culture , volume =

    Kramsch, Claire , year =. Language and Culture , volume =. doi:10.1075/aila.27.02kra , journal =

  20. [28]

    arXiv preprint arXiv:2402.15018 , year=

    Unintended Impacts of LLM Alignment on Global Representation , author=. arXiv preprint arXiv:2402.15018 , year=

  21. [29]

    arXiv preprint arXiv:2306.16388 , year=

    Towards measuring the representation of subjective global opinions in language models , author=. arXiv preprint arXiv:2306.16388 , year=

  22. [30]

    arXiv preprint arXiv:2402.09369 , year=

    Massively multi-cultural knowledge acquisition & lm benchmarking , author=. arXiv preprint arXiv:2402.09369 , year=

  23. [31]

    Proceedings of the ACM Web Conference 2023 , pages =

    Nguyen, Tuan-Phong and Razniewski, Simon and Varde, Aparna and Weikum, Gerhard , title =. Proceedings of the ACM Web Conference 2023 , pages =. 2023 , isbn =. doi:10.1145/3543507.3583535 , abstract =

  24. [32]

    arXiv preprint arXiv:2404.15238 , year=

    CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies , author=. arXiv preprint arXiv:2404.15238 , year=

  25. [33]

    arXiv preprint arXiv:2309.16609 , year=

    Qwen Technical Report , author=. arXiv preprint arXiv:2309.16609 , year=

  26. [34]

    arXiv preprint arXiv:2402.07827 , year=

    Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model , author=. arXiv preprint arXiv:2402.07827 , year=

  27. [35]

    arXiv preprint arXiv:2404.01954 , year=

    HyperCLOVA X Technical Report , author=. arXiv preprint arXiv:2404.01954 , year=

  28. [36]

    Can Common Sense uncover cultural differences in computer applications?

    Anacleto, Junia and Lieberman, Henry and Tsutsumi, Marie and Neris, V \^a nia and Carvalho, Aparecido and Espinosa, Jose and Godoi, Muriel and Zem-Mascarenhas, Silvia. Can Common Sense uncover cultural differences in computer applications?. Artificial Intelligence in Theory an...

  29. [37]

    arXiv preprint arXiv:2401.15585 , year=

    Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting , author=. arXiv preprint arXiv:2401.15585 , year=

  30. [38]

    ACM Journal of Data and Information Quality , volume=

    Biases in large language models: origins, inventory, and discussion , author=. ACM Journal of Data and Information Quality , volume=. 2023 , publisher=

  31. [39]

    arXiv preprint arXiv:2404.01854 , year=

    IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces , author=. arXiv preprint arXiv:2404.01854 , year=

  32. [40]

    arXiv preprint arXiv:2312.00738 , url=

    Xuan-Phi Nguyen and Wenxuan Zhang and Xin Li and Mahani Aljunied and Qingyu Tan and Liying Cheng and Guanzheng Chen and Yue Deng and Sen Yang and Chaoqun Liu and Hang Zhang and Lidong Bing , title =. arXiv preprint arXiv:2312.00738 , url=

  33. [41]

    2024 , eprint=

    Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis , author=. 2024 , eprint=

  34. [42]

    HAE - RAE Bench: Evaluation of K orean Knowledge in Language Models

    Son, Guijin and Lee, Hanwool and Kim, Suwan and Kim, Huiseo and Lee, Jae cheol and Yeom, Je Won and Jung, Jihyu and Kim, Jung woo and Kim, Songseong. HAE - RAE Bench: Evaluation of K orean Knowledge in Language Models. Proceedings of the 2024 Joint International Conference on ...

  35. [43]

    CLI c K : A Benchmark Dataset of Cultural and Linguistic Intelligence in K orean

    Kim, Eunsu and Suk, Juyoung and Oh, Philhoon and Yoo, Haneul and Thorne, James and Oh, Alice. CLI c K : A Benchmark Dataset of Cultural and Linguistic Intelligence in K orean. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resourc...

  36. [44]

    2024 , eprint=

    COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances , author=. 2024 , eprint=

  37. [45]

    2024 , eprint=

    Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese , author=. 2024 , eprint=

  38. [46]

    Communications of the ACM , volume=

    Building a multilingual Wikipedia , author=. Communications of the ACM , volume=. 2021 , publisher=

  39. [47]

    , author=

    qalsadi, Arabic mophological analyzer Library for python. , author=

  40. [48]

    and Patrick Burns and John Stewart and Todd Cook , title =

    Johnson, Kyle P. and Patrick Burns and John Stewart and Todd Cook , title =

  41. [49]

    ACM Trans

    Setiawan, Irwan and Kao, Hung-Yu , title =. ACM Trans. Asian Low-Resour. Lang. Inf. Process. , month =. 2024 , publisher =. doi:10.1145/3656342 , abstract =

  42. [50]

    The IndicNLP Library

    Anoop Kunchukuttan. The IndicNLP Library. 2020

  43. [51]

    Stemming Hausa text: using affix-stripping rules and reference look-up , volume =

    Bimba, Andrew and Idris, Norisma and Khamis, Norazlina and Noor, Nurul , year =. Stemming Hausa text: using affix-stripping rules and reference look-up , volume =. Language Resources and Evaluation , doi =

  44. [52]

    Proceedings of the international AAAI conference on web and social media , volume=

    Big questions for social media big data: Representativeness, validity and other methodological pitfalls , author=. Proceedings of the international AAAI conference on web and social media , volume=

  45. [53]

    2024 , eprint=

    Gemini: A Family of Highly Capable Multimodal Models , author=. 2024 , eprint=

  46. [54]

    2024 , eprint=

    GPT-4 Technical Report , author=. 2024 , eprint=

  47. [55]

    2023 , eprint=

    PaLM 2 Technical Report , author=. 2023 , eprint=

  48. [56]

    BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages , url =

    Myung, Junho and Lee, Nayeon and Zhou, Yi and Jin, Jiho and Putri, Rifki Afina and Antypas, Dimosthenis and Borkakoty, Hsuvas and Kim, Eunsu and Perez-Almendros, Carla and Ayele, Abinew Ali and Guti\'. BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and L...

  49. [57]

    2024 , eprint=

    Survey of Cultural Awareness in Language Models: Text and Beyond , author=. 2024 , eprint=

  50. [58]

    2025 , eprint=

    Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art , author=. 2025 , eprint=

  51. [59]

    Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

    CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean , author=. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

  52. [60]

    Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

    HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models , author=. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

  53. [61]

    Proceedings of the 20th Workshop of Young Researchers' Roundtable on Spoken Dialogue Systems. 2024

  54. [62]

    Conversational XAI and Explanation Dialogues

    Feldhus, Nils. Conversational XAI and Explanation Dialogues. 2024

  55. [63]

    Enhancing Emotion Recognition in Spoken Dialogue Systems through Multimodal Integration and Personalization

    Kaneko, Takumasa. Enhancing Emotion Recognition in Spoken Dialogue Systems through Multimodal Integration and Personalization. 2024

  56. [64]

    Towards Personalisation of User Support Systems

    Higuchi, Tomoya. Towards Personalisation of User Support Systems. 2024

  57. [65]

    Social Agents for Positively Influencing Human Psychological States

    Baihaqi, Muhammad Yeza. Social Agents for Positively Influencing Human Psychological States. 2024

  58. [66]

    Personalized Topic Transition for Dialogue System

    Yoshida, Kai. Personalized Topic Transition for Dialogue System. 2024

  59. [67]

    Elucidation of Psychotherapy and Development of New Treatment Methods Using AI

    Maeda, Shio. Elucidation of Psychotherapy and Development of New Treatment Methods Using AI. 2024

  60. [68]

    Assessing Interactional Competence with Multimodal Dialog Systems

    Saeki, Mao. Assessing Interactional Competence with Multimodal Dialog Systems. 2024

  61. [69]

    Faithfulness of Natural Language Generation

    Schmidtova, Patricia. Faithfulness of Natural Language Generation. 2024

  62. [70]

    Knowledge-Grounded Dialogue Systems for Generating Interesting and Engaging Responses

    Onozeki, Hiroki. Knowledge-Grounded Dialogue Systems for Generating Interesting and Engaging Responses. 2024

  63. [71]

    Towards a Dialogue System That Can Take Interlocutors' Values into Account

    Zenimoto, Yuki. Towards a Dialogue System That Can Take Interlocutors' Values into Account. 2024

  64. [72]

    Multimodal Spoken Dialogue System with Biosignals

    Katada, Shun. Multimodal Spoken Dialogue System with Biosignals. 2024

  65. [73]

    Timing Sensitive Turn-Taking in Spoken Dialogue Systems Based on User Satisfaction

    Yoshikawa, Sadahiro. Timing Sensitive Turn-Taking in Spoken Dialogue Systems Based on User Satisfaction. 2024

  66. [74]

    Towards Robust and Multilingual Task-Oriented Dialogue Systems

    Ohashi, Atsumoto. Towards Robust and Multilingual Task-Oriented Dialogue Systems. 2024

  67. [75]

    Toward Faithful Dialogs: Evaluating and Improving the Faithfulness of Dialog Systems

    Huang, Sicong. Toward Faithful Dialogs: Evaluating and Improving the Faithfulness of Dialog Systems. 2024

  68. [76]

    Cognitive Model of Listener Response Generation and Its Application to Dialogue Systems

    Mori, Taiga. Cognitive Model of Listener Response Generation and Its Application to Dialogue Systems. 2024

  69. [77]

    Topological Deep Learning for Term Extraction

    Ruppik, Benjamin Matthias. Topological Deep Learning for Term Extraction. 2024

  70. [78]

    Dialogue Management with Graph-structured Knowledge

    Walker, Nicholas Thomas. Dialogue Management with Graph-structured Knowledge. 2024

  71. [79]

    Towards a Co-creation Dialogue System

    Zhou, Xulin. Towards a Co-creation Dialogue System. 2024

  72. [80]

    Enhancing Decision-Making with AI Assistance

    Tanaka, Yoshiki. Enhancing Decision-Making with AI Assistance. 2024

  73. [81]

    Ontology Construction for Task-oriented Dialogue

    Vukovic, Renato. Ontology Construction for Task-oriented Dialogue. 2024

  74. [82]

    Generalized Visual-Language Grounding with Complex Language Context

    Hemanthage, Bhathiya. Generalized Visual-Language Grounding with Complex Language Context. 2024

  75. [83]

    Towards a Real-Time Multimodal Emotion Estimation Model for Dialogue Systems

    Jiang, Jingjing. Towards a Real-Time Multimodal Emotion Estimation Model for Dialogue Systems. 2024

  76. [84]

    Exploring Explainability and Interpretability in Generative AI

    Huang, Shiyuan. Exploring Explainability and Interpretability in Generative AI. 2024

  77. [85]

    Innovative Approaches to Enhancing Safety and Ethical AI Interactions in Digital Environments

    Yang, Zachary. Innovative Approaches to Enhancing Safety and Ethical AI Interactions in Digital Environments. 2024

  78. [86]

    Leveraging Linguistic Structural Information for Improving the Model`s Semantic Understanding Ability

    Lee, Sangmyeong. Leveraging Linguistic Structural Information for Improving the Model`s Semantic Understanding Ability. 2024

  79. [87]

    Multi-User Dialogue Systems and Controllable Language Generation

    Wagner, Nicolas. Multi-User Dialogue Systems and Controllable Language Generation. 2024

  80. [88]

    Enhancing Role-Playing Capabilities in Persona Dialogue Systems through Corpus Construction and Evaluation Methods

    Uehara, Ryuichi. Enhancing Role-Playing Capabilities in Persona Dialogue Systems through Corpus Construction and Evaluation Methods. 2024

  81. [89]

    Character Expression and User Adaptation for Spoken Dialogue Systems

    Yamamoto, Kenta. Character Expression and User Adaptation for Spoken Dialogue Systems. 2024

  82. [90]

    Interactive Explanations through Dialogue Systems

    Feustel, Isabel. Interactive Explanations through Dialogue Systems. 2024

  83. [91]

    Towards Emotion-aware Task-oriented Dialogue Systems in the Era of Large Language Models

    Feng, Shutong. Towards Emotion-aware Task-oriented Dialogue Systems in the Era of Large Language Models. 2024

  84. [92]

    Utilizing Large Language Models for Customized Dialogue Data Augmentation and Psychological Counseling

    Qi, Zhiyang. Utilizing Large Language Models for Customized Dialogue Data Augmentation and Psychological Counseling. 2024

  85. [93]

    Toward More Human-like SDS s: Advancing Emotional and Social Engagement in Embodied Conversational Agents

    Pang, Zi Haur. Toward More Human-like SDS s: Advancing Emotional and Social Engagement in Embodied Conversational Agents. 2024

  86. [94]

    Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024). 2024

  87. [95]

    Investigating radicalisation indicators in online extremist communities

    De Kock, Christine and Hovy, Eduard. Investigating radicalisation indicators in online extremist communities. 2024. doi:10.18653/v1/2024.woah-1.1

  88. [96]

    Detection of Conspiracy Theories Beyond Keyword Bias in G erman-Language Telegram Using Large Language Models

    Pustet, Milena and Steffen, Elisabeth and Mihaljevic, Helena. Detection of Conspiracy Theories Beyond Keyword Bias in G erman-Language Telegram Using Large Language Models. 2024. doi:10.18653/v1/2024.woah-1.2

  89. [97]

    E ko H ate: Abusive Language and Hate Speech Detection for Code-switched Political Discussions on N igerian T witter

    Ilevbare, Comfort and Alabi, Jesujoba and Adelani, David Ifeoluwa and Bakare, Firdous and Abiola, Oluwatoyin and Adeyemo, Oluwaseyi. E ko H ate: Abusive Language and Hate Speech Detection for Code-switched Political Discussions on N igerian T witter. 2024. doi:10.18653/v1/2024...

  90. [98]

    A Study of the Class Imbalance Problem in Abusive Language Detection

    Zhang, Yaqi and Hangya, Viktor and Fraser, Alexander. A Study of the Class Imbalance Problem in Abusive Language Detection. 2024. doi:10.18653/v1/2024.woah-1.4

  91. [99]

    H ausa H ate: An Expert Annotated Corpus for H ausa Hate Speech Detection

    Vargas, Francielle and Guimar \ a es, Samuel and Muhammad, Shamsuddeen Hassan and Alves, Diego and Ahmad, Ibrahim Said and Abdulmumin, Idris and Mohamed, Diallo and Pardo, Thiago and Benevenuto, Fabr \'i cio. H ausa H ate: An Expert Annotated Corpus for H ausa Hate Speech Dete...

  92. [100]

    VIDA : The Visual Incel Data Archive

    Anastasi, Selenia and Schneider, Florian and Biemann, Chris and Fischer, Tim. VIDA : The Visual Incel Data Archive. A Theory-oriented Annotated Dataset To Enhance Hate Detection Through Visual Culture. 2024. doi:10.18653/v1/2024.woah-1.6

  93. [101]

    Towards a Unified Framework for Adaptable Problematic Content Detection via Continual Learning

    Omrani, Ali and Salkhordeh Ziabari, Alireza and Golazizian, Preni and Sorensen, Jeffrey and Dehghani, Morteza. Towards a Unified Framework for Adaptable Problematic Content Detection via Continual Learning. 2024. doi:10.18653/v1/2024.woah-1.7

  94. [102]

    From Linguistics to Practice: a Case Study of Offensive Language Taxonomy in H ebrew

    Liebeskind, Chaya and Litvak, Marina and Vanetik, Natalia. From Linguistics to Practice: a Case Study of Offensive Language Taxonomy in H ebrew. 2024. doi:10.18653/v1/2024.woah-1.8

  95. [103]

    Estimating the Emotion of Disgust in G reek Parliament Records

    Lislevand, Vanessa and Pavlopoulos, John and Louridas, Panos and Dritsa, Konstantina. Estimating the Emotion of Disgust in G reek Parliament Records. 2024. doi:10.18653/v1/2024.woah-1.9

  96. [104]

    Simple LLM based Approach to Counter Algospeak

    Fillies, Jan and Paschke, Adrian. Simple LLM based Approach to Counter Algospeak. 2024. doi:10.18653/v1/2024.woah-1.10

  97. [105]

    Harnessing Personalization Methods to Identify and Predict Unreliable Information Spreader Behavior

    Ashraf, Shaina and Gruschka, Fabio and Flek, Lucie and Welch, Charles. Harnessing Personalization Methods to Identify and Predict Unreliable Information Spreader Behavior. 2024. doi:10.18653/v1/2024.woah-1.11

  98. [106]

    Robust Safety Classifier Against Jailbreaking Attacks: Adversarial Prompt Shield

    Kim, Jinhwa and Derakhshan, Ali and Harris, Ian. Robust Safety Classifier Against Jailbreaking Attacks: Adversarial Prompt Shield. 2024. doi:10.18653/v1/2024.woah-1.12

  99. [107]

    Improving aggressiveness detection using a data augmentation technique based on a Diffusion Language Model

    Reyes-Ram \'i rez, Antonio and Arag \'o n, Mario and S \'a nchez-Vega, Fernando and L \'o pez-Monroy, Adrian. Improving aggressiveness detection using a data augmentation technique based on a Diffusion Language Model. 2024. doi:10.18653/v1/2024.woah-1.13

  100. [108]

    The M exican Gayze: A Computational Analysis of the Attitudes towards the LGBT + Population in M exico on Social Media Across a Decade

    Andersen, Scott and Ojeda-Trueba, Segio-Luis and V \'a squez, Juan and Bel-Enguix, Gemma. The M exican Gayze: A Computational Analysis of the Attitudes towards the LGBT + Population in M exico on Social Media Across a Decade. 2024. doi:10.18653/v1/2024.woah-1.14

  101. [109]

    X -posing Free Speech: Examining the Impact of Moderation Relaxation on Online Social Networks

    Arun, Arvindh and Chhatani, Saurav and An, Jisun and Kumaraguru, Ponnurangam. X -posing Free Speech: Examining the Impact of Moderation Relaxation on Online Social Networks. 2024. doi:10.18653/v1/2024.woah-1.15

  102. [110]

    The Uli Dataset: An Exercise in Experience Led Annotation of o GBV

    Arora, Arnav and Jinadoss, Maha and Arora, Cheshta and George, Denny and Brindaalakshmi and Khan, Haseena and Rawat, Kirti and Div and Ritash and Mathur, Seema. The Uli Dataset: An Exercise in Experience Led Annotation of o GBV. 2024. doi:10.18653/v1/2024.woah-1.16

  103. [111]

    Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales

    Nirmal, Ayushi and Bhattacharjee, Amrita and Sheth, Paras and Liu, Huan. Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales. 2024. doi:10.18653/v1/2024.woah-1.17

  104. [112]

    A B ayesian Quantification of Aporophobia and the Aggravating Effect of Low -- Wealth Contexts on Stigmatization

    Brate, Ryan and Van Erp, Marieke and Van Den Bosch, Antal. A B ayesian Quantification of Aporophobia and the Aggravating Effect of Low -- Wealth Contexts on Stigmatization. 2024. doi:10.18653/v1/2024.woah-1.18

  105. [113]

    Toxicity Classification in U krainian

    Dementieva, Daryna and Khylenko, Valeriia and Babakov, Nikolay and Groh, Georg. Toxicity Classification in U krainian. 2024. doi:10.18653/v1/2024.woah-1.19

  106. [114]

    A Strategy Labelled Dataset of Counterspeech

    Poudhar, Aashima and Konstas, Ioannis and Abercrombie, Gavin. A Strategy Labelled Dataset of Counterspeech. 2024. doi:10.18653/v1/2024.woah-1.20

  107. [115]

    Improving Covert Toxicity Detection by Retrieving and Generating References

    Lee, Dong-Ho and Cho, Hyundong and Jin, Woojeong and Moon, Jihyung and Park, Sungjoon and R. Improving Covert Toxicity Detection by Retrieving and Generating References. 2024. doi:10.18653/v1/2024.woah-1.21

  108. [116]

    Subjective Isms? On the Danger of Conflating Hate and Offence in Abusive Language Detection

    Cercas Curry, Amanda and Abercrombie, Gavin and Talat, Zeerak. Subjective Isms? On the Danger of Conflating Hate and Offence in Abusive Language Detection. 2024. doi:10.18653/v1/2024.woah-1.22

  109. [117]

    From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets

    Tonneau, Manuel and Liu, Diyi and Fraiberger, Samuel and Schroeder, Ralph and Hale, Scott and R. From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets. 2024. doi:10.18653/v1/2024.woah-1.23

  110. [118]

    SGH ate C heck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of S ingapore

    Ng, Ri Chi and Prakash, Nirmalendu and Hee, Ming Shan and Choo, Kenny Tsu Wei and Lee, Roy Ka-wei. SGH ate C heck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of S ingapore. 2024. doi:10.18653/v1/2024.woah-1.24

  111. [119]

    Proceedings of the Ninth Workshop on Noisy and User-generated Text (W-NUT 2024). 2024

  112. [120]

    Correcting Challenging F innish Learner Texts With Claude, GPT -3.5 and GPT -4 Large Language Models

    Creutz, Mathias. Correcting Challenging F innish Learner Texts With Claude, GPT -3.5 and GPT -4 Large Language Models. 2024

  113. [121]

    Context-aware Adversarial Attack on Named Entity Recognition

    Chen, Shuguang and Neves, Leonardo and Solorio, Thamar. Context-aware Adversarial Attack on Named Entity Recognition. 2024

  114. [122]

    Effects of different types of noise in user-generated reviews on human and machine translations including C hat GPT

    Popovic, Maja and Lapshinova-Koltunski, Ekaterina and Koponen, Maarit. Effects of different types of noise in user-generated reviews on human and machine translations including C hat GPT. 2024

  115. [123]

    Stanceosaurus 2.0 - Classifying Stance Towards R ussian and S panish Misinformation

    Lavrouk, Anton and Ligon, Ian and Zheng, Jonathan and Naous, Tarek and Xu, Wei and Ritter, Alan. Stanceosaurus 2.0 - Classifying Stance Towards R ussian and S panish Misinformation. 2024

  116. [124]

    and Shahariar, G

    Elahi, Kazi and Rahman, Tasnuva and Shahriar, Shakil and Sarker, Samir and Shawon, Md. and Shahariar, G. M. A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy B angla Texts. 2024

  117. [125]

    Label Supervised Contrastive Learning for Imbalanced Text Classification in E uclidean and Hyperbolic Embedding Spaces

    Khalid, Baber and Dai, Shuyang and Taghavi, Tara and Lee, Sungjin. Label Supervised Contrastive Learning for Imbalanced Text Classification in E uclidean and Hyperbolic Embedding Spaces. 2024

  118. [126]

    M aint N orm: A corpus and benchmark model for lexical normalisation and masking of industrial maintenance short text

    Bikaun, Tyler and Hodkiewicz, Melinda and Liu, Wei. M aint N orm: A corpus and benchmark model for lexical normalisation and masking of industrial maintenance short text. 2024

  119. [127]

    The Effects of Data Quality on Named Entity Recognition

    Bhadauria, Divya and Sierra M \'u nera, Alejandro and Krestel, Ralf. The Effects of Data Quality on Named Entity Recognition. 2024

  120. [128]

    Topic Bias in Emotion Classification

    Wegge, Maximilian and Klinger, Roman. Topic Bias in Emotion Classification. 2024

  121. [129]

    Stars Are All You Need: A Distantly Supervised Pyramid Network for Unified Sentiment Analysis

    Li, Wenchang and Chen, Yixing and Zheng, Shuang and Wang, Lei and Lalor, John. Stars Are All You Need: A Distantly Supervised Pyramid Network for Unified Sentiment Analysis. 2024

  122. [130]

    Proceedings of the The 6th Workshop on Narrative Understanding. 2024. doi:10.18653/v1/2024.wnu-1.0

  123. [131]

    Narration as Functions: from Events to Narratives

    Huang, Junbo and Usbeck, Ricardo. Narration as Functions: from Events to Narratives. 2024. doi:10.18653/v1/2024.wnu-1.1

  124. [132]

    How to tame your plotline: A framework for goal-driven interactive fairy tale generation

    Ermolaeva, Marina and Shakhmatova, Anastasia and Nepomnyashchikh, Alina and Fenogenova, Alena. How to tame your plotline: A framework for goal-driven interactive fairy tale generation. 2024. doi:10.18653/v1/2024.wnu-1.2

  125. [133]

    Understanding Transmedia Storytelling: Reception and Narrative Comprehension in Bill Willingham`s Fables Franchise

    Lagrange, Victoria. Understanding Transmedia Storytelling: Reception and Narrative Comprehension in Bill Willingham`s Fables Franchise. 2024. doi:10.18653/v1/2024.wnu-1.3

  126. [134]

    Using Large Language Models for Understanding Narrative Discourse

    Piper, Andrew and Bagga, Sunyam. Using Large Language Models for Understanding Narrative Discourse. 2024. doi:10.18653/v1/2024.wnu-1.4

  127. [135]

    Is It Safe to Tell Your Story? Towards Achieving Privacy for Sensitive Narratives

    Shokri, Mohammad and Bishop, Allison and Levitan, Sarah Ita. Is It Safe to Tell Your Story? Towards Achieving Privacy for Sensitive Narratives. 2024. doi:10.18653/v1/2024.wnu-1.7

  128. [136]

    Annotating Mystery Novels: Guidelines and Adaptations

    Heyns, Nuette and Van Zaanen, Menno. Annotating Mystery Novels: Guidelines and Adaptations. 2024. doi:10.18653/v1/2024.wnu-1.9

  129. [137]

    Causal Micro-Narratives

    Heddaya, Mourad and Zeng, Qingcheng and Zentefis, Alexander and Voigt, Rob and Tan, Chenhao. Causal Micro-Narratives. 2024. doi:10.18653/v1/2024.wnu-1.12

  130. [138]

    Media Framing through the Lens of Event-Centric Narratives

    Das, Rohan and Chandra, Aditya and Lee, I-Ta and Pacheco, Maria Leonor. Media Framing through the Lens of Event-Centric Narratives. 2024. doi:10.18653/v1/2024.wnu-1.15

  131. [139]

    BERT -based Annotation of Oral Texts Elicited via Multilingual Assessment Instrument for Narratives

    Baumann, Timo and Eller, Korbinian and Gagarina, Natalia. BERT -based Annotation of Oral Texts Elicited via Multilingual Assessment Instrument for Narratives. 2024. doi:10.18653/v1/2024.wnu-1.16

  132. [140]

    Proceedings of the Ninth Conference on Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.0

  133. [141]

    Findings of the WMT 24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet

    Kocmi, Tom and Avramidis, Eleftherios and Bawden, Rachel and Bojar, Ond r ej and Dvorkovich, Anton and Federmann, Christian and Fishel, Mark and Freitag, Markus and Gowda, Thamme and Grundkiewicz, Roman and Haddow, Barry and Karpinska, Marzena and Koehn, Philipp and Marie, Ben...

  134. [142]

    Are LLM s Breaking MT Metrics? Results of the WMT 24 Metrics Shared Task

    Freitag, Markus and Mathur, Nitika and Deutsch, Daniel and Lo, Chi-Kiu and Avramidis, Eleftherios and Rei, Ricardo and Thompson, Brian and Blain, Frederic and Kocmi, Tom and Wang, Jiayi and Adelani, David Ifeoluwa and Buchicchio, Marianna and Zerva, Chrysoula and Lavie, Alon. ...

  135. [143]

    De Souza, Jos \'e G

    Zerva, Chrysoula and Blain, Frederic and C. De Souza, Jos \'e G. and Kanojia, Diptesh and Deoghare, Sourabh and Guerreiro, Nuno M. and Attanasio, Giuseppe and Rei, Ricardo and Orasan, Constantin and Negri, Matteo and Turchi, Marco and Chatterjee, Rajen and Bhattacharyya, Pushp...

  136. [144]

    Findings of the WMT 2024 Shared Task of the Open Language Data Initiative

    Maillard, Jean and Burchell, Laurie and Anastasopoulos, Antonios and Federmann, Christian and Koehn, Philipp and Wang, Skyler. Findings of the WMT 2024 Shared Task of the Open Language Data Initiative. 2024. doi:10.18653/v1/2024.wmt-1.4

  137. [145]

    Results of the WAT / WMT 2024 Shared Task on Patent Translation

    Higashiyama, Shohei. Results of the WAT / WMT 2024 Shared Task on Patent Translation. 2024. doi:10.18653/v1/2024.wmt-1.5

  138. [146]

    Findings of the WMT 2024 Biomedical Translation Shared Task: Test Sets on Abstract Level

    Neves, Mariana and Grozea, Cristian and Thomas, Philippe and Roller, Roland and Bawden, Rachel and N \'e v \'e ol, Aur \'e lie and Castle, Steffen and Bonato, Vanessa and Di Nunzio, Giorgio Maria and Vezzani, Federica and Vicente Navarro, Maika and Yeganova, Lana and Jimeno Ye...

  139. [147]

    MSLC 24 Submissions to the General Machine Translation Task

    Larkin, Samuel and Lo, Chi-Kiu and Knowles, Rebecca. MSLC 24 Submissions to the General Machine Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.7

  140. [148]

    IOL Research Machine Translation Systems for WMT 24 General Machine Translation Shared Task

    Zhang, Wenbo. IOL Research Machine Translation Systems for WMT 24 General Machine Translation Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.8

  141. [149]

    Choose the Final Translation from NMT and LLM Hypotheses Using MBR Decoding: HW - TSC `s Submission to the WMT 24 General MT Shared Task

    Wu, Zhanglin and Wei, Daimeng and Li, Zongyao and Shang, Hengchao and Guo, Jiaxin and Li, Shaojun and Rao, Zhiqiang and Luo, Yuanchang and Xie, Ning and Yang, Hao. Choose the Final Translation from NMT and LLM Hypotheses Using MBR Decoding: HW - TSC `s Submission to the WMT 24...

  142. [150]

    C ycle GN : A Cycle Consistent Approach for Neural Machine Translation

    Dreano, S. C ycle GN : A Cycle Consistent Approach for Neural Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.10

  143. [151]

    U v A - MT `s Participation in the WMT 24 General Translation Shared Task

    Tan, Shaomu and Stap, David and Aycock, Seth and Monz, Christof and Wu, Di. U v A - MT `s Participation in the WMT 24 General Translation Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.11

  144. [152]

    Rei, Ricardo and Pombal, Jose and Guerreiro, Nuno M. and Alves, Jo \ a o and Martins, Pedro Henrique and Fernandes, Patrick and Wu, Helena and Vaz, Tania and Alves, Duarte and Farajian, Amin and Agrawal, Sweta and Farinhas, Antonio and C. De Souza, Jos \'e G. and Martins, Andr...

  145. [153]

    TSU HITS `s Submissions to the WMT 2024 General Machine Translation Shared Task

    Mynka, Vladimir and Mikhaylovskiy, Nikolay. TSU HITS `s Submissions to the WMT 2024 General Machine Translation Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.13

  146. [154]

    Document-level Translation with LLM Reranking: Team- J at WMT 2024 General Translation Task

    Kudo, Keito and Deguchi, Hiroyuki and Morishita, Makoto and Fujii, Ryo and Ito, Takumi and Ozaki, Shintaro and Natsumi, Koki and Sato, Kai and Yano, Kazuki and Takahashi, Ryosuke and Kimura, Subaru and Hara, Tomomasa and Sakai, Yusuke and Suzuki, Jun. Document-level Translatio...

  147. [155]

    DLUT and GTCOM `s Neural Machine Translation Systems for WMT 24

    Zong, Hao and Bei, Chao and Liu, Huan and Yuan, Conghu and Chen, Wentao and Huang, Degen. DLUT and GTCOM `s Neural Machine Translation Systems for WMT 24. 2024. doi:10.18653/v1/2024.wmt-1.15

  148. [156]

    CUNI at WMT 24 General Translation Task: LLM s, ( Q ) L o RA , CPO and Model Merging

    Hrabal, Miroslav and Jon, Josef and Popel, Martin and Luu, Nam and Semin, Danil and Bojar, Ond r ej. CUNI at WMT 24 General Translation Task: LLM s, ( Q ) L o RA , CPO and Model Merging. 2024. doi:10.18653/v1/2024.wmt-1.16

  149. [157]

    From General LLM to Translation: How We Dramatically Improve Translation Quality Using Human Evaluation Data for LLM Finetuning

    Elshin, Denis and Karpachev, Nikolay and Gruzdev, Boris and Golovanov, Ilya and Ivanov, Georgy and Antonov, Alexander and Skachkov, Nickolay and Latypova, Ekaterina and Layner, Vladimir and Enikeeva, Ekaterina and Popov, Dmitry and Chekashev, Anton and Negodin, Vladislav and F...

  150. [158]

    Cogs in a Machine, Doing What They`re Meant to Do -- the AMI Submission to the WMT 24 General Translation Task

    Jasonarson, Atli and Hafsteinsson, Hinrik and \'A rmannsson, Bjarki and Steingr \'i msson, Steinth \'o r. Cogs in a Machine, Doing What They`re Meant to Do -- the AMI Submission to the WMT 24 General Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.18

  151. [159]

    IKUN for WMT 24 General MT Task: LLM s Are Here for Multilingual Machine Translation

    Liao, Baohao and Herold, Christian and Khadivi, Shahram and Monz, Christof. IKUN for WMT 24 General MT Task: LLM s Are Here for Multilingual Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.19

  152. [160]

    NTTSU at WMT 2024 General Translation Task

    Kondo, Minato and Fukuda, Ryo and Wang, Xiaotian and Chousa, Katsuki and Nishimura, Masato and Buma, Kosei and Kano, Takatomo and Utsuro, Takehito. NTTSU at WMT 2024 General Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.20

  153. [161]

    SCIR - MT `s Submission for WMT 24 General Machine Translation Task

    Li, Baohang and Ye, Zekai and Huang, Yichong and Feng, Xiaocheng and Qin, Bing. SCIR - MT `s Submission for WMT 24 General Machine Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.21

  154. [162]

    AIST AIRC Systems for the WMT 2024 Shared Tasks

    Rikters, Matiss and Miwa, Makoto. AIST AIRC Systems for the WMT 2024 Shared Tasks. 2024. doi:10.18653/v1/2024.wmt-1.22

  155. [163]

    Occiglot at WMT 24: E uropean Open-source Large Language Models Evaluated on Translation

    Avramidis, Eleftherios and Gr. Occiglot at WMT 24: E uropean Open-source Large Language Models Evaluated on Translation. 2024. doi:10.18653/v1/2024.wmt-1.23

  156. [164]

    C o ST of breaking the LLM s

    Mukherjee, Ananya and Yadav, Saumitra and Shrivastava, Manish. C o ST of breaking the LLM s. 2024. doi:10.18653/v1/2024.wmt-1.24

  157. [165]

    WMT 24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles

    Dawkins, Hillary and Nejadgholi, Isar and Lo, Chi-Kiu. WMT 24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles. 2024. doi:10.18653/v1/2024.wmt-1.25

  158. [166]

    The G ender Q ueer Test Suite

    Friidhriksd \'o ttir, Steinunn Rut. The G ender Q ueer Test Suite. 2024. doi:10.18653/v1/2024.wmt-1.26

  159. [167]

    Domain Dynamics: Evaluating Large Language Models in E nglish- H indi Translation

    Bhattacharjee, Soham and Gain, Baban and Ekbal, Asif. Domain Dynamics: Evaluating Large Language Models in E nglish- H indi Translation. 2024. doi:10.18653/v1/2024.wmt-1.27

  160. [168]

    Investigating the Linguistic Performance of Large Language Models in Machine Translation

    Manakhimova, Shushen and Macketanz, Vivien and Avramidis, Eleftherios and Lapshinova-Koltunski, Ekaterina and Bagdasarov, Sergei and M. Investigating the Linguistic Performance of Large Language Models in Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.28

  161. [169]

    I so C hrono M eter: A Simple and Effective Isochronic Translation Evaluation Metric

    Rozanov, Nikolai and Pankov, Vikentiy and Mukhutdinov, Dmitrii and Vypirailenko, Dima. I so C hrono M eter: A Simple and Effective Isochronic Translation Evaluation Metric. 2024. doi:10.18653/v1/2024.wmt-1.29

  162. [170]

    A Test Suite of Prompt Injection Attacks for LLM -based Machine Translation

    Miceli Barone, Antonio Valerio and Sun, Zhifan. A Test Suite of Prompt Injection Attacks for LLM -based Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.30

  163. [171]

    Killing Two Flies with One Stone: An Attempt to Break LLM s Using E nglish- I celandic Idioms and Proper Names

    \'A rmannsson, Bjarki and Hafsteinsson, Hinrik and Jasonarson, Atli and Steingrimsson, Steinthor. Killing Two Flies with One Stone: An Attempt to Break LLM s Using E nglish- I celandic Idioms and Proper Names. 2024. doi:10.18653/v1/2024.wmt-1.31

  164. [172]

    M eta M etrics- MT : Tuning Meta-Metrics for Machine Translation via Human Preference Calibration

    Anugraha, David and Kuwanto, Garry and Susanto, Lucky and Wijaya, Derry Tanti and Winata, Genta. M eta M etrics- MT : Tuning Meta-Metrics for Machine Translation via Human Preference Calibration. 2024. doi:10.18653/v1/2024.wmt-1.32

  165. [173]

    chr F - S : Semantics Is All You Need

    Mukherjee, Ananya and Shrivastava, Manish. chr F - S : Semantics Is All You Need. 2024. doi:10.18653/v1/2024.wmt-1.33

  166. [174]

    MSLC 24: Further Challenges for Metrics on a Wide Landscape of Translation Quality

    Knowles, Rebecca and Larkin, Samuel and Lo, Chi-Kiu. MSLC 24: Further Challenges for Metrics on a Wide Landscape of Translation Quality. 2024. doi:10.18653/v1/2024.wmt-1.34

  167. [175]

    M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task

    Juraska, Juraj and Deutsch, Daniel and Finkelstein, Mara and Freitag, Markus. M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.35

  168. [176]

    Evaluating WMT 2024 Metrics Shared Task Submissions on A fri MTE (the A frican Challenge Set)

    Wang, Jiayi and Adelani, David Ifeoluwa and Stenetorp, Pontus. Evaluating WMT 2024 Metrics Shared Task Submissions on A fri MTE (the A frican Challenge Set). 2024. doi:10.18653/v1/2024.wmt-1.36

  169. [177]

    Machine Translation Metrics Are Better in Evaluating Linguistic Errors on LLM s than on Encoder-Decoder Systems

    Avramidis, Eleftherios and Manakhimova, Shushen and Macketanz, Vivien and M. Machine Translation Metrics Are Better in Evaluating Linguistic Errors on LLM s than on Encoder-Decoder Systems. 2024. doi:10.18653/v1/2024.wmt-1.37

  170. [178]

    TMU - HIT `s Submission for the WMT 24 Quality Estimation Shared Task: Is GPT -4 a Good Evaluator for Machine Translation?

    Sato, Ayako and Nakajima, Kyotaro and Kim, Hwichan and Chen, Zhousi and Komachi, Mamoru. TMU - HIT `s Submission for the WMT 24 Quality Estimation Shared Task: Is GPT -4 a Good Evaluator for Machine Translation?. 2024. doi:10.18653/v1/2024.wmt-1.38

  171. [179]

    HW - TSC 2024 Submission for the Quality Estimation Shared Task

    Shan, Weiqiao and Zhu, Ming and Li, Yuang and Piao, Mengyao and Zhao, Xiaofeng and Su, Chang and Zhang, Min and Yang, Hao and Jiang, Yanfei. HW - TSC 2024 Submission for the Quality Estimation Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.39

  172. [180]

    HW - TSC `s Participation in the WMT 2024 QEAPE Task

    Yu, Jiawei and Zhao, Xiaofeng and Zhang, Min and Yanqing, Zhao and Li, Yuang and Chang, Su and Qiao, Xiaosong and Miaomiao, Ma and Yang, Hao. HW - TSC `s Participation in the WMT 2024 QEAPE Task. 2024. doi:10.18653/v1/2024.wmt-1.40

  173. [181]

    Perez-Ortiz, Juan Antonio and S \'a nchez-Mart \'i nez, Felipe and S \'a nchez-Cartagena, V \'i ctor M. and Espl \`a -Gomis, Miquel and Galiano Jimenez, Aaron and Oliver, Antoni and Avent \'i n-Boya, Claudi and Pardos, Alejandro and Vald \'e s, Cristina and Sans Socasau, Jus \...

  174. [182]

    The B angla/ B engali Seed Dataset Submission to the WMT 24 Open Language Data Initiative Shared Task

    Ahmed, Firoz and Venkateswaran, Nitin and Moeller, Sarah. The B angla/ B engali Seed Dataset Submission to the WMT 24 Open Language Data Initiative Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.42

  175. [183]

    A High-quality Seed Dataset for I talian Machine Translation

    Ferrante, Edoardo. A High-quality Seed Dataset for I talian Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.43

  176. [184]

    Correcting FLORES Evaluation Dataset for Four A frican Languages

    Abdulmumin, Idris and Mkhwanazi, Sthembiso and Mbooi, Mahlatse and Muhammad, Shamsuddeen Hassan and Ahmad, Ibrahim Said and Putini, Neo and Mathebula, Miehleketo and Shingange, Matimba and Gwadabe, Tajuddeen and Marivate, Vukosi. Correcting FLORES Evaluation Dataset for Four A...

  177. [185]

    Expanding FLORES + Benchmark for More Low-Resource Settings: P ortuguese-Emakhuwa Machine Translation Evaluation

    Ali, Felermino Dario Mario and Lopes Cardoso, Henrique and Sousa-Silva, Rui. Expanding FLORES + Benchmark for More Low-Resource Settings: P ortuguese-Emakhuwa Machine Translation Evaluation. 2024. doi:10.18653/v1/2024.wmt-1.45

  178. [186]

    Enhancing Tuvan Language Resources through the FLORES Dataset

    Kuzhuget, Ali and Mongush, Airana and Oorzhak, Nachyn-Enkhedorzhu. Enhancing Tuvan Language Resources through the FLORES Dataset. 2024. doi:10.18653/v1/2024.wmt-1.46

  179. [187]

    Machine Translation Evaluation Benchmark for W u C hinese: Workflow and Analysis

    Yu, Hongjian and Shi, Yiming and Zhou, Zherui and Haberland, Christopher. Machine Translation Evaluation Benchmark for W u C hinese: Workflow and Analysis. 2024. doi:10.18653/v1/2024.wmt-1.47

  180. [188]

    Open Language Data Initiative: Advancing Low-Resource Machine Translation for K arakalpak

    Mamasaidov, Mukhammadsaid and Shopulatov, Abror. Open Language Data Initiative: Advancing Low-Resource Machine Translation for K arakalpak. 2024. doi:10.18653/v1/2024.wmt-1.48

  181. [189]

    FLORES + Translation and Machine Translation Evaluation for the E rzya Language

    Gordeev, Isai and Kuldin, Sergey and Dale, David. FLORES + Translation and Machine Translation Evaluation for the E rzya Language. 2024. doi:10.18653/v1/2024.wmt-1.49

  182. [190]

    S panish Corpus and Provenance with Computer-Aided Translation for the WMT 24 OLDI Shared Task

    Cols, Jose. S panish Corpus and Provenance with Computer-Aided Translation for the WMT 24 OLDI Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.50

  183. [191]

    Efficient Terminology Integration for LLM -based Translation in Specialized Domains

    Kim, Sejoon and Sung, Mingi and Lee, Jeonghwan and Lim, Hyunkuk and Gimenez Perez, Jorge. Efficient Terminology Integration for LLM -based Translation in Specialized Domains. 2024. doi:10.18653/v1/2024.wmt-1.51

  184. [192]

    Rakuten`s Participation in WMT 2024 Patent Translation Task

    Htun, Ohnmar and Poncelas, Alberto. Rakuten`s Participation in WMT 2024 Patent Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.52

  185. [193]

    The SETU - ADAPT Submission for WMT 24 Biomedical Shared Task

    Castaldo, Antonio and Zafar, Maria and Nayak, Prashanth and Haque, Rejwanul and Way, Andy and Monti, Johanna. The SETU - ADAPT Submission for WMT 24 Biomedical Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.53

  186. [194]

    Findings of WMT 2024 Shared Task on Low-Resource I ndic Languages Translation

    Pakray, Partha and Pal, Santanu and Vetagiri, Advaitha and Krishna, Reddi and Maji, Arnab Kumar and Dash, Sandeep and Laitonjam, Lenin and Sarah, Lyngdoh and Manna, Riyanka. Findings of WMT 2024 Shared Task on Low-Resource I ndic Languages Translation. 2024. doi:10.18653/v1/20...

  187. [195]

    Findings of WMT 2024`s M ulti I ndic22 MT Shared Task for Machine Translation of 22 I ndian Languages

    Dabre, Raj and Kunchukuttan, Anoop. Findings of WMT 2024`s M ulti I ndic22 MT Shared Task for Machine Translation of 22 I ndian Languages. 2024. doi:10.18653/v1/2024.wmt-1.55

  188. [196]

    Findings of WMT 2024 E nglish-to-Low Resource Multimodal Translation Task

    Parida, Shantipriya and Bojar, Ond r ej and Abdulmumin, Idris and Muhammad, Shamsuddeen Hassan and Ahmad, Ibrahim Said. Findings of WMT 2024 E nglish-to-Low Resource Multimodal Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.56

  189. [197]

    Findings of the WMT 2024 Shared Task Translation into Low-Resource Languages of S pain: Blending Rule-Based and Neural Systems

    S \'a nchez-Mart \'i nez, Felipe and Perez-Ortiz, Juan Antonio and Galiano Jimenez, Aaron and Oliver, Antoni. Findings of the WMT 2024 Shared Task Translation into Low-Resource Languages of S pain: Blending Rule-Based and Neural Systems. 2024. doi:10.18653/v1/2024.wmt-1.57

  190. [198]

    Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation

    Wang, Longyue and Liu, Siyou and Lyu, Chenyang and Jiao, Wenxiang and Wang, Xing and Xu, Jiahao and Tu, Zhaopeng and Gu, Yan and Chen, Weiyu and Wu, Minghao and Zhou, Liting and Koehn, Philipp and Way, Andy and Yuan, Yulin. Findings of the WMT 2024 Shared Task on Discourse-Lev...

  191. [199]

    De Souza, Jos \'e G

    Mohammed, Wafaa and Agrawal, Sweta and Farajian, Amin and Cabarr \ a o, Vera and Eikema, Bryan and Farinha, Ana C and C. De Souza, Jos \'e G. Findings of the WMT 2024 Shared Task on Chat Translation. 2024. doi:10.18653/v1/2024.wmt-1.59

  192. [200]

    Findings of the WMT 2024 Shared Task on Non-Repetitive Translation

    Kinugawa, Kazutaka and Mino, Hideya and Goto, Isao and Shirai, Naoto. Findings of the WMT 2024 Shared Task on Non-Repetitive Translation. 2024. doi:10.18653/v1/2024.wmt-1.60

  193. [201]

    A3-108 Controlling Token Generation in Low Resource Machine Translation Systems

    Yadav, Saumitra and Mukherjee, Ananya and Shrivastava, Manish. A3-108 Controlling Token Generation in Low Resource Machine Translation Systems. 2024. doi:10.18653/v1/2024.wmt-1.61

  194. [202]

    S amsung R & D Institute P hilippines @ WMT 2024 I ndic MT Task

    Roque, Matthew Theodore and Catalan, Carlos Rafael and Velasco, Dan John and Rufino, Manuel Antonio and Cruz, Jan Christian Blaise. S amsung R & D Institute P hilippines @ WMT 2024 I ndic MT Task. 2024. doi:10.18653/v1/2024.wmt-1.62

  195. [203]

    DLUT - NLP Machine Translation Systems for WMT 24 Low-Resource I ndic Language Translation

    Ju, Chenfei and Liu, Junpeng and Huang, Kaiyu and Huang, Degen. DLUT - NLP Machine Translation Systems for WMT 24 Low-Resource I ndic Language Translation. 2024. doi:10.18653/v1/2024.wmt-1.63

  196. [204]

    SRIB - NMT `s Submission to the I ndic MT Shared Task in WMT 2024

    Patil, Pranamya and Hr, Raghavendra and Raghuwanshi, Aditya and Verma, Kushal. SRIB - NMT `s Submission to the I ndic MT Shared Task in WMT 2024. 2024. doi:10.18653/v1/2024.wmt-1.64

  197. [205]

    MTNLP - IIITH : Machine Translation for Low-Resource I ndic Languages

    P M, Abhinav and Shetye, Ketaki and Krishnamurthy, Parameswari. MTNLP - IIITH : Machine Translation for Low-Resource I ndic Languages. 2024. doi:10.18653/v1/2024.wmt-1.65

  198. [206]

    Exploration of the C ycle GN Framework for Low-Resource Languages

    Dreano, S. Exploration of the C ycle GN Framework for Low-Resource Languages. 2024. doi:10.18653/v1/2024.wmt-1.66

  199. [207]

    The SETU - ADAPT Submissions to the WMT 24 Low-Resource I ndic Language Translation Task

    Gajakos, Neha and Nayak, Prashanth and Haque, Rejwanul and Way, Andy. The SETU - ADAPT Submissions to the WMT 24 Low-Resource I ndic Language Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.67

  200. [208]

    SPRING Lab IITM `s Submission to Low Resource I ndic Language Translation Shared Task

    Sayed, Hamees and Joglekar, Advait and Umesh, Srinivasan. SPRING Lab IITM `s Submission to Low Resource I ndic Language Translation Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.68

  201. [209]

    Machine Translation Advancements of Low-Resource I ndian Languages by Transfer Learning

    Wei, Bin and Jiawei, Zheng and Li, Zongyao and Wu, Zhanglin and Guo, Jiaxin and Wei, Daimeng and Rao, Zhiqiang and Li, Shaojun and Luo, Yuanchang and Shang, Hengchao and Yang, Jinlong and Xie, Yuhao and Yang, Hao. Machine Translation Advancements of Low-Resource I ndian Langua...

  202. [210]

    NLIP \_ L ab- IITH Low-Resource MT System for WMT 24 I ndic MT Shared Task

    Sahoo, Pramit and Brahma, Maharaj and Desarkar, Maunendra Sankar. NLIP \_ L ab- IITH Low-Resource MT System for WMT 24 I ndic MT Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.70

  203. [211]

    Yes- MT `s Submission to the Low-Resource I ndic Language Translation Shared Task in WMT 2024

    Bhaskar, Yash and Krishnamurthy, Parameswari. Yes- MT `s Submission to the Low-Resource I ndic Language Translation Shared Task in WMT 2024. 2024. doi:10.18653/v1/2024.wmt-1.71

  204. [212]

    System Description of BV - SLP for S indhi- E nglish Machine Translation in M ulti I ndic22 MT 2024 Shared Task

    Joshi, Nisheeth and Katyayan, Pragya and Arora, Palak and Nathani, Bharti. System Description of BV - SLP for S indhi- E nglish Machine Translation in M ulti I ndic22 MT 2024 Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.72

  205. [213]

    WMT 24 System Description for the M ulti I ndic22 MT Shared Task on M anipuri Language

    Singh, Ningthoujam Justwant and Singh, Kshetrimayum Boynao and Singh, Ningthoujam Avichandra and Phijam, Sanjita and Singh, Thoudam Doren. WMT 24 System Description for the M ulti I ndic22 MT Shared Task on M anipuri Language. 2024. doi:10.18653/v1/2024.wmt-1.73

  206. [214]

    NLIP -Lab- IITH Multilingual MT System for WAT 24 MT Shared Task

    Brahma, Maharaj and Sahoo, Pramit and Desarkar, Maunendra Sankar. NLIP -Lab- IITH Multilingual MT System for WAT 24 MT Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.74

  207. [215]

    DCU ADAPT at WMT 24: E nglish to Low-resource Multi-Modal Translation Task

    Haq, Sami and Huidrom, Rudali and Castilho, Sheila. DCU ADAPT at WMT 24: E nglish to Low-resource Multi-Modal Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.75

  208. [216]

    E nglish-to-Low-Resource Translation: A Multimodal Approach for H indi, M alayalam, B engali, and H ausa

    Hatami, Ali and Banerjee, Shubhanker and Arcan, Mihael and Buitelaar, Paul and Philip McCrae, John. E nglish-to-Low-Resource Translation: A Multimodal Approach for H indi, M alayalam, B engali, and H ausa. 2024. doi:10.18653/v1/2024.wmt-1.76

  209. [217]

    O dia G en AI `s Participation in WMT 2024 E nglish-to-Low Resource Multimodal Translation Task

    Parida, Shantipriya and Sahoo, Shashikanta and Sekhar, Sambit and Jena, Upendra and Jena, Sushovan and Lata, Kusum. O dia G en AI `s Participation in WMT 2024 E nglish-to-Low Resource Multimodal Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.77

  210. [218]

    Arewa NLP `s Participation at WMT 24

    Ahmad, Mahmoud and Khalid, Auwal and Aliyu, Lukman and Sani, Babangida and Abdullahi, Mariya. Arewa NLP `s Participation at WMT 24. 2024. doi:10.18653/v1/2024.wmt-1.78

  211. [219]

    Multimodal Machine Translation for Low-Resource I ndic Languages: A Chain-of-Thought Approach Using Large Language Models

    Rajpoot, Pawan and Bhat, Nagaraj and Shrivastava, Ashish. Multimodal Machine Translation for Low-Resource I ndic Languages: A Chain-of-Thought Approach Using Large Language Models. 2024. doi:10.18653/v1/2024.wmt-1.79

  212. [220]

    Chitranuvad: Adapting Multi-lingual LLM s for Multimodal Translation

    Khan, Shaharukh and Tarun, Ayush and Faraz, Ali and Kamble, Palash and Dahiya, Vivek and Pokala, Praveen and Kulkarni, Ashish and Khatri, Chandra and Ravi, Abhinav and Agarwal, Shubham. Chitranuvad: Adapting Multi-lingual LLM s for Multimodal Translation. 2024. doi:10.18653/v1...

  213. [221]

    Brotherhood at WMT 2024: Leveraging LLM -Generated Contextual Conversations for Cross-Lingual Image Captioning

    Betala, Siddharth and Chokshi, Ishan. Brotherhood at WMT 2024: Leveraging LLM -Generated Contextual Conversations for Cross-Lingual Image Captioning. 2024. doi:10.18653/v1/2024.wmt-1.81

  214. [222]

    TIM - UNIGE Translation into Low-Resource Languages of S pain for WMT 24

    Mutal, Jonathan and Ormaechea, Luc \'i a. TIM - UNIGE Translation into Low-Resource Languages of S pain for WMT 24. 2024. doi:10.18653/v1/2024.wmt-1.82

  215. [223]

    TAN - IBE Participation in the Shared Task: Translation into Low-Resource Languages of S pain

    Oliver, Antoni. TAN - IBE Participation in the Shared Task: Translation into Low-Resource Languages of S pain. 2024. doi:10.18653/v1/2024.wmt-1.83

  216. [224]

    Enhaced Apertium System: Translation into Low-Resource Languages of S pain S panish -- A sturian

    Garc \'i a, Sof \'i a. Enhaced Apertium System: Translation into Low-Resource Languages of S pain S panish -- A sturian. 2024. doi:10.18653/v1/2024.wmt-1.84

  217. [225]

    and Perez-Ortiz, Juan Antonio and S \'a nchez-Mart \'i nez, Felipe

    Galiano Jimenez, Aaron and S \'a nchez-Cartagena, V \'i ctor M. and Perez-Ortiz, Juan Antonio and S \'a nchez-Mart \'i nez, Felipe. U niversitat d`Alacant`s Submission to the WMT 2024 Shared Task on Translation into Low-Resource Languages of S pain. 2024. doi:10.18653/v1/2024.wmt-1.85

  218. [226]

    S amsung R & D Institute P hilippines @ WMT 2024 Low-resource Languages of S pain Shared Task

    Velasco, Dan John and Rufino, Manuel Antonio and Cruz, Jan Christian Blaise. S amsung R & D Institute P hilippines @ WMT 2024 Low-resource Languages of S pain Shared Task. 2024. doi:10.18653/v1/2024.wmt-1.86

  219. [227]

    Back to the Stats: Rescuing Low Resource Neural Machine Translation with Statistical Methods

    Velayuthan, Menan and Jayakody, Dilith and De Silva, Nisansa and Fernando, Aloka and Ranathunga, Surangika. Back to the Stats: Rescuing Low Resource Neural Machine Translation with Statistical Methods. 2024. doi:10.18653/v1/2024.wmt-1.87

  220. [228]

    Hybrid Distillation from RBMT and NMT : H elsinki- NLP `s Submission to the Shared Task on Translation into Low-Resource Languages of S pain

    De Gibert, Ona and Aulamo, Mikko and Scherrer, Yves and Tiedemann, J. Hybrid Distillation from RBMT and NMT : H elsinki- NLP `s Submission to the Shared Task on Translation into Low-Resource Languages of S pain. 2024. doi:10.18653/v1/2024.wmt-1.88

  221. [229]

    Robustness of Fine-Tuned Models for Machine Translation with Varying Noise Levels: Insights for A sturian, A ragonese and Aranese

    B. Robustness of Fine-Tuned Models for Machine Translation with Varying Noise Levels: Insights for A sturian, A ragonese and Aranese. 2024. doi:10.18653/v1/2024.wmt-1.89

  222. [230]

    Training and Fine-Tuning NMT Models for Low-Resource Languages Using Apertium-Based Synthetic Corpora

    Sant, Aleix and Bardanca, Daniel and Pichel Campos, Jos \'e Ramom and De Luca Fornaciari, Francesca and Escolano, Carlos and Garcia Gilabert, Javier and Gamallo, Pablo and Mash, Audrey and Liao, Xixian and Melero, Maite. Training and Fine-Tuning NMT Models for Low-Resource Lan...

  223. [231]

    Vicomtech@ WMT 2024: Shared Task on Translation into Low-Resource Languages of S pain

    Ponce, David and Gete, Harritxu and Etchegoyhen, Thierry. Vicomtech@ WMT 2024: Shared Task on Translation into Low-Resource Languages of S pain. 2024. doi:10.18653/v1/2024.wmt-1.91

  224. [232]

    SJTU System Description for the WMT 24 Low-Resource Languages of S pain Task

    Hu, Tianxiang and Sun, Haoxiang and Gao, Ruize and Tang, Jialong and Zhang, Pei and Yang, Baosong and Wang, Rui. SJTU System Description for the WMT 24 Low-Resource Languages of S pain Task. 2024. doi:10.18653/v1/2024.wmt-1.92

  225. [233]

    Multilingual Transfer and Domain Adaptation for Low-Resource Languages of S pain

    Luo, Yuanchang and Wu, Zhanglin and Wei, Daimeng and Shang, Hengchao and Li, Zongyao and Guo, Jiaxin and Rao, Zhiqiang and Li, Shaojun and Yang, Jinlong and Xie, Yuhao and Jiawei, Zheng and Wei, Bin and Yang, Hao. Multilingual Transfer and Domain Adaptation for Low-Resource La...

  226. [234]

    TRIBBLE - TR anslating IB erian languages Based on Limited E -resources

    Kuzmin, Igor and Przyby a, Piotr and Mcgill, Euan and Saggion, Horacio. TRIBBLE - TR anslating IB erian languages Based on Limited E -resources. 2024. doi:10.18653/v1/2024.wmt-1.94

  227. [235]

    C loud S heep System for WMT 24 Discourse-Level Literary Translation

    Liu, Lisa and Liu, Ryan and Tsai, Angela and Shang, Jingbo. C loud S heep System for WMT 24 Discourse-Level Literary Translation. 2024. doi:10.18653/v1/2024.wmt-1.95

  228. [236]

    Final Submission of SJTUL ove F iction to Literary Task

    Sun, Haoxiang and Hu, Tianxiang and Gao, Ruize and Tang, Jialong and Zhang, Pei and Yang, Baosong and Wang, Rui. Final Submission of SJTUL ove F iction to Literary Task. 2024. doi:10.18653/v1/2024.wmt-1.96

  229. [237]

    Context-aware and Style-related Incremental Decoding Framework for Discourse-Level Literary Translation

    Luo, Yuanchang and Guo, Jiaxin and Wei, Daimeng and Shang, Hengchao and Li, Zongyao and Wu, Zhanglin and Rao, Zhiqiang and Li, Shaojun and Yang, Jinlong and Yang, Hao. Context-aware and Style-related Incremental Decoding Framework for Discourse-Level Literary Translation. 2024...

  230. [238]

    N ovel T rans: System for WMT 24 Discourse-Level Literary Translation

    Liu, Yuchen and Yao, Yutong and Zhan, Runzhe and Lin, Yuchu and Wong, Derek F. N ovel T rans: System for WMT 24 Discourse-Level Literary Translation. 2024. doi:10.18653/v1/2024.wmt-1.98

  231. [239]

    L in C hance- NTU for Unconstrained WMT 2024 Literary Translation

    Li, Kechen and Tao, Yaotian and Huang, Hongyi and Ji, Tianbo. L in C hance- NTU for Unconstrained WMT 2024 Literary Translation. 2024. doi:10.18653/v1/2024.wmt-1.99

  232. [240]

    Improving Context Usage for Translating Bilingual Customer Support Chat with Large Language Models

    Pombal, Jose and Agrawal, Sweta and Martins, Andr \'e. Improving Context Usage for Translating Bilingual Customer Support Chat with Large Language Models. 2024. doi:10.18653/v1/2024.wmt-1.100

  233. [241]

    Optimising LLM -Driven Machine Translation with Context-Aware Sliding Windows

    Yang, Xinye and Mu, Yida and Bontcheva, Kalina and Song, Xingyi. Optimising LLM -Driven Machine Translation with Context-Aware Sliding Windows. 2024. doi:10.18653/v1/2024.wmt-1.101

  234. [242]

    Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History

    Sung, Mingi and Lee, Seungmin and Kim, Jiwon and Kim, Sejoon. Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History. 2024. doi:10.18653/v1/2024.wmt-1.102

  235. [243]

    Enhancing Translation Quality: A Comparative Study of Fine-Tuning and Prompt Engineering in Dialog-Oriented Machine Translation Systems

    Zhu, Lichao and Zimina, Maria and Namdarzadeh, Behnoosh and Ballier, Nicolas and Yun \`e s, Jean-Baptiste. Enhancing Translation Quality: A Comparative Study of Fine-Tuning and Prompt Engineering in Dialog-Oriented Machine Translation Systems. Insights from the MULTITAN - GML ...

  236. [244]

    The SETU - ADAPT Submissions to WMT 2024 Chat Translation Tasks

    Zafar, Maria and Castaldo, Antonio and Nayak, Prashanth and Haque, Rejwanul and Way, Andy. The SETU - ADAPT Submissions to WMT 2024 Chat Translation Tasks. 2024. doi:10.18653/v1/2024.wmt-1.104

  237. [245]

    Exploring the Traditional NMT Model and Large Language Model for Chat Translation

    Yang, Jinlong and Shang, Hengchao and Wei, Daimeng and Guo, Jiaxin and Li, Zongyao and Wu, Zhanglin and Rao, Zhiqiang and Li, Shaojun and Xie, Yuhao and Luo, Yuanchang and Jiawei, Zheng and Wei, Bin and Yang, Hao. Exploring the Traditional NMT Model and Large Language Model fo...

  238. [246]

    Graph Representations for Machine Translation in Dialogue Settings

    Krause, Lea and Baez Santamaria, Selene and Kalo, Jan-Christoph. Graph Representations for Machine Translation in Dialogue Settings. 2024. doi:10.18653/v1/2024.wmt-1.106

  239. [247]

    Reducing Redundancy in J apanese-to- E nglish Translation: A Multi-Pipeline Approach for Translating Repeated Elements in J apanese

    Wang, Qiao and Huang, Yixuan and Yuan, Zheng. Reducing Redundancy in J apanese-to- E nglish Translation: A Multi-Pipeline Approach for Translating Repeated Elements in J apanese. 2024. doi:10.18653/v1/2024.wmt-1.107

  240. [248]

    SYSTRAN @ WMT 24 Non-Repetitive Translation Task

    Avila, Marko and Crego, Josep. SYSTRAN @ WMT 24 Non-Repetitive Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.108

  241. [249]

    Mitigating Metric Bias in Minimum B ayes Risk Decoding

    Kovacs, Geza and Deutsch, Daniel and Freitag, Markus. Mitigating Metric Bias in Minimum B ayes Risk Decoding. 2024. doi:10.18653/v1/2024.wmt-1.109

  242. [250]

    Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data

    Liu, Zhongtao and Riley, Parker and Deutsch, Daniel and Lui, Alison and Niu, Mengmeng and Shah, Apurva and Freitag, Markus. Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data. 2024. doi:10.18653/v1/2024.wmt-1.110

  243. [251]

    How Effective Are State Space Models for Machine Translation?

    Pitorro, Hugo and Vasylenko, Pavlo and Treviso, Marcos and Martins, Andr \'e. How Effective Are State Space Models for Machine Translation?. 2024. doi:10.18653/v1/2024.wmt-1.111

  244. [252]

    Evaluation and Large-scale Training for Contextual Machine Translation

    Post, Matt and Junczys-Dowmunt, Marcin. Evaluation and Large-scale Training for Contextual Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.112

  245. [253]

    A Multi-task Learning Framework for Evaluating Machine Translation of Emotion-loaded User-generated Content

    Qian, Shenbin and Orasan, Constantin and Kanojia, Diptesh and Do Carmo, F \'e lix. A Multi-task Learning Framework for Evaluating Machine Translation of Emotion-loaded User-generated Content. 2024. doi:10.18653/v1/2024.wmt-1.113

  246. [254]

    On Instruction-Finetuning Neural Machine Translation Models

    Raunak, Vikas and Grundkiewicz, Roman and Junczys-Dowmunt, Marcin. On Instruction-Finetuning Neural Machine Translation Models. 2024. doi:10.18653/v1/2024.wmt-1.114

  247. [255]

    Benchmarking Visually-Situated Translation of Text in Natural Images

    Salesky, Elizabeth and Koehn, Philipp and Post, Matt. Benchmarking Visually-Situated Translation of Text in Natural Images. 2024. doi:10.18653/v1/2024.wmt-1.115

  248. [256]

    Analysing Translation Artifacts: A Comparative Study of LLM s, NMT s, and Human Translations

    Sizov, Fedor and Espa \ n a-Bonet, Cristina and Van Genabith, Josef and Xie, Roy and Dutta Chowdhury, Koel. Analysing Translation Artifacts: A Comparative Study of LLM s, NMT s, and Human Translations. 2024. doi:10.18653/v1/2024.wmt-1.116

  249. [257]

    How Grammatical Features Impact Machine Translation: A New Test Suite for C hinese- E nglish MT Evaluation

    Song, Huacheng and Li, Yi and Wu, Yiwen and Liu, Yu and Lin, Jingxia and Xu, Hongzhi. How Grammatical Features Impact Machine Translation: A New Test Suite for C hinese- E nglish MT Evaluation. 2024. doi:10.18653/v1/2024.wmt-1.117

  250. [258]

    Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy

    Thompson, Brian and Mathur, Nitika and Deutsch, Daniel and Khayrallah, Huda. Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy. 2024. doi:10.18653/v1/2024.wmt-1.118

  251. [259]

    Speech Is More than Words: Do Speech-to-Text Translation Systems Leverage Prosody?

    Tsiamas, Ioannis and Sperber, Matthias and Finch, Andrew and Garg, Sarthak. Speech Is More than Words: Do Speech-to-Text Translation Systems Leverage Prosody?. 2024. doi:10.18653/v1/2024.wmt-1.119

  252. [260]

    Cultural Adaptation of Menus: A Fine-Grained Approach

    Zhang, Zhonghe and He, Xiaoyu and Iyer, Vivek and Birch, Alexandra. Cultural Adaptation of Menus: A Fine-Grained Approach. 2024. doi:10.18653/v1/2024.wmt-1.120

  253. [261]

    Pitfalls and Outlooks in Using COMET

    Zouhar, Vil \'e m and Chen, Pinzhen and Lam, Tsz Kin and Moghe, Nikita and Haddow, Barry. Pitfalls and Outlooks in Using COMET. 2024. doi:10.18653/v1/2024.wmt-1.121

  254. [262]

    Post-edits Are Preferences Too

    Berger, Nathaniel and Riezler, Stefan and Exel, Miriam and Huck, Matthias. Post-edits Are Preferences Too. 2024. doi:10.18653/v1/2024.wmt-1.122

  255. [263]

    Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts

    Briakou, Eleftheria and Luo, Jiaming and Cherry, Colin and Freitag, Markus. Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts. 2024. doi:10.18653/v1/2024.wmt-1.123

  256. [264]

    Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task

    Caillaut, Ga. Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task. 2024. doi:10.18653/v1/2024.wmt-1.124

  257. [265]

    Shortcomings of LLM s for Low-Resource Translation: Retrieval and Understanding Are Both the Problem

    Court, Sara and Elsner, Micha. Shortcomings of LLM s for Low-Resource Translation: Retrieval and Understanding Are Both the Problem. 2024. doi:10.18653/v1/2024.wmt-1.125

  258. [266]

    Introducing the N ews P a LM MBR and QE Dataset: LLM -Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data

    Finkelstein, Mara and Vilar, David and Freitag, Markus. Introducing the N ews P a LM MBR and QE Dataset: LLM -Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data. 2024. doi:10.18653/v1/2024.wmt-1.126

  259. [267]

    Is Preference Alignment Always the Best Option to Enhance LLM -Based Translation? An Empirical Analysis

    Gisserot-Boukhlef, Hippolyte and Rei, Ricardo and Malherbe, Emmanuel and Hudelot, C \'e line and Colombo, Pierre and Guerreiro, Nuno M. Is Preference Alignment Always the Best Option to Enhance LLM -Based Translation? An Empirical Analysis. 2024. doi:10.18653/v1/2024.wmt-1.127

  260. [268]

    Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation

    Iyer, Vivek and Malik, Bhavitvya and Stepachev, Pavel and Chen, Pinzhen and Haddow, Barry and Birch, Alexandra. Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation. 2024. doi:10.18653/v1/2024.wmt-1.128

  261. [269]

    Efficient Technical Term Translation: A Knowledge Distillation Approach for Parenthetical Terminology Translation

    Jiyoon, Myung and Park, Jihyeon and Son, Jungki and Lee, Kyungro and Han, Joohyung. Efficient Technical Term Translation: A Knowledge Distillation Approach for Parenthetical Terminology Translation. 2024. doi:10.18653/v1/2024.wmt-1.129

  262. [270]

    Assessing the Role of Imagery in Multimodal Machine Translation

    Kashani Motlagh, Nicholas and Davis, Jim and Gwinnup, Jeremy and Erdmann, Grant and Anderson, Tim. Assessing the Role of Imagery in Multimodal Machine Translation. 2024. doi:10.18653/v1/2024.wmt-1.130

  263. [271]

    Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation

    Kocmi, Tom and Zouhar, Vil \'e m and Avramidis, Eleftherios and Grundkiewicz, Roman and Karpinska, Marzena and Popovi \'c , Maja and Sachan, Mrinmaya and Shmatova, Mariya. Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation. 2024. doi:10.1865...

  264. [272]

    Neural Methods for Aligning Large-Scale Parallel Corpora from the Web for South and E ast A sian Languages

    Koehn, Philipp. Neural Methods for Aligning Large-Scale Parallel Corpora from the Web for South and E ast A sian Languages. 2024. doi:10.18653/v1/2024.wmt-1.132

  265. [273]

    Plug, Play, and Fuse: Zero-Shot Joint Decoding via Word-Level Re-ranking across Diverse Vocabularies

    Koneru, Sai and Huck, Matthias and Exel, Miriam and Niehues, Jan. Plug, Play, and Fuse: Zero-Shot Joint Decoding via Word-Level Re-ranking across Diverse Vocabularies. 2024. doi:10.18653/v1/2024.wmt-1.133

  266. [274]

    Proceedings of the Eighth Widening NLP Workshop. 2024. doi:10.18653/v1/2024.winlp-1.0

  267. [275]

    Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation. 2024

  268. [276]

    Towards Disfluency Annotated Corpora for I ndian Languages

    Kochar, Chayan and Mujadia, Vandan Vasantlal and Mishra, Pruthwik and Sharma, Dipti Misra. Towards Disfluency Annotated Corpora for I ndian Languages. 2024

  269. [277]

    E mo M ix-3 L : A Code-Mixed Dataset for B angla- E nglish- H indi for Emotion Detection

    Raihan, Nishat and Goswami, Dhiman and Mahmud, Antara and Anastasopoulos, Antonios and Zampieri, Marcos. E mo M ix-3 L : A Code-Mixed Dataset for B angla- E nglish- H indi for Emotion Detection. 2024

  270. [278]

    and Buitelaar, Paul and McCrae, John P

    Rani, Priya and Negi, Gaurav and Jha, Saroj and Suryawanshi, Shardul and Ojha, Atul Kr. and Buitelaar, Paul and McCrae, John P. Findings of the WILDRE Shared Task on Code-mixed Less-resourced Sentiment Analysis for I ndo- A ryan Languages. 2024

  271. [279]

    Multilingual Bias Detection and Mitigation for I ndian Languages

    Maity, Ankita and Sharma, Anubhav and Dhar, Rudra and Abhishek, Tushar and Gupta, Manish and Varma, Vasudeva. Multilingual Bias Detection and Mitigation for I ndian Languages. 2024

  272. [280]

    Dharma \'s \= a stra Informatics: Concept Mining System for Socio-Cultural Facet in A ncient I ndia

    Nigam, Arooshi and Chandra, Subhash. Dharma \'s \= a stra Informatics: Concept Mining System for Socio-Cultural Facet in A ncient I ndia. 2024

  273. [281]

    Exploring News Summarization and Enrichment in a Highly Resource-Scarce I ndian Language: A Case Study of Mizo

    Bala, Abhinaba and Urlana, Ashok and Mishra, Rahul and Krishnamurthy, Parameswari. Exploring News Summarization and Enrichment in a Highly Resource-Scarce I ndian Language: A Case Study of Mizo. 2024

  274. [282]

    Finding the Causality of an Event in News Articles

    Lalitha Devi, Sobha and RK Rao, Pattabhi. Finding the Causality of an Event in News Articles. 2024

  275. [283]

    Creating Corpus of Low Resource I ndian Languages for Natural Language Processing: Challenges and Opportunities

    Dongare, Pratibha. Creating Corpus of Low Resource I ndian Languages for Natural Language Processing: Challenges and Opportunities. 2024

  276. [284]

    FZZG at WILDRE -7: Fine-tuning Pre-trained Models for Code-mixed, Less-resourced Sentiment Analysis

    Thakkar, Gaurish and Tadi \'c , Marko and Mikelic Preradovic, Nives. FZZG at WILDRE -7: Fine-tuning Pre-trained Models for Code-mixed, Less-resourced Sentiment Analysis. 2024

  277. [285]

    MLI nitiative@ WILDRE 7: Hybrid Approaches with Large Language Models for Enhanced Sentiment Analysis in Code-Switched and Code-Mixed Texts

    Veeramani, Hariram and Thapa, Surendrabikram and Naseem, Usman. MLI nitiative@ WILDRE 7: Hybrid Approaches with Large Language Models for Enhanced Sentiment Analysis in Code-Switched and Code-Mixed Texts. 2024

  278. [286]

    Aalamaram: A Large-Scale Linguistically Annotated Treebank for the T amil Language

    Abirami, A M and Leong, Wei Qi and Rengarajan, Hamsawardhini and Anitha, D and Suganya, R and Singh, Himanshu and Sarveswaran, Kengatharaiyer and Tjhi, William Chandra and Shah, Rajiv Ratn. Aalamaram: A Large-Scale Linguistically Annotated Treebank for the T amil Language. 2024

  279. [287]

    Proceedings of the First Workshop on Advancing Natural Language Processing for Wikipedia. 2024. doi:10.18653/v1/2024.wikinlp-1.0

  280. [288]

    B ord IR lines: A Dataset for Evaluating Cross-lingual Retrieval Augmented Generation

    Li, Bryan and Haider, Samar and Luo, Fiona and Agashe, Adwait and Callison-Burch, Chris. B ord IR lines: A Dataset for Evaluating Cross-lingual Retrieval Augmented Generation. 2024. doi:10.18653/v1/2024.wikinlp-1.3

  281. [289]

    Multi-Label Field Classification for Scientific Documents using Expert and Crowd-sourced Knowledge

    Gelles, Rebecca and Dunham, James. Multi-Label Field Classification for Scientific Documents using Expert and Crowd-sourced Knowledge. 2024. doi:10.18653/v1/2024.wikinlp-1.7

  282. [290]

    Uncovering Differences in Persuasive Language in R ussian versus E nglish W ikipedia

    Li, Bryan and Panasyuk, Aleksey and Callison-Burch, Chris. Uncovering Differences in Persuasive Language in R ussian versus E nglish W ikipedia. 2024. doi:10.18653/v1/2024.wikinlp-1.8

  283. [291]

    and Srinivasan, Krishna and Clinchant, St \'e phane and Lin, Jimmy

    Yang, Jheng-Hong and Lassance, Carlos and Rezende, Rafael S. and Srinivasan, Krishna and Clinchant, St \'e phane and Lin, Jimmy. Retrieval Evaluation for Long-Form and Knowledge-Intensive Image -- Text Article Composition. 2024. doi:10.18653/v1/2024.wikinlp-1.9

  284. [292]

    and Lopez-Ponce, Francisco Fernando and Ojeda-Trueba, Sergio-Luis and Bel-Enguix, Gemma

    Salas-Jimenez, K. and Lopez-Ponce, Francisco Fernando and Ojeda-Trueba, Sergio-Luis and Bel-Enguix, Gemma. W iki B ias as an Extrapolation Corpus for Bias Detection. 2024. doi:10.18653/v1/2024.wikinlp-1.10

  285. [293]

    HOAXPEDIA : A Unified W ikipedia Hoax Articles Dataset

    Borkakoty, Hsuvas and Espinosa-Anke, Luis. HOAXPEDIA : A Unified W ikipedia Hoax Articles Dataset. 2024. doi:10.18653/v1/2024.wikinlp-1.11

  286. [294]

    The Rise of AI -Generated Content in W ikipedia

    Brooks, Creston and Eggert, Samuel and Peskoff, Denis. The Rise of AI -Generated Content in W ikipedia. 2024. doi:10.18653/v1/2024.wikinlp-1.12

  287. [295]

    Embedded Topic Models Enhanced by Wikification

    Shibuya, Takashi and Utsuro, Takehito. Embedded Topic Models Enhanced by Wikification. 2024. doi:10.18653/v1/2024.wikinlp-1.13

  288. [296]

    Wikimedia data for AI : a review of Wikimedia datasets for NLP tasks and AI -assisted editing

    Johnson, Isaac and Kaffee, Lucie-Aim \'e e and Redi, Miriam. Wikimedia data for AI : a review of Wikimedia datasets for NLP tasks and AI -assisted editing. 2024. doi:10.18653/v1/2024.wikinlp-1.14

  289. [297]

    Blocks Architecture ( B lo A rk): Efficient, Cost-Effective, and Incremental Dataset Architecture for W ikipedia Revision History

    Li, Lingxi and Yao, Zonghai and Kwon, Sunjae and Yu, Hong. Blocks Architecture ( B lo A rk): Efficient, Cost-Effective, and Incremental Dataset Architecture for W ikipedia Revision History. 2024. doi:10.18653/v1/2024.wikinlp-1.16

  290. [298]

    ARMADA : Attribute-Based Multimodal Data Augmentation

    Jin, Xiaomeng and Kim, Jeonghwan and Zhou, Yu and Huang, Kuan-Hao and Wu, Te-Lin and Peng, Nanyun and Ji, Heng. ARMADA : Attribute-Based Multimodal Data Augmentation. 2024. doi:10.18653/v1/2024.wikinlp-1.17

  291. [299]

    Summarization-Based Document ID s for Generative Retrieval with Language Models

    Li, Alan and Cheng, Daniel and Keung, Phillip and Kasai, Jungo and Smith, Noah A. Summarization-Based Document ID s for Generative Retrieval with Language Models. 2024. doi:10.18653/v1/2024.wikinlp-1.18

  292. [300]

    Proceedings of the Eleventh Workshop on Asian Translation (WAT 2024). 2024. doi:10.18653/v1/2024.wat-1.0

Pith tools

Reviewed May 9, 2026 · model on record in the stance chip above.