REVIEW 1 major objections 2 minor 299 references
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
T0 review · 1 major / 2 minor · reviewed 2026-05-09 · grok-4.3
Pith's one-line read Combining differential privacy with LLM redaction improves the privacy-utility trade-off for Dutch clinical notes.
desk verdict Hybrid DP plus LLM preprocessing beats standalone DP on Dutch clinical notes, but the utility gains are measured on tasks too close to the de-identification steps themselves. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Hybrid pipelines that apply linguistic preprocessing (NER or LLM redaction) before differential privacy mechanisms.
What would settle it
A follow-up evaluation that applies the same de-identified notes to an actual secondary research task such as outcome prediction and finds no utility advantage for the hybrid methods over pure DP.
Extended reading notes
Core claim
The authors show that differential privacy mechanisms applied alone to Dutch clinical text cause large drops in utility on entity and relation classification tasks, while hybrid strategies that first redact protected information using NER or especially LLM-based methods before applying DP deliver markedly better privacy-utility trade-offs as measured by both leakage metrics and downstream task performance.
Load-bearing premise
That performance on entity and relation classification tasks accurately reflects the real-world usefulness of the de-identified notes for secondary healthcare research.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents the first comparative evaluation of differentially private (DP) mechanisms, named entity recognition (NER), and large language model (LLM)-based methods for de-identifying Dutch clinical notes. It examines standalone approaches as well as hybrid pipelines that apply NER or LLM preprocessing before DP, measuring performance via privacy leakage metrics and extrinsic utility on entity and relation classification tasks. The central finding is that DP alone substantially degrades utility, whereas hybrid strategies—particularly LLM-based redaction followed by DP—yield a meaningfully better privacy-utility trade-off.
Significance. If the empirical results hold under scrutiny, the work supplies timely, language-specific evidence on practical de-identification strategies for Dutch clinical text, an under-studied setting relative to English. The hybrid LLM-DP approach is shown to mitigate the utility penalty of pure DP while retaining formal privacy guarantees, which could directly inform GDPR-compliant secondary-use pipelines in healthcare. The inclusion of extrinsic downstream tasks adds relevance beyond intrinsic privacy metrics, though the paper's own evaluation design limits the strength of claims about broader clinical utility.
major comments (1)
- Evaluation section (extrinsic tasks): The central claim that LLM preprocessing improves the privacy-utility trade-off rests on entity and relation classification performance. These tasks are semantically close to the NER/LLM redaction step itself, so measured gains may reflect task alignment rather than preserved semantic content for secondary clinical uses (e.g., cohort studies or outcome modeling). No results are reported on more distant tasks such as diagnosis prediction or temporal event extraction, leaving the generalizability of the improvement untested and weakening support for the headline conclusion.
minor comments (2)
- Abstract: The abstract states the evaluation approach and main finding but provides no quantitative results (e.g., specific privacy leakage rates, F1 scores, or DP parameters such as ε), making it difficult for readers to gauge the magnitude of the reported improvements without reading the full results section.
- Dataset and experimental details: The manuscript would benefit from an explicit table or subsection listing the Dutch clinical corpus size, number of notes, protected entity types, and the exact DP mechanisms and privacy budgets (ε, δ) used in each condition.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address the single major comment below.
read point-by-point responses
-
Referee: Evaluation section (extrinsic tasks): The central claim that LLM preprocessing improves the privacy-utility trade-off rests on entity and relation classification performance. These tasks are semantically close to the NER/LLM redaction step itself, so measured gains may reflect task alignment rather than preserved semantic content for secondary clinical uses (e.g., cohort studies or outcome modeling). No results are reported on more distant tasks such as diagnosis prediction or temporal event extraction, leaving the generalizability of the improvement untested and weakening support for the headline conclusion.
Authors: We thank the referee for highlighting this important consideration. Entity and relation classification were chosen as extrinsic tasks because they are standard benchmarks in clinical NLP literature for assessing de-identification utility and because Dutch-annotated datasets are available for them, enabling direct comparison across methods. The relation classification task requires contextual inference and semantic linking beyond entity detection alone, offering evidence that hybrid LLM-DP approaches preserve more than surface-level information. We agree, however, that more distant tasks such as diagnosis prediction or temporal event extraction would better demonstrate generalizability to broader secondary uses. Such evaluations would require additional annotated data and resources beyond the current study scope. In the revised manuscript we will add an explicit limitations paragraph in the Discussion section that acknowledges the scope of the chosen tasks, qualifies the headline claims accordingly, and identifies these more distant tasks as valuable directions for future work. revision: partial
Circularity Check
No circularity: purely empirical comparative evaluation with independent benchmarks
full rationale
The paper conducts a comparative study of DP, NER, and LLM-based de-identification methods on Dutch clinical notes, measuring privacy leakage and utility via standard extrinsic tasks (entity and relation classification). No mathematical derivations, equations, fitted parameters, or predictions are present. Central claims rest on direct experimental results rather than any self-referential reduction, self-citation chains, or ansatz smuggling. Evaluations use established metrics and tasks that do not reduce to the preprocessing steps by construction. This is a standard empirical setup with no load-bearing circular elements.
Assumptions & free parameters
assumptions (2)
- domain assumption Differential privacy mechanisms provide formal privacy guarantees when correctly implemented.
- domain assumption NER and LLM models can reliably identify protected health information in clinical text.
Cite this review
Pith. "Pith review of Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation." pith.science (2026). https://pith.science/paper/2604.21421
@misc{pith2026260421421,
author = {Pith},
title = {Pith review of: Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/2604.21421}},
note = {Machine review of arXiv:2604.21421}
}
read the original abstract
Protecting patient privacy in clinical narratives is essential for enabling secondary use of healthcare data under regulations such as GDPR and HIPAA. While manual de-identification remains the gold standard, it is costly and slow, motivating the need for automated methods that combine privacy guarantees with high utility. Most automated text de-identification pipelines employed named entity recognition (NER) to identify protected entities for redaction. Although methods based on differential privacy (DP) provide formal privacy guarantees, more recently also large language models (LLMs) are increasingly used for text de-identification in the clinical domain. In this work, we present the first comparative study of DP, NER, and LLMs for Dutch clinical text de-identification. We investigate these methods separately as well as hybrid strategies that apply NER or LLM preprocessing prior to DP, and assess performance in terms of privacy leakage and extrinsic evaluation (entity and relation classification). We show that DP mechanisms alone degrade utility substantially, but combining them with linguistic preprocessing, especially LLM-based redaction, significantly improves the privacy-utility trade-off.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
- [1]
- [2]
-
[3]
The OrienTel Moroccan MCA (Modern Colloquial Arabic) database
Khalid Choukri and Niklas Paullson. The OrienTel Moroccan MCA (Modern Colloquial Arabic) database. 2004
work page 2004
-
[4]
Roventini, Adriana and Marinelli, Rita and Bertagna, Francesca. ItalWordNet v.2
-
[5]
Dwork, Cynthia , title =. Proceedings of the 33rd International Conference on Automata, Languages and Programming - Volume Part II , pages =. 2006 , isbn =. doi:10.1007/11787006_1 , abstract =
-
[6]
Dwork, Cynthia and Roth, Aaron , title =. Found. Trends Theor. Comput. Sci. , month = aug, pages =. 2014 , issue_date =. doi:10.1561/0400000042 , abstract =
-
[7]
Meystre, Stephane M. and Friedlin, F. Jeffrey and South, Brett R. and Shen, Shuying and Samore, Matthew H. , title=. BMC Medical Research Methodology , year=. doi:10.1186/1471-2288-10-70 , url=
-
[8]
International Symposium on Privacy Enhancing Technologies , year=
Broadening the Scope of Differential Privacy Using Metrics , author=. International Symposium on Privacy Enhancing Technologies , year=
Show all 299 references
-
[9]
2019 IEEE International Conference on Data Mining (ICDM) , year=
Leveraging Hierarchical Representations for Preserving Privacy and Utility in Text , author=. 2019 IEEE International Conference on Data Mining (ICDM) , year=
2019
-
[10]
BART : Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, Mike and Liu, Yinhan and Goyal, Naman and Ghazvininejad, Marjan and Mohamed, Abdelrahman and Levy, Omer and Stoyanov, Veselin and Zettlemoyer, Luke. BART : Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. Proce...
2020 doi
-
[11]
Conference of the European Chapter of the Association for Computational Linguistics , year=
ADePT: Auto-encoder based Differentially Private Text Transformation , author=. Conference of the European Chapter of the Association for Computational Linguistics , year=
-
[12]
2025 , eprint=
InferDPT: Privacy-Preserving Inference for Black-box Large Language Model , author=. 2025 , eprint=
2025
-
[13]
Yue, Xiang and Du, Minxin and Wang, Tianhao and Li, Yaliang and Sun, Huan and Chow, Sherman S. M. Differential Privacy for Text Analytics via Natural Text Sanitization. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021. doi:10.18653/v1/2021.findi...
2021 doi
-
[14]
A Customized Text Sanitization Mechanism with Differential Privacy
Chen, Sai and Mo, Fengran and Wang, Yanhao and Chen, Cen and Nie, Jian-Yun and Wang, Chengyu and Cui, Jamie. A Customized Text Sanitization Mechanism with Differential Privacy. Findings of the Association for Computational Linguistics: ACL 2023. 2023. doi:10.18653/v1/2023.find...
2023 doi
-
[15]
2023 , eprint=
DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4 , author=. 2023 , eprint=
2023
-
[16]
2020 , eprint=
Language Models are Few-Shot Learners , author=. 2020 , eprint=
2020
-
[17]
Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus
Stubbs, Amber and Uzuner, \"O zlem. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. J Biomed Inform
2014
-
[18]
Department of Health and Human Services
U.S. Department of Health and Human Services. 45 CFR § 164.514 – de-identification of health information
-
[19]
European Parliament and Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Dat...
2016
-
[20]
2019 , eprint=
Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations , author=. 2019 , eprint=
2019
-
[21]
Language Resources and Evaluation , pages=
Creation of a gold standard Dutch corpus of clinical notes for adverse drug event detection: the Dutch ADE corpus , author=. Language Resources and Evaluation , pages=. 2025 , publisher=
2025
-
[22]
arXiv preprint arXiv:2211.01147 , year=
An easy-to-use and robust approach for the differentially private de-identification of clinical textual documents , author=. arXiv preprint arXiv:2211.01147 , year=
-
[23]
Digital Health , volume=
Data privacy in healthcare: Global challenges and solutions , author=. Digital Health , volume=. 2025 , publisher=
2025
-
[24]
Journal of medical Internet research , volume=
Use and understanding of anonymization and de-identification in the biomedical literature: scoping review , author=. Journal of medical Internet research , volume=. 2019 , publisher=
2019
-
[25]
arXiv preprint arXiv:1912.09582 , year=
Bertje: A dutch bert model , author=. arXiv preprint arXiv:1912.09582 , year=
1912
-
[26]
nl: a language model for Dutch electronic health records , author=
MedRoBERTa. nl: a language model for Dutch electronic health records , author=. Computational Linguistics in the Netherlands , volume=. 2021 , organization=
2021
-
[27]
How to successfully recycle English GPT-2 to make models for other languages , author=
As good as new. How to successfully recycle English GPT-2 to make models for other languages , author=. 2020 , eprint=
2020
-
[28]
De-identification of patient notes with recurrent neural networks
Dernoncourt, Franck and Lee, Ji Young and Uzuner, Ozlem and Szolovits, Peter. De-identification of patient notes with recurrent neural networks. Journal of the American Medical Informatics Association (JAMIA)
-
[29]
The International FLAIRS Conference Proceedings , author=
De-identification of Emergency Medical Records in French: Survey and Comparison of State-of-the-Art Automated Systems , volume=. The International FLAIRS Conference Proceedings , author=. 2021 , month=. doi:10.32473/flairs.v34i1.128480 , abstractNote=
2021 doi
-
[30]
An Efficient Method for Deidentifying Protected Health Information in Chinese Electronic Health Records: Algorithm Development and Validation
Wang, Peng and Li, Yong and Yang, Liang and Li, Simin and Li, Linfeng and Zhao, Zehan and Long, Shaopei and Wang, Fei and Wang, Hongqian and Li, Ying and Wang, Chengliang. An Efficient Method for Deidentifying Protected Health Information in Chinese Electronic Health Records: ...
-
[32]
Enhancing text anonymization via re-identification risk-based explainability , journal =
Benet Manzanares-Salor and David Sánchez , keywords =. Enhancing text anonymization via re-identification risk-based explainability , journal =. 2025 , issn =. doi:https://doi.org/10.1016/j.knosys.2024.112945 , url =
2025 doi
-
[33]
C-sanitized: A privacy model for document redaction and sanitization , year =
S\'. C-sanitized: A privacy model for document redaction and sanitization , year =. J. Assoc. Inf. Sci. Technol. , month = jan, pages =. doi:10.1002/asi.23363 , abstract =
-
[34]
Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe
Yue, Xiang and Inan, Huseyin and Li, Xuechen and Kumar, Girish and McAnallen, Julia and Shajari, Hoda and Sun, Huan and Levitan, David and Sim, Robert. Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe. Proceedings of the 61st Annual Meeting of...
2023 doi
-
[35]
International Conference on Learning Representations , year=
Differentially Private Fine-tuning of Language Models , author=. International Conference on Learning Representations , year=
-
[36]
Edward J Hu and yelong shen and Phillip Wallis and Zeyuan Allen-Zhu and Yuanzhi Li and Shean Wang and Lu Wang and Weizhu Chen , booktitle=. Lo. 2022 , url=
2022
-
[37]
ACM Trans
Liu, Xiao-Yang and Zhu, Rongyi and Zha, Daochen and Gao, Jiechao and Zhong, Shan and White, Matt and Qiu, Meikang , title =. ACM Trans. Manage. Inf. Syst. , month = aug, keywords =. 2024 , publisher =. doi:10.1145/3682068 , abstract =
2024 doi
-
[38]
arXiv preprint arXiv:2209.09631 , year=
De-identification of French unstructured clinical notes for machine learning tasks , author=. arXiv preprint arXiv:2209.09631 , year=
-
[39]
arXiv preprint arXiv:2507.19396 , year=
Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study , author=. arXiv preprint arXiv:2507.19396 , year=
-
[40]
Robust Utility-Preserving Text Anonymization Based on Large Language Models
Yang, Tianyu and Zhu, Xiaodan and Gurevych, Iryna. Robust Utility-Preserving Text Anonymization Based on Large Language Models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.1404
2025 doi
-
[41]
arXiv preprint arXiv:2303.08774 , year=
Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=
-
[42]
arXiv preprint arXiv:2501.12948 , year=
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , author=. arXiv preprint arXiv:2501.12948 , year=
-
[43]
arXiv e-prints , pages=
The llama 3 herd of models , author=. arXiv e-prints , pages=
-
[44]
arXiv preprint arXiv:2507.05201 , year=
Medgemma technical report , author=. arXiv preprint arXiv:2507.05201 , year=
-
[45]
, author=
De-identification of Personal Information:. , author=. 2015 , publisher=
2015
-
[46]
Health (San Francisco) , volume=
Simple demographics often identify people uniquely , author=. Health (San Francisco) , volume=
-
[47]
Mamma Mia! Where ' s My Name? De-Identifying I talian Clinical Notes with Large Language Models
Miranda, Michele and Brati \`e res, S \'e bastien and Patarnello, Stefano and Lilli, Livia. Mamma Mia! Where ' s My Name? De-Identifying I talian Clinical Notes with Large Language Models. Proceedings of the Eleventh Italian Conference on Computational Linguistics (CLiC-it 2025). 2025
2025
-
[48]
Proceedings of the First Workshop on Writing Aids at the Crossroads of AI, Cognitive Science and NLP (WRAICOGS 2025). 2025
2025
-
[49]
Chain-of- M eta W riting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts
Buhnila, Ioana and Cislaru, Georgeta and Todirascu, Amalia. Chain-of- M eta W riting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts. 2025
2025
-
[50]
Semantic Masking in a Needle-in-a-haystack Test for Evaluating Large Language Model Long-Text Capabilities
Shi, Ken and Penn, Gerald. Semantic Masking in a Needle-in-a-haystack Test for Evaluating Large Language Model Long-Text Capabilities. 2025
2025
-
[51]
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
Khallaf, Nouran and Eugeni, Carlo and Sharoff, Serge. Reading Between the Lines: A dataset and a study on why some texts are tougher than others. 2025
2025
-
[52]
P ara R ev : Building a dataset for Scientific Paragraph Revision annotated with revision instruction
Jourdan, L \'e ane and Boudin, Florian and Dufour, Richard and Hernandez, Nicolas and Aizawa, Akiko. P ara R ev : Building a dataset for Scientific Paragraph Revision annotated with revision instruction. 2025
2025
-
[53]
Towards an operative definition of creative writing: a preliminary assessment of creativeness in AI and human texts
Maggi, Chiara and Vitaletti, Andrea. Towards an operative definition of creative writing: a preliminary assessment of creativeness in AI and human texts. 2025
2025
-
[54]
Decoding Semantic Representations in the Brain Under Language Stimuli with Large Language Models
Sato, Anna and Kobayashi, Ichiro. Decoding Semantic Representations in the Brain Under Language Stimuli with Large Language Models. 2025
2025
-
[55]
Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4). 2025
2025
-
[56]
A rabic S ense: A Benchmark for Evaluating Commonsense Reasoning in A rabic with Large Language Models
Lamsiyah, Salima and Zeinalipour, Kamyar and El amrany, Samir and Brust, Matthias and Maggini, Marco and Bouvry, Pascal and Schommer, Christoph. A rabic S ense: A Benchmark for Evaluating Commonsense Reasoning in A rabic with Large Language Models. 2025
2025
-
[57]
Lahjawi: A rabic Cross-Dialect Translator
Hamed, Mohamed Motasim and Hreden, Muhammad and Hennara, Khalil and Aldallal, Zeina and Chrouf, Sara and AlModhayan, Safwan. Lahjawi: A rabic Cross-Dialect Translator. 2025
2025
-
[58]
Lost in Variation: An Unsupervised Methodology for Mining Lexico-syntactic Patterns in Middle A rabic Texts
Bezan. Lost in Variation: An Unsupervised Methodology for Mining Lexico-syntactic Patterns in Middle A rabic Texts. 2025
2025
-
[59]
SADSL y C : A Corpus for Saudi A rabian Multi-dialect Identification through Song Lyrics
Alahmari, Salwa Saad. SADSL y C : A Corpus for Saudi A rabian Multi-dialect Identification through Song Lyrics. 2025
2025
-
[60]
Enhancing Dialectal A rabic Intent Detection through Cross-Dialect Multilingual Input Augmentation
Hossain, Shehenaz and Shammary, Fouad and Shammary, Bahaulddin and Afli, Haithem. Enhancing Dialectal A rabic Intent Detection through Cross-Dialect Multilingual Input Augmentation. 2025
2025
-
[61]
D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic
Khered, Abdullah and Benkhedda, Youcef and Batista-Navarro, Riza. D ial2 MSA -Verified: A Multi-Dialect A rabic Social Media Dataset for Neural Machine Translation to M odern S tandard A rabic. 2025
2025
-
[62]
Web-Based Corpus Compilation of the Emirati A rabic Dialect
El-Ghawi, Yousra A. Web-Based Corpus Compilation of the Emirati A rabic Dialect. 2025
2025
-
[63]
Evaluating Calibration of A rabic Pre-trained Language Models on Dialectal Text
Al-Laith, Ali and Kebdani, Rachida. Evaluating Calibration of A rabic Pre-trained Language Models on Dialectal Text. 2025
2025
-
[64]
Empirical Evaluation of Pre-trained Language Models for Summarizing M oroccan D arija News Articles
Aftiss, Azzedine and Lamsiyah, Salima and Schommer, Christoph and El Alaoui, Said Ouatik. Empirical Evaluation of Pre-trained Language Models for Summarizing M oroccan D arija News Articles. 2025
2025
-
[65]
D ialect2 SQL : A Novel Text-to- SQL Dataset for A rabic Dialects with a Focus on M oroccan D arija
Chafik, Salmane and Ezzini, Saad and Berrada, Ismail. D ialect2 SQL : A Novel Text-to- SQL Dataset for A rabic Dialects with a Focus on M oroccan D arija. 2025
2025
-
[66]
A ra S im: Optimizing A rabic Dialect Translation in Children`s Literature with LLM s and Similarity Scores
Bouomar, Alaa Hassan and Abbas, Noorhan. A ra S im: Optimizing A rabic Dialect Translation in Children`s Literature with LLM s and Similarity Scores. 2025
2025
-
[67]
Navigating Dialectal Bias and Ethical Complexities in L evantine A rabic Hate Speech Detection
Haj Ahmed, Ahmed and Yew, Rui-Jie and Minocher, Xerxes and Venkatasubramanian, Suresh. Navigating Dialectal Bias and Ethical Complexities in L evantine A rabic Hate Speech Detection. 2025
2025
-
[68]
Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects. 2025
2025
-
[69]
Findings of the V ar D ial Evaluation Campaign 2025: The N or SID Shared Task on N orwegian Slot, Intent and Dialect Identification
Scherrer, Yves and van der Goot, Rob and M hlum, Petter. Findings of the V ar D ial Evaluation Campaign 2025: The N or SID Shared Task on N orwegian Slot, Intent and Dialect Identification. 2025
2025
-
[70]
Information Theory and Linguistic Variation: A Study of B razilian and E uropean P ortuguese
Alves, Diego. Information Theory and Linguistic Variation: A Study of B razilian and E uropean P ortuguese. 2025
2025
-
[71]
Leveraging Open-Source Large Language Models for Native Language Identification
Ng, Yee Man and Markov, Ilia. Leveraging Open-Source Large Language Models for Native Language Identification. 2025
2025
-
[72]
and Tayyar Madabushi, Harish
Torgbi, Melissa and Clayman, Andrew and Speight, Jordan J. and Tayyar Madabushi, Harish. Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the U nited K ingdom. 2025
2025
-
[73]
Large Language Models as a Normalizer for Transliteration and Dialectal Translation
Alam, Md Mahfuz Ibn and Anastasopoulos, Antonios. Large Language Models as a Normalizer for Transliteration and Dialectal Translation. 2025
2025
-
[74]
Testing the Boundaries of LLM s: Dialectal and Language-Variety Tasks
Faisal, Fahim and Anastasopoulos, Antonios. Testing the Boundaries of LLM s: Dialectal and Language-Variety Tasks. 2025
2025
-
[75]
Text Generation Models for L uxembourgish with Limited Data: A Balanced Multilingual Strategy
Plum, Alistair and Ranasinghe, Tharindu and Purschke, Christoph. Text Generation Models for L uxembourgish with Limited Data: A Balanced Multilingual Strategy. 2025
2025
-
[76]
Retrieval of Parallelizable Texts Across C hurch S lavic Variants
Lendvai, Piroska and Reichel, Uwe and Jouravel, Anna and Rabus, Achim and Renje, Elena. Retrieval of Parallelizable Texts Across C hurch S lavic Variants. 2025
2025
-
[77]
Neural Text Normalization for L uxembourgish Using Real-Life Variation Data
Lutgen, Anne-Marie and Plum, Alistair and Purschke, Christoph and Plank, Barbara. Neural Text Normalization for L uxembourgish Using Real-Life Variation Data. 2025
2025
-
[78]
Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal B avarian Case Study
Kr. Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal B avarian Case Study. 2025
2025
-
[79]
Regional Distribution of the /el/-/ l/ Merger in A ustralian E nglish
Coats, Steven and Diskin-Holdaway, Chlo \'e and Loakes, Debbie. Regional Distribution of the /el/-/ l/ Merger in A ustralian E nglish. 2025
2025
-
[80]
Learning Cross-Dialectal Morphophonology with Syllable Structure Constraints
Khalifa, Salam and Qaddoumi, Abdelrahim and Kodner, Jordan and Rambow, Owen. Learning Cross-Dialectal Morphophonology with Syllable Structure Constraints. 2025
2025
-
[81]
and Riabi, Arij and Seddah, Djam \'e
Lopetegui, Javier A. and Riabi, Arij and Seddah, Djam \'e. Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in S panish Varieties. 2025
2025
-
[82]
Add Noise, Tasks, or Layers? M ai NLP at the V ar D ial 2025 Shared Task on N orwegian Dialectal Slot and Intent Detection
Blaschke, Verena and K. Add Noise, Tasks, or Layers? M ai NLP at the V ar D ial 2025 Shared Task on N orwegian Dialectal Slot and Intent Detection. 2025
2025
-
[83]
LTG at V ar D ial 2025 N or SID : More and Better Training Data for Slot and Intent Detection
Midtgaard, Marthe and M hlum, Petter and Scherrer, Yves. LTG at V ar D ial 2025 N or SID : More and Better Training Data for Slot and Intent Detection. 2025
2025
-
[84]
H i TZ at V ar D ial 2025 N or SID : Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation
Bengoetxea, Jaione and Zubillaga, Mikel and Azurmendi, Ekhi and Heredia, Maite and Etxaniz, Julen and Ferro, Markel and Barnes, Jeremy. H i TZ at V ar D ial 2025 N or SID : Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation. 2025
2025
-
[85]
CUFE @ V ar D ial 2025 N or SID : Multilingual BERT for N orwegian Dialect Identification and Intent Detection
Ibrahim, Michael. CUFE @ V ar D ial 2025 N or SID : Multilingual BERT for N orwegian Dialect Identification and Intent Detection. 2025
2025
-
[86]
Dolomites: Domain-Specific Long-Form Methodical Tasks
Malaviya, Chaitanya and Agrawal, Priyanka and Ganchev, Kuzman and Srinivasan, Pranesh and Huot, Fantine and Berant, Jonathan and Yatskar, Mark and Das, Dipanjan and Lapata, Mirella and Alberti, Chris. Dolomites: Domain-Specific Long-Form Methodical Tasks. Transactions of the A...
2025 doi
-
[87]
Nguyen, Tu Anh and Muller, Benjamin and Yu, Bokai and Costa-jussa, Marta R. and Elbayad, Maha and Popuri, Sravya and Ropers, Christophe and Duquenne, Paul-Ambroise and Algayres, Robin and Mavlyutov, Ruslan and Gat, Itai and Williamson, Mary and Synnaeve, Gabriel and Pino, Juan...
2025 doi
-
[88]
CLAP nq: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
Rosenthal, Sara and Sil, Avirup and Florian, Radu and Roukos, Salim. CLAP nq: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems. Transactions of the Association for Computational Linguistics. 2025. doi:10.1162/tacl_a_00729
2025 doi
-
[89]
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
Pang, Jianhui and Ye, Fanghua and Wong, Derek Fai and Yu, Dian and Shi, Shuming and Tu, Zhaopeng and Wang, Longyue. Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models. Transactions of the Association for Computational Linguisti...
2025 doi
-
[90]
Investigating Critical Period Effects in Language Acquisition through Neural Language Models
Constantinescu, Ionut and Pimentel, Tiago and Cotterell, Ryan and Warstadt, Alex. Investigating Critical Period Effects in Language Acquisition through Neural Language Models. Transactions of the Association for Computational Linguistics. 2025. doi:10.1162/tacl_a_00725
2025 doi
-
[91]
and Goyal, Navin and Tsvetkov, Yulia
Ahuja, Kabir and Balachandran, Vidhisha and Panwar, Madhur and He, Tianxing and Smith, Noah A. and Goyal, Navin and Tsvetkov, Yulia. Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers. Transactions of the Association for Computati...
2025 doi
-
[92]
Proceedings of the Second Workshop on Scaling Up Multilingual & Multi-Cultural Evaluation. 2025
2025
-
[93]
The First Multilingual Model For The Detection of Suicide Texts
Zevallos, Rodolfo Joel and Schoene, Annika Marie and Ortega, John E. The First Multilingual Model For The Detection of Suicide Texts. 2025
2025
-
[94]
C ross I n: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment
Lin, Geyu and Wang, Bin and Liu, Zhengyuan and Chen, Nancy F. C ross I n: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment. 2025
2025
-
[95]
Evaluating Dialect Robustness of Language Models via Conversation Understanding
Srirag, Dipankar and Sahoo, Nihar Ranjan and Joshi, Aditya. Evaluating Dialect Robustness of Language Models via Conversation Understanding. 2025
2025
-
[96]
Cross-Lingual Document Recommendations with Transformer-Based Representations: Evaluating Multilingual Models and Mapping Techniques
Tashu, Tsegaye Misikir and Kontos, Eduard-Raul and Sabatelli, Matthia and Valdenegro-Toro, Matias. Cross-Lingual Document Recommendations with Transformer-Based Representations: Evaluating Multilingual Models and Mapping Techniques. 2025
2025
-
[97]
VRCP : Vocabulary Replacement Continued Pretraining for Efficient Multilingual Language Models
Nozaki, Yuta and Nakashima, Dai and Sato, Ryo and Asaba, Naoki and Kawamura, Shintaro. VRCP : Vocabulary Replacement Continued Pretraining for Efficient Multilingual Language Models. 2025
2025
-
[98]
Proceedings of the Second Workshop in South East Asian Language Processing. 2025
2025
-
[99]
and Estuar, Maria Regina Justina E
Bernardo, Jacob Simon D. and Estuar, Maria Regina Justina E. b AI -b AI : A Context-Aware Transliteration System for Baybayin Scripts. 2025
2025
-
[100]
N usa BERT : Teaching I ndo BERT to be Multilingual and Multicultural
Wongso, Wilson and Setiawan, David Samuel and Limcorn, Steven and Joyoadikusumo, Ananto. N usa BERT : Teaching I ndo BERT to be Multilingual and Multicultural. 2025
2025
-
[101]
Evaluating Sampling Strategies for Similarity-Based Short Answer Scoring: a Case Study in T hailand
Boonsarngsuk, Pachara and Arpanantikul, Pacharapon and Hiranwipas, Supakorn and Watcharakajorn, Wipu and Chuangsuwanich, Ekapol. Evaluating Sampling Strategies for Similarity-Based Short Answer Scoring: a Case Study in T hailand. 2025
2025
-
[102]
T hai W inograd Schemas: A Benchmark for T hai Commonsense Reasoning
Artkaew, Phakphum. T hai W inograd Schemas: A Benchmark for T hai Commonsense Reasoning. 2025
2025
-
[103]
Anak Baik: A Low-Cost Approach to Curate I ndonesian Ethical and Unethical Instructions
Hakim, Sulthan Abiyyu and Perdana, Rizal Setya and Fatyanosa, Tirana Noor. Anak Baik: A Low-Cost Approach to Curate I ndonesian Ethical and Unethical Instructions. 2025
2025
-
[104]
I ndonesian Speech Content De-Identification in Low Resource Transcripts
Abdjul, Rifqi Naufal and Puji Lestari, Dessi and Purwarianti, Ayu and Mawalim, Candy Olivia and Sakti, Sakriani and Unoki, Masashi. I ndonesian Speech Content De-Identification in Low Resource Transcripts. 2025
2025
-
[105]
I ndo M orph: a Morphology Engine for I ndonesian
Kamajaya, Ian and Moeljadi, David. I ndo M orph: a Morphology Engine for I ndonesian. 2025
2025
-
[106]
N usa D ialogue: Dialogue Summarization and Generation for Underrepresented and Extremely Low-Resource Languages
Purwarianti, Ayu and Adhista, Dea and Baptiso, Agung and Mahfuzh, Miftahul and Sabila, Yusrina and Adila, Aulia and Cahyawijaya, Samuel and Aji, Alham Fikri. N usa D ialogue: Dialogue Summarization and Generation for Underrepresented and Extremely Low-Resource Languages. 2025
2025
-
[107]
Proceedings of the 1st Regulatory NLP Workshop (RegNLP 2025). 2025
2025
-
[108]
Shared Task RIRAG -2025: Regulatory Information Retrieval and Answer Generation
Gokhan, Tuba and Wang, Kexin and Gurevych, Iryna and Briscoe, Ted. Shared Task RIRAG -2025: Regulatory Information Retrieval and Answer Generation. 2025
2025
-
[109]
Challenges in Technical Regulatory Text Variation Detection
Chikati, Shriya Vaagdevi and Larkin, Samuel and Minicola, David and Lo, Chi-kiu. Challenges in Technical Regulatory Text Variation Detection. 2025
2025
-
[110]
Bilingual BSARD : Extending Statutory Article Retrieval to D utch
Lotfi, Ehsan and Banar, Nikolay and Yuzbashyan, Nerses and Daelemans, Walter. Bilingual BSARD : Extending Statutory Article Retrieval to D utch. 2025
2025
-
[111]
Unifying Large Language Models and Knowledge Graphs for efficient Regulatory Information Retrieval and Answer Generation
Vanapalli, Kishore and Kilaru, Aravind and Shafiq, Omair and Khan, Shahzad. Unifying Large Language Models and Knowledge Graphs for efficient Regulatory Information Retrieval and Answer Generation. 2025
2025
-
[112]
A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts
Rayo Mosquera, Jhon Stewar and De La Rosa Peredo, Carlos Raul and Garrido Cordoba, Mario. A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts. 2025
2025
-
[113]
1-800- SHARED - TASKS at R eg NLP : Lexical Reranking of Semantic Retrieval ( L e S e R ) for Regulatory Question Answering
Purbey, Jebish and Sharma, Drishti and Gupta, Siddhant and Murad, Khawaja and Pullakhandam, Siddartha and Kadiyala, Ram Mohan Rao. 1-800- SHARED - TASKS at R eg NLP : Lexical Reranking of Semantic Retrieval ( L e S e R ) for Regulatory Question Answering. 2025
2025
-
[114]
MST - R : Multi-Stage Tuning for Retrieval Systems and Metric Evaluation
Malviya, Yash and Dhingra, Karan and Singh, Maneesh. MST - R : Multi-Stage Tuning for Retrieval Systems and Metric Evaluation. 2025
2025
-
[115]
and Androutsopoulos, Ion
Chasandras, Ioannis and Chlapanis, Odysseas S. and Androutsopoulos, Ion. AUEB -Archimedes at RIRAG -2025: Is Obligation concatenation really all you need?. 2025
2025
-
[116]
Structured Tender Entities Extraction from Complex Tables with Few-short Learning
Abbas, Asim and Lee, Mark and Shanavas, Niloofer and Kovatchev, Venelin and Ali, Mubashir. Structured Tender Entities Extraction from Complex Tables with Few-short Learning. 2025
2025
-
[117]
A Two-Stage LLM System for Enhanced Regulatory Information Retrieval and Answer Generation
Sun, Fengzhao and Yu, Jun and Hou, Jiaming and Lin, Yutong and Liu, Tianyu. A Two-Stage LLM System for Enhanced Regulatory Information Retrieval and Answer Generation. 2025
2025
-
[118]
NUST Nova at RIRAG 2025: A Hybrid Framework for Regulatory Information Retrieval and Question Answering
Khan, Mariam Babar and Ameer, Huma and Latif, Seemab and Fatima, Mehwish. NUST Nova at RIRAG 2025: A Hybrid Framework for Regulatory Information Retrieval and Question Answering. 2025
2025
-
[119]
NUST Alpha at RIRAG 2025: Fusion RAG for Bridging Lexical and Semantic Retrieval and Question Answering
Faisal, Muhammad Rouhan and Abdullah, Muhammad and Shah, Faizyaab Ali and Riaz, Shalina and Ameer, Huma and Latif, Seemab and Fatima, Mehwish. NUST Alpha at RIRAG 2025: Fusion RAG for Bridging Lexical and Semantic Retrieval and Question Answering. 2025
2025
-
[120]
NUST Omega at RIRAG 2025: Investigating Context-aware Retrieval and Answer Generations-Lessons and Challenges
Ameer, Huma and Akram, Muhammad Hannan and Latif, Seemab and Fatima, Mehwish. NUST Omega at RIRAG 2025: Investigating Context-aware Retrieval and Answer Generations-Lessons and Challenges. 2025
2025
-
[121]
Enhancing Regulatory Compliance Through Automated Retrieval, Reranking, and Answer Generation
Umar, K. Enhancing Regulatory Compliance Through Automated Retrieval, Reranking, and Answer Generation. 2025
2025
-
[122]
A REGNLP Framework: Developing Retrieval-Augmented Generation for Regulatory Document Analysis
Bayer, Ozan and Ulu, Elif Nehir and Sark. A REGNLP Framework: Developing Retrieval-Augmented Generation for Regulatory Document Analysis. 2025
2025
-
[123]
and Yousfi, Iman and Pudota, Nirmala and Bhattacharya, Sanmitra
Quinn, Devin and Pai, Sumit P. and Yousfi, Iman and Pudota, Nirmala and Bhattacharya, Sanmitra. Regulatory Question-Answering using Generative AI. 2025
2025
-
[124]
RIRAG : A Bi-Directional Retrieval-Enhanced Framework for Financial Legal QA in O bli QA Shared Task
Zhang, Xinyan and Feng, Xiaobing and Xu, Xiujuan and Zheng, Zhiliang and Wu, Kai. RIRAG : A Bi-Directional Retrieval-Enhanced Framework for Financial Legal QA in O bli QA Shared Task. 2025
2025
-
[125]
RAG ulator: Effective RAG for Regulatory Question Answering
Aushev, Islam and Kratkov, Egor and Nikolaev, Evgenii and Glinskii, Andrei and Krikunov, Vasilii and Panchenko, Alexander and Konovalov, Vasily and Belikova, Julia. RAG ulator: Effective RAG for Regulatory Question Answering. 2025
2025
-
[126]
Proceedings of Bridging Neurons and Symbols for Natural Language Processing and Knowledge Graphs Reasoning @ COLING 2025. 2025
2025
-
[127]
Chain of Knowledge Graph: Information-Preserving Multi-Document Summarization for Noisy Documents
Lee, Kangil and Jang, Jinwoo and Lim, Youngjin and Shin, Minsu. Chain of Knowledge Graph: Information-Preserving Multi-Document Summarization for Noisy Documents. 2025
2025
-
[128]
CEGRL - TKGR : A Causal Enhanced Graph Representation Learning Framework for Temporal Knowledge Graph Reasoning
Sun, Jinze and Sheng, Yongpan and He, Lirong and Qin, Yongbin and Liu, Ming and Jia, Tao. CEGRL - TKGR : A Causal Enhanced Graph Representation Learning Framework for Temporal Knowledge Graph Reasoning. 2025
2025
-
[129]
Reasoning Knowledge Filter for Logical Table-to-Text Generation
Bai, Yu and Liu, Baoqiang and Xue, Shuang and Cai, Fang and Ye, Na and Zhang, Guiping. Reasoning Knowledge Filter for Logical Table-to-Text Generation. 2025
2025
-
[130]
From Chain to Tree: Refining Chain-like Rules into Tree-like Rules on Knowledge Graphs
Sun, Wangtao and He, Shizhu and Zhao, Jun and Liu, Kang. From Chain to Tree: Refining Chain-like Rules into Tree-like Rules on Knowledge Graphs. 2025
2025
-
[131]
LAB - KG : A Retrieval-Augmented Generation Method with Knowledge Graphs for Medical Lab Test Interpretation
Guo, Rui and Devereux, Barry and Farnan, Greg and McLaughlin, Niall. LAB - KG : A Retrieval-Augmented Generation Method with Knowledge Graphs for Medical Lab Test Interpretation. 2025
2025
-
[132]
Bridging Language and Scenes through Explicit 3- D Model Construction
Dong, Tiansi and Das, Writwick and Sifa, Rafet. Bridging Language and Scenes through Explicit 3- D Model Construction. 2025
2025
-
[133]
VCRMNER : Visual Cue Refinement in Multimodal NER using CLIP Prompts
Bai, Yu and Wang, Lianji and Liu, Xiang and Chi, Haifeng and Zhang, Guiping. VCRMNER : Visual Cue Refinement in Multimodal NER using CLIP Prompts. 2025
2025
-
[134]
Neuro-Conceptual Artificial Intelligence: Integrating OPM with Deep Learning to Enhance Question Answering Quality
Kang, Xin and Shteyngardt, Veronika and Wang, Yuhan and Dori, Dov. Neuro-Conceptual Artificial Intelligence: Integrating OPM with Deep Learning to Enhance Question Answering Quality. 2025
2025
-
[135]
Emergence of symbolic abstraction heads for in-context learning in large language models
Al-Saeedi, Ali and Harma, Aki. Emergence of symbolic abstraction heads for in-context learning in large language models. 2025
2025
-
[136]
Linking language model predictions to human behaviour on scalar implicatures
Zinova, Yulia and Arps, David and Spalek, Katharina and Romoli, Jacopo. Linking language model predictions to human behaviour on scalar implicatures. 2025
2025
-
[137]
Generative F rame N et: Scalable and Adaptive Frames for Interpretable Knowledge Storage and Retrieval for LLM s Powered by LLM s
Tayyar Madabushi, Harish and Hudson, Taylor and Bonial, Claire. Generative F rame N et: Scalable and Adaptive Frames for Interpretable Knowledge Storage and Retrieval for LLM s Powered by LLM s. 2025
2025
-
[138]
Proceedings of the first International Workshop on Nakba Narratives as Language Resources. 2025
2025
-
[139]
Deciphering Implicatures: On NLP and Oral Testimonies
Sabra, Zainab. Deciphering Implicatures: On NLP and Oral Testimonies. 2025
2025
-
[140]
A cultural shift in Western perceptions of P alestine
Regier, Terry and Khalidi, Muhammad Ali. A cultural shift in Western perceptions of P alestine. 2025
2025
-
[141]
and Castle, Rick and Chappell, Carissa and Schoinoplokaki, Emmanouela and Seet, Allene M
Lamar, Annie K. and Castle, Rick and Chappell, Carissa and Schoinoplokaki, Emmanouela and Seet, Allene M. and Shilo, Amit and Nahas, Chloe. Cognitive Geographies of Catastrophe Narratives: Georeferenced Interview Transcriptions as Language Resource for Models of Forced Displac...
2025
-
[142]
Sentiment Analysis of Nakba Oral Histories: A Critical Study of Large Language Models
Ashqar, Huthaifa I. Sentiment Analysis of Nakba Oral Histories: A Critical Study of Large Language Models. 2025
2025
-
[143]
The Nakba Lexicon: Building a Comprehensive Dataset from Palestinian Literature
AbuHaija, Izza and Al Mandhari, Salim and El-Haj, Mo and Sibony, Jonas and Rayson, Paul. The Nakba Lexicon: Building a Comprehensive Dataset from Palestinian Literature. 2025
2025
-
[144]
A rabic Topic Classification Corpus of the Nakba Short Stories
Hamed, Osama and Zaidkilani, Nadeem. A rabic Topic Classification Corpus of the Nakba Short Stories. 2025
2025
-
[145]
Exploring Author Style in Nakba Short Stories: A Comparative Study of Transformer-Based Models
Hamed, Osama and Zaidkilani, Nadeem. Exploring Author Style in Nakba Short Stories: A Comparative Study of Transformer-Based Models. 2025
2025
-
[146]
Detecting Inconsistencies in Narrative Elements of Cross Lingual Nakba Texts
Hamarsheh, Nada and Elabour, Zahia and Murra, Aya and Yahya, Adnan. Detecting Inconsistencies in Narrative Elements of Cross Lingual Nakba Texts. 2025
2025
-
[147]
Multilingual Propaganda Detection: Exploring Transformer-Based Models m BERT , XLM - R o BERT a, and m T 5
Ragab, Mohamed Ibrahim and Mohamed, Ensaf Hussein and Medhat, Walaa. Multilingual Propaganda Detection: Exploring Transformer-Based Models m BERT , XLM - R o BERT a, and m T 5. 2025
2025
-
[148]
and Rayan, Tamara N
Awad, Ghadir A. and Rayan, Tamara N. and Dunagan, Lavinia and Gamba, David. Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in L ebanon. 2025
2025
-
[149]
The Missing Cause: An Analysis of Causal Attributions in Reporting on P alestine
Garcia Corral, Paulina and Bechara, Hannah and Manohara, Krishnamoorthy and Jankin, Slava. The Missing Cause: An Analysis of Causal Attributions in Reporting on P alestine. 2025
2025
-
[150]
Bias Detection in Media: Traditional Models vs
Mohammed, Marryam Yahya and Mohamed, Esraa Ismail and Esmat, Mariam Nabil and Nagib, Yomna Ashraf and Radwan, Nada Ahmed and Elshaer, Ziad Mohamed and Mohamed, Ensaf Hussein. Bias Detection in Media: Traditional Models vs. Transformers in Analyzing Social Media Coverage of the...
2025
-
[151]
N akba TR : A T urkish NER Dataset for Nakba Narratives
Bilgin Tasdemir, Esma Fat. N akba TR : A T urkish NER Dataset for Nakba Narratives. 2025
2025
-
[152]
Integrating Argumentation Features for Enhanced Propaganda Detection in A rabic Narratives on the Israeli War on G aza
Nabhani, Sara and Borg, Claudia and Micallef, Kurt and Al-Khatib, Khalid. Integrating Argumentation Features for Enhanced Propaganda Detection in A rabic Narratives on the Israeli War on G aza. 2025
2025
-
[153]
Proceedings of the First Workshop on Multilingual Counterspeech Generation. 2025
2025
-
[154]
PANDA - Paired Anti-hate Narratives Dataset from A sia: Using an LLM -as-a-Judge to Create the First C hinese Counterspeech Dataset
Bennie, Michael and Zhang, Demi and Xiao, Bushi and Cao, Jing and Liu, Chryseis Xinyi and Meng, Jian and Tripp, Alayo. PANDA - Paired Anti-hate Narratives Dataset from A sia: Using an LLM -as-a-Judge to Create the First C hinese Counterspeech Dataset. 2025
2025
-
[155]
RSSN at Multilingual Counterspeech Generation: Leveraging Lightweight Transformers for Efficient and Context-Aware Counter-Narrative Generation
V, Ravindran. RSSN at Multilingual Counterspeech Generation: Leveraging Lightweight Transformers for Efficient and Context-Aware Counter-Narrative Generation. 2025
2025
-
[156]
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
Wadhwa, Sahil and Xu, Chengtian and Chen, Haoming and Mahalingam, Aakash and Kar, Akankshya and Chaudhary, Divya. Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization. 2025
2025
-
[157]
NLP @ IIMAS - CLTL at Multilingual Counterspeech Generation: Combating Hate Speech Using Contextualized Knowledge Graph Representations and LLM s
Preciado M \'a rquez, David Salvador and G \'o mez Adorno, Helena and Markov, Ilia and Baez Santamaria, Selene. NLP @ IIMAS - CLTL at Multilingual Counterspeech Generation: Combating Hate Speech Using Contextualized Knowledge Graph Representations and LLM s. 2025
2025
-
[158]
CODEOFCONDUCT at Multilingual Counterspeech Generation: A Context-Aware Model for Robust Counterspeech Generation in Low-Resource Languages
Bennie, Michael and Xiao, Bushi and Liu, Chryseis Xinyi and Zhang, Demi and Meng, Jian and Tripp, Alayo. CODEOFCONDUCT at Multilingual Counterspeech Generation: A Context-Aware Model for Robust Counterspeech Generation in Low-Resource Languages. 2025
2025
-
[159]
HW - TSC at Multilingual Counterspeech Generation
Lyu, Xinglin and Wang, Haolin and Zhang, Min and Yang, Hao. HW - TSC at Multilingual Counterspeech Generation. 2025
2025
-
[160]
MilaNLP @Multilingual Counterspeech Generation: Evaluating Translation and Background Knowledge Filtering
Moscato, Emanuele and Muti, Arianna and Nozza, Debora. MilaNLP @Multilingual Counterspeech Generation: Evaluating Translation and Background Knowledge Filtering. 2025
2025
-
[161]
Hyderabadi Pearls at Multilingual Counterspeech Generation : HALT : Hate Speech Alleviation using Large Language Models and Transformers
Farhan, Md Shariq. Hyderabadi Pearls at Multilingual Counterspeech Generation : HALT : Hate Speech Alleviation using Large Language Models and Transformers. 2025
2025
-
[162]
T ren T eam at Multilingual Counterspeech Generation: Multilingual Passage Re-Ranking Approaches for Knowledge-Driven Counterspeech Generation Against Hate
Russo, Daniel. T ren T eam at Multilingual Counterspeech Generation: Multilingual Passage Re-Ranking Approaches for Knowledge-Driven Counterspeech Generation Against Hate. 2025
2025
-
[163]
The First Workshop on Multilingual Counterspeech Generation at COLING 2025: Overview of the Shared Task
Bonaldi, Helena and Vallecillo-Rodr \'i guez, Mar \'i a Estrella and Zubiaga, Irune and Montejo-Raez, Arturo and Soroa, Aitor and Mart \'i n-Valdivia, Mar \'i a-Teresa and Guerini, Marco and Agerri, Rodrigo. The First Workshop on Multilingual Counterspeech Generation at COLING...
2025
-
[164]
Proceedings of the First Workshop on Language Models for Low-Resource Languages. 2025
2025
-
[165]
Overview of the First Workshop on Language Models for Low-Resource Languages ( L o R es LM 2025)
Hettiarachchi, Hansi and Ranasinghe, Tharindu and Rayson, Paul and Mitkov, Ruslan and Gaber, Mohamed and Premasiri, Damith and Tan, Fiona Anting and Uyangodage, Lasitha Randunu Chandrakantha. Overview of the First Workshop on Language Models for Low-Resource Languages ( L o R ...
2025
-
[166]
Atlas-Chat: Adapting Large Language Models for Low-Resource M oroccan A rabic Dialect
Shang, Guokan and Abdine, Hadi and Khoubrane, Yousef and Mohamed, Amr and Abbahaddou, Yassine and Ennadir, Sofiane and Momayiz, Imane and Ren, Xuguang and Moulines, Eric and Nakov, Preslav and Vazirgiannis, Michalis and Xing, Eric. Atlas-Chat: Adapting Large Language Models fo...
2025
-
[167]
Empowering P ersian LLM s for Instruction Following: A Novel Dataset and Training Approach
Mokhtarabadi, Hojjat and Zamani, Ziba and Maazallahi, Abbas and Manshaei, Mohammad Hossein. Empowering P ersian LLM s for Instruction Following: A Novel Dataset and Training Approach. 2025
2025
-
[168]
B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis
Alam, Sadia and Ishmam, Md Farhan and Alvee, Navid Hasin and Siddique, Md Shahnewaz and Hossain, Md Azam and Kamal, Abu Raihan Mostofa. B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis. 2025
2025
-
[169]
Using Language Models for assessment of users' satisfaction with their partner in P ersian
Habibzadeh, Zahra and Asadpour, Masoud. Using Language Models for assessment of users' satisfaction with their partner in P ersian. 2025
2025
-
[170]
Enhancing Plagiarism Detection in M arathi with a Weighted Ensemble of TF - IDF and BERT Embeddings for Low-Resource Language Processing
Mutsaddi, Atharva and Choudhary, Aditya Prashant. Enhancing Plagiarism Detection in M arathi with a Weighted Ensemble of TF - IDF and BERT Embeddings for Low-Resource Language Processing. 2025
2025
-
[171]
Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in H ausa Language Using A fri BERT a
Sani, Sani Abdullahi and Muhammad, Shamsuddeen Hassan and Jarvis, Devon. Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in H ausa Language Using A fri BERT a. 2025
2025
-
[172]
and Gipp, Bela
Zhukova, Anastasia and Matt, Christian E. and Gipp, Bela. Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language. 2025
2025
-
[173]
F ilipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from S outheast A sia
Gamboa, Lance Calvin Lim and Lee, Mark. F ilipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from S outheast A sia. 2025
2025
-
[174]
Exploiting Word Sense Disambiguation in Large Language Models for Machine Translation
Tran, Van-Hien and Dabre, Raj and Kaing, Hour and Song, Haiyue and Tanaka, Hideki and Utiyama, Masao. Exploiting Word Sense Disambiguation in Large Language Models for Machine Translation. 2025
2025
-
[175]
Low-Resource Interlinear Translation: Morphology-Enhanced Neural Models for A ncient G reek
Rapacz, Maciej and Smywi \'n ski-Pohl, Aleksander. Low-Resource Interlinear Translation: Morphology-Enhanced Neural Models for A ncient G reek. 2025
2025
-
[176]
Language ver Y Rare for All
Merad, Ibrahim and Wolf, Amos and Mazzawi, Ziad and L \'e o, Yannick. Language ver Y Rare for All. 2025
2025
-
[177]
and Doh, Joon Young and Rodan, Eid and Zhu, Kevin and O ' Brien, Sean
Donthi, Sundesh and Spencer, Maximilian and Patel, Om B. and Doh, Joon Young and Rodan, Eid and Zhu, Kevin and O ' Brien, Sean. Improving LLM Abilities in Idiomatic Translation. 2025
2025
-
[178]
A Comparative Study of Static and Contextual Embeddings for Analyzing Semantic Changes in Medieval L atin Charters
Liu, Yifan and Tilahun, Gelila and Gao, Xinxiang and Wen, Qianfeng and Gervers, Michael. A Comparative Study of Static and Contextual Embeddings for Analyzing Semantic Changes in Medieval L atin Charters. 2025
2025
-
[179]
Bridging Literacy Gaps in A frican Informal Business Management with Low-Resource Conversational Agents
Ouattara, Maimouna and Kabor \'e , Abdoul Kader and Klein, Jacques and Bissyand \'e , Tegawend \'e F. Bridging Literacy Gaps in A frican Informal Business Management with Low-Resource Conversational Agents. 2025
2025
-
[180]
Social Bias in Large Language Models For B angla: An Empirical Study on Gender and Religious Bias
Sadhu, Jayanta and Saha, Maneesha Rani and Shahriyar, Rifat. Social Bias in Large Language Models For B angla: An Empirical Study on Gender and Religious Bias. 2025
2025
-
[181]
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
Cruz, Jan Christian Blaise. Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation. 2025
2025
-
[182]
Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models
Bagheri Nezhad, Sina and Agrawal, Ameeta and Pokharel, Rhitabrat. Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models. 2025
2025
-
[183]
B aby LM s for isi X hosa: Data-Efficient Language Modelling in a Low-Resource Context
Matzopoulos, Alexis and Hendriks, Charl and Mahomed, Hishaam and Meyer, Francois. B aby LM s for isi X hosa: Data-Efficient Language Modelling in a Low-Resource Context. 2025
2025
-
[184]
Mapping Cross-Lingual Sentence Representations for Low-Resource Language Pairs Using Pre-trained Language Models
Tudor, Andreea Ioana and Tashu, Tsegaye Misikir. Mapping Cross-Lingual Sentence Representations for Low-Resource Language Pairs Using Pre-trained Language Models. 2025
2025
-
[185]
How to age BERT Well: Continuous Training for Historical Language Adaptation
Harju, Anika and van der Goot, Rob. How to age BERT Well: Continuous Training for Historical Language Adaptation. 2025
2025
-
[186]
Exploiting Task Reversibility of DRS Parsing and Generation: Challenges and Insights from a Multi-lingual Perspective
Amin, Muhammad Saad and Anselma, Luca and Mazzei, Alessandro. Exploiting Task Reversibility of DRS Parsing and Generation: Challenges and Insights from a Multi-lingual Perspective. 2025
2025
-
[187]
BBPOS : BERT -based Part-of-Speech Tagging for U zbek
Bobojonova, Latofat and Akhundjanova, Arofat and Ostheimer, Phil Sidney and Fellenz, Sophie. BBPOS : BERT -based Part-of-Speech Tagging for U zbek. 2025
2025
-
[188]
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
Dewangan, Vikrant and S, Bharath Raj and Suri, Garvit and Sonavane, Raghav. When Every Token Counts: Optimal Segmentation for Low-Resource Language Models. 2025
2025
-
[189]
Recent Advancements and Challenges of T urkic C entral A sian Language Processing
Veitsman, Yana and Hartmann, Mareike. Recent Advancements and Challenges of T urkic C entral A sian Language Processing. 2025
2025
-
[190]
C a LQ uest
Lasheras, Uriel Anderson and Pinheiro, Vladia. C a LQ uest. PT : Towards the Collection and Evaluation of Natural Causal Ladder Questions in P ortuguese for AI Agents. 2025
2025
-
[191]
P ersian MCQ -Instruct: A Comprehensive Resource for Generating Multiple-Choice Questions in P ersian
Zeinalipour, Kamyar and Jamshidi, Neda and Akbari, Fahimeh and Maggini, Marco and Bianchini, Monica and Gori, Marco. P ersian MCQ -Instruct: A Comprehensive Resource for Generating Multiple-Choice Questions in P ersian. 2025
2025
-
[192]
Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss
Turumtaev, Galim. Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss. 2025
2025
-
[193]
and Alsehibani, Arwa and Qandos, Nour and Elshehy, Omar and Abdelkader, Mohamed and Koubaa, Anis
Nacar, Omer and Sibaee, Serry Taiseer and Ahmed, Samar and Ben Atitallah, Safa and Ammar, Adel and Alhabashi, Yasser and Al-Batati, Abdulrahman S. and Alsehibani, Arwa and Qandos, Nour and Elshehy, Omar and Abdelkader, Mohamed and Koubaa, Anis. Towards Inclusive A rabic LLM s:...
2025
-
[194]
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models
Kryvosheieva, Daria and Levy, Roger. Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models. 2025
2025
-
[195]
Evaluating Large Language Models for In-Context Learning of Linguistic Patterns In Unseen Low Resource Languages
Zhu, Hongpu and Liang, Yuqi and Xu, Wenjing and Xu, Hongzhi. Evaluating Large Language Models for In-Context Learning of Linguistic Patterns In Unseen Low Resource Languages. 2025
2025
-
[196]
Next-Level C antonese-to- M andarin Translation: Fine-Tuning and Post-Processing with LLM s
Dai, Yuqian and Chan, Chun Fai and Wong, Ying Ki and Pun, Tsz Ho. Next-Level C antonese-to- M andarin Translation: Fine-Tuning and Post-Processing with LLM s. 2025
2025
-
[197]
When LLM s Struggle: Reference-less Translation Evaluation for Low-resource Languages
Sindhujan, Archchana and Kanojia, Diptesh and Orasan, Constantin and Qian, Shenbin. When LLM s Struggle: Reference-less Translation Evaluation for Low-resource Languages. 2025
2025
-
[198]
Does Machine Translation Impact Offensive Language Identification? The Case of I ndo- A ryan Languages
Dmonte, Alphaeus and Satapara, Shrey and Alsudais, Rehab and Ranasinghe, Tharindu and Zampieri, Marcos. Does Machine Translation Impact Offensive Language Identification? The Case of I ndo- A ryan Languages. 2025
2025
-
[199]
Maria and Sayed, Imaan and Van Der Leek, Alexander
Mahlaza, Zola and Keet, C. Maria and Sayed, Imaan and Van Der Leek, Alexander. I si Z ulu noun classification based on replicating the ensemble approach for R unyankore. 2025
2025
-
[200]
From A rabic Text to Puzzles: LLM -Driven Development of A rabic Educational Crosswords
Zeinalipour, Kamyar and Saad, Moahmmad and Maggini, Marco and Gori, Marco. From A rabic Text to Puzzles: LLM -Driven Development of A rabic Educational Crosswords. 2025
2025
-
[201]
Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages. 2025
2025
-
[202]
H indi Reading Comprehension: Do Large Language Models Exhibit Semantic Understanding?
Lal, Daisy Monika and Rayson, Paul and El-Haj, Mo. H indi Reading Comprehension: Do Large Language Models Exhibit Semantic Understanding?. 2025
2025
-
[203]
Machine Translation and Transliteration for I ndo- A ryan Languages: A Systematic Review
Perera, Sandun Sameera and Sumanathilaka, Deshan Koshala. Machine Translation and Transliteration for I ndo- A ryan Languages: A Systematic Review. 2025
2025
-
[204]
BERT opic for Topic Modeling of H indi Short Texts: A Comparative Study
Mutsaddi, Atharva and Jamkhande, Anvi and Thakre, Aryan Shirish and Haribhakta, Yashodhara. BERT opic for Topic Modeling of H indi Short Texts: A Comparative Study. 2025
2025
-
[205]
Evaluating Structural and Linguistic Quality in U rdu DRS Parsing and Generation through Bidirectional Evaluation
Amin, Muhammad Saad and Anselma, Luca and Mazzei, Alessandro. Evaluating Structural and Linguistic Quality in U rdu DRS Parsing and Generation through Bidirectional Evaluation. 2025
2025
-
[206]
Studying the Effect of H indi Tokenizer Performance on Downstream Tasks
Goel, Rashi and Sadat, Fatiha. Studying the Effect of H indi Tokenizer Performance on Downstream Tasks. 2025
2025
-
[207]
Adapting Multilingual LLM s to Low-Resource Languages using Continued Pre-training and Synthetic Corpus: A Case Study for H indi LLM s
Joshi, Raviraj and Singla, Kanishk and Kamath, Anusha and Kalani, Raunak and Paul, Rakesh and Vaidya, Utkarsh and Chauhan, Sanjay Singh and Wartikar, Niranjan and Long, Eileen. Adapting Multilingual LLM s to Low-Resource Languages using Continued Pre-training and Synthetic Cor...
2025
-
[208]
OVQA : A Dataset for Visual Question Answering and Multimodal Research in O dia Language
Parida, Shantipriya and Sahoo, Shashikanta and Sekhar, Sambit and Sahoo, Kalyanamalini and Kotwal, Ketan and Khosla, Sonal and Dash, Satya Ranjan and Bose, Aneesh and Kohli, Guneet Singh and Lenka, Smruti Smita and Bojar, Ond r ej. OVQA : A Dataset for Visual Question Answerin...
2025
-
[209]
Advancing Multilingual Speaker Identification and Verification for I ndo- A ryan and D ravidian Languages
Sritharan, Braveenan and Thayasivam, Uthayasanker. Advancing Multilingual Speaker Identification and Verification for I ndo- A ryan and D ravidian Languages. 2025
2025
-
[210]
Sentiment Analysis of S inhala News Comments Using Transformers
Bandaranayake, Isuru and Usoof, Hakim. Sentiment Analysis of S inhala News Comments Using Transformers. 2025
2025
-
[211]
E x M ute: A Context-Enriched Multimodal Dataset for Hateful Memes
Debnath, Riddhiman Swanan and Firuj, Nahian Beente and Shakib, Abdul Wadud and Sultana, Sadia and Islam, Md Saiful. E x M ute: A Context-Enriched Multimodal Dataset for Hateful Memes. 2025
2025
-
[212]
Studying the capabilities of Large Language Models in solving Combinatorics Problems posed in H indi
Kumar, Yash and Roy, Subhajit. Studying the capabilities of Large Language Models in solving Combinatorics Problems posed in H indi. 2025
2025
-
[213]
Sumon and Sami, Nasrullah and Chowdhury, Mahruba Sharmin and Islam, Md Saiful
Shibu, Hrithik Majumdar and Datta, Shrestha and Miah, Md. Sumon and Sami, Nasrullah and Chowdhury, Mahruba Sharmin and Islam, Md Saiful. From Scarcity to Capability: Empowering Fake News Detection in Low-Resource Languages with LLM s. 2025
2025
-
[214]
Enhancing Participatory Development Research in S outh A sia through LLM Agents System: An Empirically-Grounded Methodological Initiative from Field Evidence in S ri L ankan
Zhao, Xinjie and Wang, Hao and Sriwarnasinghe, Shyaman Maduranga and Tang, Jiacheng and Wang, Shiyun and Sugiyama, Sayaka and Morikawa, So. Enhancing Participatory Development Research in S outh A sia through LLM Agents System: An Empirically-Grounded Methodological Initiative...
2025
-
[215]
Identifying Aggression and Offensive Language in Code-Mixed Tweets: A Multi-Task Transfer Learning Approach
Kancharla, Bharath and Singh, Prabhjot and Kancharla, Lohith Bhagavan and Chama, Yashita and Sharma, Raksha. Identifying Aggression and Offensive Language in Code-Mixed Tweets: A Multi-Task Transfer Learning Approach. 2025
2025
-
[216]
Team I ndi D ata M iner at I ndo NLP 2025: H indi Back Transliteration - R oman to D evanagari using LL a M a
Kumar, Saurabh and Kakadiya, Dhruvkumar Babubhai and Singh, Sanasam Ranbir. Team I ndi D ata M iner at I ndo NLP 2025: H indi Back Transliteration - R oman to D evanagari using LL a M a. 2025
2025
-
[217]
I ndo NLP 2025 Shared Task: R omanized S inhala to S inhala Reverse Transliteration Using BERT
Perera, Sandun Sameera and Jayakodi, Lahiru Prabhath and Sumanathilaka, Deshan Koshala and Anuradha, Isuri. I ndo NLP 2025 Shared Task: R omanized S inhala to S inhala Reverse Transliteration Using BERT. 2025
2025
-
[218]
Crossing Language Boundaries: Evaluation of Large Language Models on U rdu- E nglish Question Answering
Kazi, Samreen and Rahim, Maria and Khoja, Shakeel Ahmed. Crossing Language Boundaries: Evaluation of Large Language Models on U rdu- E nglish Question Answering. 2025
2025
-
[219]
Investigating the Effect of Backtranslation for I ndic Languages
Das, Sudhansu Bala and Choudhury, Samujjal and Mishra, Dr Tapas Kumar and Patra, Dr Bidyut Kr. Investigating the Effect of Backtranslation for I ndic Languages. 2025
2025
-
[220]
S inhala Transliteration: A Comparative Analysis Between Rule-based and S eq2 S eq Approaches
De Mel, Yomal and Wickramasinghe, Kasun and de Silva, Nisansa and Ranathunga, Surangika. S inhala Transliteration: A Comparative Analysis Between Rule-based and S eq2 S eq Approaches. 2025
2025
-
[221]
and Sherly, Elizabeth
Baiju, Bajiyo and Manohar, Kavya and Pillai, Leena G. and Sherly, Elizabeth. R omanized to Native M alayalam Script Transliteration Using an Encoder-Decoder Framework. 2025
2025
-
[222]
Proceedings of the Workshop on Generative AI and Knowledge Graphs (GenAIK). 2025
2025
-
[223]
Effective Modeling of Generative Framework for Document-level Relational Triple Extraction
Saini, Pratik and Nayak, Tapas. Effective Modeling of Generative Framework for Document-level Relational Triple Extraction. 2025
2025
-
[224]
Learn Together: Joint Multitask Finetuning of Pretrained KG -enhanced LLM for Downstream Tasks
Martynova, Anastasia and Tishin, Vladislav and Semenova, Natalia. Learn Together: Joint Multitask Finetuning of Pretrained KG -enhanced LLM for Downstream Tasks. 2025
2025
-
[225]
GNET - QG : Graph Network for Multi-hop Question Generation
Jamshidi, Samin and Chali, Yllias. GNET - QG : Graph Network for Multi-hop Question Generation. 2025
2025
-
[226]
SKETCH : Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
Mahalingam, Aakash and Gande, Vinesh Kumar and Chadha, Aman and Jain, Vinija and Chaudhary, Divya. SKETCH : Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval. 2025
2025
-
[227]
On Reducing Factual Hallucinations in Graph-to-Text Generation Using Large Language Models
Iarosh, Dmitrii and Panchenko, Alexander and Salnikov, Mikhail. On Reducing Factual Hallucinations in Graph-to-Text Generation Using Large Language Models. 2025
2025
-
[228]
G raph RAG : Leveraging Graph-Based Efficiency to Minimize Hallucinations in LLM -Driven RAG for Finance Data
Barry, Mariam and Caillaut, Gaetan and Halftermeyer, Pierre and Qader, Raheel and Mouayad, Mehdi and Le Deit, Fabrice and Cariolaro, Dimitri and Gesnouin, Joseph. G raph RAG : Leveraging Graph-Based Efficiency to Minimize Hallucinations in LLM -Driven RAG for Finance Data. 2025
2025
-
[229]
Structured Knowledge meets G en AI : A Framework for Logic-Driven Language Models
Eldessouky, Farida Helmy and Ehab, Nourhan and Schindler, Carolin and Abuelkheir, Mervat and Minker, Wolfgang. Structured Knowledge meets G en AI : A Framework for Logic-Driven Language Models. 2025
2025
-
[230]
Performance and Limitations of Fine-Tuned LLM s in SPARQL Query Generation
Mecharnia, Thamer and d ' Aquin, Mathieu. Performance and Limitations of Fine-Tuned LLM s in SPARQL Query Generation. 2025
2025
-
[231]
Refining Noisy Knowledge Graph with Large Language Models
Dong, Na and Kertkeidkachorn, Natthawut and Liu, Xin and Shirai, Kiyoaki. Refining Noisy Knowledge Graph with Large Language Models. 2025
2025
-
[232]
Can LLM s be Knowledge Graph Curators for Validating Triple Insertions?
Regino, Andr \'e Gomes and dos Reis, Julio Cesar. Can LLM s be Knowledge Graph Curators for Validating Triple Insertions?. 2025
2025
-
[233]
T ext2 C ypher: Bridging Natural Language and Graph Databases
Ozsoy, Makbule Gulcin and Messallem, Leila and Besga, Jon and Minneci, Gianandrea. T ext2 C ypher: Bridging Natural Language and Graph Databases. 2025
2025
-
[234]
KGF ake N et: A Knowledge Graph-Enhanced Model for Fake News Detection
Kumar, Anuj and Kumar, Pardeep and Yadav, Abhishek and Ahlawat, Satyadev and Prasad, Yamuna. KGF ake N et: A Knowledge Graph-Enhanced Model for Fake News Detection. 2025
2025
-
[235]
Style Knowledge Graph: Augmenting Text Style Transfer with Knowledge Graphs
Toshevska, Martina and Kalajdziski, Slobodan and Gievska, Sonja. Style Knowledge Graph: Augmenting Text Style Transfer with Knowledge Graphs. 2025
2025
-
[236]
Entity Quality Enhancement in Knowledge Graphs through LLM -based Question Answering
Kamaladdini Ezzabady, Morteza and Benamara, Farah. Entity Quality Enhancement in Knowledge Graphs through LLM -based Question Answering. 2025
2025
-
[237]
Multilingual Skill Extraction for Job Vacancy -- Job Seeker Matching in Knowledge Graphs
Kavas, Hamit and Serra-Vidal, Marc and Wanner, Leo. Multilingual Skill Extraction for Job Vacancy -- Job Seeker Matching in Knowledge Graphs. 2025
2025
-
[238]
Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect). 2025
2025
-
[239]
S ilver S peak: Evading AI -Generated Text Detectors using Homoglyphs
Creo, Aldan and Pudasaini, Shushanta. S ilver S peak: Evading AI -Generated Text Detectors using Homoglyphs. 2025
2025
-
[240]
Human vs
Moe ner, Philipp and Adel, Heike. Human vs. AI : A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts. 2025
2025
-
[241]
Mirror Minds : An Empirical Study on Detecting LLM -Generated Text via LLM s
Baradia, Josh and Gupta, Shubham and Kundu, Suman. Mirror Minds : An Empirical Study on Detecting LLM -Generated Text via LLM s. 2025
2025
-
[242]
Benchmarking AI Text Detection: Assessing Detectors Against New Datasets, Evasion Tactics, and Enhanced LLM s
Pudasaini, Shushanta and Miralles, Luis and Lillis, David and Salvador, Marisa Llorens. Benchmarking AI Text Detection: Assessing Detectors Against New Datasets, Evasion Tactics, and Enhanced LLM s. 2025
2025
-
[243]
Charbel N
Kindji, G. Charbel N. and Rojas Barahona, Lina M. and Fromont, Elisa and Urvoy, Tanguy. Cross-table Synthetic Tabular Data Detection. 2025
2025
-
[244]
Your Large Language Models are Leaving Fingerprints
McGovern, Hope Elizabeth and Stureborg, Rickard and Suhara, Yoshi and Alikaniotis, Dimitris. Your Large Language Models are Leaving Fingerprints. 2025
2025
-
[245]
and Taylor, Sydney and Bergen, Benjamin and Jones, Cameron
Rathi, Ishika M. and Taylor, Sydney and Bergen, Benjamin and Jones, Cameron. GPT -4 is Judged More Human than Humans in Displaced and Inverted T uring Tests. 2025
2025
-
[246]
and Schwartz, H
Varadarajan, Vasudha and Giorgi, Salvatore and Mangalik, Siddharth and Soni, Nikita and Markowitz, Dave M. and Schwartz, H. Andrew. The Consistent Lack of Variance of Psychological Factors Expressed by LLM s and Spambots. 2025
2025
-
[247]
and Spero, Max
Masrour, Elyas and Emi, Bradley N. and Spero, Max. DAMAGE : Detecting Adversarially Modified AI Generated Text. 2025
2025
-
[248]
Text Graph Neural Networks for Detecting AI -Generated Content
Valdez-Valenzuela, Andric and G \'o mez-Adorno, Helena and Montes-y-G \'o mez, Manuel. Text Graph Neural Networks for Detecting AI -Generated Content. 2025
2025
-
[249]
I Know You Did Not Write That! A Sampling Based Watermarking Method for Identifying Machine Generated Text
Kele. I Know You Did Not Write That! A Sampling Based Watermarking Method for Identifying Machine Generated Text. 2025
2025
-
[250]
DCBU at G en AI Detection Task 1: Enhancing Machine-Generated Text Detection with Semantic and Probabilistic Features
Zhang, Zhaowen and Chen, Songhao and Liu, Bingquan. DCBU at G en AI Detection Task 1: Enhancing Machine-Generated Text Detection with Semantic and Probabilistic Features. 2025
2025
-
[251]
L3i++ at G en AI Detection Task 1: Can Label-Supervised LL a MA Detect Machine-Generated Text?
Tran, Hanh Thi Hong and Nam, Nguyen Tien. L3i++ at G en AI Detection Task 1: Can Label-Supervised LL a MA Detect Machine-Generated Text?. 2025
2025
-
[252]
T ech E xperts( IPN ) at G en AI Detection Task 1: Detecting AI -Generated Text in E nglish and Multilingual Contexts
Mehak, Gull and Qasim, Amna and Meque, Abdul Gafar Manuel and Hussain, Nisar and Sidorov, Grigori and Gelbukh, Alexander. T ech E xperts( IPN ) at G en AI Detection Task 1: Detecting AI -Generated Text in E nglish and Multilingual Contexts. 2025
2025
-
[253]
S zeged AI at G en AI Detection Task 1: Beyond Binary - Soft-Voting Multi-Class Classification for Binary Machine-Generated Text Detection Across Diverse Language Models
Kiss, Mihaly and Berend, G \'a bor. S zeged AI at G en AI Detection Task 1: Beyond Binary - Soft-Voting Multi-Class Classification for Binary Machine-Generated Text Detection Across Diverse Language Models. 2025
2025
-
[254]
Team U nibuc - NLP at G en AI Detection Task 1: Qwen it detect machine-generated text?
Creanga, Claudiu and Marchitan, Teodor-George and Dinu, Liviu P. Team U nibuc - NLP at G en AI Detection Task 1: Qwen it detect machine-generated text?. 2025
2025
-
[255]
Fraunhofer SIT at G en AI Detection Task 1: Adapter Fusion for AI -generated Text Detection
Schaefer, Karla and Steinebach, Martin. Fraunhofer SIT at G en AI Detection Task 1: Adapter Fusion for AI -generated Text Detection. 2025
2025
-
[256]
OSINT at G en AI Detection Task 1: Multilingual MGT Detection: Leveraging Cross-Lingual Adaptation for Robust LLM s Text Identification
Agrahari, Shifali and Ranbir Singh, Sanasam. OSINT at G en AI Detection Task 1: Multilingual MGT Detection: Leveraging Cross-Lingual Adaptation for Robust LLM s Text Identification. 2025
2025
-
[257]
Nota AI at G en AI Detection Task 1: Unseen Language-Aware Detection System for Multilingual Machine-Generated Text
Park, Hancheol and Kim, Jaeyeon and Kim, Geonmin and Kim, Tae-Ho. Nota AI at G en AI Detection Task 1: Unseen Language-Aware Detection System for Multilingual Machine-Generated Text. 2025
2025
-
[258]
CNLP - NITS - PP at G en AI Detection Task 1: AI -Generated Text Using Transformer-Based Approaches
Yadagiri, Annepaka and Lekkala, Sai Teja and Vardhan, Mandadoddi Srikar and Pakray, Partha and Krishna, Reddi Mohana. CNLP - NITS - PP at G en AI Detection Task 1: AI -Generated Text Using Transformer-Based Approaches. 2025
2025
-
[259]
Kamrujjaman and Islam, Md Saiful
Mobin, MD. Kamrujjaman and Islam, Md Saiful. L ux V eri at G en AI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI -Generated Text across E nglish and Multilingual Contexts. 2025
2025
-
[260]
Grape at G en AI Detection Task 1: Leveraging Compact Models and Linguistic Features for Robust Machine-Generated Text Detection
Doan, Nhi Hoai and Inui, Kentaro. Grape at G en AI Detection Task 1: Leveraging Compact Models and Linguistic Features for Robust Machine-Generated Text Detection. 2025
2025
-
[261]
AAIG at G en AI Detection Task 1: Exploring Syntactically-Aware, Resource-Efficient Small Autoregressive Decoders for AI Content Detection
Bhandarkar, Avanti and Wilson, Ronald and Woodard, Damon. AAIG at G en AI Detection Task 1: Exploring Syntactically-Aware, Resource-Efficient Small Autoregressive Decoders for AI Content Detection. 2025
2025
-
[262]
T ur QU az at G en AI Detection Task 1:Dr
Kele s , Kaan Efe and Kutlu, Mucahid. T ur QU az at G en AI Detection Task 1:Dr. Perplexity or: How I Learned to Stop Worrying and Love the Finetuning. 2025
2025
-
[263]
AI -Monitors at G en AI Detection Task 1: Fast and Scalable Machine Generated Text Detection
Singh, Azad and Tripathi, Vishnu and Pandey, Ravindra Kumar and Saho, Pragyanand and Joshi, Prakhar and Mani, Neel and Alagh, Richa and Mishra, Pallaw and Arora, Piyush. AI -Monitors at G en AI Detection Task 1: Fast and Scalable Machine Generated Text Detection. 2025
2025
-
[264]
Advacheck at G en AI Detection Task 1: AI Detection Powered by Domain-Aware Multi-Tasking
Gritsai, German and Voznyuk, Anastasia and Khabutdinov, Ildar and Grabovoy, Andrey. Advacheck at G en AI Detection Task 1: AI Detection Powered by Domain-Aware Multi-Tasking. 2025
2025
-
[265]
G en AI Content Detection Task 1: E nglish and Multilingual Machine-Generated Text Detection: AI vs
Wang, Yuxia and Shelmanov, Artem and Mansurov, Jonibek and Tsvigun, Akim and Mikhailov, Vladislav and Xing, Rui and Xie, Zhuohan and Geng, Jiahui and Puccetti, Giovanni and Artemova, Ekaterina and Su, Jinyan and Ta, Minh Ngoc and Abassy, Mervat and Elozeiri, Kareem Ashraf and ...
2025
-
[266]
CIC - NLP at G en AI Detection Task 1: Advancing Multilingual Machine-Generated Text Detection
Abiola, Tolulope Olalekan and Bizuneh, Tewodros Achamaleh and Uroosa, Fatima and Hafeez, Nida and Sidorov, Grigori and Kolesnikova, Olga and Ojo, Olumide Ebenezer. CIC - NLP at G en AI Detection Task 1: Advancing Multilingual Machine-Generated Text Detection. 2025
2025
-
[267]
CIC - NLP at G en AI Detection Task 1: Leveraging D istil BERT for Detecting Machine-Generated Text in E nglish
Abiola, Tolulope Olalekan and Bizuneh, Tewodros Achamaleh and Abiola, Oluwatobi Joseph and Oladepo, Temitope Olasunkanmi and Ojo, Olumide Ebenezer and Sidorov, Grigori and Kolesnikova, Olga. CIC - NLP at G en AI Detection Task 1: Leveraging D istil BERT for Detecting Machine-G...
2025
-
[268]
nits \_ teja \_ srikar at G en AI Detection Task 2: Distinguishing Human and AI -Generated Essays Using Machine Learning and Transformer Models
Lekkala, Sai Teja and Yadagiri, Annepaka and Vardhan, Mangadoddi Srikar and Pakray, Partha. nits \_ teja \_ srikar at G en AI Detection Task 2: Distinguishing Human and AI -Generated Essays Using Machine Learning and Transformer Models. 2025
2025
-
[269]
I ntegrity AI at G en AI Detection Task 2: Detecting Machine-Generated Academic Essays in E nglish and A rabic Using ELECTRA and Stylometry
AL-Smadi, Mohammad. I ntegrity AI at G en AI Detection Task 2: Detecting Machine-Generated Academic Essays in E nglish and A rabic Using ELECTRA and Stylometry. 2025
2025
-
[270]
CMI - AIGCX at G en AI Detection Task 2: Leveraging Multilingual Proxy LLM s for Machine-Generated Text Detection in Academic Essays
Jiao, Kaijie and Yao, Xingyu and Ma, Shixuan and Fang, Sifan and Guo, Zikang and Xu, Benfeng and Zhang, Licheng and Wang, Quan and Zhang, Yongdong and Mao, Zhendong. CMI - AIGCX at G en AI Detection Task 2: Leveraging Multilingual Proxy LLM s for Machine-Generated Text Detecti...
2025
-
[271]
E ssay D etect at G en AI Detection Task 2: Guardians of Academic Integrity: Multilingual Detection of AI -Generated Essays
Agrahari, Shifali and Jayant, Subhashi and Kumar, Saurabh and Ranbir Singh, Sanasam. E ssay D etect at G en AI Detection Task 2: Guardians of Academic Integrity: Multilingual Detection of AI -Generated Essays. 2025
2025
-
[272]
CNLP - NITS - PP at G en AI Detection Task 2: Leveraging D istil BERT and XLM - R o BERT a for Multilingual AI -Generated Text Detection
Yadagiri, Annepaka and Krishna, Reddi Mohana and Pakray, Partha. CNLP - NITS - PP at G en AI Detection Task 2: Leveraging D istil BERT and XLM - R o BERT a for Multilingual AI -Generated Text Detection. 2025
2025
-
[273]
RA at G en AI Detection Task 2: Fine-tuned Language Models For Detection of Academic Authenticity, Results and Thoughts
Gharib, Rana and Elgendy, Ahmed. RA at G en AI Detection Task 2: Fine-tuned Language Models For Detection of Academic Authenticity, Results and Thoughts. 2025
2025
-
[274]
Tesla at G en AI Detection Task 2: Fast and Scalable Method for Detection of Academic Essay Authenticity
Indurthi, Vijayasaradhi and Varma, Vasudeva. Tesla at G en AI Detection Task 2: Fast and Scalable Method for Detection of Academic Essay Authenticity. 2025
2025
-
[275]
G en AI Content Detection Task 2: AI vs
Chowdhury, Shammur Absar and Almerekhi, Hind and Kutlu, Mucahid and Kele s , Kaan Efe and Ahmad, Fatema and Mohiuddin, Tasnim and Mikros, George and Alam, Firoj. G en AI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge. 2025
2025
-
[276]
CNLP - NITS - PP at G en AI Detection Task 3: Cross-Domain Machine-Generated Text Detection Using D istil BERT Techniques
Lekkala, Sai Teja and Yadagiri, Annepaka and Vardhan, Mangadoddi Srikar and Pakray, Partha. CNLP - NITS - PP at G en AI Detection Task 3: Cross-Domain Machine-Generated Text Detection Using D istil BERT Techniques. 2025
2025
-
[277]
and Katsios, Gregorios A
Edikala, Abishek R. and Katsios, Gregorios A. and Creaghe, Noelie and Yu, Ning. Leidos at G en AI Detection Task 3: A Weight-Balanced Transformer Approach for AI Generated Text Detection Across Domains. 2025
2025
-
[278]
and Spero, Max and Masrour, Elyas
Emi, Bradley N. and Spero, Max and Masrour, Elyas. Pangram at G en AI Detection Task 3: An Active Learning Approach to Machine-Generated Text Detection. 2025
2025
-
[279]
Kamrujjaman and Islam, Md Saiful
Mobin, MD. Kamrujjaman and Islam, Md Saiful. L ux V eri at G en AI Detection Task 3: Cross-Domain Detection of AI -Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models. 2025
2025
-
[280]
Kandula, Hemanth and Li, Chak Fai and Qiu, Haoling and Karakos, Damianos and Man, Hieu and Nguyen, Thien Huu and Ulicny, Brian. BBN - U . O regon`s ALERT system at G en AI Content Detection Task 3: Robust Authorship Style Representations for Cross-Domain Machine-Generated Text...
2025
-
[281]
Random at G en AI Detection Task 3: A Hybrid Approach to Cross-Domain Detection of Machine-Generated Text with Adversarial Attack Mitigation
Agrahari, Shifali and Mishra, Prabhat and Kumar, Sujit. Random at G en AI Detection Task 3: A Hybrid Approach to Cross-Domain Detection of Machine-Generated Text with Adversarial Attack Mitigation. 2025
2025
-
[282]
MOSAIC at GENAI Detection Task 3 : Zero-Shot Detection Using an Ensemble of Models
Dubois, Matthieu and Yvon, Fran c ois and Piantanida, Pablo. MOSAIC at GENAI Detection Task 3 : Zero-Shot Detection Using an Ensemble of Models. 2025
2025
-
[283]
G en AI Content Detection Task 3: Cross-Domain Machine Generated Text Detection Challenge
Dugan, Liam and Zhu, Andrew and Alam, Firoj and Nakov, Preslav and Apidianaki, Marianna and Callison-Burch, Chris. G en AI Content Detection Task 3: Cross-Domain Machine Generated Text Detection Challenge. 2025
2025
-
[284]
Proceedings of the Joint Workshop of the 9th Financial Technology and Natural Language Processing (FinNLP), the 6th Financial Narrative Processing (FNP), and the 1st Workshop on Large Language Models for Finance and Legal (LLMFinLegal). 2025
2025
-
[285]
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
Biancotti, Claudia and Camassa, Carolina and Coletta, Andrea and Giudice, Oliver and Glielmo, Aldo. Chat Bankman-Fried: an Exploration of LLM Alignment in Finance. 2025
2025
-
[286]
G raph RAG Analysis for Financial Narrative Summarization and A Framework for Optimizing Domain Adaptation
Shukla, Neelesh Kumar and Prabhakar, Prabhat and Thangaraj, Sakthivel and Singh, Sandeep and Sun, Weiyi and Venkatesan, C Prasanna and Krishnamurthy, Viji. G raph RAG Analysis for Financial Narrative Summarization and A Framework for Optimizing Domain Adaptation. 2025
2025
-
[287]
Wang, Dongsheng and Zmigrod, Ran and Sibue, Mathieu J. and Pei, Yulong and Babkin, Petr and Brugere, Ivan and Liu, Xiaomo and Navarro, Nacho and Papadimitriou, Antony and Watson, William and Ma, Zhiqiang and Nourbakhsh, Armineh and Shah, Sameena. B u DDIE : A Business Document...
2025
-
[288]
F in M o E : A M o E -based Large C hinese Financial Language Model
Zhang, Xuanyu and Yang, Qing. F in M o E : A M o E -based Large C hinese Financial Language Model. 2025
2025
-
[289]
Bridging the Gap: Efficient Cross-Lingual NER in Low-Resource Financial Domain
Kumar, Sunisth and ElKholy, Mohammed and Liu, Davide and Boulenger, Alexandre. Bridging the Gap: Efficient Cross-Lingual NER in Low-Resource Financial Domain. 2025
2025
-
[290]
Evaluating Financial Literacy of Large Language Models through Domain Specific Languages for Plain Text Accounting
Figueroa Rosero, Alexei Gustavo and Grundmann, Paul and Freidank, Julius and Nejdl, Wolfgang and Loeser, Alexander. Evaluating Financial Literacy of Large Language Models through Domain Specific Languages for Plain Text Accounting. 2025
2025
-
[291]
Synthetic Data Generation Using Large Language Models for Financial Question Answering
Harsha, Chetan and Phogat, Karmvir Singh and Dasaratha, Sridhar and Puranam, Sai Akhil and Ramakrishna, Shashishekar. Synthetic Data Generation Using Large Language Models for Financial Question Answering. 2025
2025
-
[292]
Concept-Based RAG Models: A High-Accuracy Fact Retrieval Approach
Lin, Cheng-Yu and Jang, Jyh-Shing. Concept-Based RAG Models: A High-Accuracy Fact Retrieval Approach. 2025
2025
-
[293]
Training L ayout LM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain
Uthayasooriyar, Benno and Ly, Antoine and Vermet, Franck and Corro, Caio. Training L ayout LM from Scratch for Efficient Named-Entity Recognition in the Insurance Domain. 2025
2025
-
[294]
A veni B ench: Accessible and Versatile Evaluation of Finance Intelligence
Klimaszewski, Mateusz and Chen, Pinzhen and Guillou, Liane and Papaioannou, Ioannis and Haddow, Barry and Birch, Alexandra. A veni B ench: Accessible and Versatile Evaluation of Finance Intelligence. 2025
2025
-
[295]
and Zohren, Stefan
Drinkall, Felix and Pierrehumbert, Janet B. and Zohren, Stefan. Forecasting Credit Ratings: A Case Study where Traditional Methods Outperform Generative LLM s. 2025
2025
-
[296]
Investigating the effectiveness of length based rewards in DPO for building Conversational Financial Question Answering Systems
Yadav, Anushka and Rallabandi, Sai Krishna and Dakle, Parag Pravin and Raghavan, Preethi. Investigating the effectiveness of length based rewards in DPO for building Conversational Financial Question Answering Systems. 2025
2025
-
[297]
C redit LLM : Constructing Financial AI Assistant for Credit Products using Financial LLM and Few Data
Yan, Sixing and Zhu, Ting. C redit LLM : Constructing Financial AI Assistant for Credit Products using Financial LLM and Few Data. 2025
2025
-
[298]
Modeling Interactions Between Stocks Using LLM -Enhanced Graphs for Volume Prediction
Xu, Zhiyu and Liu, Yi and Wang, Yuchi and Bao, Ruihan and Harimoto, Keiko and Sun, Xu. Modeling Interactions Between Stocks Using LLM -Enhanced Graphs for Volume Prediction. 2025
2025
-
[299]
Financial Named Entity Recognition: How Far Can LLM Go?
Lu, Yi-Te and Huo, Yintong. Financial Named Entity Recognition: How Far Can LLM Go?. 2025
2025
-
[300]
Proxy Tuning for Financial Sentiment Analysis: Overcoming Data Scarcity and Computational Barriers
Wang, Yuxiang and Wang, Yuchi and Liu, Yi and Bao, Ruihan and Harimoto, Keiko and Sun, Xu. Proxy Tuning for Financial Sentiment Analysis: Overcoming Data Scarcity and Computational Barriers. 2025
2025
Reviewed May 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.