REVIEW 4 major objections 3 minor 38 references
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
T0 review · 4 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read PalmX 2025, the first shared task for benchmarking LLMs on Arabic and Islamic culture, finds that task-specific fine-tuning substantially improves cultural QA, with top systems reaching 72.15% on Arabic culture and 84.22% on Islamic…
desk verdict Useful first benchmark for Arabic and Islamic cultural QA, but the fine-tuning claim is overstated because baseline comparisons mix model sizes and families. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the PalmX 2025 benchmark itself: two sets of multiple-choice questions in Modern Standard Arabic, one on general Arabic culture and one on Islamic culture, each with four options and one human-reviewed correct answer, split into training, dev, and held-out test sets. The mechanism that turns the benchmark into a repeatable measurement is a likelihood-based scoring protocol: for each item, a model receives the question and the four labeled options, and its answer is the label with the highest normalized log-probability as a continuation of the prompt. Participation rules require submitting model weights rather than retrieval systems, cap models at 13 billion parameters, and keep the test set private until evaluation, so the reported accuracies are meant to reflect internalized knowledge plus whatever fine-tuning the teams applied.
What would settle it
Re-annotate a random sample of, say, 300 test questions per subtask with new professional linguists who have not seen the official key; if agreement with the official answers is materially below perfect, or if model rankings change when only independently confirmed items are scored, the benchmark's central claim is weakened. A cheaper version is to isolate the items the original two reviewers disagreed about before consolidation and compare scores on that subset.
Extended reading notes
Core claim
The paper's central claim is that task-specific fine-tuning substantially improves performance over baseline models on both subtasks, as stated in its analysis of results. Against a zero-shot NileChat-3B baseline of 67.55% on general Arabic culture and 75.12% on Islamic culture, the top teams reached 72.15% and 84.22%, respectively, gains of 4.6 and 9.1 points. The same result pattern also supports a methodological sub-claim: parameter-efficient fine-tuning, especially LoRA, is the dominant and most effective strategy, while data augmentation helps in the Islamic domain but did not help the same team on cultural questions. The authors interpret the higher Islamic scores as evidence that Islamic knowledge has more structured, canonical answers than the broader cultural domain.
Load-bearing premise
The rankings all rest on one assumption: that the human-reviewed answer key is correct and unambiguous, and the authors themselves flag that LLM-generated or reformulated questions could carry subtle artifacts, so systematic errors in the key would shift every reported accuracy.
Editorial extensions
If this is right
- Fine-tuning on PalmX's human-reviewed training questions is a reliable route to improving Arabic cultural QA, lifting the top culture score 4.6 points and the top Islamic score 9.1 points over the zero-shot baseline.
- Smaller Arabic-centric models can win: the 3-billion-parameter NileChat-3B took first place in the culture subtask through fine-tuning, while larger 7B and 9B models finished close behind, suggesting scale is not the main lever.
- Parameter-efficient LoRA fine-tuning is competitive with full fine-tuning, so substantial cultural gains do not require retraining all weights.
- Data augmentation is domain-dependent: the AYA team found paraphrase augmentation decisive for Islamic questions but not helpful on the cultural development set.
- Because even winning systems answer only 72.15% of cultural and 84.22% of Islamic questions correctly, stable headroom remains for future culturally grounded Arabic LLMs.
Reading between the lines
- Beyond the paper, a country-stratified accuracy score would test whether the overall numbers hide weak performance on under-represented countries such as Iraq and Algeria, an imbalance the authors acknowledge.
- Beyond the paper, the label-only likelihood scorer may favor models with well-calibrated Arabic next-token distributions; a free-text generation variant is a direct robustness check.
- Beyond the paper, if the gains replicate in future rounds, the PalmX training split could serve as reusable instruction data for grounding Arabic LLMs culturally, not just as an evaluation set.
- Beyond the paper, the 'Islamic knowledge is more structured' explanation predicts that inter-annotator agreement and answer determinism should be measurably higher on Islamic questions; that is testable with the dataset's own review records.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents PalmX 2025, which the authors describe as the first shared task for benchmarking LLM cultural competence in Arabic and Islamic domains. The task consists of two MSA multiple-choice subtasks, one on general Arabic culture and one on general Islamic culture, with public training and development splits and a held-out test split. The paper describes the data collection pipeline, which combines Palm-derived items, web-crawled content, and LLM-assisted generation, followed by two-linguist human review; it also presents the likelihood-based evaluation protocol and the results of the nine and six valid submissions to the two subtasks. The central empirical claims are that task-specific fine-tuning substantially improves over zero-shot baselines, that LoRA-style parameter-efficient fine-tuning is the dominant and most effective approach, and that data augmentation helps in the Islamic domain but not in the general cultural domain. The authors report top accuracies of 72.15% for Subtask 1 and 84.22% for Subtask 2, and they release the dataset and evaluation code publicly.
Significance. If the benchmark data are of high quality, this is a useful community resource: it is the first standardized Arabic and Islamic cultural MCQ benchmark that I am aware of, it uses a transparent likelihood-based evaluation harness, it reports results from 15 independently developed systems, and it makes the data and evaluation code publicly available. The human-review design, the private test split, and the explicit accuracy metric are sensible choices, and the participation statistics indicate genuine community interest. The paper's strongest contribution is the resource itself and the documented baseline results. The more general methodological conclusion about task-specific fine-tuning is weaker than the wording suggests, because the baseline is a single small model and most submissions use different base models; the controlled comparisons requested below are needed before that conclusion can be treated as established.
major comments (4)
- [Section 5.1, Tables 4-5] The claim that 'task-specific fine-tuning significantly improves performance over baseline models' is underdetermined by the reported comparisons. The only baseline is NileChat-3B in zero-shot mode, while most submissions use different base models. In Table 5, the top two Islamic-subtask systems fine-tune ALLaM-7B-Instruct and no zero-shot ALLaM-7B score is reported, so the 9.1% improvement over the baseline could be attributable to base-model choice; moreover, two fine-tuned Qwen2.5 submissions (74.13% and 70.83%) fall below the NileChat-3B baseline. In Table 4, only the first- and second-place systems share the baseline base model, while MarsadLabM ties the baseline and Hamyaria and Star finish below it. Please add same-base zero-shot comparisons for ALLaM-7B and Fanar-9B, or explicitly restrict the conclusion to NileChat-3B and the first two Subtask-1 systems.
- [Section 3.1 / Table 4] The integrity of the evaluation is put in doubt by two entries. First, CultranAI's dataset column in Table 4 lists 'PalmX (test)' among the training sources, which directly contradicts the statement in Section 3.1 that the test set was private and held out; if this is literal, the submission should be disqualified or re-evaluated without the test split. Second, the ISL row describes a 'Retrieval-augmented (Gemini)' approach, and Section 5.4 praises their 'external knowledge retrieval,' which appears to conflict with the rule prohibiting systems with RAG or live internet access; please clarify whether retrieval was used only to construct training data and not at inference time. These points must be resolved because they bear directly on the validity of the leaderboard.
- [Sections 4.3 and 5.1] No uncertainty quantification or significance testing is provided for any of the accuracy differences. The top four Subtask-1 systems fall within 1.0 percentage point (72.15, 71.65, 71.45, 71.35) on a 2,000-item test, so the ordering may not be stable; for the 1,000-item Islamic test the gap between the top two is also small (84.22 vs. 83.82). The word 'significantly' in the Abstract and Section 5.1 should be supported by exact binomial confidence intervals or a pairwise significance test, or replaced by a weaker formulation.
- [Section 2.1.1 and Limitations] Because the benchmark's value depends on correct and unambiguous ground-truth answers, the human-review process should be documented quantitatively. The manuscript states that two linguists independently reviewed all questions and that discrepancies were consolidated, but it reports no inter-annotator agreement rate and no post-hoc audit of the final test set; the Limitations section acknowledges that 'small sections' of the dataset were LLM-generated or reformulated and could contain subtle artifacts. Please report the number and percentage of LLM-generated items in each split, the agreement between annotators, and the outcome of discrepancy resolution, so that readers can assess the risk of systematic answer-key errors.
minor comments (3)
- [Section 5.2] The statement that parameter-efficient fine-tuning was 'the predominant and most effective approach' is only partially supported: LoRA was predominant, but the overall winner of Subtask 1 used full fine-tuning, so 'most effective' should be softened or justified with a same-base comparison.
- [Section 3.1, footnote 9] The sentence 'Test data was shared only after the leaderboard announcement' is confusing, since evaluation requires the organizers to use the test set before announcing the leaderboard; please clarify whether this refers to the public release of the test set after the announcement.
- [Table 4] The MarsadLabM row is ranked 7th with the same accuracy (67.55%) as the baseline row; an explicit tie-breaking rule would help readers interpret the rankings.
Circularity Check
No circularity: the benchmark results are empirical measurements, not derivations from fitted inputs.
full rationale
The paper's central claim—that task-specific fine-tuning improves over baseline models—is an observed comparison of independently submitted systems on a held-out MCQ test set, scored by likelihood-based accuracy (Section 3.2, Eq. 1). The authors use their own Palm dataset to source some questions and NileChat-3B as the baseline, but these are data and model provenance choices, not circular reductions. The ground-truth answers were independently reviewed by professional linguists, the test split was held out, and submissions came from external teams. No equation defines the finding in terms of a fitted parameter, and no load-bearing uniqueness theorem or ansatz is imported via self-citation. The skeptic's concern that top systems in Subtask 2 use different base models (e.g., ALLaM-7B) than the NileChat baseline is a potential confound in the strength of the generalization claim, but confounding is a validity/correctness issue, not circularity. The paper reports a benchmark outcome rather than deriving a result from its own assumptions, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Multiple-choice questions in Modern Standard Arabic are a valid proxy for cultural competence of LLMs in the Arab world.
- domain assumption LLM-generated MCQs reviewed by two professional linguists are factually correct and culturally appropriate.
- domain assumption The test set stayed private during the competition, so reported results measure generalization rather than memorization.
- domain assumption The sources used (Palm, web pages, Islamic competitions, Mawdoo3) are representative of relevant Arabic and Islamic cultural knowledge.
Cite this review
Pith. "Pith review of PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture." pith.science (2026). https://pith.science/paper/4QNDO3M4
@misc{pith2026250902550,
author = {Pith},
title = {Pith review of: PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture},
year = {2026},
howpublished = {\url{https://pith.science/paper/4QNDO3M4}},
note = {Machine review of arXiv:2509.02550}
}
read the original abstract
Large Language Models (LLMs) inherently reflect the vast data distributions they encounter during their pre-training phase. As this data is predominantly sourced from the web, there is a high chance it will be skewed towards high-resourced languages and cultures, such as those of the West. Consequently, LLMs often exhibit a diminished understanding of certain communities, a gap that is particularly evident in their knowledge of Arabic and Islamic cultures. This issue becomes even more pronounced with increasingly under-represented topics. To address this critical challenge, we introduce PalmX 2025, the first shared task designed to benchmark the cultural competence of LLMs in these specific domains. The task is composed of two subtasks featuring multiple-choice questions (MCQs) in Modern Standard Arabic (MSA): General Arabic Culture and General Islamic Culture. These subtasks cover a wide range of topics, including traditions, food, history, religious practices, and language expressions from across 22 Arab countries. The initiative drew considerable interest, with 26 teams registering for Subtask 1 and 19 for Subtask 2, culminating in nine and six valid submissions, respectively. Our findings reveal that task-specific fine-tuning substantially boosts performance over baseline models. The top-performing systems achieved an accuracy of 72.15% on cultural questions and 84.22% on Islamic knowledge. Parameter-efficient fine-tuning emerged as the predominant and most effective approach among participants, while the utility of data augmentation was found to be domain-dependent.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Muhammad Abdul-Mageed, Abdelrahim Elmadany, Alcides Inciarte, Md Tawkat Islam Khondaker, and 1 others. 2023. Jasmine: Arabic gpt models for few-shot learning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 16721--16744
work page 2023
-
[4]
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, and Monojit Choudhury. 2024. https://arxiv.org/abs/2403.15412 Towards measuring and modeling" culture" in llms: A survey . arXiv preprint arXiv:2403.15412
arXiv 2024
-
[5]
Emran Al-Buraihy, Dan Wang, Tariq Hussain, Razaz Waheeb Attar, Ahmad Ali AlZubi, Khalid Zaman, and Zengkang Gan. 2025. https://www.nature.com/articles/s41598-025-02894-z Aratraditions10k bridging cultures with a comprehensive dataset for enhanced cross lingual image annotation retrieval and tagging . Scientific Reports, 15(1):19624
work page 2025
-
[6]
Walid Al-Dhabyani and Hamzah A. Alsayadi. 2025. Leveraging Large Language Models to Improve Arabic Multiple-Choice Questions in Cultural and Islamic Domains . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Computational Linguistics. Co-located with EMNLP 2025, November 5--9
work page 2025
-
[7]
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab. 2024. https://arxiv.org/abs/2402.13231 Investigating cultural alignment of large language models . arXiv preprint arXiv:2402.13231
arXiv 2024
-
[8]
Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, and 1 others. 2025 a . https://doi.org/10.18653/v1/2025.acl-long.1579 Palm: A culturally inclusive and linguistically diverse dataset for A rabic LLM s . In Proceedings of the 63rd Annual ...
Show all 38 references
-
[9]
Fakhraddin Alwajih, Samar Mohamed Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, and 1 others. 2025 b . https://arxiv.org/pdf/2505.21979 Pearl: A multimodal culturally-aware ...
2025
-
[10]
Fakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdelrahman Mohamed, and Muhammad Abdul-Mageed. 2024. https://doi.org/10.18653/v1/2024.acl-long.689 Peacock: A family of A rabic multimodal large language models and benchmarks . In Proceedings of the 62nd Annual Meet...
2024 doi
-
[11]
Zaid Alyafeai, Khalid Almubarak, Ahmed Ashraf, Deema Alnuhait, Saied Alshahrani, Gubran AQ Abdulrahman, Gamil Ahmed, Qais Gawah, Zead Saleh, Mustafa Ghaleb, and 1 others. 2024. https://arxiv.org/pdf/2402.03177 Cidar: Culturally relevant instruction dataset for arabic . arXiv p...
2024 arXiv
-
[12]
Houdaifa Atou, Issam Ait Yahia, and Ismail Berrada. 2025. Phoenix at Palmx: Exploring Data Augmentation for Arabic Cultural Question Answering . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Computati...
2025
-
[13]
Lama Ayash, Hassan Alhuzali, Ashwag Alasmari, and Sultan Aloufi. 2025. https://arxiv.org/abs/2503.17485 Saudiculture: A benchmark for evaluating large language models’ cultural competence within saudi arabia . Journal of King Saud University Computer and Information Sciences, ...
2025 arXiv
-
[14]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, and 1 others. 2022. https://arxiv.org/pdf/2204.05862 Training a helpful and harmless assistant with reinforcement learning from human feedbac...
2022 arXiv
-
[15]
M Saiful Bari, Yazeed Alnumay, Norah A Alzahrani, Nouf M Alotaibi, Hisham A Alyahya, Sultan AlRashed, Faisal A Mirza, Shaykhah Z Alsubaie, Hassan A Alahmed, Ghadah Alabduljabbar, and 1 others. 2024. https://arxiv.org/abs/2407.15390 Allam: Large language models for arabic and e...
2024 arXiv
-
[16]
Lee, Haonan Li, and 11 others
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal...
2024 arXiv
-
[17]
Rafiul Biswas, Shimaa Ibrahim, Kais Attia, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani
Md. Rafiul Biswas, Shimaa Ibrahim, Kais Attia, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani. 2025. MarsadLab at PalmX2025: An LLM Benchmark for Arabic Culture and Islamic Civilization . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNL...
2025
-
[18]
Pulkit Chatwal and Santosh Kumar Mishra. 2025. Cultura-Arabica: Probing and Enhancing Arabic Cultural Awareness in Large Language Models via LORA . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Comput...
2025
-
[19]
Rochelle Choenni and Ekaterina Shutova. 2024. https://arxiv.org/pdf/2408.16482 Self-alignment: Improving alignment of cultural values in llms via in-context learning . arXiv preprint arXiv:2408.16482
2024 arXiv
-
[20]
Eman Elrefai, Esraa Khaled, and Alhassan Ehab. 2025. Star at PalmX 2025: Arabic Cultural Understanding via Targeted Pretraining and Lightweight Fine-tuning . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association ...
2025
-
[21]
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, and Rao Muhammad Anwer. 2025. https://doi.org/10.18653/v1/2025.findings-naacl.105 CAMEL -bench: A comprehensive A rab...
2025 doi
-
[22]
Mohamed Gomaa and Noureldin Elmadany. 2025. Retrieval-Augmented Fine-Tuning for Arabic Cultural Question Answering . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Computational Linguistics. Co-located...
2025
-
[23]
Shehenaz Hossain and Haithem Afli. 2025. ADAPT–MTU HAI at PalmX 2025: Leveraging Full and Parameter‑Efficient LLM Fine‑Tuning for Arabic Cultural QA . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Com...
2025
-
[24]
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, and Jinchao Xu. 2024. https://doi.org/10.18653/v1/2024.na...
2024 doi
-
[25]
Rebecca L Johnson, Giada Pistilli, Natalia Men \'e dez-Gonz \'a lez, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022. https://arxiv.org/pdf/2203.07785 The ghost in the machine has an american accent: value conflict in gpt-3 . arXiv pre...
2022 arXiv
-
[26]
Cheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/9a16935bf54c4af233e25d998b7f4a2c-Paper-Conference.pdf Culturellm: Incorporating cultural differences into large language models . In Advances...
2024
-
[27]
Zhaoming Liu. 2025. https://doi.org/10.1515/jtc-2023-0019 Cultural bias in large language models: A comprehensive analysis and mitigation strategies . Journal of Transcultural Communication, 3(2):224--244
2025 doi
-
[28]
Samar Mohamed Magdy, Sang Yun Kwon, Fakhraddin Alwajih, Safaa Taher Abdelfadil, Shady Shehata, and Muhammad Abdul-Mageed. 2025. https://doi.org/10.18653/v1/2025.naacl-long.613 JAWAHER : A multidialectal dataset of A rabic proverbs for LLM benchmarking . In Proceedings of the 2...
2025 doi
-
[29]
Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, and Muhammad Abdul-Mageed. 2025. https://arxiv.org/pdf/2505.18383 Nilechat: Towards linguistically diverse and culturally aware llms for local communities . arXiv preprint arXiv:2505.18383
2025
-
[30]
Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam
Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam. 2025. https://aclanthology.org/2025.coling-main.283/ A ra D i CE : Benchmarks for dialectal and cultural capabilities in LLM s . In P...
2025
-
[31]
Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2023. https://arxiv.org/abs/2305.14456 Having beer after prayer? measuring cultural bias in large language models . arXiv preprint arXiv:2305.14456
2023 arXiv
-
[32]
Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, and Isabelle Augenstein. 2025. https://arxiv.org/abs/2411.00860 Survey of cultural awareness in language models: Text and beyond . Computational Lin...
2025 arXiv
-
[33]
Abdelrahman Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri, Farah Atif, Chatrine Qwaider, Karima Kadaoui, Sara Shatnawi, Yaser Alesh, and Fajri Koto. 2025. https://doi.org/10.18653/v1/2025.acl-long.380 Commonsense reasoning in A rab culture . In Proceedings of...
2025 doi
-
[34]
Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, and 1 others. 2023. https://arxiv.org/abs/2308.16149 Jais and jais-chat: Arabic-centric foundation and instruction-tuned open gen...
2023 arXiv
-
[35]
Jannatul Tajrin, Bir Ballav Roy, and Firoj Alam. 2025. AYA at PalmX 2025: Modeling Cultural and Islamic Knowledge in LLMs . In Proceedings of the Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou, China. Association for Computational Linguistics. Co-...
2025
-
[36]
Yan Tao, Olga Viberg, Ryan S Baker, and René F Kizilcec. 2024. https://doi.org/10.1093/pnasnexus/pgae346 Cultural bias and cultural alignment of large language models . PNAS Nexus, 3(9):pgae346
2024 doi
-
[37]
Fanar Team, Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, and 1 others. 2025. https://arxiv.org/pdf/2501.13944 Fanar: An arabic-centric multimodal generative ai platform ....
2025 arXiv
-
[38]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. https://arxiv.org/pdf/2505.09388 Qwen3 technical report . arXiv preprint arXiv:2505.09388
2025 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.