REVIEW 4 major objections 4 minor 52 references
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Fann or Flop, the first Arabic poetry benchmark spanning 12 eras and 14 genres, finds that state-of-the-art LLMs, strong on standard Arabic tasks, fall short on poetry's interpretive depth.
desk verdict Era-label error undermines the multiera claim, but the benchmark resource and task are real; fix labels and evaluation details. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is the benchmark itself: a taxonomy of twelve historical poetic eras and fourteen genres, matched to 6,984 curated poems each carrying a verse-by-verse explanation in formal Arabic. The explanations are the crux — they integrate literal meaning with figurative and rhetorical analysis (naming devices such as metaphor, simile, and paronomasia) and with cultural-historical context, which is what forces models to go beyond surface statistics. Around this corpus sits a multi-tier evaluation stack: BLEU and chrF(++) for lexical overlap, AraBERT-based BERTScore and mDeBERTaV3-based textual entailment for semantic alignment, GPT-4o as an automated judge of faithfulness and fluency, and a rubric-based human score for interpretive depth. The stack is designed so that the benchmark can separate fluent paraphrase from genuine understanding.
What would settle it
Have independent Arabic-literature experts re-annotate a random sample of the 6,984 pairs without seeing the gold material: assign era and genre labels, and write fresh verse explanations. If pairwise expert agreement falls below roughly 70 percent, the human references themselves are not stable enough to rank models, and the struggle finding would be measuring disagreement among humans as much as model failure. In the opposite direction, the claim would soften if an open-weight model scored at or above the human-reference level on interpretive depth in a blind human evaluation.
Extended reading notes
Core claim
The paper's central claim is that Fann or Flop exposes a real gap that standard Arabic benchmarks hide: most LLMs perform well on ordinary Arabic tasks but consistently fall short when asked to interpret Arabic poetry. On the benchmark's 6,984 poem-explanation pairs, fourteen state-of-the-art models are asked to produce verse-by-verse explanations in formal Arabic, scored against human-written references by lexical overlap (BLEU, chrF(++)), semantic similarity (BERTScore), textual entailment, an LLM judge, and human experts rating interpretive depth on a 0–10 rubric. The strongest model, GPT-4o, reaches about 0.64 BERTScore and about 7.5 for human-judged interpretive depth; several open and Arabic-centric models land well below that, and every model family scores lower on pre-Islamic and Umayyad verse than on modern poetry. The authors read this as evidence that current models have absorbed Modern Standard Arabic but not the layered, figurative, and historically embedded language of classical Arabic poetry, which is why they propose poetry comprehension as a sharper test of cultural and linguistic depth.
Load-bearing premise
The ranking stands on the gold explanations and the era and genre labels being trustworthy references, but the paper never reports how the reference explanations were authored, how many experts wrote them, or how often experts agree, so inconsistent or unrepresentative references would distort every model score.
Editorial extensions
If this is right
- Standard Arabic benchmarks overstate real competence: high scores on sentiment, question answering, and named-entity tasks coexist with weak poetic interpretation, so a model cannot be certified as deeply Arabic-capable on surface tasks alone.
- The era-wise breakdown pinpoints the weakness as classical Arabic: every model family scores lower on pre-Islamic and Umayyad poetry than on modern poetry, giving trainers a specific register to target.
- Fluency and understanding are separable: several models produce fluent Arabic explanations that human or automated judges find shallow or unfaithful, so evaluation suites should measure interpretive depth independently of grammatical quality.
- The open-source release makes progress measurable: future Arabic LLMs can be compared against the same human references, turning poetic comprehension from anecdote into a reproducible diagnostic.
Reading between the lines
- The paper does not run this, but its 6,984 gold explanations could serve as instruction-tuning data: fine-tuning an open model on them and testing on a held-out split would reveal whether the poetry deficit is missing knowledge of classical Arabic or a failure to apply interpretive reasoning on demand.
- A multiple-choice comprehension variant derived from the same poems would separate understanding from generation: if models that fail open-ended explanation pass forced-choice questions, the bottleneck is articulation or evaluation rather than comprehension itself.
- The pattern likely generalizes across heritage languages: if Arabic poetry shows this cliff against prose tasks, comparable benchmarks for classical Persian, Sanskrit, or Chinese would probably show a similar gap, making shallow literary understanding a general property of LLMs rather than an Arabic-specific one.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Fann or Flop, a benchmark for Arabic poetry understanding consisting of 6,984 poems with expert-written explanations, spanning 12 historical eras and 14 genres. The authors evaluate 14 open and closed LLMs by prompting them to produce verse-by-verse Arabic explanations and comparing with automatic metrics (BLEU, chrF++, BERTScore), textual entailment, GPT-4o-based faithfulness/fluency judgments, and a human interpretive-depth score. The main finding is that most models score relatively low on poetic interpretation compared with their typical performance on standard Arabic benchmarks, with GPT-4o and Gemini-2.5-Flash leading. The dataset and evaluation code are released.
Significance. The benchmark addresses a real gap: existing Arabic NLP benchmarks focus on prose, and poetry understanding is largely untested. The historical and genre coverage, if accurate, is substantially broader than prior poetry resources such as Ashaar. The release of data and code is a strength, as is the multi-metric evaluation. However, the central claim that the benchmark is a reliable multiera instrument depends on the correctness of the era/genre labels and on the quality and consistency of the gold explanations, both of which have unresolved issues detailed below.
major comments (4)
- [Section 2.3, Tables 2 and 11] The expert-validation claim is contradicted by a concrete factual error: Table 2 lists Bashar ibn Burd (d. c. 783 CE) under 'Between the Two Dynasties,' which the same table dates to 1258–1517 CE, and Table 11 assigns 321 of his poems to that era. This is a misplacement of roughly five centuries; Bashar ibn Burd is a well-known late Umayyad/early Abbasid poet. Because the same taxonomy-driven pipeline labels the entire dataset, this error raises doubts about the reliability of the era labels used in Tables 4 and 17, and it undermines the statement in Section 2.3 that 'all genre and era annotations were reviewed by Arabic language and literature experts.' The authors should audit all era assignments, especially for poets whose dates are well known, and report the audit results.
- [Section 3, Table 3] The human evaluation component is not adequately documented. The paper reports interpretive-depth scores (e.g., 7.52 for GPT-4o) but does not state how many poems or outputs were annotated, the number of annotators, their expertise or linguistic background, or inter-annotator agreement (e.g., Cohen's kappa). Without these details, the human scores cannot be reproduced or used to compare models, and the standard deviations given for other metrics are not matched by corresponding statistics for human scores. This is a load-bearing issue because the 'most models struggle' finding is partly based on the human interpretive-depth column.
- [Section 2.2/2.3] There is no description of how the gold explanations were created. The text states that each sample is 'manually verified by native Arabic speakers with domain knowledge' (Section 1) and that 'all genre and era annotations were reviewed' (Section 2.3), but it never says who wrote the verse-by-verse explanations, whether each poem has multiple independent references, how many experts were involved, or how often they agreed. Since BLEU, chrF++, BERTScore, textual entailment, and GPT-4o-based faithfulness scores all compare model outputs against these references, the authors must provide this information and ideally measure reference-explanation diversity.
- [Section 3, Evaluation Metric] GPT-4o is used both as an automatic judge for faithfulness/fluency and as one of the evaluated models. This creates a potential self-preference bias for the GPT-4o results in Table 3, and it also makes the lexical and semantic metrics the only 'neutral' comparison. The authors should either use an independent judge (e.g., a different LLM or human raters for the whole set) or report a calibration/agreement study between GPT-4o and humans on these two dimensions.
minor comments (4)
- [Section 3, Table 3] BLEU scores near 0.04 are all near zero and are not informative for open-ended explanation generation; the paper itself notes the limitation, so these columns should be either removed from the main table or replaced with a more appropriate measure such as a retrieval-based semantic match.
- [Section 5] The limitations section acknowledges that 'poetry often invites multiple valid interpretations' but does not connect this to the evaluation design; a short discussion of how the chosen references handle ambiguity would strengthen the paper.
- [References] Several citations are to Wikipedia or personal blogs (e.g., 'wikipedia, 2025', 'oussama, 2024', 'alsharekh, 2019'); these should be replaced by scholarly sources or primary references where possible.
- [Figures 2 and 9] The Arabic text in the example figures appears in a small and partially garbled rendering; please ensure the vectorized text is legible in the final PDF.
Circularity Check
No circularity: the benchmark construction and model evaluation are empirical and externally referenced; no fitted parameter or derivation loop is present.
full rationale
Fann or Flop is a dataset-and-evaluation paper rather than a derivational model, so the circularity patterns that apply to fitted-parameter or self-citation chains do not arise. The central claim that LLMs struggle with Arabic poetry is measured by comparing model outputs against human-authored reference explanations via BLEU, chrF(++), BERTScore, textual entailment, GPT-4o-based faithfulness/fluency scoring, and human interpretive-depth annotation. No parameter is fitted to a subset of the data and then renamed a prediction; the reference explanations are external human-authored texts, and the era/genre taxonomy is sourced from an external archive and then expert-reviewed. The only self-referential element is the use of GPT-4o as the LLM judge for faithfulness and fluency while GPT-4o is also among the evaluated models; this is a methodological confound that could bias those two columns, but it does not force the central finding, which is additionally supported by BERTScore, textual entailment, and human interpretive-depth scores that do not depend on GPT-4o. The concern about Table 11 assigning Bashar ibn Burd to the 1258-1517 'Between the Two Dynasties' era is a factual-accuracy and label-validity issue, not a circular derivation: the paper's era-wise results inherit any label errors, but the results are not equivalent to the labels by construction. The paper's own Limitations section acknowledges that poetry invites multiple valid interpretations that current metrics may not fully capture even with expert-curated references, which is an honest validity caveat rather than a circular step. No load-bearing argument reduces to a self-citation chain, and no claim is defined in terms of the quantity it purports to predict. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption The arabic-poetry.net archive is authoritative and representative enough to support claims about the broader Arabic poetic tradition.
- domain assumption Experts correctly validated genre and era annotations for all 6,984 poems.
- domain assumption The human reference explanations are valid gold interpretations of the poems.
- domain assumption Arabic-pretrained AraBERT and mDeBERTaV3 produce meaningful semantic and entailment scores for classical Arabic poetry.
- domain assumption GPT-4o is a reliable judge of faithfulness and fluency for Arabic poetic explanations.
Cite this review
Pith. "Pith review of Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs." pith.science (2026). https://pith.science/paper/4HLIIGGX
@misc{pith2026250518152,
author = {Pith},
title = {Pith review of: Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/4HLIIGGX}},
note = {Machine review of arXiv:2505.18152}
}
read the original abstract
Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages and tasks, their ability to understand Arabic poetry remains largely unexplored. In this work, we introduce \emph{Fann or Flop}, the first benchmark designed to assess the comprehension of Arabic poetry by LLMs in 12 historical eras, covering 14 core poetic genres and a variety of metrical forms, from classical structures to contemporary free verse. The benchmark comprises a curated corpus of poems with explanations that assess semantic understanding, metaphor interpretation, prosodic awareness, and cultural context. We argue that poetic comprehension offers a strong indicator for testing how good the LLM understands classical Arabic through Arabic poetry. Unlike surface-level tasks, this domain demands deeper interpretive reasoning and cultural sensitivity. Our evaluation of state-of-the-art LLMs shows that most models struggle with poetic understanding despite strong results on standard Arabic benchmarks. We release "Fann or Flop" along with the evaluation suite as an open-source resource to enable rigorous evaluation and advancement for Arabic language models. Code is available at: https://github.com/mbzuai-oryx/FannOrFlop.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Muhammad Abdul-Mageed, Shady Elbassuoni, Jad Doughman, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Yorgo Zoughby, Ahmad Shaher, Iskander Gaba, Ahmed Helal, and Mohammed El-Razzaz. 2021. https://www.aclweb.org/anthology/2021.wanlp-1.2 D ia L ex: A benchmark for evaluating multidialectal A rabic word embeddings . In Proceedings of the Sixth Arabic Natura...
work page 2021
-
[4]
Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2020. Arbert & marbert: Deep bidirectional transformers for arabic. arXiv preprint arXiv:2101.01785
arXiv 2020
-
[5]
Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, \"O mer \"O zsoy, Alexander Mehler, and Gemma Roig. 2022. https://aclanthology.org/2022.coling-1.330/ Tafsir dataset: A novel multi-task benchmark for named entity recognition and topic modeling in classical A rabic literature . In Proceedings of the 29th International Conference on Computation...
work page 2022
-
[6]
Google AI. 2025 a . https://ai.google.dev/gemini-api/docs/models/gemini#gemini-2.0-flash Gemini 2.0 flash . Large language model, accessed May 20, 2025
work page 2025
-
[7]
Google AI. 2025 b . https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-flash Gemini 2.5 flash . Large language model (Preview), accessed May 20, 2025
work page 2025
-
[8]
Abu Nasr al Jawhari. 10th Century. Taj al-lugha wa sihah al-arabiya - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Abu_Nasr_al-Jawhari. [Accessed 06-05-2025]
work page 2025
Show all 52 references
-
[9]
alsharekh. 2019. Al-mujam al muaser-- lexicon.alsharekh.org. https://lexicon.alsharekh.org/. [Accessed 06-05-2025]
2019
-
[10]
15th Century
AlSuyuti. 15th Century. Al-mizhar fi eulum allughat wa'anwaeiha - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Al-Suyuti. [Accessed 06-05-2025]
2025
-
[11]
Zaid Alyafeai, Maged S Al-Shaibani, and Moataz Ahmed. 2023. Ashaar: Automatic analysis and generation of arabic poetry using deep learning approaches. arXiv preprint arXiv:2307.06218
2023 arXiv
-
[12]
Toni Andrews. 2024. I s A rabic T he R ichest L anguage I n W ords? - I nterpreters & T ranslators, I nc. --- ititranslates.com. https://ititranslates.com/is-arabic-the-richest-language-in-words/. [Accessed 06-05-2025]
2024
-
[13]
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020. Arabert: Transformer-based model for arabic language understanding. arXiv preprint arXiv:2003.00104
2020 arXiv
-
[14]
Mujam Ar-Riyadh. 2025. Mujam ar-riyadh--- dictionary.ksaa.gov.sa. https://dictionary.ksaa.gov.sa/. [Accessed 06-05-2025]
2025
-
[15]
Alzahrani, Nouf M
M Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham A. Alyahya, Sultan AlRashed, Faisal A. Mirza, Shaykhah Z. Alsubaie, Hassan A. Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih, Raneem Alnajim, Salman Alsubaihi, Maryam Al Mansour, ...
2024 arXiv
-
[16]
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432--7439
2020
-
[17]
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, et al. 2018. The madar arabic dialect corpus and lexicon. In Proceedings of the eleventh international conference on lang...
2018
-
[18]
Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, and Smaranda Muresan. 2022. Flute: Figurative language understanding through textual explanations. arXiv preprint arXiv:2205.12404
2022 arXiv
-
[19]
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2025. Sharegpt4v: Improving large multi-modal models with better captions. In European Conference on Computer Vision, pages 370--387. Springer
2025
-
[20]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, et al. 2024. Aya expanse: Combining research breakthroughs for a new multilingual frontier. arXiv preprint arXiv:2412.04261
2024 arXiv
-
[21]
FIGLANG202. 2024. F ig L ang2024 --- sites.google.com. https://sites.google.com/view/figlang2024. [Accessed 07-05-2025]
2024
-
[22]
Giuseppe Gallipoli and Luca Cagliero. 2025. It is not a piece of cake for gpt: Explaining textual entailment recognition in the presence of figurative language. In Proceedings of the 31st International Conference on Computational Linguistics, pages 9656--9674
2025
-
[23]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[24]
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543
2021 arXiv
-
[25]
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Dingjie Song, Zhihong Chen, Abdulmohsen Alharthi, Bang An, Juncai He, et al. 2023. Acegpt, localizing large language models in arabic. arXiv preprint arXiv:2309.12053
2023 arXiv
-
[26]
Salam Khalifa, Nizar Habash, Fadhl Eryani, Ossama Obeid, Dana Abdulrahim, and Meera Al Kaabi. 2018. https://aclanthology.org/L18-1607/ A morphologically annotated corpus of emirati A rabic . In Proceedings of the Eleventh International Conference on Language Resources and Eval...
2018
-
[27]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
2024 arXiv
-
[28]
Emmy Liu, Chen Cui, Kenneth Zheng, and Graham Neubig. 2022. Testing the ability of language models to interpret figurative language. arXiv preprint arXiv:2204.12632
2022 arXiv
-
[29]
Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru, Mugariya Farooq, Giulia Campesan, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Maksim Velikanov, Basma El Amel Boussaha, et al. 2024. Falcon2-11b technical report. arXiv preprint arXiv:2407.14885
2024 arXiv
-
[30]
14th Century
Ibn Manzur. 14th Century. Lisan al-arab --- en.wikipedia.org. https://en.wikipedia.org/wiki/Ibn_Manzur. [Accessed 06-05-2025]
2025
-
[31]
Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015. Machine translation experiments on padic: A parallel arabic dialect corpus. In Proceedings of the 29th Pacific Asia conference on language, information and computation, pages 26--34
2015
-
[32]
Meta AI . 2024. https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
2024
-
[33]
Behrang Mohit, Nathan Schneider, Rishav Bhowmick, Kemal Oflazer, and Noah A Smith. 2012. Recall-oriented learning of named entities in arabic wikipedia. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pages 162--173
2012
-
[34]
Hussein Mozannar, Karl El Hajal, Elie Maamary, and Hazem Hajj. 2019. Neural arabic question answering. arXiv preprint arXiv:1906.05394
2019 arXiv
-
[35]
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. 2020. Camel tools: An open source python toolkit for arabic natural language processing. In Proceedings of the twelfth language resou...
2020
-
[36]
Susanna Olivero. 2024. Figurative Language Understanding based on Large Language Models. Ph.D. thesis, Politecnico di Torino
2024
-
[37]
OpenAI . 2024. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence
2024
-
[38]
OpenAI. 2024. https://arxiv.org/abs/2410.21276 Gpt-4o system card . Preprint, arXiv:2410.21276
2024 arXiv
-
[39]
oussama. 2024. M odern S tandard A rabic – T he M issing G lossary - --- blog.jarrousse.org. https://blog.jarrousse.org/2024/03/27/modern-standard-arabic-the-missing-glossary/. [Accessed 06-05-2025]
2024
-
[40]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311--318
2002
-
[41]
Maja Popovi \'c . 2017. https://doi.org/10.18653/v1/W17-4770 chr F ++: words helping character n-grams . In Proceedings of the Second Conference on Machine Translation, pages 612--618, Copenhagen, Denmark. Association for Computational Linguistics
2017 doi
-
[42]
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, and et al. 2024. https://arxiv.org/abs/2403.05530 Gemini 1.5: Unlocking multimodal understanding acro...
2024 arXiv
-
[43]
Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020. Arabench: Benchmarking dialectal arabic-english machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5094--5107
2020
-
[44]
Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, et al. 2023. Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. arXiv pre...
2023 arXiv
-
[45]
Fanar Team, Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, Fahim Dalvi, Kareem Darwish, Nadir Durrani, Mohamed Elfeky, Ahmed Elmagarmid, Mohamed Eltabakh, Masoomali Fatehki...
2025 arXiv
-
[46]
Qwen Team. 2025. https://qwenlm.github.io/blog/qwen3/ Qwen3
2025
-
[47]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[48]
wikipedia. 2025. ar.wikipedia.org. https://ar.wikipedia.org/wiki/ [Accessed 06-05-2025]
2025
-
[49]
wikipediaArabic. 2025. V arieties of A rabic - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Varieties_of_Arabic. [Accessed 06-05-2025]
2025
-
[50]
Taha Zerrouki and Amar Balla. 2017. Tashkeela: Novel corpus of arabic vocalized texts, data for auto-diacritization systems. Data in brief, 11:147
2017
-
[51]
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675
2019 arXiv
-
[52]
Cheng Zhao, Bin Wang, and Zhen Wang. 2024. Understanding literary texts by llms: A case study of ancient chinese poetry. arXiv preprint arXiv:2409.00060
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.