Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Fann or Flop, the first Arabic poetry benchmark spanning 12 eras and 14 genres, finds that state-of-the-art LLMs, strong on standard Arabic tasks, fall short on poetry's interpretive depth.

desk verdict Era-label error undermines the multiera claim, but the benchmark resource and task are real; fix labels and evaluation details. read the letter →

arxiv 2505.18152 v2 pith:4HLIIGGX submitted 2025-05-23 cs.CL

classification cs.CL
keywords ArabicpoetryunderstandingLLMbenchmarkclassicalfigurativelanguagehistoricaleraspoeticgenresinterpretivereasoningNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Fann or Flop sets out to measure something existing Arabic language benchmarks do not: whether large language models can actually understand Arabic poetry in the literary sense, rather than merely process Arabic text. The paper assembles 6,984 poems with expert-verified historical era and genre labels, each paired with a human-authored verse-by-verse explanation in formal Arabic, spanning twelve eras from the pre-Islamic period to the present and fourteen genres from praise and satire to love and elegy. Fourteen state-of-the-art closed and open models are asked to produce their own explanations, which are scored against the human references by lexical overlap, semantic similarity, entailment, an automated judge, and human experts rating interpretive depth on a ten-point scale. The consistent finding is that models do well on ordinary Arabic tasks yet fall short on poetry, missing metaphors, rhetorical devices, and cultural context. The paper argues that poetic comprehension is a strong indicator of whether a model has truly internalized classical Arabic, and it is releasing Fann or Flop as an open-source diagnostic.

What carries the argument

The load-bearing instrument is the benchmark itself: a taxonomy of twelve historical poetic eras and fourteen genres, matched to 6,984 curated poems each carrying a verse-by-verse explanation in formal Arabic. The explanations are the crux — they integrate literal meaning with figurative and rhetorical analysis (naming devices such as metaphor, simile, and paronomasia) and with cultural-historical context, which is what forces models to go beyond surface statistics. Around this corpus sits a multi-tier evaluation stack: BLEU and chrF(++) for lexical overlap, AraBERT-based BERTScore and mDeBERTaV3-based textual entailment for semantic alignment, GPT-4o as an automated judge of faithfulness and fluency, and a rubric-based human score for interpretive depth. The stack is designed so that the benchmark can separate fluent paraphrase from genuine understanding.

What would settle it

Have independent Arabic-literature experts re-annotate a random sample of the 6,984 pairs without seeing the gold material: assign era and genre labels, and write fresh verse explanations. If pairwise expert agreement falls below roughly 70 percent, the human references themselves are not stable enough to rank models, and the struggle finding would be measuring disagreement among humans as much as model failure. In the opposite direction, the claim would soften if an open-weight model scored at or above the human-reference level on interpretive depth in a blind human evaluation.

Watch

Extended reading notes

Core claim

The paper's central claim is that Fann or Flop exposes a real gap that standard Arabic benchmarks hide: most LLMs perform well on ordinary Arabic tasks but consistently fall short when asked to interpret Arabic poetry. On the benchmark's 6,984 poem-explanation pairs, fourteen state-of-the-art models are asked to produce verse-by-verse explanations in formal Arabic, scored against human-written references by lexical overlap (BLEU, chrF(++)), semantic similarity (BERTScore), textual entailment, an LLM judge, and human experts rating interpretive depth on a 0–10 rubric. The strongest model, GPT-4o, reaches about 0.64 BERTScore and about 7.5 for human-judged interpretive depth; several open and Arabic-centric models land well below that, and every model family scores lower on pre-Islamic and Umayyad verse than on modern poetry. The authors read this as evidence that current models have absorbed Modern Standard Arabic but not the layered, figurative, and historically embedded language of classical Arabic poetry, which is why they propose poetry comprehension as a sharper test of cultural and linguistic depth.

Load-bearing premise

The ranking stands on the gold explanations and the era and genre labels being trustworthy references, but the paper never reports how the reference explanations were authored, how many experts wrote them, or how often experts agree, so inconsistent or unrepresentative references would distort every model score.

Editorial extensions

If this is right

  • Standard Arabic benchmarks overstate real competence: high scores on sentiment, question answering, and named-entity tasks coexist with weak poetic interpretation, so a model cannot be certified as deeply Arabic-capable on surface tasks alone.
  • The era-wise breakdown pinpoints the weakness as classical Arabic: every model family scores lower on pre-Islamic and Umayyad poetry than on modern poetry, giving trainers a specific register to target.
  • Fluency and understanding are separable: several models produce fluent Arabic explanations that human or automated judges find shallow or unfaithful, so evaluation suites should measure interpretive depth independently of grammatical quality.
  • The open-source release makes progress measurable: future Arabic LLMs can be compared against the same human references, turning poetic comprehension from anecdote into a reproducible diagnostic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not run this, but its 6,984 gold explanations could serve as instruction-tuning data: fine-tuning an open model on them and testing on a held-out split would reveal whether the poetry deficit is missing knowledge of classical Arabic or a failure to apply interpretive reasoning on demand.
  • A multiple-choice comprehension variant derived from the same poems would separate understanding from generation: if models that fail open-ended explanation pass forced-choice questions, the bottleneck is articulation or evaluation rather than comprehension itself.
  • The pattern likely generalizes across heritage languages: if Arabic poetry shows this cliff against prose tasks, comparable benchmarks for classical Persian, Sanskrit, or Chinese would probably show a similar gap, making shallow literary understanding a general property of LLMs rather than an Arabic-specific one.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Fann or Flop, a benchmark for Arabic poetry understanding consisting of 6,984 poems with expert-written explanations, spanning 12 historical eras and 14 genres. The authors evaluate 14 open and closed LLMs by prompting them to produce verse-by-verse Arabic explanations and comparing with automatic metrics (BLEU, chrF++, BERTScore), textual entailment, GPT-4o-based faithfulness/fluency judgments, and a human interpretive-depth score. The main finding is that most models score relatively low on poetic interpretation compared with their typical performance on standard Arabic benchmarks, with GPT-4o and Gemini-2.5-Flash leading. The dataset and evaluation code are released.

Significance. The benchmark addresses a real gap: existing Arabic NLP benchmarks focus on prose, and poetry understanding is largely untested. The historical and genre coverage, if accurate, is substantially broader than prior poetry resources such as Ashaar. The release of data and code is a strength, as is the multi-metric evaluation. However, the central claim that the benchmark is a reliable multiera instrument depends on the correctness of the era/genre labels and on the quality and consistency of the gold explanations, both of which have unresolved issues detailed below.

major comments (4)
  1. [Section 2.3, Tables 2 and 11] The expert-validation claim is contradicted by a concrete factual error: Table 2 lists Bashar ibn Burd (d. c. 783 CE) under 'Between the Two Dynasties,' which the same table dates to 1258–1517 CE, and Table 11 assigns 321 of his poems to that era. This is a misplacement of roughly five centuries; Bashar ibn Burd is a well-known late Umayyad/early Abbasid poet. Because the same taxonomy-driven pipeline labels the entire dataset, this error raises doubts about the reliability of the era labels used in Tables 4 and 17, and it undermines the statement in Section 2.3 that 'all genre and era annotations were reviewed by Arabic language and literature experts.' The authors should audit all era assignments, especially for poets whose dates are well known, and report the audit results.
  2. [Section 3, Table 3] The human evaluation component is not adequately documented. The paper reports interpretive-depth scores (e.g., 7.52 for GPT-4o) but does not state how many poems or outputs were annotated, the number of annotators, their expertise or linguistic background, or inter-annotator agreement (e.g., Cohen's kappa). Without these details, the human scores cannot be reproduced or used to compare models, and the standard deviations given for other metrics are not matched by corresponding statistics for human scores. This is a load-bearing issue because the 'most models struggle' finding is partly based on the human interpretive-depth column.
  3. [Section 2.2/2.3] There is no description of how the gold explanations were created. The text states that each sample is 'manually verified by native Arabic speakers with domain knowledge' (Section 1) and that 'all genre and era annotations were reviewed' (Section 2.3), but it never says who wrote the verse-by-verse explanations, whether each poem has multiple independent references, how many experts were involved, or how often they agreed. Since BLEU, chrF++, BERTScore, textual entailment, and GPT-4o-based faithfulness scores all compare model outputs against these references, the authors must provide this information and ideally measure reference-explanation diversity.
  4. [Section 3, Evaluation Metric] GPT-4o is used both as an automatic judge for faithfulness/fluency and as one of the evaluated models. This creates a potential self-preference bias for the GPT-4o results in Table 3, and it also makes the lexical and semantic metrics the only 'neutral' comparison. The authors should either use an independent judge (e.g., a different LLM or human raters for the whole set) or report a calibration/agreement study between GPT-4o and humans on these two dimensions.
minor comments (4)
  1. [Section 3, Table 3] BLEU scores near 0.04 are all near zero and are not informative for open-ended explanation generation; the paper itself notes the limitation, so these columns should be either removed from the main table or replaced with a more appropriate measure such as a retrieval-based semantic match.
  2. [Section 5] The limitations section acknowledges that 'poetry often invites multiple valid interpretations' but does not connect this to the evaluation design; a short discussion of how the chosen references handle ambiguity would strengthen the paper.
  3. [References] Several citations are to Wikipedia or personal blogs (e.g., 'wikipedia, 2025', 'oussama, 2024', 'alsharekh, 2019'); these should be replaced by scholarly sources or primary references where possible.
  4. [Figures 2 and 9] The Arabic text in the example figures appears in a small and partially garbled rendering; please ensure the vectorized text is legible in the final PDF.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the benchmark construction and model evaluation are empirical and externally referenced; no fitted parameter or derivation loop is present.

full rationale

Fann or Flop is a dataset-and-evaluation paper rather than a derivational model, so the circularity patterns that apply to fitted-parameter or self-citation chains do not arise. The central claim that LLMs struggle with Arabic poetry is measured by comparing model outputs against human-authored reference explanations via BLEU, chrF(++), BERTScore, textual entailment, GPT-4o-based faithfulness/fluency scoring, and human interpretive-depth annotation. No parameter is fitted to a subset of the data and then renamed a prediction; the reference explanations are external human-authored texts, and the era/genre taxonomy is sourced from an external archive and then expert-reviewed. The only self-referential element is the use of GPT-4o as the LLM judge for faithfulness and fluency while GPT-4o is also among the evaluated models; this is a methodological confound that could bias those two columns, but it does not force the central finding, which is additionally supported by BERTScore, textual entailment, and human interpretive-depth scores that do not depend on GPT-4o. The concern about Table 11 assigning Bashar ibn Burd to the 1258-1517 'Between the Two Dynasties' era is a factual-accuracy and label-validity issue, not a circular derivation: the paper's era-wise results inherit any label errors, but the results are not equivalent to the labels by construction. The paper's own Limitations section acknowledges that poetry invites multiple valid interpretations that current metrics may not fully capture even with expert-curated references, which is an honest validity caveat rather than a circular step. No load-bearing argument reduces to a self-citation chain, and no claim is defined in terms of the quantity it purports to predict. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the quality of the human annotations and on the validity of the evaluation metrics, not on free parameters or invented entities. No numerical constants are fitted to data, but several domain assumptions about archive representativeness, expert ground truth, and metric suitability are load-bearing.

assumptions (5)
  • domain assumption The arabic-poetry.net archive is authoritative and representative enough to support claims about the broader Arabic poetic tradition.
    Data collection in Section 2.2 scrapes poems from this single archive; if the archive is biased, era and genre coverage and the labels inherit that bias.
  • domain assumption Experts correctly validated genre and era annotations for all 6,984 poems.
    Sections 2.1 and 2.3 state that experts reviewed the taxonomy and labels, but no inter-annotator agreement or validation statistics are reported, and these labels are the ground truth for evaluations.
  • domain assumption The human reference explanations are valid gold interpretations of the poems.
    Section 3 compares all model outputs to human-authored references, but the paper does not describe how the reference explanations were authored or independently validated.
  • domain assumption Arabic-pretrained AraBERT and mDeBERTaV3 produce meaningful semantic and entailment scores for classical Arabic poetry.
    Section 3 uses BERTScore and textual entailment to measure model outputs; these models are largely trained on MSA and modern text, and no validation is provided for classical poetic language.
  • domain assumption GPT-4o is a reliable judge of faithfulness and fluency for Arabic poetic explanations.
    Section 3 uses GPT-4o as an LLM-as-judge without auditing judge reliability or bias, and GPT-4o is also one of the evaluated models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs." pith.science (2026). https://pith.science/paper/4HLIIGGX

@misc{pith2026250518152,
  author       = {Pith},
  title        = {Pith review of: Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4HLIIGGX}},
  note         = {Machine review of arXiv:2505.18152}
}
read the original abstract

Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages and tasks, their ability to understand Arabic poetry remains largely unexplored. In this work, we introduce \emph{Fann or Flop}, the first benchmark designed to assess the comprehension of Arabic poetry by LLMs in 12 historical eras, covering 14 core poetic genres and a variety of metrical forms, from classical structures to contemporary free verse. The benchmark comprises a curated corpus of poems with explanations that assess semantic understanding, metaphor interpretation, prosodic awareness, and cultural context. We argue that poetic comprehension offers a strong indicator for testing how good the LLM understands classical Arabic through Arabic poetry. Unlike surface-level tasks, this domain demands deeper interpretive reasoning and cultural sensitivity. Our evaluation of state-of-the-art LLMs shows that most models struggle with poetic understanding despite strong results on standard Arabic benchmarks. We release "Fann or Flop" along with the evaluation suite as an open-source resource to enable rigorous evaluation and advancement for Arabic language models. Code is available at: https://github.com/mbzuai-oryx/FannOrFlop.

Figures

Figures reproduced from arXiv: 2505.18152 by the authors.

Figure 1
Figure 1. Chronological Wheel of Arabic Poetic Eras. This circular taxonomy visualizes the evolution of Ara￾bic poetry across 12 major historical eras, from the Pre￾Islamic and Transitional periods through the Abbasid, Andalusian, and Mamluk dynasties, up to the Modern era. The layout reflects both temporal flow and the rich cultural shifts that shaped poetic expression. Detailed taxonomy by genre, meter, and notable poets pr… view at source ↗
Figure 2
Figure 2. Representative Poetic Samples Across Arabic Literary Eras. This figure presents curated excerpts from Arabic poems spanning key historical eras, illustrating the evolution of language, themes, and stylistic expression. The Pre-Islamic sample reflects tribal valor and rhetorical precision; the Umayyad excerpt captures satire and social commentary; the Abbasid example highlights philosophical reflection and refined me… view at source ↗
Figure 3
Figure 3. Fann or Flop Pipeline. Fann or Flop is built out of the multi-stage pipeline. It begins with scraping Arabic poems from a trusted online archive using a custom web scraper. Extracted poems are matched to an initial expert-verified taxonomy and filtered to remove duplicates, ambiguous metadata, and invalid entries. The filtered texts then undergo normalization (e.g., unifying diacritics, punctuation, and letter forms… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Qualitative Comparison of Model-Generated Explanations for a Single Arabic Poem. This figure presents a representative Arabic poem alongside its original human-written explanation and corresponding verse￾by-verse explanations generated by four different language models…
Figure 5
Figure 5. Figure 5: Era and Genre Statistics. Subfigure (a) displays the distribution of poems across historical eras, while subfig￾ure (b) shows the overall genre distribution across the dataset. Complementing these visualizations, we also in￾clude detailed per-era tables listing the mos…
Figure 6
Figure 6. Figure 6: Genre distribution across historical eras. This stacked bar chart illustrates how poetic themes evolved across different dynasties. It highlights patterns such as the prominence of Praise and Satire during the Abbasid and Umayyad eras, and the diverse thematic expressi…
Figure 7
Figure 7. Figure 7: The verse-level explanation prompt used for evaluation. This prompt instructs the model to produce de￾tailed verse-by-verse explanations in Arabic. It guides the model to integrate both literal and figurative interpretations, explicitly name rhetorical devices (e.g., m…
Figure 8
Figure 8. Figure 8: System prompt used for LLM-Judge evaluation of verse-by-verse poem explanations. LLM ( [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Translated Samples. This figure presents English translations of the Arabic samples shown in [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]
Figure 10
Figure 10. Figure 10: Fann or Flop Samples by Genre. Additional representative examples from the Fann or Flop benchmark, illustrating the diversity of genres covered, including Love (Ghazal), Praise (Madh. ), Wisdom (Hikma), Satire (Hija’), Elegy (Rith ¯ a’), ¯ Reproach (’Itab), Political …

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

52 extracted references · 26 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Muhammad Abdul-Mageed, Shady Elbassuoni, Jad Doughman, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Yorgo Zoughby, Ahmad Shaher, Iskander Gaba, Ahmed Helal, and Mohammed El-Razzaz. 2021. https://www.aclweb.org/anthology/2021.wanlp-1.2 D ia L ex: A benchmark for evaluating multidialectal A rabic word embeddings . In Proceedings of the Sixth Arabic Natura...

  4. [4]

    Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2020. Arbert & marbert: Deep bidirectional transformers for arabic. arXiv preprint arXiv:2101.01785

  5. [5]

    Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, \"O mer \"O zsoy, Alexander Mehler, and Gemma Roig. 2022. https://aclanthology.org/2022.coling-1.330/ Tafsir dataset: A novel multi-task benchmark for named entity recognition and topic modeling in classical A rabic literature . In Proceedings of the 29th International Conference on Computation...

  6. [6]

    Google AI. 2025 a . https://ai.google.dev/gemini-api/docs/models/gemini#gemini-2.0-flash Gemini 2.0 flash . Large language model, accessed May 20, 2025

  7. [7]

    Google AI. 2025 b . https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-flash Gemini 2.5 flash . Large language model (Preview), accessed May 20, 2025

  8. [8]

    10th Century

    Abu Nasr al Jawhari. 10th Century. Taj al-lugha wa sihah al-arabiya - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Abu_Nasr_al-Jawhari. [Accessed 06-05-2025]

Show all 52 references
  1. [9]

    alsharekh. 2019. Al-mujam al muaser-- lexicon.alsharekh.org. https://lexicon.alsharekh.org/. [Accessed 06-05-2025]

  2. [10]

    15th Century

    AlSuyuti. 15th Century. Al-mizhar fi eulum allughat wa'anwaeiha - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Al-Suyuti. [Accessed 06-05-2025]

  3. [11]

    Zaid Alyafeai, Maged S Al-Shaibani, and Moataz Ahmed. 2023. Ashaar: Automatic analysis and generation of arabic poetry using deep learning approaches. arXiv preprint arXiv:2307.06218

  4. [12]

    Toni Andrews. 2024. I s A rabic T he R ichest L anguage I n W ords? - I nterpreters & T ranslators, I nc. --- ititranslates.com. https://ititranslates.com/is-arabic-the-richest-language-in-words/. [Accessed 06-05-2025]

  5. [13]

    Wissam Antoun, Fady Baly, and Hazem Hajj. 2020. Arabert: Transformer-based model for arabic language understanding. arXiv preprint arXiv:2003.00104

  6. [14]

    Mujam Ar-Riyadh. 2025. Mujam ar-riyadh--- dictionary.ksaa.gov.sa. https://dictionary.ksaa.gov.sa/. [Accessed 06-05-2025]

  7. [15]

    Alzahrani, Nouf M

    M Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham A. Alyahya, Sultan AlRashed, Faisal A. Mirza, Shaykhah Z. Alsubaie, Hassan A. Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih, Raneem Alnajim, Salman Alsubaihi, Maryam Al Mansour, ...

  8. [16]

    Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432--7439

  9. [17]

    Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, et al. 2018. The madar arabic dialect corpus and lexicon. In Proceedings of the eleventh international conference on lang...

  10. [18]

    Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, and Smaranda Muresan. 2022. Flute: Figurative language understanding through textual explanations. arXiv preprint arXiv:2205.12404

  11. [19]

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2025. Sharegpt4v: Improving large multi-modal models with better captions. In European Conference on Computer Vision, pages 370--387. Springer

  12. [20]

    John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, et al. 2024. Aya expanse: Combining research breakthroughs for a new multilingual frontier. arXiv preprint arXiv:2412.04261

  13. [21]

    FIGLANG202. 2024. F ig L ang2024 --- sites.google.com. https://sites.google.com/view/figlang2024. [Accessed 07-05-2025]

  14. [22]

    Giuseppe Gallipoli and Luca Cagliero. 2025. It is not a piece of cake for gpt: Explaining textual entailment recognition in the presence of figurative language. In Proceedings of the 31st International Conference on Computational Linguistics, pages 9656--9674

  15. [23]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  16. [24]

    Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543

  17. [25]

    Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Dingjie Song, Zhihong Chen, Abdulmohsen Alharthi, Bang An, Juncai He, et al. 2023. Acegpt, localizing large language models in arabic. arXiv preprint arXiv:2309.12053

  18. [26]

    Salam Khalifa, Nizar Habash, Fadhl Eryani, Ossama Obeid, Dana Abdulrahim, and Meera Al Kaabi. 2018. https://aclanthology.org/L18-1607/ A morphologically annotated corpus of emirati A rabic . In Proceedings of the Eleventh International Conference on Language Resources and Eval...

  19. [27]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437

  20. [28]

    Emmy Liu, Chen Cui, Kenneth Zheng, and Graham Neubig. 2022. Testing the ability of language models to interpret figurative language. arXiv preprint arXiv:2204.12632

  21. [29]

    Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru, Mugariya Farooq, Giulia Campesan, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Maksim Velikanov, Basma El Amel Boussaha, et al. 2024. Falcon2-11b technical report. arXiv preprint arXiv:2407.14885

  22. [30]

    14th Century

    Ibn Manzur. 14th Century. Lisan al-arab --- en.wikipedia.org. https://en.wikipedia.org/wiki/Ibn_Manzur. [Accessed 06-05-2025]

  23. [31]

    Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015. Machine translation experiments on padic: A parallel arabic dialect corpus. In Proceedings of the 29th Pacific Asia conference on language, information and computation, pages 26--34

  24. [32]

    Meta AI . 2024. https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

  25. [33]

    Behrang Mohit, Nathan Schneider, Rishav Bhowmick, Kemal Oflazer, and Noah A Smith. 2012. Recall-oriented learning of named entities in arabic wikipedia. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, pages 162--173

  26. [34]

    Hussein Mozannar, Karl El Hajal, Elie Maamary, and Hazem Hajj. 2019. Neural arabic question answering. arXiv preprint arXiv:1906.05394

  27. [35]

    Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. 2020. Camel tools: An open source python toolkit for arabic natural language processing. In Proceedings of the twelfth language resou...

  28. [36]

    Susanna Olivero. 2024. Figurative Language Understanding based on Large Language Models. Ph.D. thesis, Politecnico di Torino

  29. [37]

    OpenAI . 2024. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence

  30. [38]

    OpenAI. 2024. https://arxiv.org/abs/2410.21276 Gpt-4o system card . Preprint, arXiv:2410.21276

  31. [39]

    oussama. 2024. M odern S tandard A rabic – T he M issing G lossary - --- blog.jarrousse.org. https://blog.jarrousse.org/2024/03/27/modern-standard-arabic-the-missing-glossary/. [Accessed 06-05-2025]

  32. [40]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311--318

  33. [41]

    Maja Popovi \'c . 2017. https://doi.org/10.18653/v1/W17-4770 chr F ++: words helping character n-grams . In Proceedings of the Second Conference on Machine Translation, pages 612--618, Copenhagen, Denmark. Association for Computational Linguistics

  34. [42]

    Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, and et al. 2024. https://arxiv.org/abs/2403.05530 Gemini 1.5: Unlocking multimodal understanding acro...

  35. [43]

    Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020. Arabench: Benchmarking dialectal arabic-english machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5094--5107

  36. [44]

    Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, et al. 2023. Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. arXiv pre...

  37. [45]

    Fanar Team, Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, Fahim Dalvi, Kareem Darwish, Nadir Durrani, Mohamed Elfeky, Ahmed Elmagarmid, Mohamed Eltabakh, Masoomali Fatehki...

  38. [46]

    Qwen Team. 2025. https://qwenlm.github.io/blog/qwen3/ Qwen3

  39. [47]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971

  40. [48]

    wikipedia. 2025. ar.wikipedia.org. https://ar.wikipedia.org/wiki/ [Accessed 06-05-2025]

  41. [49]

    wikipediaArabic. 2025. V arieties of A rabic - W ikipedia --- en.wikipedia.org. https://en.wikipedia.org/wiki/Varieties_of_Arabic. [Accessed 06-05-2025]

  42. [50]

    Taha Zerrouki and Amar Balla. 2017. Tashkeela: Novel corpus of arabic vocalized texts, data for auto-diacritization systems. Data in brief, 11:147

  43. [51]

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675

  44. [52]

    Cheng Zhao, Bin Wang, and Zhen Wang. 2024. Understanding literary texts by llms: A case study of ancient chinese poetry. arXiv preprint arXiv:2409.00060

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.