Pith. sign in

REVIEW 4 major objections 4 minor 42 references

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models

T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read FarsEval-PKBETS, a new 4,000-question Persian benchmark, reports sub-50% accuracy for three current large language models, concluding that Persian and Iranian cultural competence is still largely unsolved.

desk verdict A genuinely useful Persian benchmark resource whose headline claim about models being 'far from solving' it is undercut by the absence of a human baseline and a released evaluation protocol. read the letter →

arxiv 2504.14690 v1 pith:FQGHVF23 submitted 2025-04-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords FarsEval-PKBETSPersianlanguagemodelsLLMbenchmarkquestionansweringculturalknowledgeethicsandbiastextgenerationhumanevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FarsEval-PKBETS is a new Persian-language evaluation benchmark built from 4,000 questions that were human-generated or human-collected and then human-reviewed, covering medicine, law, religion, Persian language, encyclopedic knowledge, human preferences, social knowledge, ethics and bias, toxicity, respect for others' rights, and text generation. The paper's central claim is that existing Persian evaluation resources are limited by translation, exam-style data, and a near-total reliance on multiple-choice questions, whereas this benchmark is designed around Persian linguistic and Iranian cultural context and includes short-answer and descriptive formats that test production, not just selection. To demonstrate that the questions are genuinely hard, the authors evaluated Llama3-70B, PersianMind, and Dorna and report fully correct answer rates of 47%, 19%, and 30%, respectively. The intended upshot is that current language models remain far from solving Persian-language tasks that depend on local knowledge, style, social judgment, and generation.

What carries the argument

The central object is FarsEval-PKBETS itself: a dataset of 4,000 records, each containing a Persian question, a reference answer, and metadata such as category, label, and a Reference field for sourced items, with 1,200 descriptive, 500 short-answer, and 2,300 multiple-choice items distributed across twelve head categories. The paper's construction machinery is a two-reviewer workflow on the Saba annotation platform, where submissions are approved or rejected and returned for revision, plus domain-expert supervision for medicine and law. For evaluation, outputs are scored as Correct, Wrong, or Semi-correct, with the semi-correct category catching cases like correct answer text attached to a distractor's identifier. This combination is what lets the benchmark claim both content validity through human review and cultural grounding, and diagnostic power through format diversity and fine-grained scoring.

What would settle it

Have several independent Persian-speaking annotators re-answer a random sample of the Ethics, Bias and Morality, Human Preferences, and Respecting Others' Rights items without seeing the reference answers; if per-item majority agreement is low or frequent disputes arise, the reference answers are not stable enough to support the accuracy claims.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that a carefully reviewed, culturally grounded Persian QA set can expose large gaps in current LLMs. Across 4,000 items, Llama3-70B averages 0.47 full-correct accuracy, PersianMind 0.19, and Dorna 0.30; even counting partially correct responses, the best model reaches only 0.54. The benchmark also reveals failure modes that multiple-choice-only evaluation misses: in a 100-question probe, models gave a correct choice but an incorrect justification 43% of the time, and models sometimes attach the right answer text to the wrong option identifier. The authors take these results as evidence that FarsEval-PKBETS is challenging and that Persian and Iranian socio-cultural competence is not yet achieved by publicly available models.

Load-bearing premise

The reference answers are treated as the ground truth, especially in ethical, moral, and rights-related questions where answers were set according to Iranian cultural norms; if those answers are contested or inconsistent, the reported accuracy numbers are not a clean measure of model capability.

Editorial extensions

If this is right

  • A reusable Persian benchmark now exists that can be rerun as models improve; sub-50% accuracy gives a concrete baseline for progress.
  • MCQ-only scores overstate capability; because semi-correct and unjustified answers are counted separately, benchmark users can see where selection outperforms reasoning.
  • Persian-specific and Iranian-cultural categories such as empathy, irony, respecting others' rights, formal register, and poems define dimensions that translated English benchmarks do not measure.
  • Domain results are uneven, such as 0.65 lexical semantics but 0.14 emergency medicine for Llama3-70B, so the benchmark can guide targeted improvement by category.
  • The balanced distribution of correct option positions reduces position-bias artifacts in multiple-choice scoring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because reference answers in ethics, bias, morality, and respecting others' rights encode Iranian cultural norms, future uses should separate knowledge accuracy from value alignment; identical item scores may mean different things across cultural backgrounds.
  • A direct test of the cultural-grounding claim would be to translate the same 4,000 items into English and evaluate English-native models: if accuracy rises substantially, the bottleneck is Persian language and local context rather than general reasoning.
  • The 43% justification-failure finding suggests a training signal: models may first improve in giving correct rationales before improving final answers, so reporting both could reveal earlier progress.
  • Re-annotating a sample of the single-annotator categories such as Emotion and Irony with multiple independent Persian-speaking annotators would estimate how much of the accuracy gap is label ambiguity rather than model deficit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces FarsEval-PKBETS, a Persian-language benchmark of 4,000 human-generated or collected question-answer samples covering 22 subcategories, three response formats (multiple-choice, short answer, descriptive), and culturally specific topics such as Iranian law, religion, ethics, and social knowledge. The authors describe the Saba annotation platform, the category design, and an evaluation of three models: Llama3-70B, PersianMind, and Dorna. They report average accuracies of 47%, 19%, and 30% respectively, and conclude that current language models are still far from being able to solve the benchmark.

Significance. If the dataset is released with proper validation, FarsEval-PKBETS would be a valuable resource for Persian LLM evaluation: it is human-generated, covers diverse formats and local cultural knowledge, involves expert supervision in medicine and law, and addresses gaps in existing Persian benchmarks that rely heavily on translated or exam-derived multiple-choice questions. The paper also documents a useful annotation platform. However, the headline claim that sub-50% accuracy shows models are 'far from being able to solve' the benchmark currently rests on unvalidated reference answers and a selection procedure aimed at making questions hard for Llama3-70B; the missing human baseline and missing evaluation protocol are load-bearing gaps.

major comments (4)
  1. [Technical Validation / Table 3] The central claim that models are 'far from being able to solve' FarsEval-PKBETS requires a human performance ceiling, which is not reported. Without a human baseline or inter-annotator agreement statistics, sub-50% scores could reflect ambiguous or contested reference answers rather than model deficiency. This is especially consequential in the normative categories (Ethics, Bias & Morality; Empathy, Intimacy & Trust; Respecting Others' Rights; Human Preferences; Religion), where the paper states that reference answers were 'determined with careful consideration of Iranian cultural norms and societal conventions.' The authors should report human accuracy on a stratified sample, and agreement statistics among annotators or external judges, before interpreting low model scores as model failure.
  2. [Methods / Design Principles] The benchmark was explicitly designed to be challenging for Llama3-70B, and the text states that questions were selected with that goal in mind. This makes the low Llama3-70B score partly a selection artifact rather than an independent measure of capability. To support the general claim that current LLMs are far from solving the benchmark, the paper needs a human ceiling: competent Persian-speaking humans should score substantially above the models, ideally on a per-category basis. Without this, the 'far from solving' conclusion is not established by the reported numbers.
  3. [Technical Validation / Table 3] The evaluation protocol for Table 3 is underspecified. The paper does not state who assigned the Correct/Wrong/Semi-correct labels (human evaluators or automatic methods), how many evaluators scored each response, what instructions or rubrics were used for descriptive answers, or what prompting setup and decoding parameters were used for the three models. These details are necessary for reproducibility and for interpreting the accuracy numbers, particularly for open-ended generation categories where scoring is subjective.
  4. [Data Availability (throughout)] No dataset link, repository, or release plan is provided, and the Saba platform is described only through screenshots. For a benchmark paper in this venue, the data must be made available (with appropriate metadata and a clear license). The paper also does not define the exact schema of the 'Reference' and 'Label' fields or explain how the Chain-of-Thought labels are used in evaluation; this information should be included in the data description.
minor comments (4)
  1. [Abstract] There are typos in the abstract: 'bechmark' should be 'benchmark' and the sentence 'This bechmark incorporates linguistics, cultural, and local considerations' should be 'This benchmark incorporates linguistic, cultural, and local considerations.'
  2. [Table 3] The paper reports an overall 'Average' row but does not specify whether the average is weighted by the number of samples per subcategory or is an unweighted mean of the subcategory accuracies. Since the subcategories vary in size from 50 to 500 samples, the authors should clarify this and report both if needed.
  3. [References] Reference [31] is cited as the source of the Persian GATS dataset, but the reference list entry for [31] is 'HmBlogs: A big general Persian corpus.' Please verify and correct this citation or provide the actual GATS dataset reference.
  4. [Domains and Categories / VII.d] The example in the Empathy, Intimacy & Trust section states that answering 'talk to the thief and let them go if they show regret' is incorrect. This is presented as a moral judgment without supporting evidence or discussion of alternative ethical views; given the paper's reliance on such items, the authors should either provide a principled justification or acknowledge the normative nature of these keys explicitly.

Circularity Check

1 steps flagged · score 3.0 of 10

Difficulty for Llama3-70B was a curation input, so its low score is partly a selection artifact.

  1. fitted input called prediction [Technical Validation, p. 19 (penultimate paragraph) and Abstract]
    "In designing this benchmark, efforts were made to ensure that the questions are diverse and challenging. At the time of planning this benchmark, the Llama3-70B model was among the newest and most powerful models available. Efforts were made to design and select questions that are challenging for this model."

    The paper explicitly selected and designed questions to be hard for Llama3-70B, then reports Llama3-70B's average accuracy (0.47) and uses below-50% scores to conclude that 'current language models are still far from being able to solve this benchmark.' For Llama3-70B, the observed difficulty is an input to item selection, not an independent result: a set curated to defeat a model is expected to defeat it. The low number therefore re-states the curation criterion rather than providing independent evidence about the model's general competence. Independent support would require out-of-sample items or a human accuracy ceiling; neither is reported.

full rationale

FarsEval-PKBETS is primarily a benchmark resource rather than a derived mathematical claim, so most of the paper is self-contained. No load-bearing self-citation chain is present: the only self-citations (HmBlogs, ParsMap) are used as data sources and are not invoked to forbid alternatives. The main circular element is the design-to-evaluation loop for Llama3-70B: because the questions were deliberately curated to be challenging for that model, its below-50% score is partly a selection artifact rather than a clean falsification. PersianMind and Dorna were not used in item selection, so their low scores retain independent empirical content, and the benchmark itself remains a useful, externally reusable resource. Missing human baselines and inter-annotator reliability statistics are validity concerns about the reference answers, not circularity in the derivation. The score of 3 reflects one partial selection-driven reduction while acknowledging that the central artifact and the other model measurements do not reduce to their inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim is empirical and artifact-based rather than mathematical. It depends on the correctness of reference answers, the representativeness of the three tested models, the reliability of the scoring rubric, and the quality of source datasets. None of these is independently verified in the paper. There are no fitted numerical parameters or invented theoretical entities.

assumptions (4)
  • domain assumption Human-authored reference answers are correct and uncontroversial, including normative ethics and social knowledge items.
    The paper relies on in-house experts and two supervisors, but reports no inter-annotator agreement or external validation. If reference answers are wrong or contested, all accuracy figures lose meaning.
  • domain assumption The three evaluated models are representative enough of current LLMs to support the 'far from being able' conclusion.
    The sample includes one large general model (Llama3-70B) and two smaller Persian-tuned models; no API or frontier models are tested, and no human baseline is provided.
  • domain assumption The Correct/Wrong/Semi-correct scoring rubric was applied consistently and reliably.
    The rubric is described informally; the paper does not state whether labels came from automatic scripts, human judges, or both, and does not report multiple runs or agreement.
  • domain assumption Source datasets used as raw material are accurate enough after revision.
    Items are drawn from medical exams, MirasIrony, ArmanEmo, ETHICS, ParsMap, and driving tests; the paper notes revisions but gives no detailed error analysis or statistics on how many labels were changed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models." pith.science (2026). https://pith.science/paper/FQGHVF23

@misc{pith2026250414690,
  author       = {Pith},
  title        = {Pith review of: FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FQGHVF23}},
  note         = {Machine review of arXiv:2504.14690}
}
read the original abstract

Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attention. This paper introduces FarsEval-PKBETS benchmark, a subset of FarsEval project for evaluating large language models in Persian. This benchmark consists of 4000 questions and answers in various formats, including multiple choice, short answer and descriptive responses. It covers a wide range of domains and tasks,including medicine, law, religion, Persian language, encyclopedic knowledge, human preferences, social knowledge, ethics and bias, text generation, and respecting others' rights. This bechmark incorporates linguistics, cultural, and local considerations relevant to the Persian language and Iran. To ensure the questions are challenging for current LLMs, three models -- Llama3-70B, PersianMind, and Dorna -- were evaluated using this benchmark. Their average accuracy was below 50%, meaning they provided fully correct answers to fewer than half of the questions. These results indicate that current language models are still far from being able to solve this benchmark

Figures

Figures reproduced from arXiv: 2504.14690 by the authors.

Figure 1
Figure 1. Head categories of FarsEval-PKBETS with the number of samples. Methods Design Principles Each record of FarsEval-PKBETS contains a question and its correct (reference) answer (if applicable). General rules have been considered to construct FarsEval-PKBETS [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Example of each question/answer format in FarsEval [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Screenshots of the annotation platform (Saba); Top left: Panel for evaluating model responses to questions. Top right: Data submission panel. Bottom: Main page for viewing and entering data in a reviewer account. In summary, Saba includes the following features: o Management of task categories o Submission of multiple-choice, short-answer, and descriptive questions o Recording details and revision history for each q… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 25 canonical work pages

  1. [1]

    Chen, Y. et al. See What LLMs Cannot Answer: A Self -Challenge Framework for Uncovering LLM Weaknesses. Preprint at http://arxiv.org/abs/2408.08978 (2024)

  2. [2]

    Laskar, M. T. R. et al. A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations. in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing 13785–13816 (Association for Computational Linguistics, Miami, Florida, USA, 2024). doi:10.18653/v1/2024.emnlp-main.764. 21

  3. [3]

    Wang, A. et al. GLUE: A Multi -Task Benchmark and Analys is Platform for Natural Language Understanding. in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP 353–355 (Association for Computational Linguistics, Brussels, Belgium, 2018). doi:10.18653/v1/W18-5446

  4. [4]

    Wang, A. et al. SuperGLUE: a stickier benchmark for general -purpose language understanding systems. in Proceedings of the 33rd International Conference on Neural Information Processing Systems (Curran Associates Inc., Red Hook, NY, USA, 2019)

  5. [5]

    & Choi, Y

    Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. & Choi, Y. HellaSwag: Can a Machine Really Finish Your Sentence? in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics 4791–4800 (Association for Computational Linguistics, F lorence, Italy, 2019). doi:10.18653/v1/P19-1472

  6. [6]

    Hendrycks, D. et al. Measuring Massive Multitask Language Understanding. in Proceedings of the International Conference on Learning Representations (ICLR) (2021)

  7. [7]

    Wang, Y. et al. MMLU-Pro: A More Robust and Challenging Multi -Task Language Understanding Benchmark. Preprint at https://doi.org/10.48550/arXiv.2406.01574 (2024)

  8. [8]

    & Evans, O

    Lin, S., Hilton, J. & Evans, O. TruthfulQA: Measuring How Models Mimic Human Falsehoods. in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 3214–3252 (Association for Computational Linguistics, Dublin, Ireland, 2022). doi:10.18653/v1/2022.acl-long.229

Show all 42 references
  1. [9]

    Srivastava, A. et al. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research (2023)

  2. [10]

    Suzgun, M. et al. Challenging BIG-Bench Tasks and Whether Chain -of-Thought Can Solve Them. arXiv preprint arXiv:2210.09261 (2022)

  3. [11]

    Rein, D. et al. GPQA: A Graduate -Level Google -Proof Q&A Benchmark. in First Conference on Language Modeling (2024)

  4. [12]

    Hu, J. et al. XTREME: a massively multilingual multi -task benchmark for evaluating cross -lingual generalization. in Proceedings of the 37th International Conference on Machine Learning (JMLR.org, 2020)

  5. [13]

    Hendrycks, D. et al. Aligning {AI} With Shared Human Values. in International Conference on Learning Representations (2021)

  6. [14]

    & Sun, D

    Kotek, H., Dockum, R. & Sun, D. Gender bias and stereotypes in Large Language Models. in Proceedings of The ACM Collective Intelligence Conference 12–24 (ACM, Delft Netherlands, 2023). doi:10.1145/3582269.3615599

  7. [15]

    Zhang, B. et al. A dataset for evaluating clinical research claims in large language models. Sci Data 12, 86 (2025). 22

  8. [16]

    Niu, Z. et al. PharmaBench: Enhancing ADMET benchmarks with large language models. Sci Data 11, 985 (2024)

  9. [17]

    & Atarodi, A

    Shojaee-Mend, H., Mohebbati, R., Amiri, M. & Atarodi, A. Evaluating the strengths and weaknesses of large language models in answering neurophysiology questions. Sci Rep 14, 10785 (2024)

  10. [18]

    Khorshidi, H. et al. Application of ChatGPT in multilingual medical education: How does ChatGPT fare in 2023 ’s Iranian residency entrance examination. Informatics in Medicine Unlocked 41, 101314 (2023)

  11. [19]

    Abaskohi, A. et al. Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT. in Proceedings of the 2024 Joint International C onference on Computational Linguistics, Language Resources and Evaluation (LREC -COLING 2024) (eds. Calzolari, N. et al...

  12. [20]

    Hendrycks, D. et al. Measuring Mathematical Problem Solving With the MATH Dataset. in Thirty- fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) (2021)

  13. [21]

    https://github.com/ParsBench/ParsBench (2024)

    ParsBench. https://github.com/ParsBench/ParsBench (2024)

  14. [22]

    Khashabi, D. et al. ParsiNLU: A Suite of Language Understanding Challenges for Persian. Preprint at https://doi.org/10.48550/arXiv.2012.06154 (2021)

  15. [23]

    Singh, S. et al. Aya Dataset: An Open -Access Collection for Multilingual Instruction Tuning. in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds. Ku, L. -W., Martins, A. & Srikumar, V.) 11521 –11567 (Associat...

  16. [24]

    https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned- llm

    Free Dolly: Introducing the World’s First Truly Open Inst ruction-Tuned LLM. https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned- llm

  17. [25]

    Ghahroodi, O. et al. Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language? in First Conference on Language Modeling (2024)

  18. [26]

    Myung, J. et al. BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages. in The Thirty -eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024)

  19. [27]

    Bandarkar, L. et al. The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants. in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 749–775 (Association for Computational Lingu istic...

  20. [28]

    https://huggingface.co/spaces/PartAI/open-persian-llm- leaderboard (2024)

    Open Persian LLM Leaderboard (v1.0.0). https://huggingface.co/spaces/PartAI/open-persian-llm- leaderboard (2024). 23

  21. [29]

    Advancing Persian LLM Evaluation

    Anonymous. Advancing Persian LLM Evaluation. in The 2025 Annual Conference of the Nations of the Americas Chapter of the ACL (2025)

  22. [30]

    Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. in Proceedings of the 36th International Conference on Neural Information Processing Systems (Curran Associates Inc., Red Hook, NY, USA, 2022)

  23. [31]

    Khansari, H. M. & Shamsfard, M. HmBlogs: A big general Persian corpus. Preprint at https://doi.org/10.48550/arXiv.2111.02362 (2021)

  24. [32]

    Mirzaee, H., Peymanfard, J., Moshtaghin, H. H. & Zeinali, H. ArmanEmo: A Persian Dataset for Text-based Emotion Detection. Preprint at https://doi.org/10.48550/arXiv.2207.11808 (2022)

  25. [33]

    Golazizian, P. et al. Irony Detection in Persian Language: A Transfer Learning Approach Using Emoji Prediction. in Proceedings of the Twelfth Language Resources and Evaluation Conference (eds. Calzolari, N. et al.) 2839–2845 (European Language Resources Association, Marseille,...

  26. [34]

    Zeng, Q. & Li, A. -R. A Survey in Automatic Irony Processing: Linguistic, Cognitive, and Multi -X Perspectives. in Proceedings of the 29th International Conference on Computational Linguistics (eds. Calzolari, N. et al.) 824 –836 (International Committee on Computational Lingu...

  27. [35]

    Sorin, V. et al. Large Language Models and Empathy: Systematic Review. J Med Internet Res 26, e52597 (2024)

  28. [36]

    & Shamsfard, M

    Tajalli, V., Kalantari, F. & Shamsfard, M. Developing an Informal -Formal Persian Corpus. Preprint at https://doi.org/10.48550/arXiv.2308.05336 (2023)

  29. [37]

    & Dixon, L

    Wulczyn, E., Thain, N. & Dixon, L. Ex Machina: Personal Attacks Seen at Scale. in Proceedings of the 26th International Conference on World Wide Web 1391–1399 (International World Wide Web Conferences Steering Committee, Perth Australia, 2017). doi:10.1145/3038912.3052591

  30. [38]

    & Vasserman, L

    Borkan, D., Dixon, L., Sorensen, J., Thain, N. & Vasserman, L. Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification. in Companion Proceedings of The 2019 World Wide Web Conference 491–500 (ACM, San Francisco USA, 2019). doi:10.1145/3308560.3317593

  31. [39]

    Villate-Castillo, G., Ser, J. D. & Urquijo, B. S. A Systematic Review of Toxicity in Large Language Models: Definitions, Datasets, Detectors, Detoxification Methods and Challenges. Preprint at https://doi.org/10.21203/rs.3.rs-4621646/v1 (2024)

  32. [40]

    Grattafiori, A. et al. The Llama 3 Herd of Models. Preprint at https://doi.org/10.48550/arXiv.2407.21783 (2024)

  33. [41]

    & Dousti, M

    Rostami, P., Salemi, A. & Dousti, M. J. PersianMind: A Cross -Lingual Persian -English Large Language Model. Preprint at https://doi.org/10.48550/arXiv.2401.06466 (2024)

  34. [42]

    https://huggingface.co/PartAI/Dorna-Llama3-8B-Instruct (2024)

    Dorna-Llama3-8B-Instruct. https://huggingface.co/PartAI/Dorna-Llama3-8B-Instruct (2024). 24

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.