REVIEW 4 major objections 4 minor 1 cited by
ELAB: Extensive LLM Alignment Benchmark in Persian Language
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A unified benchmark scores Persian LLMs on safety, fairness, and social norms.
desk verdict New Persian alignment datasets are welcome, but the leaderboard rests on an unvalidated GPT-4o-mini judge — treat scores as provisional, not definitive. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the benchmark suite itself: three data modes (translated, generated, collected) aggregated across five new Persian datasets (ProhibiBench-fa, SafeBench-fa, FairBench-fa, SocialBench-fa, GuardBench-fa) plus four translated benchmarks (Anthropic-fa, AdvBench-fa, HarmBench-fa, DecodingTrust-fa). Scoring is done by an LLM-as-a-judge protocol in which GPT-4o-mini assigns 0–10 scores using four handcrafted system prompts (safety, fairness, social norms, and a cultural-harmlessness variant); the mean score across questions becomes each model's alignment score for the leaderboard.
What would settle it
A human-annotation study on a stratified sample of, say, 200 responses per model, comparing human alignment ratings to GPT-4o-mini's scores; if the judge's scores diverge systematically from human judgments (e.g., ranking models differently), the leaderboard claims would fail.
Extended reading notes
Core claim
The central claim is that Persian LLM alignment can be measured along three interdependent axes—safety, fairness, and social norms—using a unified benchmark suite that combines translated, synthetic, and naturally collected Persian data. The paper asserts that this is the first large-scale structured framework of its kind for Persian, and that cultural specifics such as 'taarof' (deference rituals) and 'aberoo' (social dignity) require indigenous datasets rather than mere translation. The reported results show Gemma-2-9B-it ranking highest across most benchmarks, with Aya-Expanse-8B close behind, while Qwen2.5 variants lag on fairness and social-norm compliance.
Load-bearing premise
The whole evaluation rests on the premise that GPT-4o-mini, given a hand-written prompt, scores Persian responses on a 0–10 scale in a way that matches what a careful Persian-speaking human would judge to be safe, fair, and socially appropriate.
Editorial extensions
If this is right
- If the framework holds, Persian LLM developers can compare models on safety, fairness, and social norms with a single public leaderboard.
- The divergence in embedding distributions between translated and native data (Figure 2) implies translated benchmarks alone are insufficient for assessing cultural alignment.
- The framework provides a template for building alignment benchmarks for other underrepresented languages by mixing translation, generation, and natural collection.
- Reported scores indicate Gemma-2-9B-it is the strongest aligned open-weight Persian model under 10B parameters, and Qwen2.5 models lag on fairness and norms.
Reading between the lines
- Because GPT-4o-mini both labels generated data and scores responses, the benchmark may partly measure the judge model's own safety and fairness preferences; a different judge or a human panel could change rankings.
- ProhibiBench-fa was produced with the Do Anything Now (DAN) jailbreak, so the suite tests robustness to that specific attack family; models may still fail on other adversarial strategies not represented.
- GuardBench-fa's social-norm scores are strikingly low for several models (e.g., around 40–50 for the 2B and 3B models), suggesting the collected data is the hardest axis; this difference could serve as a proxy for cultural alignment difficulty.
- The framework could be extended to other Persian varieties (Dari, Tajik) or neighboring languages to test whether cultural constructs travel across dialects.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ELAB, a Persian-language alignment benchmark that combines translated versions of English safety/fairness benchmarks (Anthropic, AdvBench, HarmBench, DecodingTrust), newly generated Persian datasets (SafeBench-fa, FairBench-fa, SocialBench-fa, ProhibiBench-fa), and a naturally collected dataset (GuardBench-fa). The authors evaluate seven open-weight models under 10B parameters using GPT-4o-mini as an LLM judge, report scores per dataset and category, and present a public leaderboard. The central claim is that ELAB is the first large-scale, structured framework for evaluating the alignment of Persian LLMs, with safety, fairness, and social norms as the three alignment dimensions.
Significance. If the benchmark is valid, it addresses a genuine gap: there is little Persian-specific alignment evaluation infrastructure, and the paper contributes several translated and culturally tailored datasets, a unified categorization, and a public leaderboard that can be reused. The translation pipeline includes back-translation and native-speaker checks, and the generated datasets were reviewed by at least one human annotator, which are strengths. However, the significance is currently conditional on the validity of the GPT-4o-mini judge and on the robustness of the reported scores; the paper's own Limitations section acknowledges reliance on LLM-as-a-judge without providing validation evidence. The claimed 'first large-scale' status is plausible but is weakened by small category sizes and by the absence of any human or alternative-judge validation of the scoring.
major comments (4)
- [Section 4, step 3; Appendix A] The entire leaderboard rests on a single unvalidated judge model, GPT-4o-mini, with hand-written prompts. No human evaluation, inter-annotator agreement, or comparison with an alternative judge is reported. Because the same model is also used to translate the English datasets (Section 3.1), classify generated items (Section 3.2), and label GuardBench-fa (Section 3.3), any systematic Persian-language blind spot in GPT-4o-mini would propagate through both dataset construction and evaluation. The claim that strong-model judges are reliable, stated in the Limitations section, is an assumption rather than evidence. Please provide at least a sample-based human evaluation or a second-judge comparison on a subset, and report agreement metrics, before using the leaderboard to rank models.
- [Table 1, HarmBench-fa row; Section 3.2.3] Several reported scores are based on extremely small samples. Table 1 lists only 3 items for HarmBench-fa Social Norm and 6 for HarmBench-fa Fairness, and Section 3.2.3 reports a final SocialBench-fa size of 17 while Table 1 lists 16. The Table 3 scores for these categories, including values of 100, are therefore not statistically meaningful and cannot support comparative claims such as 'Gemma-2-9B-it consistently achieves the highest scores.' Please either aggregate these small categories, report confidence intervals, or explicitly flag them as illustrative only.
- [Appendix A, Safety/Fairness/Social Norm evaluation prompts] The judge prompts assign a score of 0 when 'the answer is in english or any other language except Persian.' This conflates language with alignment and likely penalizes multilingual models that code-switch, such as Aya-Expanse-8B. Since the goal is to measure alignment, not language adherence, this design choice can distort rankings. Please report how many responses per model were penalized for non-Persian language, or adjust the judge to score content alignment separately from language correctness.
- [Table 3] No error bars, confidence intervals, or significance tests are reported for any of the scores. The differences between adjacent models in Table 3 are often a few points (e.g., Anthropic-fa Safety: 79.40 for Qwen2.5-7B-Instruct vs. 80.06 for Ministral-8B), and without variance estimates these differences cannot be distinguished from noise. The conclusion that certain models 'consistently achieve the highest scores' needs statistical support, especially given the small item counts in several categories.
minor comments (4)
- [Table 1] The dataset name 'HarmBanch-fa' appears to be a typo for 'HarmBench-fa'; please fix it.
- [Section 3.2.3 and Table 1] The SocialBench-fa size is stated as 17 in the text and as 16 in Table 1; please reconcile the counts.
- [Section 3.3.1] The text says 6,146 offensive entries plus 505 swear-related entries, but Table 1 lists 6,651 for GuardBench-fa. The arithmetic is consistent (6,146 + 505 = 6,651), but the mismatch in phrasing (offensive vs. swear) should be clarified.
- [Abstract and Section 3.4] The claim of being the 'first large-scale, structured framework' for Persian LLM alignment would benefit from a comparison with any earlier Persian safety or bias evaluation efforts in the related-work section, to make the novelty precise.
Circularity Check
No significant circularity: the benchmark construction is self-contained and no prediction reduces to its inputs by construction.
full rationale
The paper is a dataset-and-evaluation benchmark; there is no derivation chain in which an output quantity is defined in terms of the quantity it is supposed to predict. The translated datasets are adapted from external English benchmarks, the generated datasets are produced by Command-R Plus and reviewed by a human annotator, and GuardBench-fa is collected from social media and manually reviewed. The leaderboard scores are computed by GPT-4o-mini as an external judge on the responses of the seven evaluated open-source models; the judge is not a parameter fitted to the data, and no equation makes the scores equal to an input. The paper's use of GPT-4o-mini for translation, classification, and evaluation is a measurement-choice assumption whose validity is not demonstrated, and the lack of human evaluation or inter-annotator agreement is a correctness and validity risk, not a circularity. Self-citations, such as the mention of PersianMMLU in related work, are not load-bearing. Therefore no circular step can be exhibited under the required standard.
Assumptions & free parameters
assumptions (4)
- domain assumption GPT-4o-mini translation preserves semantic meaning and cultural nuance after back-translation and native speaker review.
- domain assumption GPT-4o-mini-as-a-judge provides valid alignment scores for Persian on a 0 to 10 scale.
- domain assumption Original English benchmarks (Anthropic, AdvBench, HarmBench, DecodingTrust) are valid source constructs for Persian safety, fairness, and social norms after translation.
- domain assumption Command-R Plus and GPT-4o-mini classified generated items into categories accurately enough to support benchmark construction.
Cite this review
Pith. "Pith review of ELAB: Extensive LLM Alignment Benchmark in Persian Language." pith.science (2026). https://pith.science/paper/UUS2554C
@misc{pith2026250412553,
author = {Pith},
title = {Pith review of: ELAB: Extensive LLM Alignment Benchmark in Persian Language},
year = {2026},
howpublished = {\url{https://pith.science/paper/UUS2554C}},
note = {Machine review of arXiv:2504.12553}
}
read the original abstract
This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing LLM evaluation frameworks by adapting them to Persian linguistic and cultural contexts. This benchmark creates three types of Persian-language benchmarks: (i) translated data, (ii) new data generated synthetically, and (iii) new naturally collected data. We translate Anthropic Red Teaming data, AdvBench, HarmBench, and DecodingTrust into Persian. Furthermore, we create ProhibiBench-fa, SafeBench-fa, FairBench-fa, and SocialBench-fa as new datasets to address harmful and prohibited content in indigenous culture. Moreover, we collect extensive dataset as GuardBench-fa to consider Persian cultural norms. By combining these datasets, our work establishes a unified framework for evaluating Persian LLMs, offering a new approach to culturally grounded alignment evaluation. A systematic evaluation of Persian LLMs is performed across the three alignment aspects: safety (avoiding harmful content), fairness (mitigating biases), and social norms (adhering to culturally accepted behaviors). We present a publicly available leaderboard that benchmarks Persian LLMs with respect to safety, fairness, and social norms at: https://huggingface.co/spaces/MCILAB/LLM_Alignment_Evaluation.
Figures
Forward citations
Cited by 1 Pith paper
-
We Politely Insist: Your LLM Must Learn the Persian Art of Taarof
A new 450-scenario benchmark shows that LLMs lag native Persian speakers by 40 to 48 points on taarof-expected interactions, and that fine-tuning on the benchmark narrows the gap.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Cohere For AI . 2024. https://doi.org/10.57967/hf/3135 c4ai-command-r-plus-08-2024
doi:10.57967/hf/3135 2024
-
[4]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, and 26 others. 2024. https://arxiv.org/abs/2412.04261 Ay...
arXiv 2024
-
[5]
Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, and Mohammad Manthouri. 2021. https://doi.org/10.1007/s11063-021-10528-4 Parsbert: Transformer-based model for persian language understanding . Neural Processing Letters, 53(6):3831–3847
-
[6]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, and 1 others. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858
arXiv 2022
-
[7]
Omid Ghahroodi, Marzia Nouri, Mohammad Vali Sanian, Alireza Sahebi, Doratossadat Dastgheib, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah, and Mohammad Hossein Rohban. 2024. https://openreview.net/forum?id=yIEyHP7AvH Khayyam challenge (persian MMLU ): Is your LLM truly wise to the persian language? In First Conference on Language Modeling
2024
-
[8]
Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabagdi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, and 6 others. 2021. https://d...
Show all 22 references
-
[9]
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024. https://doi.org/10.18653/v1/2024.findings-acl.235 SALAD -bench: A hierarchical and comprehensive safety benchmark for large language models . In Findings of the Association for ...
2024 doi
-
[11]
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023 b . Trustworthy llms: a survey and guideline for evaluating large language models' alignment. arXiv preprint arXiv:2308.05374
2023 arXiv
-
[13]
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, and 1 others. 2024 b . Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249
2024 arXiv
-
[14]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. https://doi.org/10.18653/v1/2022.findings-acl.165 BBQ : A hand-built bias benchmark for question answering . In Findings of the Association for ...
2022 doi
-
[15]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwe...
2025 arXiv
-
[16]
Paul R \"o ttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. https://doi.org/10.18653/v1/2024.naacl-long.301 XST est: A test suite for identifying exaggerated safety behaviours in large language models . In Proceedings of the 2024 Co...
2024 doi
-
[17]
do anything now
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang. 2024 a . “do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security, pages 1671--1685
2024
-
[18]
do anything now
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024 b . " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, p...
2024
-
[19]
Ehsan Taher, Seyed Abbas Hoseini, and Mehrnoush Shamsfard. 2020. https://arxiv.org/abs/2003.08875 Beheshti-ner: Persian named entity recognition using bert . Preprint, arXiv:2003.08875
2020 arXiv
-
[20]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[21]
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, and et al. 2023 a . Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In NeurIPS
2023
-
[22]
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, and 1 others. 2023 b . Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In NeurIPS
2023
-
[23]
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. https://arxiv.org/abs/2309.07045 Safetybench: Evaluating the safety of large language models . Preprint, arXiv:2309.07045
2024 arXiv
-
[25]
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 b . Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.