{"id":"c2d1d65a-23be-42fa-935f-e7cc31add553","arxiv_id":"2504.12553","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ELAB provides the first Persian LLM alignment benchmark uniting translated, synthetic, and collected datasets for safety, fairness, and social norms evaluation.","lead":"This paper introduces ELAB, a Persian-language benchmark suite that translates existing English safety and bias datasets and adds new Persian-specific datasets to measure LLM safety, fairness, and social norms. It evaluates seven open-source models with GPT-4o-mini as judge and publishes a public leaderboard.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The entire leaderboard rests on unvalidated GPT-4o-mini judge scores for Persian; if those scores are biased, the framework's central claim and model rankings lack support.","rationale":"The reader identified the unvalidated GPT-4o-mini judge as the weakest assumption, and my stress-test pass converges on the same point. The paper's strongest claim is that it provides a large-scale, structured framework for Persian LLM alignment evaluation. That framework is operationalized entirely through LLM-as-a-judge scoring; the datasets, category labels, and leaderboard numbers all flow from GPT-4o-mini. There is no internal evidence that this judge produces valid Persian alignment scores: no human agreement study, no comparison with another judge, and no analysis of prompt sensitivity. The judge prompts themselves introduce a potentially severe scoring rule (0 for any non-Persian response) that is not discussed and could distort cross-model comparisons. Because this is a measurement-validity problem rather than a mere artifact-availability issue, it is genuinely load-bearing. However, I do not think the concern warrants rejection: the datasets may still be useful resources, and the paper explicitly discloses its reliance on LLM-as-a-judge in the Limitations section. The appropriate verdict is conditional on releasing the judge outputs and providing human-validation evidence, which matches the reader's CONDITIONAL verdict. No change to the reader's verdict is needed.","tokens_in":11795,"tokens_out":3248,"duration_ms":36943,"concrete_test":"Sample approximately 300 responses stratified across all seven models and the main dataset families, have three native Persian-speaking annotators rate each response on the same 0-10 rubric used by the judge, and compute inter-annotator agreement (e.g., ICC or weighted kappa) and the correlation between human and GPT-4o-mini scores. Then recompute the model rankings using human scores; if Spearman rank correlation is below about 0.7, or if the top-ranked model changes, the judge scores are not validated and the leaderboard claims are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ELAB is the first large-scale, structured framework for evaluating the alignment of Persian LLMs. Every quantitative result in Table 3, and therefore every leaderboard rank, is produced by a single judge model, GPT-4o-mini, scoring model outputs on a 0-10 scale using four hand-written prompts in Appendix A. No human evaluation, inter-annotator agreement, or alternative-judge comparison is reported. The same GPT-4o-mini model is also used to translate the English datasets, classify synthetic and collected data into safety/fairness/social-norm categories, and label GuardBench-fa, so any systematic Persian-language blind spot in this model propagates through dataset construction and evaluation alike. A concrete distortion is visible in the judge prompts: a score of 0 is assigned if the answer is in English or any language other than Persian. This penalizes a model that responds correctly but in English, and it is especially consequential for multilingual models such as Aya-Expanse, which may code-switch. The paper's Limitations section asserts that strong-model judges are reliable because this is common practice, but that is an assumption, not evidence, and it is particularly fragile for Persian, where the judge's competence is unverified. Without judge validation, the claim of providing a framework for evaluating Persian LLM alignment rests on an unvalidated measurement instrument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ELAB, a Persian-language alignment benchmark that combines translated versions of English safety/fairness benchmarks (Anthropic, AdvBench, HarmBench, DecodingTrust), newly generated Persian datasets (SafeBench-fa, FairBench-fa, SocialBench-fa, ProhibiBench-fa), and a naturally collected dataset (GuardBench-fa). The authors evaluate seven open-weight models under 10B parameters using GPT-4o-mini as an LLM judge, report scores per dataset and category, and present a public leaderboard. The central claim is that ELAB is the first large-scale, structured framework for evaluating the alignment of Persian LLMs, with safety, fairness, and social norms as the three alignment dimensions.","tokens_in":12065,"tokens_out":2167,"duration_ms":23507,"significance":"If the benchmark is valid, it addresses a genuine gap: there is little Persian-specific alignment evaluation infrastructure, and the paper contributes several translated and culturally tailored datasets, a unified categorization, and a public leaderboard that can be reused. The translation pipeline includes back-translation and native-speaker checks, and the generated datasets were reviewed by at least one human annotator, which are strengths. However, the significance is currently conditional on the validity of the GPT-4o-mini judge and on the robustness of the reported scores; the paper's own Limitations section acknowledges reliance on LLM-as-a-judge without providing validation evidence. The claimed 'first large-scale' status is plausible but is weakened by small category sizes and by the absence of any human or alternative-judge validation of the scoring.","major_comments":[{"comment":"The entire leaderboard rests on a single unvalidated judge model, GPT-4o-mini, with hand-written prompts. No human evaluation, inter-annotator agreement, or comparison with an alternative judge is reported. Because the same model is also used to translate the English datasets (Section 3.1), classify generated items (Section 3.2), and label GuardBench-fa (Section 3.3), any systematic Persian-language blind spot in GPT-4o-mini would propagate through both dataset construction and evaluation. The claim that strong-model judges are reliable, stated in the Limitations section, is an assumption rather than evidence. Please provide at least a sample-based human evaluation or a second-judge comparison on a subset, and report agreement metrics, before using the leaderboard to rank models.","section":"Section 4, step 3; Appendix A"},{"comment":"Several reported scores are based on extremely small samples. Table 1 lists only 3 items for HarmBench-fa Social Norm and 6 for HarmBench-fa Fairness, and Section 3.2.3 reports a final SocialBench-fa size of 17 while Table 1 lists 16. The Table 3 scores for these categories, including values of 100, are therefore not statistically meaningful and cannot support comparative claims such as 'Gemma-2-9B-it consistently achieves the highest scores.' Please either aggregate these small categories, report confidence intervals, or explicitly flag them as illustrative only.","section":"Table 1, HarmBench-fa row; Section 3.2.3"},{"comment":"The judge prompts assign a score of 0 when 'the answer is in english or any other language except Persian.' This conflates language with alignment and likely penalizes multilingual models that code-switch, such as Aya-Expanse-8B. Since the goal is to measure alignment, not language adherence, this design choice can distort rankings. Please report how many responses per model were penalized for non-Persian language, or adjust the judge to score content alignment separately from language correctness.","section":"Appendix A, Safety/Fairness/Social Norm evaluation prompts"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any of the scores. The differences between adjacent models in Table 3 are often a few points (e.g., Anthropic-fa Safety: 79.40 for Qwen2.5-7B-Instruct vs. 80.06 for Ministral-8B), and without variance estimates these differences cannot be distinguished from noise. The conclusion that certain models 'consistently achieve the highest scores' needs statistical support, especially given the small item counts in several categories.","section":"Table 3"}],"minor_comments":[{"comment":"The dataset name 'HarmBanch-fa' appears to be a typo for 'HarmBench-fa'; please fix it.","section":"Table 1"},{"comment":"The SocialBench-fa size is stated as 17 in the text and as 16 in Table 1; please reconcile the counts.","section":"Section 3.2.3 and Table 1"},{"comment":"The text says 6,146 offensive entries plus 505 swear-related entries, but Table 1 lists 6,651 for GuardBench-fa. The arithmetic is consistent (6,146 + 505 = 6,651), but the mismatch in phrasing (offensive vs. swear) should be clarified.","section":"Section 3.3.1"},{"comment":"The claim of being the 'first large-scale, structured framework' for Persian LLM alignment would benefit from a comparison with any earlier Persian safety or bias evaluation efforts in the related-work section, to make the novelty precise.","section":"Abstract and Section 3.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper gives Persian NLP something it didn't have — a combined safety/fairness/social-norms benchmark with genuinely new culturally grounded data. The evaluation numbers, though, are only as good as an unvalidated GPT-4o-mini judge, so I'd treat the leaderboard as provisional, not as a reliable ranking.\n\nWhat's genuinely new: five Persian datasets (ProhibiBench-fa, SafeBench-fa, FairBench-fa, SocialBench-fa, GuardBench-fa), with GuardBench-fa collected from real social media and offensive-language content. The construction is careful: back-translation verification, native-speaker cultural checks, human review of generated items. The paper is transparent about the pipeline and includes all prompts in the appendix. The SNE analysis showing a distributional gap between translated and native data makes the case that translation alone doesn't capture Persian norms — that's a strong, useful point.\n\nWhere it gets soft: the entire evaluation is scored by GPT-4o-mini using four hand-written prompts. There's no human agreement sample, no second judge, no error bars. The judge prompt assigns 0 to any response not in Persian, which penalizes a perfectly good answer that code-switches — a real concern for multilingual models like Aya-Expanse. And because the same model family translated the English datasets, classified the generated items, and judged the outputs, any Persian-specific blind spot in the model propagates through construction and evaluation alike. The Limitations section defends LLM-as-judge because it's common practice, but that's an assertion, not evidence. For a language where the judge's competence is unverified, this matters.\n\nSome of the per-category numbers are also based on very few items: SocialBench-fa has 16-17 items, HarmBench-fa social norm has 3. Those scores are essentially unstable. The paper would be stronger if it reported confidence intervals or at least flagged low-N categories.\n\nMy overall read: the resource is a solid contribution, and it deserves peer review. But it should be reframed as a dataset paper with a preliminary leaderboard, not as a definitive evaluation. The authors should release the data and judge outputs, add a small human-validated sample, and soften the claim about being 'the first large-scale structured framework' — the framework is fine, the validation is thin.\n\nWho this is for: anyone working on Persian LLM evaluation, multilingual safety, or culturally grounded alignment. I'd cite it for the datasets, not for the model rankings. A serious referee should engage with it, with the close-reading question being judge validity.","headline":"New Persian alignment datasets are welcome, but the leaderboard rests on an unvalidated GPT-4o-mini judge — treat scores as provisional, not definitive.","tokens_in":12606,"tokens_out":2837,"would_cite":true,"duration_ms":27558,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A unified benchmark scores Persian LLMs on safety, fairness, and social norms.","keywords":["Persian LLM alignment","safety evaluation","fairness evaluation","social norms","LLM-as-a-judge","benchmark suite","red teaming","culturally grounded evaluation"],"falsifier":"A human-annotation study on a stratified sample of, say, 200 responses per model, comparing human alignment ratings to GPT-4o-mini's scores; if the judge's scores diverge systematically from human judgments (e.g., ranking models differently), the leaderboard claims would fail.","tokens_in":11640,"feed_emoji":"⚖️","tokens_out":4288,"duration_ms":36825,"temperature":0.7,"pith_summary":"The paper claims to provide the first large-scale, structured framework for evaluating whether Persian-language LLMs align with safety, fairness, and social norms. It builds the framework from three data sources: translations of English alignment benchmarks, newly generated culturally specific Persian datasets, and naturally collected Persian offensive and swear-word data. The authors evaluate seven open-weight models under 10 billion parameters using an LLM-as-a-judge protocol and publish a public leaderboard. If the framework works as claimed, Persian LLM developers gain a standard way to compare models on culturally grounded alignment rather than relying on translated English tests.","feed_headline":"New Persian benchmark scores LLMs on safety, fairness, and norms","feed_subtitle":"Public leaderboard ranks seven open models on culturally grounded Persian safety, fairness, and social-norm tests.","key_machinery":"The load-bearing mechanism is the benchmark suite itself: three data modes (translated, generated, collected) aggregated across five new Persian datasets (ProhibiBench-fa, SafeBench-fa, FairBench-fa, SocialBench-fa, GuardBench-fa) plus four translated benchmarks (Anthropic-fa, AdvBench-fa, HarmBench-fa, DecodingTrust-fa). Scoring is done by an LLM-as-a-judge protocol in which GPT-4o-mini assigns 0–10 scores using four handcrafted system prompts (safety, fairness, social norms, and a cultural-harmlessness variant); the mean score across questions becomes each model's alignment score for the leaderboard.","core_discovery":"The central claim is that Persian LLM alignment can be measured along three interdependent axes—safety, fairness, and social norms—using a unified benchmark suite that combines translated, synthetic, and naturally collected Persian data. The paper asserts that this is the first large-scale structured framework of its kind for Persian, and that cultural specifics such as 'taarof' (deference rituals) and 'aberoo' (social dignity) require indigenous datasets rather than mere translation. The reported results show Gemma-2-9B-it ranking highest across most benchmarks, with Aya-Expanse-8B close behind, while Qwen2.5 variants lag on fairness and social-norm compliance.","pith_inferences":["Because GPT-4o-mini both labels generated data and scores responses, the benchmark may partly measure the judge model's own safety and fairness preferences; a different judge or a human panel could change rankings.","ProhibiBench-fa was produced with the Do Anything Now (DAN) jailbreak, so the suite tests robustness to that specific attack family; models may still fail on other adversarial strategies not represented.","GuardBench-fa's social-norm scores are strikingly low for several models (e.g., around 40–50 for the 2B and 3B models), suggesting the collected data is the hardest axis; this difference could serve as a proxy for cultural alignment difficulty.","The framework could be extended to other Persian varieties (Dari, Tajik) or neighboring languages to test whether cultural constructs travel across dialects."],"forward_implications":["If the framework holds, Persian LLM developers can compare models on safety, fairness, and social norms with a single public leaderboard.","The divergence in embedding distributions between translated and native data (Figure 2) implies translated benchmarks alone are insufficient for assessing cultural alignment.","The framework provides a template for building alignment benchmarks for other underrepresented languages by mixing translation, generation, and natural collection.","Reported scores indicate Gemma-2-9B-it is the strongest aligned open-weight Persian model under 10B parameters, and Qwen2.5 models lag on fairness and norms."],"supporting_citations":[{"why":"Source of the Anthropic red-teaming data that was translated into Anthropic-fa for safety evaluation.","marker":"Ganguli et al. (2022)"},{"why":"Source of AdvBench, translated into AdvBench-fa to evaluate harmful-behavior resistance.","marker":"Zou et al. (2023b)"},{"why":"Source of HarmBench, translated into HarmBench-fa for standardized red-teaming and refusal evaluation.","marker":"Mazeika et al. (2024b)"},{"why":"Source of DecodingTrust, translated into DecodingTrust-fa for trustworthiness and fairness assessment.","marker":"Wang et al. (2023b)"},{"why":"Provides the topic and subtopic taxonomy used to generate SafeBench-fa, FairBench-fa, and SocialBench-fa.","marker":"Liu et al. (2023b)"},{"why":"Provides the 'Do Anything Now' (DAN) jailbreak technique used to generate the adversarial ProhibiBench-fa dataset.","marker":"Shen et al. (2024b)"},{"why":"Source of Command-R Plus, the generator model used to create the synthetic benchmark questions.","marker":"Cohere For AI (2024)"},{"why":"Supplies the BBQ bias taxonomy that is extended with Persian-centric fairness dimensions.","marker":"Parrish et al. (2022)"}],"fun_headline_variants":["Persian LLM benchmark covers safety, fairness, and norms","First Persian alignment benchmark for LLMs includes cultural norms","New leaderboard: Persian LLMs scored on safety, fairness, norms","Culturally grounded benchmark evaluates Persian LLM alignment","ELAB: Persian benchmark tests LLM safety, fairness, and norms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation rests on the premise that GPT-4o-mini, given a hand-written prompt, scores Persian responses on a 0–10 scale in a way that matches what a careful Persian-speaking human would judge to be safe, fair, and socially appropriate.","fun_headline_variants_meta":{"raw":{"variants":["Persian LLM benchmark covers safety, fairness, and norms","First Persian alignment benchmark for LLMs includes cultural norms","New leaderboard: Persian LLMs scored on safety, fairness, norms","Culturally grounded benchmark evaluates Persian LLM alignment","ELAB: Persian benchmark tests LLM safety, fairness, and norms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":2001,"prompt_tokens":926,"completion_tokens":1075,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":989}},"tokens_in":542,"tokens_out":1075,"duration_ms":11529,"temperature":1.0,"reasoning_tokens":989,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:27:58.916696+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-annotation study on a stratified sample of, say, 200 responses per model, comparing human alignment ratings to GPT-4o-mini's scores; if the judge's scores diverge systematically from human judgments (e.g., ranking models differently), the leaderboard claims would fail.","supporting_citations":[],"review_version":1}