{"id":"3bd802be-8b10-421d-b77d-bdec66a8d76f","arxiv_id":"2504.14690","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FarsEval-PKBETS is a new human-reviewed Persian benchmark of 4,000 questions where three tested models score below 50% correctly.","lead":"This paper introduces FarsEval-PKBETS, a new set of 4,000 Persian-language questions for testing AI language models across medicine, law, culture, writing, and ethics. The authors report that three current models answer fewer than half of the questions correctly, suggesting Persian AI still has large gaps.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human baseline or inter-annotator reliability is reported for newly authored items, so below-50% scores could indicate ambiguous or contested reference answers rather than genuine model deficiency; the 'far from solving' conclusion needs a human-accuracy check.","rationale":"I read the paper as a benchmark/data-descriptor claim: 4,000 human-generated/collected Persian QA items, with human review, are diverse and difficult enough that current LLMs score below 50%. The strongest evidence is Table 3: three models, including Llama3-70B, average 0.47/0.54, 0.19/0.23, and 0.30/0.39 on correct/semi-correct scoring. I checked whether the unweighted average hides a weighted result; a weighted recomputation from Table 1 and Table 3 gives Llama3 roughly 48% correct, so the below-50% statement is not an artifact of equal category weighting. The remaining weak point is external validity of the reference answers, not arithmetic. The paper's own design goal—'to design questions that challenge the model as much as possible'—makes the low scores partly by construction, but that is acceptable for a benchmark only if the gold answers are known to be human-solvable. For exam-sourced and fact-based categories (medical residency exams, Constitution of IRI, driving-test right-of-way questions) human solvability is plausible. For Ethics, Bias & Morality, Empathy, Respecting Others' Rights, Human Preferences, Religion, and subjective text-generation categories, the reference answers were created/reviewed by a small group with no reported agreement statistics; the paper even gives a normative example where it declares a response incorrect. Thus the load-bearing assumption is that these gold labels are uncontested. I did not find an internal inconsistency or a mathematical error; the concern is an external validity gap. The proposed human-baseline study would directly settle it. This matches the reader's weakest_assumption at least partially; the reader focused on normative ground truth, while I frame the same issue as missing human solvability/IAA evidence across all newly authored items. Because the paper could be made sound by adding such a study and releasing data/protocol, I do not move the verdict: CONDITIONAL remains appropriate.","tokens_in":15571,"tokens_out":8226,"duration_ms":80552,"concrete_test":"Run a human-validation study on a stratified random sample of ~400 records, oversampling Ethics/Bias & Morality, Respecting Others' Rights, Empathy/Intimacy & Trust, Human Preferences, and Religion. Have 3-5 independent Persian-speaking annotators (not authors; domain experts for medicine and law) answer each sampled question without seeing the reference answer. Compute agreement with the reference answer and pairwise inter-annotator agreement (Fleiss' or Cohen's kappa). If human-answer agreement is below ~90% or kappa below ~0.7 in any sampled category, the reference answers are not uncontested ground truth; the paper should then report human accuracy per category and qualify the 'far from solving' conclusion. If human agreement is near ceiling, the concern is resolved and the headline result is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FarsEval-PKBETS is a valuable resource, but the headline inference—that average accuracies below 50% show current LLMs are 'far from being able to solve' it—requires that the reference answers are uncontroversially correct and that competent humans would agree with them. The paper reports neither a human baseline nor inter-annotator agreement. Technical Validation describes two in-house reviewers (also authors) supervising quality, and Annotator Diversity reports only the number of annotators; no section gives agreement statistics. This matters most for the roughly 650 normative/cultural records in Ethics, Bias & Morality, Respecting Others' Rights, Empathy/Intimacy & Trust, Human Preferences, Religion, and parts of Social Knowledge. For Ethics the paper explicitly says reference answers were 'determined with careful consideration of Iranian cultural norms and societal conventions'; for Empathy it asserts that answering 'talk to the thief and let them go if they show regret' is incorrect. These are moral judgments, not factual keys, and different Persian-speaking respondents can reasonably disagree. Because the design principle was to make questions as challenging as possible for Llama3-70B, low model scores were selected for, and without a human ceiling they cannot be interpreted as a model-specific deficit. If a nontrivial share of the normative items are contested, the accuracy numbers conflate model error with disagreement over ground truth, and the 'far from solving' claim is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FarsEval-PKBETS, a Persian-language benchmark of 4,000 human-generated or collected question-answer samples covering 22 subcategories, three response formats (multiple-choice, short answer, descriptive), and culturally specific topics such as Iranian law, religion, ethics, and social knowledge. The authors describe the Saba annotation platform, the category design, and an evaluation of three models: Llama3-70B, PersianMind, and Dorna. They report average accuracies of 47%, 19%, and 30% respectively, and conclude that current language models are still far from being able to solve the benchmark.","tokens_in":15765,"tokens_out":3289,"duration_ms":30267,"significance":"If the dataset is released with proper validation, FarsEval-PKBETS would be a valuable resource for Persian LLM evaluation: it is human-generated, covers diverse formats and local cultural knowledge, involves expert supervision in medicine and law, and addresses gaps in existing Persian benchmarks that rely heavily on translated or exam-derived multiple-choice questions. The paper also documents a useful annotation platform. However, the headline claim that sub-50% accuracy shows models are 'far from being able to solve' the benchmark currently rests on unvalidated reference answers and a selection procedure aimed at making questions hard for Llama3-70B; the missing human baseline and missing evaluation protocol are load-bearing gaps.","major_comments":[{"comment":"The central claim that models are 'far from being able to solve' FarsEval-PKBETS requires a human performance ceiling, which is not reported. Without a human baseline or inter-annotator agreement statistics, sub-50% scores could reflect ambiguous or contested reference answers rather than model deficiency. This is especially consequential in the normative categories (Ethics, Bias & Morality; Empathy, Intimacy & Trust; Respecting Others' Rights; Human Preferences; Religion), where the paper states that reference answers were 'determined with careful consideration of Iranian cultural norms and societal conventions.' The authors should report human accuracy on a stratified sample, and agreement statistics among annotators or external judges, before interpreting low model scores as model failure.","section":"Technical Validation / Table 3"},{"comment":"The benchmark was explicitly designed to be challenging for Llama3-70B, and the text states that questions were selected with that goal in mind. This makes the low Llama3-70B score partly a selection artifact rather than an independent measure of capability. To support the general claim that current LLMs are far from solving the benchmark, the paper needs a human ceiling: competent Persian-speaking humans should score substantially above the models, ideally on a per-category basis. Without this, the 'far from solving' conclusion is not established by the reported numbers.","section":"Methods / Design Principles"},{"comment":"The evaluation protocol for Table 3 is underspecified. The paper does not state who assigned the Correct/Wrong/Semi-correct labels (human evaluators or automatic methods), how many evaluators scored each response, what instructions or rubrics were used for descriptive answers, or what prompting setup and decoding parameters were used for the three models. These details are necessary for reproducibility and for interpreting the accuracy numbers, particularly for open-ended generation categories where scoring is subjective.","section":"Technical Validation / Table 3"},{"comment":"No dataset link, repository, or release plan is provided, and the Saba platform is described only through screenshots. For a benchmark paper in this venue, the data must be made available (with appropriate metadata and a clear license). The paper also does not define the exact schema of the 'Reference' and 'Label' fields or explain how the Chain-of-Thought labels are used in evaluation; this information should be included in the data description.","section":"Data Availability (throughout)"}],"minor_comments":[{"comment":"There are typos in the abstract: 'bechmark' should be 'benchmark' and the sentence 'This bechmark incorporates linguistics, cultural, and local considerations' should be 'This benchmark incorporates linguistic, cultural, and local considerations.'","section":"Abstract"},{"comment":"The paper reports an overall 'Average' row but does not specify whether the average is weighted by the number of samples per subcategory or is an unweighted mean of the subcategory accuracies. Since the subcategories vary in size from 50 to 500 samples, the authors should clarify this and report both if needed.","section":"Table 3"},{"comment":"Reference [31] is cited as the source of the Persian GATS dataset, but the reference list entry for [31] is 'HmBlogs: A big general Persian corpus.' Please verify and correct this citation or provide the actual GATS dataset reference.","section":"References"},{"comment":"The example in the Empathy, Intimacy & Trust section states that answering 'talk to the thief and let them go if they show regret' is incorrect. This is presented as a moral judgment without supporting evidence or discussion of alternative ethical views; given the paper's reliance on such items, the authors should either provide a principled justification or acknowledge the normative nature of these keys explicitly.","section":"Domains and Categories / VII.d"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially useful resource, but the current version is not yet ready for publication because the benchmark data are not released and the headline conclusion lacks a human baseline and a reproducible evaluation protocol. The normative ground-truth concern is substantial but addressable within the manuscript's scope by adding human evaluation and agreement statistics. I would also ask the editor to ensure that the data availability requirement is strictly enforced for this submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FarsEval-PKBETS is a new and useful Persian benchmark: 4,000 human-authored or curated items, three answer formats, a wide category spread, and explicit attention to Iranian cultural and linguistic specifics. That fills a real gap, because most existing Persian benchmarks are translated, exam-based, or multiple-choice-only. The paper is also honest about the weaknesses of MCQ-only evaluation, and Table 3 gives a clear picture: all three models average below 50% on the strict accuracy metric.\n\nCredit where due: the construction pipeline (Saba platform, two supervising reviewers, expert oversight in medicine and law, rejection loops) is standard but real, and the semi-correct response category is a sensible attempt to handle partial credit. Reuse of ArmanEmo, MirasIrony, ETHICS, and ParsMap is modest and flagged in the text.\n\nThe soft spots are real and mostly fixable. The main inference—that sub-50% accuracy means models are \"far from being able to solve\" this benchmark—rests on an unstated assumption that the reference answers are uncontroversially correct. For the normative categories (Ethics, Empathy, Respecting Others' Rights, Human Preferences, Religion), that assumption is shaky. There is no inter-annotator agreement, no human baseline, and the paper explicitly says reference answers were determined with Iranian cultural norms in mind. The \"talk to the thief\" example is a normative judgment, not a factual key. So part of the accuracy gap could be disagreement with one cultural interpretation rather than model deficiency. Additionally, the design principle was to select questions challenging for Llama3-70B, which makes that model's low score partly a selection artifact.\n\nThe paper does not release the data and does not specify the evaluation protocol in enough detail: who scored the descriptive answers, and how were correct/semi-correct judged—automatically or by humans? These omissions matter more than the missing IAA, because without the protocol the numbers in Table 3 are not independently reproducible.\n\nThe stress-test note holds up on reading. The reader's conditional verdict is fair: the resource is valuable, but the headline claim needs a human accuracy check and a narrowed interpretation.\n\nWho is this for? Persian NLP researchers and anyone building multilingual benchmarks. A serious referee should engage with it, but the referee should ask for a dataset release (at least a sample), the scoring protocol, and human baselines on a subset, especially the normative items. With those additions, this could be a solid contribution to the Persian evaluation toolkit.","headline":"A genuinely useful Persian benchmark resource whose headline claim about models being 'far from solving' it is undercut by the absence of a human baseline and a released evaluation protocol.","tokens_in":16470,"tokens_out":1676,"would_cite":true,"duration_ms":17033,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FarsEval-PKBETS, a new 4,000-question Persian benchmark, reports sub-50% accuracy for three current large language models, concluding that Persian and Iranian cultural competence is still largely unsolved.","keywords":["FarsEval-PKBETS","Persian language models","LLM benchmark","question answering","cultural knowledge","ethics and bias","text generation","human evaluation"],"falsifier":"Have several independent Persian-speaking annotators re-answer a random sample of the Ethics, Bias and Morality, Human Preferences, and Respecting Others' Rights items without seeing the reference answers; if per-item majority agreement is low or frequent disputes arise, the reference answers are not stable enough to support the accuracy claims.","tokens_in":15339,"feed_emoji":"📊","tokens_out":7428,"duration_ms":62675,"temperature":0.7,"pith_summary":"FarsEval-PKBETS is a new Persian-language evaluation benchmark built from 4,000 questions that were human-generated or human-collected and then human-reviewed, covering medicine, law, religion, Persian language, encyclopedic knowledge, human preferences, social knowledge, ethics and bias, toxicity, respect for others' rights, and text generation. The paper's central claim is that existing Persian evaluation resources are limited by translation, exam-style data, and a near-total reliance on multiple-choice questions, whereas this benchmark is designed around Persian linguistic and Iranian cultural context and includes short-answer and descriptive formats that test production, not just selection. To demonstrate that the questions are genuinely hard, the authors evaluated Llama3-70B, PersianMind, and Dorna and report fully correct answer rates of 47%, 19%, and 30%, respectively. The intended upshot is that current language models remain far from solving Persian-language tasks that depend on local knowledge, style, social judgment, and generation.","feed_headline":"New Persian benchmark stumps LLMs, which score under 50 percent","feed_subtitle":"4,000 human-reviewed questions in Persian; Llama3-70B gets 47%, Dorna 30%, PersianMind 19%.","key_machinery":"The central object is FarsEval-PKBETS itself: a dataset of 4,000 records, each containing a Persian question, a reference answer, and metadata such as category, label, and a Reference field for sourced items, with 1,200 descriptive, 500 short-answer, and 2,300 multiple-choice items distributed across twelve head categories. The paper's construction machinery is a two-reviewer workflow on the Saba annotation platform, where submissions are approved or rejected and returned for revision, plus domain-expert supervision for medicine and law. For evaluation, outputs are scored as Correct, Wrong, or Semi-correct, with the semi-correct category catching cases like correct answer text attached to a distractor's identifier. This combination is what lets the benchmark claim both content validity through human review and cultural grounding, and diagnostic power through format diversity and fine-grained scoring.","core_discovery":"On its own terms, the paper establishes that a carefully reviewed, culturally grounded Persian QA set can expose large gaps in current LLMs. Across 4,000 items, Llama3-70B averages 0.47 full-correct accuracy, PersianMind 0.19, and Dorna 0.30; even counting partially correct responses, the best model reaches only 0.54. The benchmark also reveals failure modes that multiple-choice-only evaluation misses: in a 100-question probe, models gave a correct choice but an incorrect justification 43% of the time, and models sometimes attach the right answer text to the wrong option identifier. The authors take these results as evidence that FarsEval-PKBETS is challenging and that Persian and Iranian socio-cultural competence is not yet achieved by publicly available models.","pith_inferences":["Because reference answers in ethics, bias, morality, and respecting others' rights encode Iranian cultural norms, future uses should separate knowledge accuracy from value alignment; identical item scores may mean different things across cultural backgrounds.","A direct test of the cultural-grounding claim would be to translate the same 4,000 items into English and evaluate English-native models: if accuracy rises substantially, the bottleneck is Persian language and local context rather than general reasoning.","The 43% justification-failure finding suggests a training signal: models may first improve in giving correct rationales before improving final answers, so reporting both could reveal earlier progress.","Re-annotating a sample of the single-annotator categories such as Emotion and Irony with multiple independent Persian-speaking annotators would estimate how much of the accuracy gap is label ambiguity rather than model deficit."],"forward_implications":["A reusable Persian benchmark now exists that can be rerun as models improve; sub-50% accuracy gives a concrete baseline for progress.","MCQ-only scores overstate capability; because semi-correct and unjustified answers are counted separately, benchmark users can see where selection outperforms reasoning.","Persian-specific and Iranian-cultural categories such as empathy, irony, respecting others' rights, formal register, and poems define dimensions that translated English benchmarks do not measure.","Domain results are uneven, such as 0.65 lexical semantics but 0.14 emergency medicine for Llama3-70B, so the benchmark can guide targeted improvement by category.","The balanced distribution of correct option positions reduces position-bias artifacts in multiple-choice scoring."],"supporting_citations":[{"why":"Supplies the four-part taxonomy for ethics decision-making questions and fewer than 10% of the records in that category.","marker":"[13]"},{"why":"ArmanEmo is the source of half of the emotion-classification samples, which were then revised or edited.","marker":"[32]"},{"why":"MirasIrony provides half of the irony items, with labels reviewed and sometimes changed.","marker":"[33]"},{"why":"ParsMap supplies the informal-to-formal Persian style-transfer sentences used in the Formality Style Transfer category.","marker":"[36]"},{"why":"Llama3-70B is the strongest evaluated model and the benchmark's difficulty target at planning time.","marker":"[40]"},{"why":"PersianMind is one of the two Persian-tuned models whose sub-50% accuracy supports the benchmark's challenge claim.","marker":"[41]"},{"why":"Dorna-Llama3-8B-Instruct is the other Persian-tuned evaluation model used in the results table.","marker":"[42]"}],"fun_headline_variants":["Persian benchmark exposes LLM gaps: best model 47%","New Persian QA set stumps LLMs: all scores below 50%","FarsEval-PKBETS: Top AI scores under half on Persian quiz","LLMs fail Persian test: right answer, wrong reasoning 43% of time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reference answers are treated as the ground truth, especially in ethical, moral, and rights-related questions where answers were set according to Iranian cultural norms; if those answers are contested or inconsistent, the reported accuracy numbers are not a clean measure of model capability.","fun_headline_variants_meta":{"raw":{"variants":["Persian benchmark exposes LLM gaps: best model 47%","New Persian QA set stumps LLMs: all scores below 50%","FarsEval-PKBETS: Top AI scores under half on Persian quiz","LLMs fail Persian test: right answer, wrong reasoning 43% of time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":3076,"prompt_tokens":918,"completion_tokens":2158,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":2075}},"tokens_in":534,"tokens_out":2158,"duration_ms":15776,"temperature":1.0,"reasoning_tokens":2075,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:43:02.739898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several independent Persian-speaking annotators re-answer a random sample of the Ethics, Bias and Morality, Human Preferences, and Respecting Others' Rights items without seeing the reference answers; if per-item majority agreement is low or frequent disputes arise, the reference answers are not stable enough to support the accuracy claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the four-part taxonomy for ethics decision-making questions and fewer than 10% of the records in that category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MirasIrony provides half of the irony items, with labels reviewed and sometimes changed."},{"cited_title":"Developing an Informal-Formal Persian Corpus","cited_arxiv_id":"2308.05336","evidence_quote":"ParsMap supplies the informal-to-formal Persian style-transfer sentences used in the Formality Style Transfer category."},{"cited_title":"https://huggingface.co/PartAI/Dorna-Llama3-8B-Instruct (2024)","cited_arxiv_id":null,"evidence_quote":"Dorna-Llama3-8B-Instruct is the other Persian-tuned evaluation model used in the results table."}],"review_version":1}