{"id":"66934537-8d35-4251-980b-6938bfe9da91","arxiv_id":"2505.18331","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A first Persian consumer medical QA benchmark is released, but the paper contains no evaluation results.","lead":"PerMedCQA is a new Persian-language dataset of 68,138 medical question-answer pairs gathered from online health forums, alongside an LLM-based grading rubric called MedJudge. The paper does not include the benchmarking results it promises; Section 5 is explicitly marked as under development.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central benchmarking claim is unverifiable because Section 5 is explicitly marked as under development; no model scores appear anywhere, yet the abstract and contributions assert that evaluations revealed key insights.","rationale":"The reader's verdict is REJECT and I agree with that outcome, but the reader's stated weakest_assumption (gold-reference correctness) is not the most load-bearing concern. Even if all forum answers were medically perfect, the paper still contains no evaluation results, so the benchmark's ability to rank models or produce insights is completely unverified. The reader's rationale does mention the empty Results section as a red flag, so there is partial agreement, but the formal weakest_assumption field points elsewhere. The missing results are decisive: they invalidate the abstract's central empirical claim and the third and fourth listed contributions ('comprehensively benchmark a variety of LLMs' and 'evaluate supervised fine-tuning strategies'). The dataset construction pipeline is described in detail and the dataset is publicly released, which is a genuine contribution, but the paper as submitted overclaims what was actually done. The human validation of MedJudge, while present, is too weak (kappa = 0.42, and near-zero kappa for clinical impact) to rescue the evaluation claim. A revision that either adds a full Results section or removes all performance-related claims from the abstract and contributions could make the paper publishable; as-is, REJECT is appropriate.","tokens_in":16289,"tokens_out":2671,"duration_ms":24738,"concrete_test":"Request from the authors the full results table for all 16 models on the 3,513-item test set, including MedJudge Correctness, Coverage, and Clinical Impact scores with per-model counts, and independently reproduce the MedJudge pipeline on a random 100-item subset from the released HuggingFace dataset using the provided prompts. If the table is unavailable or the reproduced scores cannot be computed from the released data, the benchmark evaluation claim remains unsupported.","verdict_should_be":"REJECT","load_bearing_attack":"The core claim of the paper is that PerMedCQA is a benchmark for evaluating LLMs and that the authors conducted extensive evaluations producing insights. However, Section 5 ('Results and Analysis and Conclusion') states: 'This section is currently under development and will be included in a future revision.' No results table, no model scores, and no comparisons are present in the main text or appendix; Figures 15–17 are captions with no accompanying data. This is not a minor omission: the benchmark's utility rests on demonstrating that MedJudge produces meaningful, non-degenerate evaluations that distinguish models. The abstract's sentence 'Our results highlight key challenges in multilingual medical QA' is therefore factually unsupported by the submitted manuscript. The only quantitative evidence for MedJudge is a 100-item human agreement study (75% collapsed accuracy, quadratic Cohen's kappa = 0.42, and kappa = 0.07 for clinical impact), which is too limited to establish a reliable evaluation framework, especially without any downstream results showing that the rubric scores discriminate between models or correlate with human preferences at scale. The dataset itself may be a useful resource, but a benchmark paper that contains no benchmark results does not deliver its central claim. Consequently, the paper should not be accepted in its current form; the authors should either add the complete Results section or revise the abstract and contribution claims to describe only the dataset and evaluation protocol, not the outcomes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PerMedCQA, a Persian-language corpus of 68,138 consumer health question-answer pairs collected from four public forums, cleaned with rule-based filtering and LLM-based PII detection, annotated with ICD-11 categories and question types, and split into stratified train/dev/test sets. It also proposes MedJudge, an LLM-based rubric grader that compares model answers against forum expert answers on Correctness, Coverage, and Clinical Impact, with a 100-item human validation study. The paper claims to benchmark 16 LLMs using zero-shot prompting, role-based prompting, pivot translation, and LoRA fine-tuning; however, Section 5, which would contain all results, is explicitly marked as under development, and no model scores or comparisons are reported anywhere in the manuscript.","tokens_in":16505,"tokens_out":6030,"duration_ms":49247,"significance":"If completed, PerMedCQA would be a valuable resource for Persian medical NLP: it is large, publicly released, derived from real consumer questions, and includes structured metadata and stratified splits. The MedJudge design is a sensible approach for open-ended QA, and the authors are transparent about the limitations of their human validation. Nevertheless, the current submission does not deliver its central benchmarking claim: the complete absence of results means the reader cannot assess whether the benchmark distinguishes models, whether MedJudge scores are stable, or whether the proposed prompting and fine-tuning methods have any measurable effect. As submitted, the paper is better described as a dataset resource than as a benchmark evaluation, and the abstract's claims about results are unsupported.","major_comments":[{"comment":"Section 5 is empty, stating 'This section is currently under development and will be included in a future revision.' The abstract and Section 1 assert that extensive evaluations were performed and that the results highlight key challenges, but no model scores, tables, or comparisons appear in the main text or appendix; Figures 15-17 are only captions with no accompanying data. This is not a presentation issue: the central claim of the paper is benchmarking, and the results are the evidence for that claim. The manuscript cannot be accepted until the full Results section is included (baseline scores, prompt-method comparisons, fine-tuning comparisons, and MedJudge reliability analyses) and the abstract is made consistent with the content actually reported.","section":"5"},{"comment":"The human validation of MedJudge is too weak to support the claim of a clinically informed evaluation framework validated by expert reviews. The primary correctness dimension has 75% collapsed agreement with quadratic Cohen's kappa = 0.42 (95% CI 0.19-0.58), and the Clinical Impact dimension has kappa = 0.07, detecting only one-third of high-impact discrepancies. Since no downstream result shows that MedJudge separates strong from weak models or correlates with human judgments at scale, the evaluation framework's reliability is not established. Please report a confusion matrix, per-label precision and recall, and a sensitivity analysis of model rankings when the grader model or rubric is changed.","section":"A.2"},{"comment":"The gold standard is assumed to be the answers posted by verified physicians on four forums, and MedJudge is explicitly instructed to judge exclusively against these answers and ignore external knowledge (A.1). If the forum answers are incomplete or contain systematic errors, every model score inherits those errors. The paper provides no evidence that the gold answers are clinically acceptable. The authors should validate a random sample of gold answers with independent expert review, or construct references through consensus, and report the resulting agreement.","section":"3.1"},{"comment":"ICD-11 and question-type labels are generated by GPT-4o-mini without any reported accuracy or human agreement. These labels are load-bearing: the test split is stratified by ICD-11 category, and the role-based prompting in Section 4.3 conditions on them. Label noise could distort both the evaluation split and the prompt conditions. Please report annotation accuracy on a human-annotated sample and, ideally, a sensitivity analysis showing that key conclusions are robust to label noise.","section":"3.3"}],"minor_comments":[{"comment":"The text says 'Figure 6 shows the distribution of ICD-11 categories,' but Figure 6 is titled 'Distribution of Question Type' and Figure 5 already shows the ICD-11 distribution; the cross-references and captions should be corrected.","section":"3.3"},{"comment":"The abstract states PerMedCQA is 'the first Persian-language benchmark,' while Section 1 uses the more careful phrase 'To the best of our knowledge.' The abstract should either match that qualification or report a systematic search for prior Persian medical QA resources.","section":"Abstract"},{"comment":"MedJudge is described in the contributions as a novel evaluation framework, but Section 4.1 states it is 'based on the criteria (Hosseini et al., 2024)'; please clarify which components are new to this work.","section":"4.1"},{"comment":"Figure 1 contains the typo 'DIFFRENTTECHNIQUES,' and Figure 13 uses 'MedJude' while the rest of the paper uses 'MedJudge'; these should be harmonized.","section":"Figure 1"},{"comment":"The reference for PaLM (Chowdhery et al., 2023) is malformed, containing an unrelated string of author names, and several URL-only entries (e.g., Claude3.5, GetZoop) lack access dates; the bibliography needs a careful cleanup.","section":"References"},{"comment":"The phrasing 'resulting the final numbers QA pairs' should be revised to 'resulting in the final number of QA pairs,' and in Section 3.3 'PerMedCQA were split' should be 'PerMedCQA was split.'","section":"3.2"}],"recommendation":"reject","confidential_remarks":"The absence of Section 5 makes this submission effectively a dataset-only paper; the title, abstract, and contributions claim benchmarking results that are not present. I would not recommend inviting a standard revision without results, because the reviewers cannot assess the benchmark's validity or the soundness of the evaluative claims. The public dataset release is a positive step, but the manuscript needs substantial new empirical content before it is suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Off the record: the PerMedCQA dataset is a genuine new resource, but this preprint is not yet a benchmark paper. Section 5 literally says \"This section is currently under development,\" and there are no model scores anywhere, yet the abstract claims \"our results highlight key challenges.\" That mismatch is the thing to know.\n\nWhat's good: the dataset construction is described carefully. 68k QA pairs from four Persian forums, two-stage cleaning, PII detection, ICD-11 and question-type annotation, public release on HuggingFace. That fills a real gap for low-resource consumer medical QA. The MedJudge rubric is defined in enough detail to reproduce, and the 100-item human agreement study is a reasonable start (75% collapsed accuracy, kappa 0.42 for correctness). Credit where due: this is not a toy dataset.\n\nThe soft spots are not subtle. The absent results are load-bearing. The abstract and contributions promise benchmarking and insights, but there is no data to support them. The human validation is also thin: kappa 0.42 is moderate at best, and the clinical-impact kappa of 0.07 means MedJudge is not reliably identifying dangerous discrepancies. The gold reference assumption—that forum physicians' answers are correct—is unexamined, and MedJudge is told to ignore external knowledge, so any bias in those answers propagates. The circularity of using GPT-4o-mini for ICD-11 labels that later condition role-based prompts is minor, but it exists.\n\nWho's this for? Persian NLP and health informatics researchers who need a consumer QA dataset. If the authors add a full results section and temper the claims, it could be a solid benchmark contribution. As is, it's a dataset release wrapped in overstated claims.\n\nMy recommendation: don't accept as is. Either the editor sends it back with a clear request for results and revised abstract, or desk-rejects with an invite to resubmit. I'd give the authors the chance to complete the work. It's not a waste of referee time once the results exist.","headline":"The PerMedCQA dataset is a genuine new resource, but the paper as submitted has no results section, so the benchmark claims are unsupported.","tokens_in":17066,"tokens_out":3184,"would_cite":true,"duration_ms":25801,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PerMedCQA introduces the first large-scale Persian benchmark for consumer medical question answering, with 68,138 QA pairs and an LLM-based judge.","keywords":["PerMedCQA","Persian medical QA","consumer health questions","benchmark dataset","LLM-as-a-judge","ICD-11","low-resource NLP","medical question answering"],"falsifier":"Take a random sample of PerMedCQA gold answers and have independent board-certified physicians, blinded to the forum posts, rate each answer's medical correctness against current guidelines; if a substantial fraction are judged inaccurate, incomplete, or outdated, the reference standard that MedJudge scores against is not reliable.","tokens_in":16087,"feed_emoji":"🩺","tokens_out":6505,"duration_ms":51575,"temperature":0.7,"pith_summary":"PerMedCQA aims to fill a gap in medical question answering: almost all consumer-oriented benchmarks are in English, and none target Persian. The paper builds a dataset of 68,138 question-answer pairs drawn from four Persian forums where verified physicians answer real patient questions, cleaned from 87,780 raw entries and annotated with ICD-11 disease categories and 25 question types. To score open-ended answers, it introduces MedJudge, an LLM grader that compares model responses with the physician answers using a rubric for correctness, coverage, and clinical impact, and reports 75% agreement with physicians on a 100-item sample. The paper also benchmarks 16 LLMs and tests prompting and fine-tuning strategies, though the results section is currently marked as under development. If the resource holds up, it would give Persian-speaking users a way to measure and improve medical AI systems in their own language.","feed_headline":"68,138 Persian medical Q&As form first consumer QA benchmark","feed_subtitle":"Real patient questions and physician answers, graded by an LLM judge, open Persian to medical AI evaluation.","key_machinery":"The load-bearing mechanism is MedJudge, a large-language-model grader prompted with a three-part rubric: Correctness (correct, partially correct, incorrect, contradictory), Coverage (equal, model subset, expert subset, no overlap), and Clinical Impact (negligible, moderate, significant, critical). MedJudge is given the patient question, the verified-physician answer as gold, and the model answer, and is explicitly forbidden to use its own medical knowledge or outside sources. The dataset pipeline is the other half of the machinery: rule-based filters remove short, duplicate, or non-textual entries; GPT-4o-mini flags personally identifiable information; the same LLM assigns ICD-11 categories; and the benchmark is split into train, evaluation, and test sets stratified by ICD-11 category.","core_discovery":"The central claim is that PerMedCQA is the first Persian-language benchmark for consumer medical question answering, and that it is large and realistic enough to support meaningful evaluation. The dataset contains 68,138 QA pairs from four public Persian medical forums, restricted to questions from real patients and answers from verified physicians, and enriched with ICD-11 labels, 25 standardized question types, patient age and sex, physician specialty, and source metadata. The paper further claims that open-ended medical answers can be reliably scored by MedJudge, an LLM-based rubric grader that is instructed to judge only against the expert answer and that reached 75% agreement with board-certified physicians on the correctness dimension of a 100-item subset. On this basis the authors argue that multilingual and instruction-tuned models vary substantially on Persian consumer health questions and that prompt-based techniques and fine-tuning can change model output quality, with the detailed results deferred to a future revision.","pith_inferences":["The reported 75% correctness agreement and quadratic Cohen's kappa of 0.42 between MedJudge and physicians are moderate; if a different judge model or prompt changes rankings, part of what is being measured is the judge, not the medical QA systems.","Because the gold answers come from forum posts and MedJudge is barred from external knowledge, any systematic error in those posts, such as outdated or overly cautious advice, becomes part of the benchmark's definition of correctness.","The text-only scope leaves out skin and visual-system questions, which are among the most common categories in the dataset; a multimodal extension would likely change model rankings.","The translation-pivot experiments imply that English-centric models may lose information in Persian; testing the same models on native Persian outputs after fine-tuning would isolate how much of the gap is language coverage rather than medical knowledge."],"forward_implications":["Persian-speaking patients can have medical QA systems evaluated on the kinds of questions they actually ask, rather than on translated exam questions.","Researchers get a public, de-identified resource for fine-tuning and for comparing models in a low-resource language.","The ICD-11 and question-type annotations allow analysis of which medical topics and question intents are hardest for current LLMs.","The reported gender and topic distributions suggest that any deployed system must handle sexual health, digestive, and skin questions well to serve Persian forum users.","MedJudge, if its agreement with physicians generalizes, offers an alternative to BLEU and ROUGE for open-ended medical evaluation in other languages."],"supporting_citations":[{"why":"Supplies the long-form medical QA benchmark approach and the rubric dimensions that MedJudge adapts for correctness, coverage, and clinical impact.","marker":"(Hosseini et al., 2024)"},{"why":"Provides the 25-category question type taxonomy used to annotate each QA pair.","marker":"(Abacha et al., 2019a)"},{"why":"Establishes ICD-11 as the classification system for the 28 disease categories assigned to the dataset.","marker":"(Khoury et al., 2017)"},{"why":"Grounds the LLM-as-a-judge evaluation method that MedJudge builds on.","marker":"(Zheng et al., 2023)"},{"why":"Supplies the LoRA parameter-efficient fine-tuning method used in the supervised fine-tuning experiments.","marker":"(Hu et al., 2022)"},{"why":"Provides MedRedQA, the consumer-oriented medical QA benchmark that motivates the need for real-world, open-ended patient questions.","marker":"(Nguyen et al., 2023)"}],"fun_headline_variants":["Persian medical QA benchmark: 68K real patient questions, LLM-graded","First Persian consumer medical QA benchmark: 68K real Q&As, LLM-ranked","68K real Persian medical Q&As get LLM-graded evaluation benchmark","LLM-graded Persian medical QA: 68K patient questions benchmark","First Persian benchmark for consumer medical QA: 68K real Q&As"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark treats answers posted by verified physicians on four public forums as the gold standard, and MedJudge is told to judge only against those answers; if those posts are medically wrong or incomplete, every model score inherits that error.","fun_headline_variants_meta":{"raw":{"variants":["Persian medical QA benchmark: 68K real patient questions, LLM-graded","First Persian consumer medical QA benchmark: 68K real Q&As, LLM-ranked","68K real Persian medical Q&As get LLM-graded evaluation benchmark","LLM-graded Persian medical QA: 68K patient questions benchmark","First Persian benchmark for consumer medical QA: 68K real Q&As"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3525,"prompt_tokens":919,"completion_tokens":2606,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":2501}},"tokens_in":535,"tokens_out":2606,"duration_ms":16508,"temperature":1.0,"reasoning_tokens":2501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:31:52.809471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of PerMedCQA gold answers and have independent board-certified physicians, blinded to the forum posts, rate each answer's medical correctness against current guidelines; if a substantial fraction are judged inaccurate, incomplete, or outdated, the reference standard that MedJudge scores against is not reliable.","supporting_citations":[],"review_version":1}