{"id":"5b7261b2-7a57-46ac-b9c7-6ea159b1ce8c","arxiv_id":"2509.07135","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MedBench-IT collects 17,410 Italian medical entrance exam questions and reports model accuracy, response consistency, ordering bias, reasoning-prompt effects, and readability correlations.","lead":"MedBench-IT is a new benchmark of 17,410 multiple-choice questions from Italian medical university entrance exams, with accuracy and robustness results for proprietary and open-source language models. It gives Italian NLP and EdTech developers a specialized testbed that did not exist before.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Answer-key correctness is unverified; the appendix's placeholder correct-answer indices suggest the labeling pipeline may be incomplete, yet every reported accuracy depends on those keys.","rationale":"The reader's weakest assumption is exactly the correctness of publisher-supplied answer keys and the representativeness of the stratified sample. My stress-test focuses on the first half of that assumption: the gold-label correctness. The manuscript provides no independent validation of the labels, and the appendix placeholders are a concrete, in-scope signal that the quality-assurance pipeline may be incomplete. This is the single most load-bearing condition because every quantitative claim in the paper—overall accuracy, per-subject rankings, reproducibility, ordering bias, reasoning-prompt effects, and readability correlations—is computed against these labels. If the labels are wrong, the central contribution (a reliable benchmark) collapses. The proposed test would settle the concern by directly measuring label accuracy on a sample; if labels are sound, the conditional acceptance stands. I therefore keep the reader's CONDITIONAL verdict unchanged: the paper should not be accepted until the authors provide evidence of label verification and data access under the sharing scheme. I do not move to REJECT because the internal consistency check of Table 4 (weighted per-subject accuracy matching overall accuracy) is a positive sign, and the placeholder issue may be an artifact of presentation rather than of the dataset itself. The verdict remains conditional on the label audit and on data release.","tokens_in":13953,"tokens_out":5370,"duration_ms":47085,"concrete_test":"Contact the corresponding author and obtain, via the data-sharing agreement, a random sample of at least 400 questions (e.g., 200 from Biology/Chemistry and 200 from Logic/Mathematics). Have two independent Italian-speaking domain experts answer the questions blind, then compare their answers against the publisher-provided keys. Compute percentage agreement and Cohen's kappa. If agreement is below 98% of questions, or if any placeholder, empty, or ambiguous correct-answer index appears in the sample, the benchmark's gold labels are not reliable and all reported accuracy scores are suspect. Additionally, verify that every example in the published appendix has a concrete correct-answer index rather than a placeholder, and that the weighted average of per-subject accuracies reproduces the reported overall accuracy for at least GPT-4o.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that MedBench-IT is a reliable benchmark of 17,410 expert-written questions with correct gold answers, supporting the reported accuracy scores. The weakest link is label correctness. The authors rely entirely on Edizioni Simone's answer keys without independent verification: there is no inter-annotator agreement study, no re-annotation of a random sample, and no audit against a second source. Internal evidence compounds this: in Appendix A, two of the four additional examples (A.1 and A.4) do not actually specify the correct answer—they contain placeholders such as '[Index for ... or similar]' and '[Index for 0.02 mol]'—and A.1 does not even list the options. For a resource paper whose only artifact is question-answer pairs, this signals that the labeling pipeline may not have been completed or checked. Since the dataset is proprietary and cannot be redistributed, reviewers and the community cannot inspect the labels. If any substantial share of gold answers is wrong, all reported accuracies, per-subject comparisons, and the readability/reasoning analyses would be measuring publisher errors or noise rather than model capability. The reproducibility and ordering-bias tests do not validate label correctness; they only demonstrate model consistency on possibly erroneous labels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"MedBench-IT is presented as the first large-scale benchmark for Italian medical university entrance examinations. The authors curate 17,410 multiple-choice questions from Edizioni Simone across six subjects and three difficulty levels, evaluate 20 proprietary and open-source LLMs under zero-shot standard and reasoning-eliciting prompts, and compute accuracy overall and by subject/difficulty. Additional analyses address reproducibility (GPT-4o, two runs), answer-ordering bias (GPT-4o and Claude 3.5 Haiku), the effect of reasoning prompts, and readability via Flesch-Vacca logistic regressions. The paper reports that top models exceed 90% accuracy, that reproducibility is 88.86% for GPT-4o, that ordering bias is small, and that readability has a statistically significant but tiny inverse relation with accuracy.","tokens_in":14155,"tokens_out":4828,"duration_ms":45633,"significance":"If the dataset and labels are reliable, MedBench-IT fills a genuine gap: native Italian, domain-specific, exam-style MCQ evaluation, with a useful model zoo spanning 0.5B to 671B parameters. The evaluation protocol is largely sound: sampling temperatures are stated, full prompt templates are provided, McNemar tests are used for two of the analyses, and the reasoning/readability decompositions are sensible. The paper also reports negative or null results (small CoT gains, mixed ordering-bias significance), which strengthens confidence in the authors' reporting. However, the benchmark's value depends on the correctness of the gold answers and on the representativeness of the 17,410-question sample, and both are currently not independently checkable because the data are proprietary. The placeholder labels in Appendix A are a concrete red flag that the answer-key pipeline may be incomplete; this must be resolved before the accuracy scores can be interpreted.","major_comments":[{"comment":"Two of the four additional examples contain placeholder correct-answer fields such as '[Index for ...]' and placeholder option lists, and the examples are presented as sample questions from the dataset. This is internal evidence that the answer-key/labeling pipeline is not fully verified. Because Eq. (1) and all accuracy scores in Tables 3 and 4 depend on gold labels, the authors should run an independent re-annotation of a random sample (e.g., 300–500 questions) against the Edizioni Simone keys, report inter-annotator agreement, and replace the placeholder examples with completed, checked items.","section":"Appendix A (A.1 and A.4)"},{"comment":"The dataset is not redistributable and cannot be inspected by reviewers, so the central claim of a 17,410-question benchmark with correct gold answers is not independently verifiable. For a resource paper of this type, the authors should at minimum release a representative public sample with gold labels and the exact prompts, report duplicate/near-duplicate checks and formatting-validation statistics, and provide a clear data-sharing agreement that lets independent researchers reproduce at least a subset. The five appendix examples are not sufficient to establish corpus quality.","section":"Section 7 (Data Availability)"},{"comment":"The contamination risk is acknowledged only as 'cannot be entirely ruled out, even if unlikely given our data source.' Given that the questions come from a commercial publisher with a known distribution channel and that the top models reach about 90% accuracy, the risk is not negligible. I ask for a concrete memorization check on a random sample of questions (e.g., prompting models to complete or answer from the stem alone) and a statement of model training cutoffs relative to the data acquisition period. Without this, the reported accuracies could partly reflect memorized publisher material rather than capability.","section":"Section 7 (Limitations) and Section 5.1"},{"comment":"Aya Expanse 8B drops from 46.7% accuracy under the standard prompt to 0.1% under the reasoning prompt, with per-subject reasoning values between 0.1% and 0.4%. This pattern strongly indicates a format-following failure rather than a genuine accuracy measurement. Reporting these values as benchmark results is misleading. The paper should either exclude such runs with an explicit non-compliance criterion or report format-success rates separately.","section":"Table 3 and Appendix B"},{"comment":"The shuffle protocol is underspecified: it is not stated whether each question received one random permutation, whether the same shuffled order was reused across models, or how many shuffled runs were averaged. The analysis also covers only two models. Since the abstract claims a 'rigorous' ordering-bias analysis, the experimental design should be described precisely and the mixed McNemar results should be accompanied by effect sizes and confidence intervals.","section":"Section 5.3 (Ordering Bias)"}],"minor_comments":[{"comment":"The reproducibility test uses temperature 1, while the main evaluation uses temperature 0. The paper should justify this choice explicitly, since a reproducibility test at temperature 0 is a different and arguably more relevant measurement of deterministic consistency.","section":"Section 4.3 and 5.2"},{"comment":"The paper does not specify how malformed model outputs (e.g., answers that are not a single number 1–5, or reasoning text without an answer) were parsed and scored. Please add a paragraph on output parsing and whether such cases are counted as incorrect.","section":"Section 3.3 and 5.1"},{"comment":"Several model identifiers are incomplete or inconsistent (e.g., 'Lexora Med. 7B', 'Maestrale v0.4', and 'Gemma 2 9B' versus the Hugging Face identifiers mentioned in the text). Provide exact model versions and hyperparameters in a reproducibility appendix.","section":"Table 3 and Appendix B"},{"comment":"The figures are referenced but their construction details are minimal; for Figure 3, the blue/red distinction should be supplemented with markers or hatching to remain accessible to color-blind readers.","section":"Figures 1–3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central asset is proprietary, and the placeholder examples in Appendix A are concerning enough that I would not recommend acceptance without an external audit of a random sample and a public data sample. The funding and data source are disclosed, but the paper reads more like a commercial-leaderboard announcement than a fully verifiable benchmark contribution; if the journal publishes resource papers, the data-access policy and label-validation protocol should be made a condition of acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: MedBench-IT is a genuinely useful resource for Italian NLP and EdTech, but the answer-key verification is a real gap and should keep you from treating the reported accuracies as established fact.\n\nWhat's new: it's the first large-scale Italian medical entrance exam benchmark, 17,410 native questions, six subjects, three difficulty levels, sourced from a real publisher. That's a real gap. The evaluation is methodologically sound in its main lines: temperature 0, standard and CoT prompts, per-subject and per-difficulty breakdowns, McNemar tests for the reproducibility and ordering analyses. The readability correlation is a nice, honest extra—small effect, statistically significant, not oversold. Related work is proper and the limitations section acknowledges data access and contamination.\n\nThe soft spot is the gold labels. The paper relies entirely on Edizioni Simone's answer keys. No re-annotation, no second source, no sample audit. That is a real problem because it's the central artifact. The appendix makes it worse: in A.1 and A.4 the options and correct-answer index are still placeholders like '[Index for ...]' rather than actual values. That looks like the labeling pipeline wasn't finished for at least those examples, or the paper was written from a template. Combined with the closed, proprietary dataset, reviewers—and the community—cannot check the keys. If even a modest share of the 17,410 keys are wrong, every accuracy number in the paper shifts, and the consistency/ordering tests don't help because they only show models are stable on the same possibly erroneous labels.\n\nI'm not claiming the numbers are wrong. There's no evidence of that. But for a resource paper, the lack of any label audit is a load-bearing gap. It's curable: release a random sample with independently re-annotated answers, or at least provide a verification protocol and evidence that a sample was checked.\n\nRecommendation: this deserves a serious referee, not a desk reject. The referee should make the audit a condition of acceptance, and the paper shouldn't be cited as an authoritative benchmark until the keys are verified. I wouldn't cite it in my own work in the next year without that.","headline":"MedBench-IT fills a real gap in Italian medical entrance exam benchmarks, but the unverified answer keys and placeholder examples in the appendix make the reported accuracies provisional until a label audit is done.","tokens_in":14678,"tokens_out":3864,"would_cite":false,"duration_ms":32451,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MedBench-IT is the first large-scale benchmark for Italian medical entrance exams, and the paper uses it to rank LLMs, finding top models near 90 percent accuracy with logic and mathematics as the persistent weak point.","keywords":["LLM evaluation","benchmark","Italian NLP","medical entrance exam","multiple-choice question answering","reproducibility","ordering bias","chain-of-thought"],"falsifier":"Audit a random sample of the benchmark questions with independent Italian medical educators and re-run the top models on the 26,115 questions from the original corpus that were excluded; non-trivial gold-answer errors or a large accuracy gap on the excluded set would invalidate the reported scores.","tokens_in":13775,"feed_emoji":"🎓","tokens_out":8377,"duration_ms":64265,"temperature":0.7,"pith_summary":"The paper introduces MedBench-IT, a benchmark of 17,410 expert-written, multiple-choice questions drawn from Italian medical university entrance examinations, spanning six subjects and three difficulty levels. Its aim is to give the Italian NLP and educational-technology communities a native-language yardstick for measuring how well large language models handle this high-stakes task. The evaluation reports a clear performance hierarchy: the best API-based models reach roughly 90 percent accuracy, resource-efficient open-source models under 30 billion parameters exceed 70 percent, and most smaller Italian fine-tunes land near 60 percent. It also finds that logic and mathematics are the hardest subjects for every model, that GPT-4o reproduces the same answer only 88.86 percent of the time across identical runs, and that requiring explicit reasoning rarely improves top-model accuracy.","feed_headline":"Italian medical exam benchmark: top AI models score near 90%","feed_subtitle":"The benchmark shows top models excel on facts but lose 15+ points on logic and math","key_machinery":"The load-bearing object is the dataset itself: 17,410 multiple-choice questions selected from a 43,525-question corpus supplied by a leading Italian preparatory publisher. The construction pipeline removes image-dependent and English-language items, strips markup, normalizes every item to a question stem with five options and one correct answer, and takes a stratified sample preserving original subject and difficulty proportions. The evaluation protocol then applies a fixed Italian prompt format under two conditions, direct answering and reasoning-eliciting, keeping temperature at zero for main runs, which allows accuracy, reproducibility, ordering-bias, and readability analyses to be compared on identical inputs.","core_discovery":"The central claim is that MedBench-IT is the first comprehensive benchmark built specifically for Italian medical entrance examinations, sourced from the publisher Edizioni Simone rather than translated from English tests. On this benchmark the paper reports model accuracies as evidence of current capability: DeepSeek-R1 scores 91.9 percent, o1-preview 89.1 percent, Claude 3.5 Sonnet 87.8 percent, and GPT-4o 83.9 percent under a direct-answer prompt, while the strongest sub-30B open models, Phi-4 and Qwen 2.5 14B, reach 76.8 and 72.6 percent. The accompanying robustness results claim that answer-order shuffling has a minimal effect on GPT-4o but a statistically significant effect on Claude 3.5 Haiku, that GPT-4o's response consistency across identical runs is 88.86 percent with subject-dependent variation, and that reasoning-eliciting prompts do not systematically improve accuracy.","pith_inferences":["If MedBench-IT becomes a standard fixture, per-subject scores could double as a diagnostic for Italian medical curricula, pinpointing whether preparation materials should emphasize reasoning practice over factual review.","Because readability has only a small inverse association with accuracy, the benchmark appears to measure domain knowledge and reasoning rather than Italian language proficiency; running the same questions in machine translation would test that interpretation directly.","The benchmark's text-only design leaves diagram-dependent items, common in real Italian medical exams, unmeasured; extending the pipeline to images would be a natural next step for evaluating multimodal models.","An independent audit of the publisher's answer keys would strengthen the benchmark's validity claims, since any systematic gold-answer errors would directly distort model rankings."],"forward_implications":["Italian EdTech developers gain a native-language benchmark for model selection, with sub-30 billion parameter models like Phi-4 and Qwen 2.5 14B already scoring above 70 percent.","Logic and mathematics appear as a consistent bottleneck across all tested models, so tutoring or admissions-support tools should be validated on those subjects separately.","The 88.86 percent reproducibility rate for GPT-4o means single-run accuracy differences of a few points may be noise; evaluations should report consistency intervals.","The weak effect of reasoning prompts suggests chain-of-thought prompting cannot be assumed to help on Italian multiple-choice medical questions, especially for strong models.","The significant ordering sensitivity found in Claude 3.5 Haiku implies answer-order robustness should be part of any acceptance test for deployed exam-answering systems."],"supporting_citations":[{"why":"Sets the comparable multiple-choice scale and format that motivated the 17,410-question stratified sample.","marker":"[6]"},{"why":"Supplies the ordering-bias testing method and the emphasis on harder, reasoning-focused questions.","marker":"[4]"},{"why":"Defines the English-language medical exam benchmark that MedBench-IT contrasts with.","marker":"[11]"},{"why":"Establishes the Italian evaluation landscape this benchmark extends with a medical-entrance focus.","marker":"[9]"},{"why":"Provides the closest Italian multi-topic MCQ collection, used to position MedBench-IT's specialized scope.","marker":"[15]"},{"why":"Introduces the chain-of-thought prompting technique used in the reasoning-eliciting condition.","marker":"[20]"}],"fun_headline_variants":["AI scores 91.9% on first Italian med-exam benchmark","DeepSeek-R1 tops Italian med entrance with 91.9%","Italian med-exam benchmark shows AI's logic gap","New benchmark: AI excels on Italian med facts, not logic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The publisher's answer keys and difficulty labels are correct, and the stratified sample of 17,410 questions faithfully represents the full 43,525-question corpus; if either assumption fails, the reported accuracies and rankings do not measure what they claim.","fun_headline_variants_meta":{"raw":{"variants":["AI scores 91.9% on first Italian med-exam benchmark","DeepSeek-R1 tops Italian med entrance with 91.9%","Italian med-exam benchmark shows AI's logic gap","New benchmark: AI excels on Italian med facts, not logic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001316,"raw_usage":{"total_tokens":5358,"prompt_tokens":939,"completion_tokens":4419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":4345}},"tokens_in":555,"tokens_out":4419,"duration_ms":29472,"temperature":1.0,"reasoning_tokens":4345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:12:45.279463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit a random sample of the benchmark questions with independent Italian medical educators and re-run the top models on the 26,115 questions from the original corpus that were excluded; non-trivial gold-answer errors or a large accuracy gap on the excluded set would invalidate the reported scores.","supporting_citations":[{"cited_title":"Hendrycks, C","cited_arxiv_id":null,"evidence_quote":"Sets the comparable multiple-choice scale and format that motivated the 17,410-question stratified sample."},{"cited_title":"Attanasio, P","cited_arxiv_id":null,"evidence_quote":"Establishes the Italian evaluation landscape this benchmark extends with a medical-entrance focus."},{"cited_title":"Rinaldi, J","cited_arxiv_id":null,"evidence_quote":"Provides the closest Italian multi-topic MCQ collection, used to position MedBench-IT's specialized scope."},{"cited_title":"Wei, et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,","cited_arxiv_id":null,"evidence_quote":"Introduces the chain-of-thought prompting technique used in the reasoning-eliciting condition."}],"review_version":2}