{"id":"0f7d7950-ce02-43bf-a7fc-8bdc8801dbef","arxiv_id":"2501.02266","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMzSzŁ is a new benchmark of almost 19,000 Polish national exam questions with evaluations of 38 language models and comparisons to human results.","lead":"Researchers built a Polish-language benchmark from almost 19,000 official national exam questions and tested dozens of open-weight AI models on it. The result is a reusable way to compare how well language models handle real Polish school and vocational tests over time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's validity rests on CKE answer keys being correct, yet Section 5 reports a known wrong key (M.42-X-18.06 Q24) without stating whether the released dataset was corrected; the unquantified label-error rate is the load-bearing risk.","rationale":"The reader's weakest assumption picks out answer-key correctness; my independent reading of the paper points to the same spot, made sharper by a detail the reader did not emphasize: Section 5 proves the existence of a wrong CKE key, and the release note does not disclose whether that example was corrected. That makes the concern concrete rather than hypothetical. I considered alternative candidates. The unsupported 'largest/most comprehensive' claim is about novelty, not validity. The lack of evaluation code is a reproducibility defect but does not by itself falsify the reported scores. The acknowledged contamination risk is real, but it affects interpretation of model comparisons, not the benchmark's ground truth. The human-correlation analysis has acknowledged limitations, but it is a secondary use of the resource. The gold-label correctness is the foundation: every accuracy number, every model ranking, and every human-correlation conclusion inherits it. Because the paper's own evidence shows the raw keys are not infallible, and because no systematic reconciliation with CKE errata or independent re-annotation is reported, the benchmark's validity should be treated as conditional on a label audit. This does not reject the contribution; a modest random audit would likely confirm a low error rate. But the paper's current wording overstates the certainty of its labels. The reader's CONDITIONAL verdict with moderate confidence is the right level, so I recommend UNCHANGED.","tokens_in":14529,"tokens_out":5483,"duration_ms":55534,"concrete_test":"Download the released dataset from HuggingFace, locate exam M.42-X-18.06, Question 24, and check whether the stored answer is 'insolvent' (the erroneous CKE key) or 'strategic' (the corrected answer). If it is 'insolvent', the known error ships in the benchmark and the Section 3.1 claim fails on its own evidence. Then independently audit a random sample of at least 400 questions stratified by exam type and domain: two native-Polish-speaking annotators, blind to the CKE key, produce gold answers; compare against dataset labels; estimate the label-error rate with a 95% confidence interval and recompute the leaderboard after removing items where the dataset disagrees with both annotators. If the CI upper bound exceeds ~1%, or if any top-model ranking changes, the benchmark's validity is conditional on a full label audit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LLMzSzŁ's almost 19,000 CKE-derived gold labels are reliable enough to support the reported accuracy scores and model comparisons. Section 3.1 asserts that CKE credibility 'minimizes the risk' of incorrect answers, but Section 5 documents a concrete counterexample: in exam M.42-X-18.06, Question 24, the expected answer was 'erroneously specified as insolvent instead of strategic.' The paper does not say whether the released HuggingFace dataset was corrected to the right label. If the erroneous label remains, the dataset contains at least one known wrong gold label, and the validity claim is contradicted by the authors' own evidence. If it was corrected, the paper omits an important detail about how labels are curated, and users cannot assume raw CKE keys were taken verbatim. Either way, the error rate of the answer-key extraction pipeline is unquantified. One error among 19,000 may be negligible for aggregate accuracy, but CKE itself publishes official errata to its answer keys; the paper does not describe reconciling the dataset with those errata. A non-negligible label-error rate (e.g., >1%) concentrated in particular domains could shift model rankings and distort the human-correlation analysis in Section 6, which is already acknowledged to compare closed-question model scores with human scores that include open questions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LLMzSzŁ, a benchmark of almost 19,000 closed-ended Polish exam questions drawn from Polish Central Examination Board (CKE) materials, covering school and vocational exams. The authors evaluate dozens of open-weight LLMs using likelihood-based accuracy, analyze performance by model size, language, release date, and instruction tuning, and investigate correlations between model scores and human examinee statistics. The dataset is released on HuggingFace, and a public leaderboard is provided.","tokens_in":14793,"tokens_out":8505,"duration_ms":76154,"significance":"If the gold labels are reliable, the dataset is a valuable evaluation resource for Polish, with per-exam timestamps that can support contamination-aware analysis. The systematic evaluation of a broad range of open-weight models is useful and reproducible. However, the paper's central novelty claim (being the largest and most comprehensive Polish LLM benchmark) is contradicted by its own cited references, and the known answer-key error documented in Section 5 is not resolved in the released dataset. The human-correlation analysis in Section 6 is statistically weak due to small sample sizes and incompatible metrics. These issues currently limit confidence in the paper's headline claims, though they are addressable in revision.","major_comments":[{"comment":"The paper claims that LLMzSzŁ is 'the largest and most comprehensive LLM benchmark developed for the Polish language that has been published to date' (Section 7). However, the cited dataset of Pokrywka et al. (2024) contains 297 tests of 120 questions each (35,640 questions), and the follow-up Łukasz Grzybowski et al. (2024) adds 144 new exams, both larger than the 'almost 19k' questions in LLMzSzŁ. Please correct or explicitly qualify this claim, for example by specifying that LLMzSzŁ is the largest multi-tier general-knowledge Polish benchmark, not the largest Polish exam-based dataset in absolute terms.","section":"Section 7 and Abstract"},{"comment":"Section 5 reports a concrete label error: in exam M.42-X-18.06, Question 24, the CKE answer key was 'erroneously specified as insolvent instead of strategic.' The paper does not state whether the released HuggingFace dataset corrects this label, nor does it describe any systematic reconciliation with CKE's official errata. Since the benchmark's gold labels are the CKE keys (Section 3.1), the existence of at least one known wrong label makes the unquantified label-error rate a load-bearing reliability issue. Please explicitly state how known errors are handled in the released dataset, document the curation process, and either provide an estimated label-error rate from a manual audit or temper the claims about answer-key reliability.","section":"Section 5 and Section 3.1"},{"comment":"The correlation analysis in Section 6 and Table 5 uses very few yearly data points (e.g., Junior High 2015–2019: n=5; 8-grade: n=5) and mixes model scores on closed questions with human scores that include open questions, while for professional exams the human values are pass rates rather than average scores. Correlations such as 0.925 (Mistral, Junior High) and 0.851 (Bielik, 8-grade) are reported without significance tests or confidence intervals. The conclusions in Section 6.4 about using LLMs to verify exam difficulty are therefore not statistically supported. Please provide p-values or confidence intervals, use comparable human metrics if available, or explicitly downgrade the strength of these conclusions.","section":"Section 6 and Table 5"},{"comment":"The abstract's claim that 'multilingual LLMs can obtain superior results over monolingual ones' is not cleanly supported by the comparisons shown. The highest-performing models are all multilingual, but there are no Polish or English models of comparable size (e.g., 70B–123B) in the evaluation; the best small model, Bielik-11B (57.52), is only directly compared with 7B multilingual models (Table 6). The comparison is thus confounded by model size and availability. Please provide a matched-size comparison (e.g., Bielik-11B against a multilingual model of similar size) or qualify the conclusion accordingly.","section":"Section 4.2 and Abstract"}],"minor_comments":[{"comment":"In Table 6, the parameter size for 'Qwen/Qwen2-1.5B' is listed as 5 (likely 1.5), and the release date for 'trurl-2-13b-academic' is given as '23-98', which is not a valid month.","section":"Table 6"},{"comment":"Table 7 contains the typo 'Phisics' for 'Physics' in several rows.","section":"Table 7"},{"comment":"The text states that for biology 'The lack of correlation may be due to the increase in difficulty of open questions,' but Table 5 does not provide a separate biology correlation; please either report the correlation or present this as a qualitative observation.","section":"Section 6.2.1"},{"comment":"The list of selected subjects is ambiguous: 'math, natural sciences, biology, physics' suggests overlapping categories; please clarify the exam-subject taxonomy used for the benchmark.","section":"Section 3.1"},{"comment":"In the Limitations section, 'questions that where published' should be 'questions that were published.'","section":"Section 8"},{"comment":"The sentence 'If this phenomenon is confirmed with more data (possibly including open questions), it will advocate for a possible use of LLMs...' should be reworded for grammatical correctness and precision.","section":"Section 6.4"}],"recommendation":"major_revision","confidential_remarks":"The claim that LLMzSzŁ is the 'largest' Polish benchmark is directly contradicted by the authors' own earlier medical-exam dataset (Pokrywka et al., 2024; Łukasz Grzybowski et al., 2024). This appears to be an oversight, but it is a factual error in the central novelty statement and should be corrected before publication. In addition, the unresolved label-error issue and the weak correlation evidence are matters that require careful revision. The dataset itself appears real and useful, and the evaluation harness is standard, so I see no reason for rejection if these concerns are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid resource paper, not a breakthrough. The dataset is real, on HuggingFace, built from CKE Polish national exams, about 19k closed questions, with publication dates and vocational coverage. That fills a genuine gap: m_mmlu and similar resources are machine-translated, and the Polish medical exam datasets are narrower. The temporal metadata is the most useful design choice, because it allows contamination-aware evaluation. Running 38 open-weight models through a likelihood-based harness is routine but competently done; the results mostly confirm known scaling and cross-lingual transfer patterns, no big surprise.\n\nWhat is genuinely good: the dataset is public, the single authoritative source avoids duplicate questions, the manual answer-matching is documented, and the Section 5 anomaly-hunting is a nice demonstration of using models to flag bad exam items. The 2018 M.42 error they found is real evidence.\n\nSoft spots, in order. First, the \"largest comprehensive benchmark for Polish\" claim is wrong on the paper's own citations: the Pokrywka medical dataset is 297 exams times 120 questions, roughly 35.6k items, nearly double this size. If they mean largest general or largest school-plus-vocational benchmark, they should say that. It is an easy fix, but it matters because it sets the wrong frame.\n\nSecond, the label-error question is not handled. The paper says CKE credibility \"minimizes the risk\" of incorrect answers, then reports a confirmed erroneous answer key in M.42-X-18.06 Q24. It never states whether the released HuggingFace dataset contains the corrected label or the raw CKE key. Since the benchmark's validity rests on gold labels, users need the correction policy and ideally the errata reconciliation. One known error among 19k probably does not shift rankings, but an unquantified extraction-pipeline error rate is exactly what a benchmark paper should report.\n\nThird, the human-correlation section is the weakest. Table 5 correlations rest on very few yearly points, the comparison mixes closed-question model scores with human scores that include open questions, and the authors read a small correlation matrix selectively. To their credit, they acknowledge the open-question limitation and do not overclaim. So this is a moderate flaw, not fatal.\n\nFourth, no evaluation code or config is shipped, only the dataset and leaderboard; reproducing exact scores requires rebuilding the harness. Minor, but noted.\n\nNo circularity concern: the models were not trained on this benchmark as far as the paper states. This is an external measurement instrument.\n\nWho this is for: people building or using Polish LLMs, and researchers interested in cross-lingual transfer or contamination-aware benchmarking. It deserves peer review; the issues are addressable and do not undercut the resource. I would send it to review with a request to fix the \"largest\" claim, state the label-correction policy, and temper the correlation analysis.","headline":"A genuinely useful new Polish exam benchmark with real public data, held back by an inflated 'largest' claim and an unquantified answer-key error rate; worth peer review with revisions.","tokens_in":15359,"tokens_out":2414,"would_cite":true,"duration_ms":24154,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark built from almost 19,000 Polish national exam questions tests open-weight LLMs against human examinees.","keywords":["LLMzSzŁ","Polish LLM benchmark","national exams","CKE","multilingual LLM evaluation","human-model correlation","answer key validation"],"falsifier":"Audit a random sample of the 19,000 answer keys against independent expert judgment; if the error rate is high enough to change the order of the top models or to erase the observed human-model correlations, the benchmark's claims about model quality and exam validation would collapse.","tokens_in":14374,"feed_emoji":"📝","tokens_out":7018,"duration_ms":64605,"temperature":0.7,"pith_summary":"The paper introduces LLMzSzŁ, a benchmark built from almost 19,000 closed-ended questions taken from Polish national school and vocational exams published by the Central Examination Board. It is designed to test whether open-weight language models can handle Polish-specific knowledge and reasoning, and to compare model performance with that of human examinees by year and exam category. The authors evaluate dozens of models and report that large multilingual models score highest, while smaller Polish-tuned models remain competitive when size is constrained. They also show that model scores correlate with human pass rates on some exams, and that low model confidence can expose genuine errors in official answer keys, including one confirmed mistake in a 2018 vocational exam.","feed_headline":"19,000 Polish exam questions rank LLMs and expose bad answers","feed_subtitle":"Multilingual models lead the scoreboard; a low-confidence flag uncovered one official answer-key error.","key_machinery":"The load-bearing object is the LLMzSzŁ dataset itself: nearly 19,000 single-choice Polish exam questions with official answer keys, stratified into middle-school, high-school, and vocational tiers, each item carrying the publication date of its exam. The evaluation procedure uses an open evaluation harness with an MMLU-style configuration: for each question the model computes the probability of each of the four answers and the highest-probability answer is scored against the gold key. A feature-level analysis using the Mann-Whitney U test identifies which question characteristics (numerical answers, words like 'wynosi' or 'oblicz', professional domains such as R.13) drive low model scores, and the per-year human score comparison provides the correlation evidence. The timestamps are the design feature that enables contamination-controlled evaluation and the year-by-year human-model comparison.","core_discovery":"The central claim is that a coherent, authoritative collection of Polish national exams can serve as a comprehensive evaluation benchmark for Polish-language LLMs, at a scale not previously available. The dataset covers four exam types across 154 domains, with per-exam publication timestamps that allow contamination-aware evaluation; all questions are closed-ended with one correct answer, and the gold labels are the official answer keys. Evaluations performed with an open evaluation harness configured in the style of MMLU show that the best overall model is Mistral-Large-Instruct-2407 (67.17% accuracy), that models below about 3 billion parameters perform near the random-guess level of 25%, and that instruction-tuned variants generally outperform their base counterparts. The authors further claim that comparing model outputs with human results can help estimate exam difficulty and validate exam questions, evidenced by the discovery of a faulty answer key in the 2018 M.42 exam when a model assigned very low probability to the expected answer.","pith_inferences":["Inference: the low-confidence detection method could be turned into a systematic audit protocol for answer keys across all exam types, rather than the single case study reported here.","Inference: if answer-key errors are not uniformly distributed across exam categories, the reported leaderboard may partly reflect how well a model matches CKE's answer conventions; re-scoring after expert correction of a random key sample would test this.","Inference: the timestamp design could serve as a template for national-exam benchmarks in other languages, offering a built-in contamination control that translated MMLU datasets lack.","Inference: the divergent human-model trends in biology (human scores falling, model scores rising) suggest that open questions, not closed ones, drive the human decline; this could be tested by scoring open responses with an LLM rubric."],"forward_implications":["Polish-language model evaluation gains a public, reproducible benchmark with a random-guess baseline of 25% accuracy.","Researchers can use the timestamps to split questions into pre- and post-release sets, reducing the effect of training-data contamination when comparing models of different release dates.","The reported results imply that for high-accuracy Polish tasks, large multilingual models are the best choice, while the 11B-parameter Polish Bielik model offers a practical alternative where size is limited.","Using LLMs to flag low-probability expected answers can serve as a screening step for national exam quality control, with one confirmed answer-key error already found.","Human-model correlations on specific exams suggest that model scores could serve as a proxy for closed-question difficulty trends, potentially separating difficulty shifts in open questions from closed questions."],"supporting_citations":[{"why":"Supplies the MMLU exam-based benchmark design and the evaluation framework that LLMzSzŁ adapts to Polish national exams.","marker":"Hendrycks et al. (2020)"},{"why":"Provides the open evaluation harness and likelihood-scoring implementation used to compute model accuracies.","marker":"Gao et al. (2024)"},{"why":"Documents errors in MMLU, motivating the paper's emphasis on ground-truth correctness and answer-key credibility.","marker":"Gema et al. (2024)"},{"why":"Contributes the feature analysis method used to identify question characteristics correlated with low model scores.","marker":"Graliński et al. (2019)"},{"why":"Previous Polish exam-based medical benchmark that frames the medical evaluation context and serves as a comparative reference.","marker":"Pokrywka et al. (2024)"},{"why":"Extends the Polish medical exam work with cross-lingual analysis, informing the interpretation of multilingual versus monolingual model performance.","marker":"Łukasz Grzybowski et al. (2024)"}],"fun_headline_variants":["Polish national exams rank LLMs, expose official answer-key error","19k Polish school questions stress-test LLMs, flag faulty key","Multilingual LLMs outscore Polish-only on 19k exam questions","19k official Polish exam Qs: LLM benchmark flags wrong answer key"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ground truth is the official answer key published by the Polish Central Examination Board; the central assumption is that those keys are correct, or at least that errors are too rare to change model rankings.","fun_headline_variants_meta":{"raw":{"variants":["Polish national exams rank LLMs, expose official answer-key error","19k Polish school questions stress-test LLMs, flag faulty key","Multilingual LLMs outscore Polish-only on 19k exam questions","19k official Polish exam Qs: LLM benchmark flags wrong answer key"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001231,"raw_usage":{"total_tokens":5033,"prompt_tokens":895,"completion_tokens":4138,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":4061}},"tokens_in":511,"tokens_out":4138,"duration_ms":24936,"temperature":1.0,"reasoning_tokens":4061,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:13:46.395569+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit a random sample of the 19,000 answer keys against independent expert judgment; if the error rate is high enough to change the order of the top models or to erase the observed human-model correlations, the benchmark's claims about model quality and exam validation would collapse.","supporting_citations":[],"review_version":1}