{"id":"479e449c-d305-414f-a520-ef1e12a7c034","arxiv_id":"2412.17933","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"BenCzechMark is a new 50-task Czech benchmark with a duel scoring system, a 320GB Czech corpus, and a leaderboard of 50 open-weight models.","lead":"This paper introduces BenCzechMark, a 50-task Czech-language benchmark for large language models with 14 newly collected tasks and a leaderboard of 50 model submissions. It also releases a 320GB cleaned Czech corpus, trains Czech-centric models of 1.2B, 1.6B and 6.7B parameters, and proposes a duel-based scoring system that ranks models by statistically significant pairwise wins.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Without FDR control, the per-duel α=0.05 tests in §4.3 allow chance wins to dominate task-level DWS; the paper's own §9 BH analysis shows a 72% increase in non-significant duels, so the 'significance-grounded' duel score is only weakly supported at task level.","rationale":"We agree with the reader that the absence of FDR control in the per-duel significance tests is the most load-bearing weakness. The paper is unusual in honestly quantifying the problem in Section 9, but that same quantification shows the default DWS is much noisier than a 'significance-grounded' score implies. The defense—that category-level rankings are stable and that the leaderboard offers a BH-corrected mode—mitigates the harm for users who want only coarse rankings, but it does not rescue the per-task DWS values or the analyses built on them, such as the task-similarity dendrogram and per-task win-score heatmaps. We therefore regard the central methodological claim as only partially supported, exactly as the reader suggests. The benchmark resource itself is valuable: 50 tasks, 14 new datasets, contamination checks, and a reproducible evaluation harness are real contributions that should be preserved. The correct verdict is conditional: accept the resource and leaderboard with the requirement that headline task-level DWS be reported with the FDR-controlled variant or that task-level false-discovery risk be clearly quantified. No verdict change is needed beyond the reader's.","tokens_in":36040,"tokens_out":6291,"duration_ms":59355,"concrete_test":"Using the released per-model, per-task outputs from the leaderboard, recompute all DWS matrices under Benjamini-Hochberg correction at α=0.05 (the leaderboard already supports this switch). Then re-generate the per-task DWS matrix (Fig. 8) and the task-similarity dendrogram (Fig. 5). If the corrected task-level vectors change such that fewer than 70% of task-level pairwise orderings are preserved, or the dendrogram's cluster membership shifts, the default task-level DWS is not robust. Additionally, run a permutation test for a few tasks to estimate the empirical false-positive rate and confirm the expected ≈5% per-duel level is inflated to a much higher rate across duels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that BenCzechMark's duel scoring mechanism is 'grounded in statistical significance theory' and yields fair per-task DWS values. This requires that per-duel significance decisions at α=0.05 are not overwhelmed by chance. In §4.3 the authors use a one-tailed paired t-test at α=0.05 for accuracy/EM, with no multiple-comparison correction. With N=50 models, each task has 1,225 duels, so under the global null roughly 61 false-positive 'wins' per task are expected. The authors' own Section 9 analysis shows that applying Benjamini-Hochberg makes 72% more duels non-significant and, for three tasks, cuts non-zero DWS values by more than half. This demonstrates that a substantial fraction of duels feeding task-level DWS are plausible noise. The paper argues category rankings remain correlated (τ=0.64–0.87), but the headline results and per-task analyses (Figures 4, 5, 8) rely on the uncorrected procedure. Thus the load-bearing premise—that chance improvements do not dominate DWS—is not met at task granularity, and task-level comparisons drawn from the default leaderboard are not reliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BenCzechMark, a Czech-language multitask benchmark with 50 tasks across 8 categories, a leaderboard with 50 model submissions, a duel scoring mechanism (DWS) based on pairwise statistical significance tests, and a category/overall win score aggregation inspired by Borda count. It also contributes BUT-LCC, a large cleaned Czech corpus, a contamination analysis pipeline, and several Czech-centric baseline models. The headline methodological claim is that the duel scoring system is 'grounded in statistical significance theory' and that it mitigates chance-driven improvements when comparing models.","tokens_in":36280,"tokens_out":2042,"duration_ms":22552,"significance":"If the methodological claims hold, BenCzechMark would be a valuable resource for Czech LLM evaluation: it is notably larger in task coverage than prior Czech benchmarks, it ships a maintained leaderboard, it includes newly collected native Czech tasks, it performs systematic contamination checks against a large corpus, and it releases the corpus and baselines. The work also proposes a concrete aggregation mechanism (DWS with Borda-style averaging) that goes beyond simple metric averaging. The paper is transparent about its own limitations, including the multiple-comparison issue discussed in Section 9, which strengthens the overall credibility of the artifact even though it weakens the central claim.","major_comments":[{"comment":"The central claim that DWS is 'grounded in statistical significance theory' and mitigates chance improvements is not supported at task granularity. Section 4.3 applies a one-tailed paired t-test at α=0.05 per duel without multiple-comparison control. With N=50 models, each task involves 1,225 pairwise duels, so under the global null roughly 61 false-positive duels per task are expected. The authors' own Section 9 analysis shows that Benjamini-Hochberg correction makes 72% more duels non-significant and, for three tasks, reduces non-zero DWS values to less than half. Since the default leaderboard and Figures 4 and 8 use the uncorrected per-duel procedure, task-level DWS values and per-task conclusions drawn from the default leaderboard are not reliable as significance-based measures. I recommend using the DWS-level corrected procedure as the primary scoring, or at least reporting all task-level analyses with corrected DWS and restricting the 'significance-grounded' claim to rank-level stability.","section":"Section 4.3 and Section 9"},{"comment":"The text states that Qwen2.5 models are excluded from evaluation, but that 'Reported DWS still captures duels with these models.' If those models are excluded from the leaderboard, their duels should not contribute to the denominator of other models' DWS values; if they are retained in the DWS computation, then the exclusion has no effect on scores and the contamination-based exclusion is purely cosmetic. Please clarify how DWS was recomputed after removing Qwen2.5, and if the reported DWS values include duels with excluded models, the leaderboard and figures should be regenerated.","section":"Section 7, 'Model Contamination'"},{"comment":"The distinction between 'duel-level' (A) and 'DWS-level' (B) statistical guarantees is important, but the paper does not state which analyses in Section 6 use which guarantee. The default leaderboard is stated to use (A), and the headline OWS/CWS results in Figures 2 and 4 appear to be based on (A). If the leaderboard allows switching to (B), the paper should specify which of the reported results are affected by the switch, and it should not present (A)-based task-level values as statistically robust.","section":"Section 9, last paragraph"}],"minor_comments":[{"comment":"The phrase 'the first comprehensive Czech language benchmark' is contestable given the concurrent CzechBench (Section 2 explicitly acknowledges it); consider softening to 'the largest' or 'the most comprehensive at time of writing.'","section":"Abstract and Introduction"},{"comment":"The hierarchical clustering dendrogram is difficult to read in its current rendering; the task labels overlap and the caption does not describe the distance metric or linkage criterion beyond 'cosine similarity' and 'hierarchical clustering.' Please make the figure legible and specify the clustering details.","section":"Figure 5"},{"comment":"The max-centered variance V_m(P_t) divides by |P_t|−1 even when a single prompt is used; clarify the handling of tasks with one prompt (e.g., LMF tasks) and state the averaging procedure over tasks explicitly.","section":"Section 4.4, Equation (1)"},{"comment":"The table lists several datasets with contamination above 50% (e.g., SumeCzech, CzechNews) that were removed, while other high-contamination datasets (e.g., Propaganda, Czech Court Decisions) were retained with a justification based on label leakage. The text is clear, but a short sentence summarizing the retention criterion in the main text would help readers who do not see the appendix.","section":"Section 7, Table 2"},{"comment":"The limitation that BCM supports only open-weight models is stated, but the consequence for the leaderboard's representativeness (no closed API models) could be discussed more; it is a design choice that also excludes models otherwise central to current LLM deployment.","section":"Section 9, first paragraph"}],"recommendation":"major_revision","confidential_remarks":"The artifact is strong and the paper is generally well-written, but the central methodological claim about significance-based scoring needs to be reconciled with the authors' own FDR analysis. I believe this is fixable within the scope of a revision, so I do not recommend rejection. The paper would also benefit from a clearer external validation of the ranking (e.g., comparison with an independent Czech benchmark) to complement the internal stability analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the artifact is real and substantial: 50 Czech tasks, 14 newly collected, a 320GB cleaned corpus, Czech-centric baselines, and a maintained leaderboard with 50 submissions. Second, the headline duel scoring method has a multiple-comparisons soft spot that the authors themselves document, and they ship the stricter variant on the leaderboard. That combination makes this a conditional accept rather than a clean one.\n\nWhat is genuinely new: the scale and care of the resource. The contamination analysis with the canary string, the manual curation of several datasets (HistoryIR, Czech SNLI, CERMAT), and the prompt-sensitivity metric (MCV) are all concrete contributions. The paper is transparent about what was removed and why, and the code and data are public. The authors also report their own negative result on HellaSwag-CZ translation quality instead of sweeping it under the rug. That is honest, reproducible work.\n\nThe soft spot is the one you already identified. In Section 4.3, per-duel tests run at alpha=0.05 without FDR control, and the authors' own Section 9 shows that BH correction makes 72% more duels non-significant and badly cuts non-zero DWS for three tasks. So task-level DWS values in Figures 4, 5, and 8 should be read as optimistic. However, the global and category-level rankings survive the correction with Kendall's tau 0.64–0.87, and the leaderboard lets users switch to the BH-corrected scores. That means the main contribution—a fair, reusable benchmark for Czech—does not collapse. The flaw is real but contained, and the authors have already mitigated it in the deployed artifact.\n\nTwo smaller concerns: the best-prompt-per-model selection can inflate scores and complicates comparisons, though the MCV analysis at least quantifies the sensitivity. And the Qwen2.5 exclusion is reasonable but the reported DWS still includes duels with those models, which is a bit inconsistent.\n\nWho this is for: anyone building benchmarks for mid-resource languages, and any Czech NLP group that needs a common ruler. It deserves a serious referee. I would recommend sending it out, with a request that the paper lead with the BH-corrected scores as the default (or at least give them equal billing), and make the task-level caveats explicit in the abstract and results.\n\nNet: cite it, use it, and review it carefully—but don't trust the per-task DWS numbers without checking the corrected leaderboard.","headline":"A genuinely useful Czech benchmark artifact with an honest but underemphasized multiple-comparisons problem in its headline scoring; deserves a real referee.","tokens_in":36937,"tokens_out":1375,"would_cite":true,"duration_ms":16308,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BenCzechMark introduces a 50-task, 8-category Czech benchmark that ranks open-weight language models by the proportion of statistically significant pairwise wins ('duels') per task.","keywords":["Czech language benchmark","duel scoring","statistical significance testing","social choice aggregation","Borda count","LLM evaluation","contamination analysis","Czech-centric language model"],"falsifier":"Split each task's test set into two halves, recompute the Duel Win Score per half, and measure the rank correlation of model orderings between halves; if Kendall τ falls well below 0.8 for a substantial fraction of tasks, the task-level scores are not reproducible and the benchmark cannot support fine-grained comparison.","tokens_in":35815,"feed_emoji":"📊","tokens_out":8177,"duration_ms":68212,"temperature":0.7,"pith_summary":"BenCzechMark (BCM) is presented as the first comprehensive Czech-language benchmark for evaluating large language models (LLMs). It combines 50 tasks spanning 8 categories, with 14 newly collected native-Czech datasets, and ranks models not by averaging accuracy but by a 'duel' mechanism: for each task, every pair of models is tested for a statistically significant improvement, and a model's Duel Win Score is the share of duels it wins. Task scores are aggregated into category and overall scores using a Borda-count-style social-choice rule that handles ties. The paper also contributes BUT-LCC, a cleaned 320GB Czech corpus used for contamination filtering and for training the first Czech-centric 7B model, and a leaderboard of 50 model submissions. The authors argue this design addresses three limitations of prior benchmarks: narrow task coverage, miscalibrated classification metrics, and fragile accuracy averaging.","feed_headline":"Czech benchmark ranks LLMs by statistically significant duels","feed_subtitle":"Fifty tasks and eight categories make Czech model comparison testable, stable, and open to new submissions.","key_machinery":"The load-bearing object is the 'duel' — a pairwise significance test between two models on a single task. For accuracy and exact-match metrics the test is a one-tailed paired t-test; for AUROC it is a Bayesian significance test with Monte-Carlo integration; for perplexity it is a bootstrap. The Duel Win Score (DWS) is the proportion of duels a model wins significantly at α=0.05, Category Win Score (CWS) averages DWS over tasks in a category, and Overall Duel Win Score (OWS) averages CWS across the 8 categories, with ties handled by a Borda-count-style rule in which each task acts as a voter and models as candidates. This mechanism is what the paper claims makes the ranking resilient to chance improvements and to changes in the set of evaluated models.","core_discovery":"On its own terms, the paper's central claim is that BenCzechMark is the first comprehensive multitask, multimetric benchmark for open-weight LLMs in Czech, and that its duel-scoring aggregation makes model comparison fairer than standard averaging. For every task and every pair of models, a one-tailed significance test (a paired t-test for accuracy and exact match, a Bayesian test for AUROC, and bootstrapping for perplexity) decides a win, loss, or tie at α=0.05. A model's Duel Win Score is the proportion of duels it wins significantly; these scores are averaged within task categories (Category Win Score) and then across categories (Overall Duel Win Score) using a Borda-count-inspired rule with ties. The paper reports that with 50 submitted models the resulting leaderboard is stable under removal of up to 20 models (Kendall τ typically high for multi-task categories), identifies a contaminated model family (Qwen2.5) via a canary string, and finds that Czech-specific tokenization gives their 7B model an advantage in language modeling perplexity but not in understanding tasks. The authors also propose max-centered variance (MCV) as a per-model and per-task measure of prompt sensitivity.","pith_inferences":["The same duel-scoring machinery could transfer to other under-resourced languages, where a few dozen native tasks plus significance-based aggregation may yield fairer rankings than translated MMLU clones.","The paper's own limitation data (72% more duels become non-significant under Benjamini-Hochberg correction, while category rankings stay moderately correlated) suggests that consumers should trust coarse category-level rankings over individual task-level DWS values.","A direct test of the stability claim would be to recompute the leaderboard after removing the historically weakest models; if the top-of-table ordering changes, the 'stable comparison' claim holds only for the full set of 50 submissions, not for smaller evaluation runs.","The superhuman model performance on the poorly-translated HellaSwag-CZ (humans ~60%, best model >70%) suggests translated benchmarks can measure translation artifacts, so future Czech benchmarks should treat automatically translated tasks with caution."],"forward_implications":["A new open-weight model can be submitted to the BCM leaderboard and compared across 50 Czech tasks, with a stability-tested Overall Duel Win Score that does not rely on a single average metric.","Czech-specific progress can be located by category (math, NLI, sentiment, NER, reading comprehension, language modeling, factual knowledge, Czech language understanding), so a model's overall win can be separated from its failure on specific task families such as Umime-to-Czech.","The prompt-sensitivity measure (MCV) tells users which models and tasks need multi-prompt evaluation, preventing one lucky prompt from being mistaken for capability.","Because the benchmark only uses open-weight models (log-likelihood access), the leaderboard provides a reproducible comparison point for the open model ecosystem rather than for API-based systems.","The release of BUT-LCC and the Czech-centric models gives other teams a cleaned corpus and baseline for continuous pretraining in Czech."],"supporting_citations":[{"why":"Supplies the paired t-test procedure used for accuracy and exact-match duels.","marker":"(Dror et al., 2018)"},{"why":"Provides the Bayesian significance test that the paper extends to AUROC duels.","marker":"(Goutte and Gaussier, 2005)"},{"why":"Provides the modern implementation of the Bayesian test used for AUROC duels.","marker":"(Morais, 2023)"},{"why":"Gives the bootstrap procedure used for perplexity duels.","marker":"(Berg-Kirkpatrick et al., 2012)"},{"why":"The social-choice voting rule that the paper's tie-handling aggregation adapts.","marker":"(Borda, 1781)"},{"why":"Showed Borda-count aggregation is more robust than averaging to model addition or removal, motivating the win-score aggregation.","marker":"(Colombo et al., 2022)"},{"why":"Provides the closest prior use of pairwise win matrices, which BCM adapts from human preferences to significance-test wins.","marker":"(Chiang et al., 2024)"},{"why":"Defines the contamination-analysis method used to filter BCM datasets against BUT-LCC.","marker":"(Brown et al., 2020)"}],"fun_headline_variants":["Czech benchmark duels LLMs for statistically fair ranking","BenCzechMark: Duel scoring for Czech LLM comparison","Czech-centric LLM benchmark uses statistical duels","Duel scoring makes Czech LLM ranking statistically sound","Benchmark duels: Czech LLMs ranked by significant wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that testing each pair of models separately at α=0.05, with no correction for the thousands of pairwise duels, is enough to keep chance improvements from distorting task-level Duel Win Scores.","fun_headline_variants_meta":{"raw":{"variants":["Czech benchmark duels LLMs for statistically fair ranking","BenCzechMark: Duel scoring for Czech LLM comparison","Czech-centric LLM benchmark uses statistical duels","Duel scoring makes Czech LLM ranking statistically sound","Benchmark duels: Czech LLMs ranked by significant wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00074,"raw_usage":{"total_tokens":3325,"prompt_tokens":990,"completion_tokens":2335,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2256}},"tokens_in":606,"tokens_out":2335,"duration_ms":17947,"temperature":1.0,"reasoning_tokens":2256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:08:03.717730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Split each task's test set into two halves, recompute the Duel Win Score per half, and measure the rank correlation of model orderings between halves; if Kendall τ falls well below 0.8 for a substantial fraction of tasks, the task-level scores are not reproducible and the benchmark cannot support fine-grained comparison.","supporting_citations":[],"review_version":1}