{"id":"f4372de2-dde5-4c2e-8d57-48247bc491cb","arxiv_id":"2608.12894","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new tri-lingual Bavarian culture benchmark shows open-weight LLMs underperform on Bavarian and source-grounded items, and that evaluation protocol materially changes accuracy and rankings.","lead":"This paper introduces BavGround, a 618-question benchmark that tests language models on Bavarian cultural knowledge across English, German, and Bavarian. It finds that open-weight models score much lower on Bavarian and source-grounded questions, and that measured scores change substantially with the evaluation method.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bavarian gap is protocol-dependent: under semantic generated-answer matching, Bavarian (55.7) is essentially tied with English (56.0), and the 'persistent difficulty' claim lacks bootstrap support outside letter scoring.","rationale":"The reader's CONDITIONAL verdict is well-founded; my independent pass lands on a different weak point. The benchmark is thoughtfully constructed: multi-parallel design, category balance, source attribution for GRD items, a full protocol suite, bootstrap CIs for the headline comparisons, and unusually candid limitation statements. The translation-authenticity concern (Section 8, sixth limitation) is real and direction-conservative: if the Bavarian items are too close to German, the measured gap underestimates the true dialect gap, and the authors say so. That makes it an item-validity concern but not one that falsifies the headline direction. The more load-bearing problem is measurement validity of the headline gap. The abstract's central empirical claim is that models show persistent difficulty on Bavarian. The only CIs provided for this gap are under letter scoring (Table 8), but the paper's own Table 16 shows that under semantic matching the Bavarian open-weight mean (55.7) is essentially tied with English (56.0) and only 4 points below German, while under hidden-state isolated alignment it is the highest language. No bootstrap intervals are reported for these alternative-protocol gaps, so the Section 5.3 robustness claim is an assertion, not a demonstrated result. Since the paper's central lesson is that protocol choice can change conclusions, the Bavarian-gap conclusion must itself be tested for protocol robustness. If the gap fails that test, the abstract overstates the finding; if it survives, the concern is resolved. This is a concrete, artifact-checkable test, and adding these CIs is a modest revision. The reader's requested translation validation remains an additional valid condition; the two concerns are complementary, and neither alone would overturn the CONDITIONAL verdict, so I recommend keeping it unchanged while adding the protocol-robustness requirement.","tokens_in":24346,"tokens_out":15792,"duration_ms":152992,"concrete_test":"From the released item-level records (results_allmodels; 111,240 rows), compute paired nonparametric bootstrap CIs (10,000 resamples over the 206 source-question IDs, matching the paper's Table 8 procedure) for the open-weight mean Bavarian-vs-English and Bavarian-vs-German gaps under three non-letter protocols: option_text_avg, semantic_embed_generated_answer, and letter_shuffled. If any 95% CI includes zero, the Section 5.3 robustness claim fails and the abstract's 'persistent difficulty' should be rephrased as a letter-scoring-specific result. Also recompute excluding GENBA-10B-it and LLaMmlein-7B to check that the gap is not carried by outlier models with extreme label priors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline empirical claim — that models show persistent difficulty on Bavarian items — is demonstrated with bootstrap CIs only for the letter-scoring protocol (Section 5.1, Table 8). Yet the paper's own protocol comparison (Section 5.3, Table 16) shows the Bavarian gap nearly vanishes under content-based evaluation: at the open-weight mean, option_text_avg gives EN 61.9 / DE 59.8 / BA 56.1 (gap 3.7–5.8 points), and semantic_embed_generated_answer gives EN 56.0 / DE 59.7 / BA 55.7, i.e. Bavarian is essentially tied with English and only 4.0 points below German. Under hidden-state isolated alignment Bavarian is actually the highest language (34.1 vs 33.6 and 33.5). No bootstrap CIs are reported for any of these alternative-protocol language gaps, so the Section 5.3 assertion that alternative protocols do not erase BAVGROUND's core difficulty structure is unsupported for the Bavarian dimension. Because the paper's central methodological lesson is that single-protocol MCQ scores can mislead, the headline gap itself must be shown to survive the same critique; otherwise the 'persistent difficulty' claim is letter-scoring-specific, not a stable dialect-competence finding. This is independent of the acknowledged translation-authenticity limitation: even with authentic items, the measured gap must be robust across meaning-based protocols.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BavGround, a multiple-choice benchmark of 206 source questions translated into English, German, and Bavarian (618 instances), spanning eight cultural domains and two item types: general-knowledge (GEN) questions and source-grounded (GRD) questions derived from regional journalism and specialist literature. Fifteen 7B–10B open-weight instruction-tuned models and one closed reference model are evaluated under a protocol-aware framework that includes letter likelihood, shuffled letters, option-text likelihood, generation parsing, semantic matching, and hidden-state diagnostics. The main findings are that, under standard letter scoring, open-weight models average 53.0% accuracy, with Bavarian items 9.8–11.7 points below English and German and GRD items 16.4 points below GEN items; that scores and rankings shift substantially across evaluation protocols; and that a longitudinal analysis of 85 GENBA-10B continued-pretraining checkpoints shows uneven gains across domains, with letter accuracy remaining low while option-text likelihood improves.","tokens_in":24565,"tokens_out":9271,"duration_ms":79888,"significance":"If the empirical claims hold, BavGround is a valuable addition to cultural NLP evaluation, targeting a genuinely underrepresented regional and dialectal setting. Strengths include the multi-parallel design (same source items in three language versions), source-linked GRD items, the use of bootstrap confidence intervals resampled over source IDs for the main letter-scoring gaps, the systematic multi-protocol comparison showing that MCQ evaluation protocols can change model rankings, and the transparent release of artifacts for reproduction. The protocol-sensitivity finding corroborates prior work by Wang et al. (2024) and others and extends it to regional cultural grounding. The checkpoint analysis, though explicitly exploratory, provides a useful demonstration of protocol-divergent learning trajectories. The main weakness is that the headline 'persistent difficulty on Bavarian items' claim is not yet shown to be robust across meaning-based protocols, which is especially consequential given the paper's own methodological message that single-protocol scores can mislead.","major_comments":[{"comment":"The assertion that 'alternative protocols do not erase BAVGROUND's core difficulty structure' is not supported for the Bavarian dimension. Under semantic_embed_generated_answer the open-weight mean is 56.0 (EN) versus 55.7 (BA), a negligible 0.3-point gap, and under option_text_avg the gap narrows to 5.8 points (61.9 versus 56.1); under hidden-state isolated alignment Bavarian is actually the highest language (34.1 versus 33.6 and 33.5). No bootstrap confidence intervals are reported for these alternative-protocol language gaps, whereas the letter-scoring gaps in Appendix Table 8 have them. Because the paper's central methodological lesson is that single-protocol scores can mislead, the headline claim of 'persistent difficulty on Bavarian items' (Abstract and Section 5.1) must be either supported by confidence intervals for meaning-based protocols, or explicitly restricted to letter scoring. The GRD gap appears more robust for GENBA (Table 18), but the Bavarian-language claim as currently stated is letter-scoring-specific.","section":"§5.3, Table 16"},{"comment":"The Bavarian translations were produced by a single native Chiemgau speaker who has resided outside Bavaria for an extended period, and the paper acknowledges likely dialect attrition, a low Levenshtein distance from German, and the absence of inter-translator agreement. This is a meaningful gap for a benchmark whose central contribution is dialect competence. The caveat that Bavarian performance is an 'upper bound' appears only in the Limitations and is absent from the Abstract and Results. I recommend moving this caveat into the main text and the abstract, and providing at least a few example Bavarian items in the appendix so that reviewers and readers can judge the authenticity of the target variety. Absent such evidence, the dialect-competence claims should be presented as provisional.","section":"§8, Limitations (sixth point)"},{"comment":"The GEN–GRD gap is reported with a bootstrap confidence interval only under letter scoring (16.4 points, CI [8.2, 24.3]). Given the paper's emphasis on protocol sensitivity, the reader cannot tell whether the 'source-grounded questions are harder' result is stable across protocols for the open-weight mean; the GENBA case study (Table 18) shows the gap persists for that model, but the open-weight mean result is missing. I recommend reporting the open-weight mean GEN–GRD gap under option_text_avg and semantic matching with confidence intervals, or explicitly stating that this cross-protocol analysis is unavailable.","section":"§5.2 and Appendix Table 8"}],"minor_comments":[{"comment":"The phrase 'Bavarian remains below English and German under most strategies' is contradicted by the hidden-state isolated row, where Bavarian (34.1) is highest; please revise to be precise, for example 'under letter and option-text scoring'.","section":"§5.3, Table 16"},{"comment":"The rows labeled 'Hidden-state contextual' and 'Hidden-state isolated' are described as 'average open-weight accuracy' but are alignment diagnostics; clarify that they are not accuracy measures.","section":"Table 16"},{"comment":"The translation process is described as performed by 'a native-speaking co-author' in the introduction and by 'an in-house expert and native speaker' in Section 3.2; clarify whether these refer to the same individual.","section":"§1 and §3.2"},{"comment":"The GEN–GRD confidence interval [8.2, 24.3] is wide; consider reporting the bootstrap distribution or a one-sided bound to give readers a better sense of precision.","section":"Appendix Table 8"},{"comment":"The phrase 'persistent difficulty with dialectal and localized cultural knowledge' could be read as implying that both components are protocol-robust; consider rewording to reflect the evidence presented in the paper.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The protocol-dependence of the Bavarian gap is the main obstacle to acceptance. If the authors can show that the Bavarian gap survives under semantic matching and option-text scoring with confidence intervals, or alternatively reframe the contribution to emphasize that Bavarian difficulty is protocol-specific, the paper could be accepted after revision. I would also encourage the editor to consider whether the single-translator Bavarian validation is sufficient for a benchmark whose central claim is dialect competence; a small validation study with additional native speakers would substantially increase confidence in the benchmark."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"BavGround is the first benchmark I know of for Bavarian regional cultural grounding, and the resource itself is the real contribution. 206 source questions across eight domains, parallel English/German/Bavarian items, a deliberate split between general and source-grounded questions, and a serious evaluation design: fifteen open-weight models, twelve protocols, bootstrap CIs resampled over source IDs, and over 111k item-level records in the anonymized repo. The GENBA checkpoint analysis is exploratory but cleanly executed. This is a solid piece of empirical work that deserves to be a community resource.\n\nThe main soft spot is the headline claim that models show persistent difficulty on Bavarian items. Under letter scoring the gap is real and bootstrapped: 9.8 points against English, 11.7 against German. But under the paper's own meaning-based protocols the gap largely disappears. Semantic matching of generated answers gives Bavarian 55.7 vs English 56.0; option-text-avg narrows the gap to 3.7-5.8 points; hidden-state isolated alignment actually puts Bavarian slightly highest. No bootstrap intervals are reported for any of these alternative protocols, and the Section 5.3 assertion that alternative protocols do not erase the core difficulty structure is therefore unsupported for the Bavarian dimension. Since the paper's central methodological lesson is that single-protocol MCQ scores mislead, the persistent-difficulty claim cannot rest on letter scoring alone. This is a load-bearing soft spot, not a minor one. The GEN-GRD gap looks more robust and I expect it would survive, but it should be proven rather than asserted.\n\nThe other weaknesses are acknowledged in Section 8 and weigh less. The Bavarian items were produced by a single translator with likely dialect attrition, so the measured Bavarian gap is an upper bound on true dialect weakness. The Claude-generated GEN questions risk contamination. The checkpoint analysis lacks control runs. All three are stated in the limitations, which I credit as honest reporting.\n\nWho is this for? Anyone doing cultural evaluation, dialect NLP, or German/Bavarian language modeling. With revision on the protocol-robustness claims, BavGround is a valuable benchmark; without it, the abstract overstates the dialect finding. I would send this to peer review, with the expectation of revision on that specific issue.","headline":"A genuinely useful regional-culture benchmark with honest limitations, but the paper's headline Bavarian difficulty claim does not survive its own protocol-sensitivity critique.","tokens_in":25145,"tokens_out":2403,"would_cite":true,"duration_ms":23380,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BavGround shows open-weight LLMs are consistently weaker on Bavarian dialect and regional knowledge than on German or English.","keywords":["Bavarian dialect","cultural grounding","LLM evaluation","multiple-choice benchmark","protocol sensitivity","regional culture","continued pretraining","multilingual evaluation"],"falsifier":"Recruit native Bavarian speakers from at least three dialect sub-regions and have them answer the 206 Bavarian items without access to the source documents; if their accuracy is near the open-weight mean of 45.9% or their agreement with the gold answers is low, BavGround is not measuring common Bavarian competence. A second check is to measure the Levenshtein distance and perceived naturalness of the Bavarian translations: if native raters judge them as near-Standard German, the dialect-gap claim loses its target.","tokens_in":24097,"feed_emoji":"🥨","tokens_out":6047,"duration_ms":58556,"temperature":0.7,"pith_summary":"The paper introduces BavGround, a multiple-choice benchmark of 206 questions about Bavarian culture, history, language, and daily life, each translated into English, German, and Bavarian, yielding 618 evaluation instances. Its central claim is that current open-weight instruction-tuned models are systematically worse on Bavarian items and on source-grounded regional questions than on general cultural facts: averaged over fifteen models, Bavarian accuracy is 45.9% versus 57.6% for German, and grounded questions fall to 46.7% versus 63.1% for general-knowledge questions. The paper further claims that the scoring protocol strongly affects measured competence, since switching from answer-letter scoring to length-normalized option-text scoring raises the open-weight mean from 53.0% to 59.3% and changes model rankings. The broader point is that regional cultural competence is not a single capability, and a single multiple-choice score can hide how much of a model's difficulty comes from dialect, local evidence, or answer-format behavior.","feed_headline":"LLMs lag on Bavarian dialect and local knowledge by 10-plus points","feed_subtitle":"A 618-question benchmark shows a Bavarian gap and that the scoring protocol changes the leaderboard.","key_machinery":"The carrying mechanism is BavGround itself: 206 source questions across eight cultural domains, split into general-knowledge (GEN) and source-grounded (GRD) items, manually translated into English, German, and Bavarian to form 618 parallel instances. Each item can be scored through several protocols—answer-letter log-probability, shuffled labels, length-normalized option-text likelihood, deterministic generation parsed back to labels, semantic embedding of generated answers, and hidden-state alignment—so the benchmark separates answer-content knowledge from label priors, option order, and format behavior. A third component is the GENBA-10B checkpoint series, 85 checkpoints from one continued-pretraining run, which lets the authors observe how cultural knowledge changes during training rather than treating it as a static property of a final model.","core_discovery":"On its own terms, the paper establishes that BavGround is a usable, protocol-sensitive instrument for measuring regional cultural grounding and dialect competence, and that current 7B–10B open-weight models have a real and consistent gap on it. Strong multilingual models cluster around 69% under standard letter scoring, but the open-weight mean is 53.0%, with Bavarian items consistently below English and German by 9.8–11.7 percentage points and grounded questions below general-knowledge items by 16.4 points. The same pattern appears in a closed-model reference, which scores 89.6% overall but still drops from 95.0% on general questions to 86.2% on grounded questions. A continued-pretraining analysis of GENBA-10B shows the same unevenness: option-text likelihood improves from 25.4% to 50.8% across checkpoints, while letter accuracy stays low and Bavarian and dialect items remain the weakest throughout. The authors' intended conclusion is that localized cultural evaluation needs to be both domain-aware and protocol-aware, because different scoring views reveal different failure modes.","pith_inferences":["Editorial inference: if Bavarian items were re-translated by native speakers from several Bavarian sub-regions, the measured Bavarian gap would likely widen, because the current single-translator version already biases toward Standard German.","Editorial inference: the label-prior diagnostics suggest that some leaderboard positions partly reflect a model's tendency to guess the benchmark's majority answer labels, so future benchmark designers should balance gold-label distributions or report label-prior-adjusted scores.","Editorial inference: the semantic-matching protocol depends on an external multilingual embedding model, so protocol comparisons under that view are partly about the embedding model's quality as well as the evaluated LLM's knowledge.","Editorial inference: applying the same protocol-aware framework to other region-versus-standard language pairs would test whether protocol sensitivity of this size is a general feature of cultural evaluation or specific to the Bavarian setup."],"forward_implications":["If the Bavarian gap is real, models deployed for Bavarian-speaking users need dialect-focused data and evaluation, not just more German-language exposure.","A single multiple-choice letter score is insufficient for cultural benchmarks, because changing the scoring protocol changes both absolute scores and model rankings.","General cultural knowledge does not transfer cleanly to source-grounded regional knowledge, so high performance on broadly accessible facts does not predict performance on regionally specific items.","Continued pretraining can improve answer-content likelihood while leaving dialect competence comparatively weak, meaning checkpoint diagnostics should separate content knowledge from answer-format behavior.","BavGround provides a template for localized evaluation below the nation-state level that could be extended to other regional and minority language communities."],"supporting_citations":[{"why":"Supplies the GENBA-10B model whose checkpoints and instruction-tuned version are evaluated.","marker":"Hoffmann et al. 2025"},{"why":"Motivates multi-prompt and protocol-aware evaluation by showing that prompt wording and answer ordering affect measured performance.","marker":"Mizrahi et al. 2024"},{"why":"Documents the mismatch between first-token probability rankings and generated text answers, grounding the option-text and generation protocols.","marker":"Wang et al. 2024"},{"why":"Shows ordering, labeling, and response-generation effects in survey-style LLM evaluations, supporting the protocol-sensitivity analysis.","marker":"Dominguez-Olmedo et al. 2024"},{"why":"Provides the localization framing and the critique of nation-state cultural proxies that BavGround operationalizes.","marker":"Zhou et al. 2025"},{"why":"Demonstrates that LLMs can disadvantage German dialect speakers, motivating the Bavarian dialect focus.","marker":"Bui et al. 2025"},{"why":"One of the primary regional sources used to construct source-grounded questions.","marker":"Liu 2021"},{"why":"Another primary regional source grounding the GRD question subset.","marker":"Merlan 2004"},{"why":"Provides the multilingual sentence-embedding model used for semantic matching of generated answers.","marker":"Reimers and Gurevych 2019"}],"fun_headline_variants":["Bavarian culture benchmark exposes LLM gaps","Scoring protocol flips rankings on Bavarian LLM benchmark","LLMs score 53% on Bavarian cultural knowledge test","Bavarian dialect and culture trip up open-weight LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Bavarian items genuinely represent Bavarian dialect and regional usage; the paper reports that all Bavarian translations came from a single native Chiemgau speaker who has lived outside Bavaria for years, so if those items are closer to Standard German than true Bavarian, the measured Bavarian gap understates the real dialect gap.","fun_headline_variants_meta":{"raw":{"variants":["Bavarian culture benchmark exposes LLM gaps","Scoring protocol flips rankings on Bavarian LLM benchmark","LLMs score 53% on Bavarian cultural knowledge test","Bavarian dialect and culture trip up open-weight LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000359,"raw_usage":{"total_tokens":1969,"prompt_tokens":997,"completion_tokens":972,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":904}},"tokens_in":613,"tokens_out":972,"duration_ms":8344,"temperature":1.0,"reasoning_tokens":904,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:06:35.970287+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit native Bavarian speakers from at least three dialect sub-regions and have them answer the 206 Bavarian items without access to the source documents; if their accuracy is near the open-weight mean of 45.9% or their agreement with the gold answers is low, BavGround is not measuring common Bavarian competence. A second check is to measure the Levenshtein distance and perceived naturalness of the Bavarian translations: if native raters judge them as near-Standard German, the dialect-gap claim loses its target.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the localization framing and the critique of nation-state cultural proxies that BavGround operationalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that LLMs can disadvantage German dialect speakers, motivating the Bavarian dialect focus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the primary regional sources used to construct source-grounded questions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Another primary regional source grounding the GRD question subset."}],"review_version":1}