{"id":"c4171a31-6c9c-4370-a62d-c1f51e8e3c05","arxiv_id":"2608.07763","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"PoVisLE, a manually constructed Polish cultural visual question-answering benchmark, shows current vision-language models reach at most 71.45% accuracy and perform worst on dialect and regionalism questions.","lead":"The authors built PoVisLE, a Polish-language benchmark with 1,117 images and 2,366 manually written visual question-answer pairs that tests whether AI models understand Polish culture, slang, dialects, and history. The best tested model, Qwen3.5-397B, reaches 71.45% accuracy, and removing the image cuts scores by 32 to 48 points, showing the test really requires looking at the picture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MCQ items can be answered without the question: question-free accuracy is 55.99% vs 65.94% full-input (Table 3), contradicting the abstract's claim that ablations rule out answer-option artefacts and weakening the benchmark's language-understanding validity.","rationale":"I read the paper in good faith. The dataset construction is careful, the annotation guidelines are detailed, and the authors report their ablations transparently, which is genuinely valuable. The central claim is that PoVisLE validly measures culturally grounded Polish multimodal understanding, and the abstract further claims that ablations confirm the benchmark cannot be solved from textual priors or answer-option artefacts alone. For this claim to hold, the question input must be necessary for the task. Table 3 shows that for multiple-choice items, question-free accuracy is 55.99% versus 65.94% with the full input, a drop of only 9.95 pp, and far above the circular-evaluation random baseline of 0.38%. Since MCQs are 36.1% of the test set, a substantial portion of the benchmark does not require processing the question. The paper's own Limitations acknowledge that closed formats may be solved by eliminating implausible options, yet the abstract's claim is worded more strongly than the data support. The reader's weakest_assumption about the single-answer requirement is a real concern, but it is less decisive here: the authors explicitly operationalize cultural correctness as shared prototypical knowledge, enforce unambiguity through structured QA and cross-validation, and the limitation is philosophical rather than demonstrated by a specific quantified failure. The MCQ artifact, by contrast, is quantified in the paper itself and directly undermines the language-understanding component of the benchmark. I therefore agree partially with the reader: the MCQ artifact is the more load-bearing issue, while the single-answer concern is a legitimate but secondary limitation. The reader's CONDITIONAL verdict already captures the need for revision, so I do not change it; the authors should be required to either demonstrate that the MCQ section is not largely solvable without the question or revise the benchmark/claim accordingly.","tokens_in":30141,"tokens_out":8867,"duration_ms":90144,"concrete_test":"Re-run the question-free condition on the 643 MCQ test items and compute, per item, the information gain of the question as full-input accuracy minus question-free accuracy, using Qwen3.5-397B-A17B Thinking or a small panel of Polish-speaking humans. If the fraction of MCQ items with near-zero gain (e.g., gain <= 2 pp) exceeds 25%, the MCQ section largely measures image-plus-option recognition rather than question comprehension; the authors should then convert or replace such items, or report a separate 'language-required' accuracy excluding them. This directly tests whether the abstract's 'cannot be solved from answer-option artefacts alone' claim survives on the MCQ subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is the question-free MCQ result in Table 3. For the best model (Qwen3.5-397B-A17B Thinking), removing the question while keeping the image and answer options yields 55.99% MCQ accuracy, only 9.95 pp below the full-input MCQ score (65.94%) and far above the theoretical random baseline of 0.38% under circular evaluation. MCQs constitute 36.1% of the test split, so a large part of the benchmark can be answered without reading the question; those items do not measure the 'language interpreted in interaction with visual context' claimed in the abstract. The paper's own Limitations concede that closed formats allow models to 'succeed by eliminating implausible options rather than demonstrating genuine understanding.' Consequently, the abstract's assertion that ablations 'confirm that the benchmark cannot be solved from textual priors or answer-option artefacts alone' is not supported by the reported data: the MCQ component is substantially solvable from image-plus-options, and the without-image condition still shows 29.08% MCQ accuracy, indicating residual textual-prior leakage. This weakens the construct validity of PoVisLE as a test of culturally grounded linguistic understanding, not merely its difficulty.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PoVisLE, a Polish vision-language benchmark with 1,117 images and 2,366 manually authored VQA pairs organized into a hierarchical taxonomy of Polish cultural and linguistic categories. The authors benchmark 16 open-weight and proprietary VLMs, reporting that Qwen3.5-397B-A17B (Thinking) achieves 71.45% macro accuracy, and run ablations that remove either the image or the question, as well as a prompt-language translation study. The central claims are that PoVisLE is the first monocultural Polish vision-language benchmark for cultural and linguistic competence and that the ablations confirm the benchmark cannot be solved from textual priors or answer-option artefacts alone.","tokens_in":30369,"tokens_out":7636,"duration_ms":66371,"significance":"If the construct validity of the benchmark holds, PoVisLE is a valuable resource for Polish cultural and linguistic multimodal evaluation. The dataset construction is unusually transparent: manual template-free annotation, cross-validation by a second annotator, an expert super-annotator, iterative LLM-based grounding checks, image-cluster bootstrap confidence intervals, detailed annotation guidelines in an appendix, and a public release. The strongest models score around 71%, leaving substantial headroom, and the per-category results identify dialect and regionalism questions as the weakest area. These properties make the resource potentially useful for the community. However, the stated interpretation of the ablation results is inconsistent with the reported data, and the answer-unambiguity requirement is in tension with the limitations discussion.","major_comments":[{"comment":"The abstract claims that the ablations 'confirm that the benchmark cannot be solved from textual priors or answer-option artefacts alone.' Table 3 shows that for the strongest model, Qwen3.5-397B-A17B (Thinking), removing the question while keeping the image and answer options yields 55.99% MCQ accuracy, only 9.95 percentage points below the full-input MCQ accuracy of 65.94%, and far above the random baseline of 0.38% under circular evaluation. Since MCQs constitute 36.1% of the test split, a large share of the benchmark can be answered without reading the question. Section 5 itself acknowledges that 'the image and answer options are often sufficient to identify the expected answer without access to the question.' This contradicts the abstract's claim. The 'without image' condition also shows residual textual-prior leakage: MCQ accuracy drops to 29.08%, still far above random. The claim as written is not supported by the reported data; the authors should either revise the abstract to describe the relative rather than absolute nature of the ablation effects, or restrict the claim to the open-ended and yes/no components, or report results on the subset of items that are genuinely not answerable without the question.","section":"Section 5, Table 3, Abstract"},{"comment":"The annotation guidelines require questions to allow 'a single clearly defined and non-debatable answer,' yet the Limitations section describes a 'perspective asymmetry' in which culturally canonical answers are privileged over 'alternative, semantically correct interpretations' and 'over-localization' where models 'may be penalized for producing correct but non-prototypical answers.' These two positions are in direct tension. If the latter effects occur in the test set, then the benchmark may measure alignment with annotator prototypes rather than objective cultural knowledge. The paper should provide evidence that the 'No ambiguity' rule was effective: for example, inter-annotator agreement on the answer key, the number of items revised or removed due to ambiguity during cross-validation, and a discussion of how many test items could admit defensible alternative answers. Without such evidence, the construct validity of PoVisLE as a measure of culturally grounded understanding remains uncertain.","section":"Section 2.3 and Limitations"},{"comment":"The LLM-based grounding check flagged MCQs as 'too easy' only when models achieved a success rate of approximately above 60% in the question-free setting. However, the final evaluation includes Qwen3.5-397B-A17B (Thinking), which achieves 55.99% question-free MCQ accuracy, just below this threshold. If the 11 LLMs used in validation were weaker than the evaluated models, the threshold would have failed to remove items that the strongest models can answer from the image and options alone. The authors should report the question-free MCQ accuracy of the validation LLMs, justify the 60% threshold, or use a model-relative or item-level criterion (e.g., flag items answerable by any model above chance). This is load-bearing because it directly affects the benchmark's guarantee that questions require the linguistic input.","section":"Section 2.4"}],"minor_comments":[{"comment":"The title contains a typo: 'Jako Takoor' should be 'Jako Tako'.","section":"Title"},{"comment":"The subcategory 'Colloqual speech and slang' is misspelled; it should be 'Colloquial speech and slang'.","section":"Table 1"},{"comment":"The model name 'LLaVA' is frequently rendered as 'LLaV A' with a stray space; this should be corrected globally.","section":"Section 4.2 and throughout"},{"comment":"The sentence 'Figure 2 present model performance' should be 'Figure 2 presents model performance'; the same issue appears in the caption of Figure 2.","section":"Section 5 and Figure 2"},{"comment":"The paragraph beginning 'To further illustrate the distribution of complexity...' is duplicated verbatim; one copy should be removed.","section":"Appendix B.2"},{"comment":"Section 2.3 states that a controlled subset was annotated by 12 auxiliary annotators, while Appendix D.5 and Table 6 report 13 auxiliary annotators; the numbers should be reconciled.","section":"Section 2.3 vs Appendix D.5"},{"comment":"The main evaluation metric is described as 'macro-averaged accuracy over question type,' but the main text does not explicitly state that overall accuracy is the unweighted mean of the three question-type accuracies (as implied by Table 3). This should be stated explicitly, with a justification for equal weighting given the different sample sizes across question types.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper evaluates LLaVA-Bielik and LLaVA-PLLuM, models developed in part by the present authors (Statkiewicz et al., 2026), and uses StyloMetrix (Okulska et al., 2023), which includes an author of this paper. These overlaps are not disclosed in the text; the authors should add an explicit statement of competing interests or self-citation disclosure. This does not affect my technical assessment but is relevant for the editor. The manuscript is otherwise thorough in its methodology and data release, and the identified issues appear addressable within a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing. PoVisLE is a useful new resource: a Polish monocultural vision-language benchmark, manually annotated, with a documented construction process and a real attempt to test culturally grounded understanding rather than surface recognition. The dataset, the taxonomy, and the release are solid. But the abstract's claim that ablations 'confirm that the benchmark cannot be solved from textual priors or answer-option artefacts alone' is not supported by Table 3. For the best model, removing the question while keeping image and options leaves MCQ accuracy at 55.99%, only about 10 points below the full-input 65.94%. A third of the test split is MCQ, so a large chunk of the benchmark does not require reading the question. The authors actually acknowledge this in the limitations, but the abstract overstates the ablation result.\n\nCredit where it is due: construction is unusually careful. Two-stage image collection, cross-validation, LLM-based grounding checks, image-cluster bootstrap intervals, detailed annotation guidelines, and demographic reporting. The ~29% annotator-provided images are a genuine plus, because they are out-of-distribution for web-trained models. The stylometric analysis gives real evidence that questions are not template-driven. The limitations section is honest about the single-answer trade-off and over-localization.\n\nThe question-free MCQ result is the main substantive issue. It weakens construct validity for the closed-format portion, because those items measure image-and-option reasoning rather than language interpreted in interaction with visual context. The fix is straightforward: report item-level leakage, revise or re-validate high-leakage items, or clearly separate MCQ results from the overall claim. The single-answer assumption is a known limitation, documented, but 'non-debatable' sits uneasily with the acknowledged perspective asymmetry. Minor point: the validation split is not sampled from the same distribution as the test split, which the authors admit; fine for illustration, not for tuning.\n\nWho is this for? Anyone building or evaluating multilingual/cultural VLMs, especially for Slavic languages. It deserves a serious refereeing round; I would conditionally accept after the MCQ issue is addressed. I would bring it to the reading group and would probably cite the dataset, with a caveat about the MCQ artifact.","headline":"A genuinely useful and transparently built Polish cultural VLM benchmark, but its own Table 3 undercuts the abstract's claim about question-free MCQ performance, so it needs revision rather than rejection.","tokens_in":30909,"tokens_out":3497,"would_cite":true,"duration_ms":35309,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces PoVisLE, the first monocultural Polish vision-language benchmark, and claims current VLMs are far from fluent in Polish cultural understanding.","keywords":["PoVisLE","vision-language benchmark","Polish cultural competence","visual question answering","dialect and regionalism","visual grounding","multilingual evaluation","cultural VQA"],"falsifier":"Take a random sample of PoVisLE test questions and have independent Polish annotators from different regions provide answers without seeing the gold standard; if agreement on the single expected answer falls well below the level needed for deterministic scoring, the no-ambiguity premise fails. A simpler check: if any future model scores near full accuracy with images removed, the visual-grounding claim is refuted.","tokens_in":29957,"feed_emoji":"🇵🇱","tokens_out":4731,"duration_ms":43254,"temperature":0.7,"pith_summary":"The paper sets out to show that Polish cultural and linguistic competence in vision-language models can be measured, and that current models are far from fluent in it. It introduces PoVisLE, a benchmark of 1,117 images and 2,366 hand-written questions in Polish that require connecting what is visible in a picture with shared cultural knowledge, dialect, slang, history, or wordplay. On this benchmark the strongest of 16 evaluated models reaches 71.45 percent accuracy, and questions about dialects and regionalisms are the hardest category. The paper argues that because removing the image cuts accuracy sharply, the benchmark really tests grounded multimodal understanding rather than text memorization.","feed_headline":"Polish vision benchmark stumps top AI models at 71 percent","feed_subtitle":"2,366 hand-written Polish questions probe dialect, slang, and visual wordplay; top models score only 71 percent.","key_machinery":"The load-bearing artifact is the PoVisLE dataset itself, organized around a hierarchical taxonomy with seven main categories: Art and entertainment, Language, Geography and nature, History and society, Culture and tradition, Image understanding, and Visual reasoning. Construction rests on three annotation rules: questions must be unambiguous with a single correct answer, must rely on visual content, and must carry strict output-format instructions. The evaluation harness adds circular option rotation for multiple-choice items, exact-match scoring with Polish diacritics and capitalization, and two ablations, removing the image and removing the question, plus Polish, English, and German prompt translation.","core_discovery":"PoVisLE is presented as the first monocultural vision-language benchmark built specifically for Polish cultural and linguistic competence. Its 2,366 question-answer pairs are manually written, template-free, and designed so that each question requires the image: even language-focused items such as the riddle that asks how a Pole answers \"how are you?\" by rhyming with the pseudonym \"Taco Hemingway\" (the answer being \"jako tako\") depend on seeing who is in the picture. The paper reports that the best model, Qwen3.5-397B-A17B in its thinking mode, scores 71.45 percent macro accuracy, that dialect and regionalism questions are the weakest category, and that removing the image lowers accuracy by 32 to 48 percentage points while removing the question leaves open-ended accuracy at zero. These results are used to claim that PoVisLE provides a valid, grounded evaluation of culturally situated Polish multimodal understanding.","pith_inferences":["If the benchmark's difficulty holds up, it could be used to track progress of Polish-specific vision-language models, which currently sit below 37 percent accuracy.","The single-answer rule is a deliberate trade-off; a future version with flexible semantic matching might reveal that models are more culturally fluent than exact-match scoring admits.","The monocultural design suggests a template for other mid-resource languages: manual, template-free questions anchored in private images could expose cultural gaps that broad multicultural benchmarks miss.","Items like the \"jako tako\" riddle indicate the benchmark tests pragmatic wordplay, so passing it requires more than object recognition or encyclopedic knowledge."],"forward_implications":["If correct, the benchmark gives a controlled yardstick for Polish cultural VQA, with a held-out test split of 1,960 pairs available for future model comparisons.","Current best models leave roughly 29 percentage points on the table, so culturally grounded Polish understanding is far from solved.","Dialects and regionalisms are the hardest category, meaning intra-language variation is a specific weakness of current VLMs.","The large accuracy drop when the image is removed confirms that high scores cannot be achieved from textual priors alone on this benchmark.","Prompt-language results suggest that stronger models exploit the original Polish formulation, while weaker models can gain up to 8.29 points from English prompts."],"supporting_citations":[{"why":"Supplies the PLCC taxonomy and text-only Polish cultural baseline that PoVisLE extends to the multimodal setting.","marker":"(Dadas et al., 2025)"},{"why":"Provides CLIP embeddings used to retrieve visually similar candidate images for question-guided augmentation.","marker":"(Radford et al., 2021)"},{"why":"Serves as a precedent for using annotator-provided private images in a culture-specific benchmark.","marker":"(Hsieh et al., 2026)"},{"why":"Motivates the question-free ablation setting, showing that LLMs can exploit multiple-choice options alone.","marker":"(Balepur et al., 2024)"},{"why":"Justifies circular evaluation for multiple-choice items to counter option-position bias.","marker":"(Zheng et al., 2024)"},{"why":"Documents the grounding gap that motivates requiring visual reference in every question.","marker":"(Ashok et al., 2025)"},{"why":"Provides reVISION, the existing Polish multimodal exam benchmark that PoVisLE contrasts with as non-culturally-focused.","marker":"(Ciesiółka and Graliński, 2025)"},{"why":"Supports the critique that prior cultural benchmarks are template-driven and surface-level.","marker":"(Yadav et al., 2025)"},{"why":"Motivates the prompt-language analysis of how language choice affects cultural performance.","marker":"(Shen et al., 2024)"}],"fun_headline_variants":["Polish culture benchmark stumps AI: top score 71%","AI struggles with Polish jokes, slang in new vision test","Vision models hit 71% on Polish cultural benchmark, need help","New Polish benchmark exposes AI cultural blind spots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each question has exactly one correct, non-debatable answer that all Poles would agree on; the paper's own limitations section concedes cases where culturally canonical answers are privileged over correct alternatives.","fun_headline_variants_meta":{"raw":{"variants":["Polish culture benchmark stumps AI: top score 71%","AI struggles with Polish jokes, slang in new vision test","Vision models hit 71% on Polish cultural benchmark, need help","New Polish benchmark exposes AI cultural blind spots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000769,"raw_usage":{"total_tokens":3380,"prompt_tokens":890,"completion_tokens":2490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2423}},"tokens_in":506,"tokens_out":2490,"duration_ms":14385,"temperature":1.0,"reasoning_tokens":2423,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:23:31.241936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of PoVisLE test questions and have independent Polish annotators from different regions provide answers without seeing the gold standard; if agreement on the single expected answer falls well below the level needed for deterministic scoring, the no-ambiguity premise fails. A simpler check: if any future model scores near full accuracy with images removed, the visual-grounding claim is refuted.","supporting_citations":[{"cited_title":"Evaluating Polish linguistic and cultural competency in large language models","cited_arxiv_id":"2503.00995","evidence_quote":"Supplies the PLCC taxonomy and text-only Polish cultural baseline that PoVisLE extends to the multimodal setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides CLIP embeddings used to retrieve visually similar candidate images for question-guided augmentation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as a precedent for using annotator-provided private images in a culture-specific benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies circular evaluation for multiple-choice items to counter option-position bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the grounding gap that motivates requiring visual reference in every question."},{"cited_title":"Evaluation of Cultural Competence of Vision-Language Models","cited_arxiv_id":"2505.22793","evidence_quote":"Supports the critique that prior cultural benchmarks are template-driven and surface-level."}],"review_version":2}