{"id":"78d9255f-942b-48b0-a743-46e3e3431037","arxiv_id":"2508.19944","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"KRETA, a 2,577-item Korean text-rich VQA benchmark, shows vision-language models recognize Korean text well but lag in multi-step reasoning, especially in open-source models.","lead":"A new benchmark, KRETA, tests AI systems on reading and reasoning about text in Korean images, with 2,577 questions across signs, menus, charts, and documents. It finds that current vision-language models can read Korean text well but fall sharply on multi-step reasoning, especially open-source models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark validity rests on LLM-generated QA without a human baseline; a human agreement study is needed.","rationale":"The reader's weakest assumption identifies the benchmark's validity as depending on the correctness and unambiguity of LLM-generated QA pairs, with no human baseline or inter-annotator agreement. I agree. The paper is a useful contribution and the pipeline is transparently described, but the empirical claims about System 2 deficiencies in open-source models are only as strong as the benchmark's construct validity. A human validation study is the natural, concrete way to settle this. The reader's CONDITIONAL verdict is therefore appropriate, and my read does not change it.","tokens_in":14805,"tokens_out":2274,"duration_ms":27231,"concrete_test":"Recruit three native Korean speakers to independently answer a random subset of 200 KRETA questions (100 System 1, 100 System 2) with the corresponding images, and have each annotator flag any question that is ambiguous or answerable without the image. Compute human accuracy and Fleiss' kappa. If human System 2 accuracy is below ~85%, or kappa is below ~0.7, or if two or more annotators flag the same question, the benchmark's validity for measuring reasoning is not established and the reported model gaps need recalibration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that KRETA is a valid measurement instrument depends on the assumption that its 2,577 QA pairs are correct, unambiguous, and require image context. Section 3.3 describes QA generation by GPT-4o-mini, Gemini-2.0-flash, and o1-mini, followed by LLM-judge selection and 'human refinement,' but no human accuracy baseline, inter-annotator agreement, or per-item rejection counts are reported. Without human performance data, we cannot tell whether the large System 1/System 2 gap reflects genuine model reasoning limitations or artifacts of the generation pipeline—e.g., questions answerable from text alone, ambiguous hard negatives, or distractors that are culturally obscure in unintended ways. The fact that the generating/evaluating models (GPT-4o-mini, Gemini-2.0-flash) are also in the evaluation pool (Table 2) further risks inflating their scores if they systematically favor their own question style. This is the weakest link in the paper's empirical argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces KRETA, a Korean text-rich VQA benchmark containing 2,577 multiple-choice QA pairs over native Korean images, organized into 15 domains and 26 image types and split into System 1 (basic text recognition, 1,426 items) and System 2 (advanced reasoning, 1,151 items). QA items are produced semi-automatically: two VLMs decompose images into structured captions, LLMs generate candidates, LLM judges score them on up to seven metrics, hard negatives are added, and human refinement is applied. The authors evaluate 18 closed- and open-source VLMs, reporting that closed models reach roughly 80--85% overall accuracy while open models lag, and that System 2 causes large drops (e.g., Qwen2.5-VL-7B from 94.5% to 36.1%). They also analyze results by domain, image type, model size, and chain-of-thought prompting, and release the code, data, prompts, and a leaderboard.","tokens_in":15020,"tokens_out":5083,"duration_ms":55486,"significance":"KRETA addresses a real gap: it is among the largest native Korean text-rich VQA datasets, with a useful dual-level reasoning split and a domain taxonomy tied to KSIC. The CSAT subsets are grounded in official examination materials, providing a valuable externally anchored component. A further strength is the release of the pipeline, prompts, and evaluation code, which supports reproducibility and adaptation to other low-resource languages. The empirical finding that open-source models collapse on System 2 items is interesting and actionable. However, the paper's central claim is that KRETA constitutes a valid measurement instrument, and the current evidence for that claim is incomplete: the benchmark's correctness and labeling rest on unquantified LLM generation/judging plus an unspecified human refinement step, and the generator/judge models also appear in the evaluation pool. These issues must be addressed before the numerical conclusions can be fully trusted.","major_comments":[{"comment":"The validity of KRETA as a measurement instrument is not established. Step 2 uses GPT-4o-mini, Gemini-2.0-flash, and o1-mini to propose QA pairs; Step 3 has GPT-4o-mini and Gemini-2.0-flash score the candidates; the human refinement step is described but no quantitative results are reported. There is no human accuracy baseline, no inter-annotator agreement, no per-item rejection rate, and no check of text-only answerability or ambiguity. Consequently, the reported System 1/System 2 gap (e.g., Qwen2.5-VL-7B 94.5 vs. 36.1) could in part be an artifact of the generation pipeline rather than a measurement of VLM reasoning. Moreover, the same models that generate and judge the data (GPT-4o-mini, Gemini-2.0-flash) are later evaluated on it in Table 2, risking inflated scores due to familiarity with their own item style. Please add human validation statistics (accuracy on a sample, agreement, e","section":null},{"comment":"Several cross-domain and image-type conclusions are drawn from very small samples without uncertainty quantification. For example, CSAT History has 60 items (Table 3), and several image types in Table 4 are near the 50-item cutoff; a 10--15 percentage point gap is close to the binomial standard error for such cells. The statement that GPT-4o 'excels' in CSAT History (93.3%) versus other models should be supported by confidence intervals or significance tests. The text should temper conclusions about specific domains and image types where the sample sizes are too small to support reliable ranking.","section":null},{"comment":"The System 1/System 2 construct validity is under-verified. The labels are assigned by the same generation prompts and LLM judges rather than by an independent protocol, and no evidence is provided that System 2 items actually require multi-step inference or that System 1 items are purely recognition-based. The Limitation section candidly notes that System 2 conflates several reasoning types, but that admission does not substitute for a validation study. Please report human agreement on System 1 vs. System 2 labels, or an analysis of the number of inference steps and answerability-without-image for a random sample of items.","section":null}],"minor_comments":[{"comment":"The reference 'Kim et al., 2025' lists an incomplete arXiv identifier (arXiv:2505.XXXXX). Please update before publication.","section":null},{"comment":"The chain-of-thought results are described with point differences (e.g., +3.7, -7.7) but no error bars or per-model tables are provided. Please include confidence intervals or a table with exact scores so the reader can assess the reliability of these differences.","section":null},{"comment":"The table as rendered in the manuscript has garbled numeric spacing (e.g., the GPT-4o row), which makes verification difficult. Please ensure the final typeset version is readable and aligned.","section":null},{"comment":"The comparison between CSAT Science (478 items) and CSAT History (60 items) should be conditioned on the large sample-size difference; statements such as 'GPT-4o excels in CSAT History' should be softened or qualified.","section":null}],"recommendation":"major_revision","confidential_remarks":"The benchmark is a useful contribution and the paper is generally well organized, but the main weakness is missing validity evidence for the LLM-generated QA. The concern raised by the reader about circularity is real: the generator/judge models are also in the evaluation pool, and the human refinement step is not quantified. I would ask the authors to run a human validation study (e.g., on 200--300 sampled items) and to report agreement, rejection rates, and a human accuracy baseline, as well as a version of Table 2 excluding GPT-4o-mini and Gemini-2.0-flash. These additions are feasible within the scope of the manuscript and would materially strengthen the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper delivers a new benchmark that fills a real gap. KRETA gives you 2,577 native Korean text-rich VQA samples across 15 domains and 26 image types, with a clean System 1 vs System 2 division. That is a solid contribution. Prior Korean sets are small or translated; KRETA is bigger than KOFFVQA and MTVQA-ko combined and built from original Korean images. The CSAT-derived items are a nice grounded addition. Releasing prompts, code, and dataset is the right move.\n\nThe empirical work is straightforward and believable up to a point: all models drop on System 2, and open-source models drop a lot. The 15.1% on CSAT Science for Qwen2.5-VL-7B is striking.\n\nThe soft spot is the one the reader flagged: the benchmark's validity is under-verified. The pipeline generates QA pairs with GPT-4o-mini, Gemini-2.0-flash, and o1-mini; selects them via LLM judges from the same families; and then those same models appear in the evaluation table. That is a circularity you can smell, even if the numbers don't show obvious favoritism (GPT-4o-mini scores 73.3 overall, below Gemini). More important: there is no human performance baseline, no inter-annotator agreement, no per-item rejection counts from the 'human refinement' step. With 2,577 items, a small-scale human pass (even 200 items) would tell you whether the System 1/System 2 gap is real or partly an artifact of how questions were written. The authors are transparent about some limits (single-image, MC-only, System 2 conflation) but don't mention this one.\n\nAlso be careful with the small domain cells: Hist has 60 samples, Med 90. Differences of 20 points there are noise. The paper shouldn't draw strong conclusions from those.\n\nNet: this is a promising resource that deserves peer review. I'd recommend accept-with-revision, with the human validation as the key missing piece. If the authors add a human agreement/baseline study and release the per-item data, the benchmark becomes much easier to trust. For a reader working on multilingual or Korean VQA, it's worth a look.","headline":"KRETA is a genuinely useful new Korean text-rich VQA benchmark with a sensible System 1/System 2 split; the main open question is whether the LLM-built questions are valid without a human performance baseline.","tokens_in":15522,"tokens_out":2343,"would_cite":true,"duration_ms":25827,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"KRETA, a 2,577-item native Korean VQA benchmark, separates reading from reasoning and finds open vision-language models collapse on reasoning.","keywords":["Korean text-rich VQA","vision-language models","System 1 / System 2 reasoning","multilingual benchmark","multiple-choice QA","semi-automated data pipeline","low-resource language evaluation"],"falsifier":"Take 200 randomly selected KRETA questions, give the image and options to native Korean speakers without exposing the gold answer, and compare their choices with the released key; also log how often a System 2 question is answerable from text alone. Low agreement, or a sizable share of text-only-answerable items, would show the measured reasoning gap is an artifact of question construction rather than model reasoning.","tokens_in":1769,"feed_emoji":"🇰🇷","tokens_out":2660,"duration_ms":93293,"temperature":0.7,"pith_summary":"KRETA gives Korean text-rich visual question answering a proper measuring instrument. It consists of 2,577 questions built from native Korean images and text, split into System 1 items that require reading text and System 2 items that require multi-step reasoning, domain knowledge, or visual math. The paper finds that current vision-language models, especially open-source ones, handle basic Korean text recognition well but collapse on System 2 reasoning; Qwen2.5-VL-7B falls from 94.5% to 36.1%. The paper attributes this to thin Korean-contextual and domain-specific training data. If the benchmark is valid, it gives researchers a way to separate reading ability from reasoning ability in Korean, and its semi-automated generation pipeline offers a template for similar benchmarks in other low-resource languages.","feed_headline":"Open vision models read Korean but fail its reasoning test","feed_subtitle":"KRETA's 2,577 native-Korean questions split reading from reasoning; top open models drop from 94.5% to 36.1%.","key_machinery":"The central object is KRETA itself: 2,577 multiple-choice QA pairs grounded in native Korean images, organized by a dual-level System 1 versus System 2 split and by a 15-domain, 26-image-type taxonomy. The argument is carried by a semi-automated pipeline: two VLMs independently decompose each image into structured captions capturing layout, text, and text-visual links; LLMs generate QA candidates; two VLMs score candidates on a seven-metric protocol; an LLM synthesizes hard distractor options; and human review refines the survivors. The System 1/System 2 distinction makes the benchmark diagnostic rather than just a leaderboard, because it lets the authors attribute performance drops to reaso","core_discovery":"The paper claims that KRETA, built only from native Korean imagery and questions rather than translated English content, is the first large-scale Korean text-rich VQA benchmark capable of evaluating both basic text recognition and advanced reasoning. Evaluation results show a 'reasoning bottleneck': closed models average about 70.5% on System 2 questions, while open-source models score as low as 29-49%. The largest drops occur in culturally rich domains like CSAT History and in cluttered image types like banners and store signs. The paper argues this gap is not mainly an OCR failure, because System 1 recognition accuracy is high, but rather a deficiency in Korean-contextual, domain-specific","pith_inferences":["Because the pipeline runs on structured captions rather than raw OCR, it could plausibly be lifted to other undersourced languages, though the paper only gestures at this transfer.","The System 2 category bundles several distinct reasoning types - multi-step deduction, math, domain knowledge - so fine-grained labels would reveal whether open models fail uniformly or mostly on knowledge-heavy items; the paper's own limitations section concedes this.","The prompt-language effect (Korean CoT boosting Qwen2.5-VL-7B while English CoT hurts it) suggests cross-lingual benchmark comparisons should control evaluation-language choice, otherwise measured gaps may partly reflect prompt language rather than model capability.","If the benchmark is valid, the 40-plus-point closed-open System 2 gaps on culturally dense domains imply that simply scaling a multilingual model without preserving low-resource language data share can dilute Korean-specific ability, a risk the paper raises in its model-size analysis."],"forward_implications":["KRETA can serve as a reusable testbed for Korean text-rich VQA across 15 domains and 26 image types, enabling targeted diagnosis of where a model's reading and reasoning skills diverge.","Because model rankings change sharply between System 1 and System 2, evaluations should report recognition and reasoning separately rather than as one aggregate score.","Open-source VLMs need targeted Korean-contextual and domain-specific training, especially for CSAT Science and History and for complex real-world layouts like banners and store signs.","Chain-of-Thought prompting is not universally helpful: it improves capable closed models but degrades small open models, so prompt effects must be part of deployment decisions.","The semi-automated pipeline can be adapted to other low-resource languages, provided prompts are written natively in the target language rather than translated from English."],"supporting_citations":[{"why":"Supplies the System 1/System 2 cognitive distinction that structures the benchmark's two reasoning levels.","marker":"Kahneman, 2011"},{"why":"TextVQA establishes the text-rich VQA task lineage that KRETA extends to Korean.","marker":"Singh et al., 2019"},{"why":"MTVQA provides the multilingual text-centric benchmark whose Korean split KRETA compares against and exceeds in scale.","marker":"Tang et al., 2024b"},{"why":"KOFFVQA is the prior small Korean VQA benchmark that KRETA positions itself against on scale, form, and text-centric focus.","marker":"Kim and Jung, 2025"},{"why":"The Korean VQA suite including K-MMB and VARCO-VISION supplies the comparison set showing prior Korean benchmarks used translated English images or were not text-centric.","marker":"Ju et al., 2024"},{"why":"MMMU-Pro provides the multiple-choice prompt format and the Chain-of-Thought evaluation setup used in KRETA's experiments.","marker":"Yue et al., 2024c"},{"why":"VLMEvalKit is the evaluation harness used to score the closed and open-source models on KRETA.","marker":"Duan et al., 2024"},{"why":"Qwen2.5-VL is a primary evaluated open-source model whose System 1-to-System 2 drop illustrates the claimed reasoning bottleneck.","marker":"Wang et al., 2024"},{"why":"The Korean Standard Industrial Classification provides the domain taxonomy that anchors KRETA's real-world industrial relevance.","marker":"Statistics Korea, 2024"}],"fun_headline_variants":["Korean VQA: Models read text but can't reason","New Korean benchmark exposes VLM reasoning gap","KRETA: Why open VLMs fail Korean reasoning","Reading vs reasoning: Open VLMs lag on Korean","Reasoning bottleneck in Korean text-rich VQA"],"cache_read_input_tokens":17408,"weakest_assumption_plain":"The benchmark's validity rests on the assumption that the automatically generated QA pairs, chosen by model judges and only lightly refined by humans, are correct, unambiguous, and not biased toward the generating models; the paper provides no human performance baseline or inter-annotator agreement to back this assumption.","fun_headline_variants_meta":{"raw":{"variants":["Korean VQA: Models read text but can't reason","New Korean benchmark exposes VLM reasoning gap","KRETA: Why open VLMs fail Korean reasoning","Reading vs reasoning: Open VLMs lag on Korean","Reasoning bottleneck in Korean text-rich VQA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1206,"prompt_tokens":761,"completion_tokens":445,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":371}},"tokens_in":505,"tokens_out":445,"duration_ms":4869,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:19:48.551553+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take 200 randomly selected KRETA questions, give the image and options to native Korean speakers without exposing the gold answer, and compare their choices with the released key; also log how often a System 2 question is answerable from text alone. Low agreement, or a sizable share of text-only-answerable items, would show the measured reasoning gap is an artifact of question construction rather than model reasoning.","supporting_citations":[],"review_version":1}