{"id":"e79147aa-dfed-41bb-a604-58a9e585c0b9","arxiv_id":"2508.21294","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An updated BLUEX benchmark with 1,422 questions and GPT-4o-generated captions that make image-based questions usable by text-only LLMs.","lead":"This paper expands the BLUEX benchmark, a set of Brazilian university entrance exam questions, with exams from 2024 and 2025 and machine-generated captions for the roughly 40% of questions that include images. It then measures how well a range of language models answer the questions using captions instead of images.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Context captions are generated with answer choices; answer leakage may inflate caption-condition accuracy and undermine the 'usable questions' claim.","rationale":"The reader's weakest assumption is that GPT-4o captions are faithful and sufficient substitutes for images. My concern is more specific and more severe: context captions are generated with access to the answer choices, so even a faithful caption can be unintentionally contaminated by the generator's knowledge of the correct answer. This does not contradict the reader's verdict—it reinforces the conditional status—but it identifies a concrete mechanism that the reader's broader fidelity concern only partially covers. The proposed test would settle whether leakage actually occurs. If leakage is found, the paper's strongest claims about caption usefulness and model reasoning would need substantial revision; if not, the conditional concerns remain as the reader stated. I recommend keeping the verdict CONDITIONAL (i.e., UNCHANGED relative to the reader's assessment), because the paper needs this validation before its conclusions can be fully trusted.","tokens_in":8398,"tokens_out":3617,"duration_ms":38725,"concrete_test":"Run a controlled caption-generation audit on a random sample of 50–100 image-based questions. For each question, generate context captions under three conditions: (a) original alternatives, (b) alternatives with the correct option replaced by a randomly selected wrong option, and (c) no alternatives at all. Then ask an independent evaluator (or a separate LLM) to determine whether each caption reveals or favors a specific answer choice. More directly, measure a text-only model's accuracy on questions using captions from condition (b) versus condition (a). If accuracy tracks the inserted answer (e.g., accuracy drops when the correct option is swapped out), answer leakage is confirmed. If accuracy is unchanged and captions do not systematically mention answer-specific content, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that captioning produces 1,422 usable questions and that context captions are an effective way to make image-based questions accessible to text-only models. Section 1 states that context captions are generated 'with access to the associated question and alternatives'; Section 3.2 says GPT-4o is 'provided with both the image and its associated question.' The correct answer is among the alternatives, so the caption generator can see it. This creates a direct leakage path: the caption may encode information that favors the correct option, either by describing the image in a way that highlights the correct alternative or by implicitly echoing answer-specific details. Table 2's large accuracy gains in the 'Context Captions' column over 'No Caption' may therefore reflect answer leakage rather than genuine visual understanding. The comparison between context and blind captions is also confounded: context captions have access to the answer set, blind captions do not. The paper's observation that context captions are shorter yet perform comparably could be explained by leakage, not by the benign 'focusing' hypothesis. If leakage is present, the '1,422 usable questions' claim is still technically true for blind captions, but the headline improvement and the conclusion that models 'leverage visual context through captions' are substantially weakened. Because the benchmark is explicitly positioned for contamination studies, answer leakage in a large fraction of the dataset would be a serious validity threat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents BLUEX Revisited, an updated version of the Portuguese university-entrance-exam benchmark BLUEX. It adds 2024 and 2025 exams and automatically generated GPT-4o captions for image-based questions under two conditions: blind captions (image only) and context captions (image plus the associated question). The authors evaluate 18 commercial and open-source models on text-only, blind-caption, and context-caption versions of the dataset, reporting large accuracy gains when captions are provided and comparing model performance with admission cutoffs at Unicamp and USP. The central claims are that captioning increases accessibility to text-only models by more than 40%, yielding 1,422 usable questions, and that context captions, despite being shorter, perform comparably to blind captions.","tokens_in":8704,"tokens_out":5810,"duration_ms":56298,"significance":"If the claims hold, the expanded dataset is a useful resource for Portuguese-language and multilingual model evaluation, especially for text-only models and for contamination-related studies. The paper's explicit strengths are its public dataset and code release, its transparent construction pipeline, and its broad model coverage. However, the central quantitative contribution depends on two assumptions that are not yet validated: that GPT-4o captions faithfully preserve the visual information needed to answer, and that context-caption results are not inflated by answer leakage when the caption generator sees the multiple-choice alternatives. The contamination claim in Section 4.2 is also asserted without probing. These issues affect the headline 'usable questions' and 'caption effectiveness' claims, so the paper is best viewed as a promising resource description with preliminary evaluations rather than a conclusive benchmark validation.","major_comments":[{"comment":"Context captions are generated by GPT-4o with access to 'the associated question' (Section 3.2). Since the question includes the multiple-choice alternatives, the correct answer is visible to the caption generator. The Context Captions condition is therefore not a clean measure of visual understanding: GPT-4o may encode cues that favor the correct alternative, and the similarity between Context and Blind Captions in Table 2 may reflect leakage compensating for shorter captions rather than the 'focusing' hypothesis stated in Section 4.1. Please add a control where context captions are generated without the correct answer (e.g., alternatives removed or replaced) and report whether the Context Captions column changes. This is load-bearing for the conclusion that context-aware captions are effective.","section":"Section 3.2; Table 2; Section 4.1"},{"comment":"Caption adequacy is never validated. The paper reports caption length distributions (Figure 3) but no human rating, no alternative captioner comparison, and no check that captions contain the information needed to answer. The entire 'usable questions' claim and all caption-condition accuracies in Table 2 assume GPT-4o captions are faithful and sufficient. A simple fix is to sample, say, 100 image questions, have human raters mark whether each caption preserves the visual information needed to answer, and report agreement; alternatively, compare model accuracy with captions vs original images. Without this, the dataset resource may contain systematic caption errors.","section":"Section 3.2; Figure 3; Table 2"},{"comment":"The paper motivates the update by contamination studies and asserts that high 2025 scores indicate genuine reasoning because 2025 exams are recent. This is asserted, not probed. No exact-match or near-duplicate search, no temporal leakage analysis, and no comparison with held-out questions are reported. Since contamination is a stated central motivation, either add a contamination probe or soften the claim. As written, the conclusion in Section 5 that performance stems from reasoning rather than memorization is unsupported.","section":"Section 4.2; Section 5"},{"comment":"All accuracies are single-run point estimates without error bars, confidence intervals, or significance tests. Many comparisons are within 1-2 points (e.g., Sabia-3 0.750 vs GPT-4o 0.754 on all questions). The model rankings and the claim that captioning improves accuracy by 'at least 10 accuracy points' in Figure 4 need repeated runs with different random seeds or bootstrap intervals. This is important because the headline contribution is quantitative.","section":"Table 2; Section 4"}],"minor_comments":[{"comment":"The abstract says 'more than 40%' and 'more than doubling the number' but the manuscript never states the original BLUEX total or the exact percentage increase. Please reconcile these numbers with Table 1's counts (812 without images, 610 with images, 1,422 total).","section":"Abstract; Table 1"},{"comment":"The caption of Figure 3 refers to 'GPT-4V' while the text consistently uses GPT-4o. Please make the terminology consistent.","section":"Figure 3"},{"comment":"The header 'Comercial models' is a typo for 'Commercial models'. Also, the table is very dense; consider separating the three partitions into sub-tables or adding clearer column separators.","section":"Table 2"},{"comment":"The phrase 'In parallel, Importantly,' is ungrammatical and should be rephrased.","section":"Section 2.2"},{"comment":"The y-axis label 'Accuracy gain' should specify 'percentage points' and clarify that the gain is computed relative to the no-caption condition. The x-axis label 'Model Size' is unclear because models are ordered by size but the axis is categorical.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The authors are affiliated with Maritaca AI, whose models Sabia-3 and Sabiazinho-3 are evaluated in the paper; this should be declared explicitly. In addition, GPT-4o is both the caption generator and one of the evaluated models, creating a same-model bootstrap that should be acknowledged. I do not view these as disqualifying, but they should be transparent to readers. The paper is more of a resource paper than a claim of new model findings; the resource itself is potentially valuable if the caption-validation and leakage concerns are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a useful resource paper. The authors extend BLUEX with 2024/2025 exams, add GPT-4o generated captions in two modes, run a broad set of open and commercial models, and release data and code. That last part matters: the dataset and harness are on HuggingFace/GitHub, so the resource is reproducible. For the Portuguese and multilingual LLM community, having a current, text-only-accessible version of a respected entrance-exam benchmark is a genuine contribution, and the breakdown of image vs. text questions by year and area is transparent.\n\nWhat I'd flag first is the leakage path in the context captions. Section 1 says context captions are generated \"with access to the associated question and alternatives\"; Section 3.2 says GPT-4o sees \"the image and its associated question.\" Since the correct answer is one of the alternatives, the caption generator can encode answer-favoring details. That doesn't kill the blind-caption arm—blind captions are question-free—but it confounds the context-vs-blind comparison. The paper's own observation that the two caption types yield similar accuracy, despite context captions being half as long, could be read as leakage canceling out, or as genuine focus; as written, the design cannot distinguish them. The 1,422 usable questions claim is safe for blind captions, but the stronger \"leverage visual context\" conclusion needs a cleaner setup or at least a leakage audit.\n\nOther soft spots: Table 2 reports single runs with no error bars or significance tests; caption fidelity is never validated by humans or by comparing caption-condition answers to image-condition answers from the same multimodal model; the anti-contamination argument (\"2025 exams are too recent\") is asserted rather than probed; and the admission-cutoff comparison is cute but not validated against actual candidate behavior. These are fixable with modest effort. The author affiliation with Maritaca, whose models are evaluated, is worth noting but not disqualifying.\n\nOverall: the resource deserves a serious referee. It's a solid dataset paper that overclaims on the analytical side. I'd send it to review with the expectation that the leakage issue be addressed, error bars or at least repeated-seed variance be added, and the caption-fidelity question be engaged.","headline":"Useful benchmark increment with a real leakage concern in the context-caption arm; the resource is worth engaging, but the analytic claims need tightening.","tokens_in":9169,"tokens_out":2457,"would_cite":true,"duration_ms":23610,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automatic captions let text-only language models handle image-based exam questions, more than doubling the number of usable items in the BLUEX benchmark.","keywords":["BLUEX","Portuguese benchmark","image captioning","LLM evaluation","multimodal reasoning","Brazilian university entrance exams","visual grounding"],"falsifier":"Take a random sample of image-based BLUEX questions, and have one group of humans answer them with only the GPT-4o context captions and another group with only the original images. If the caption-only group scores materially worse, or an annotator audit finds a recurring rate of hallucinated details such as wrong labels, invented numbers, or missing diagram elements, then the benchmark's claim that captions make these questions accessible to text-only models is contradicted.","tokens_in":8341,"feed_emoji":"🎓","tokens_out":10916,"duration_ms":98962,"temperature":0.7,"pith_summary":"This paper claims that automatically generated image captions can make an image-heavy, Portuguese-language exam benchmark usable by text-only language models. The authors extend BLUEX, a dataset of multiple-choice questions from Brazilian university entrance exams, with 2024 and 2025 exams and with GPT-4o captions for every image, produced under two conditions: blind captions written from the image alone, and context captions written with the question in view. They report 1,422 usable questions, more than double the count previously available to text-only models, and accuracy gains of at least 10 points on image-only-alternative questions for most evaluated models when context captions are supplied. If correct, this gives the Portuguese LLM community a larger, more current benchmark and a controlled way to separate visual grounding from textual reasoning.","feed_headline":"Captioning unlocks 1,422 exam questions for text-only AI models","feed_subtitle":"GPT-4o image descriptions let text-only models answer visual questions from Brazilian entrance exams.","key_machinery":"The machinery is the dual captioning pipeline applied to every image in BLUEX via GPT-4o. A blind caption is generated from the image alone; a context caption is generated from the image plus the question and answer options. In evaluation, captions are inserted at the position where the image originally appeared, yielding three conditions: no caption, blind caption, and context caption. This design isolates the effect of visual description on model accuracy and is what makes the image-based questions answerable by text-only models.","core_discovery":"The central claim is that for BLUEX, GPT-4o-generated captions are sufficient to make image-dependent questions accessible to text-only LLMs, and that context-aware captions perform as well as blind captions despite being roughly half as long. The expanded set covers exams from 2018 through 2025, with 610 image-based items; converting those items to captions raises the count of text-only-usable questions to 1,422. The paper also reports that captioning improves accuracy on questions whose answer choices are images alone, that larger models benefit most from captions, and that the top evaluated models now score at or above the admission threshold for roughly 90% of undergraduate programs in e","pith_inferences":["Our inference: the near-tie between blind and context captions suggests task-relevant visual information is usually localizable, and that shorter context captions may be the better default for cost and clarity.","Our inference: if caption fidelity holds up under human audit, the same captioning pipeline could be transferred to other image-heavy exams, expanding non-multimodal evaluation beyond this one benchmark.","Our inference: the paper does not audit captions for hallucinated details, so the next natural test is an error analysis tagging invented labels, numbers, or diagram elements; systematic hallucination would change how the caption-condition scores should be interpreted."],"forward_implications":["Text-only language models can now be benchmarked on the full BLUEX image set instead of only the text-only subset, more than doubling the available items for non-multimodal evaluation.","Researchers can compare no-caption, blind-caption, context-caption, and true-image conditions on identical questions to isolate how much of an exam question depends on visual information.","The added 2024 and 2025 exams are recent enough that high scores are less likely to come from memorized training data, strengthening their use in contamination and generalization studies.","Because context captions are shorter yet match blind captions in accuracy, future evaluation runs can use more economical prompts without sacrificing measured performance."],"supporting_citations":[{"why":"Defines the original BLUEX dataset that this paper expands; supplies the base questions, image distribution, and the comparison point for the claim of more than doubling usable items.","marker":"[Almeida et al. 2023]"},{"why":"Describes GPT-4o, the model used to generate both blind and context captions and one of the evaluated commercial models; the captioning pipeline depends on it.","marker":"[OpenAI 2024]"},{"why":"Prior result that textual transcriptions of exam images can outperform direct image input; motivates the caption-condition design.","marker":"[Pires et al. 2023]"},{"why":"Provides the evaluation harness used to run the open-source models, making the reported accuracy numbers reproducible from a single tool.","marker":"[Gao et al. 2024]"},{"why":"Defines Sabia-3, the best commercial model on text-only questions and a key comparison point across conditions.","marker":"[Abonizio et al. 2024]"},{"why":"Defines DeepSeek-V3, the strongest open-source model across most conditions in the reported comparisons.","marker":"[Liu et al. 2024a]"},{"why":"Defines the Qwen2.5 family, which supplies the strongest small open-source models and the size-scaling trend the paper observes.","marker":"[Yang et al. 2024]"}],"fun_headline_variants":["Auto captions boost text-only AI exam access by 40%","BLUEX update: 1,422 questions now open to text-only models","GPT-4o captions let text-only LLMs solve visual test items","Captioning doubles usable questions in BLUEX benchmark"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Everything caption-based rests on GPT-4o's captions being faithful, complete substitutes for the original images; if a caption omits or invents a detail the question depends on, the caption-condition accuracy numbers no longer measure the model's reasoning.","fun_headline_variants_meta":{"raw":{"variants":["Auto captions boost text-only AI exam access by 40%","BLUEX update: 1,422 questions now open to text-only models","GPT-4o captions let text-only LLMs solve visual test items","Captioning doubles usable questions in BLUEX benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1065,"prompt_tokens":633,"completion_tokens":432,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":377,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":377,"tokens_out":432,"duration_ms":4760,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:24:30.933644+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of image-based BLUEX questions, and have one group of humans answer them with only the GPT-4o context captions and another group with only the original images. If the caption-only group scores materially worse, or an annotator audit finds a recurring rate of hallucinated details such as wrong labels, invented numbers, or missing diagram elements, then the benchmark's claim that captions make these questions accessible to text-only models is contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the original BLUEX dataset that this paper expands; supplies the base questions, image distribution, and the comparison point for the claim of more than doubling usable items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the evaluation harness used to run the open-source models, making the reported accuracy numbers reproducible from a single tool."}],"review_version":1}