{"id":"656cf463-2cab-46ba-9873-2104d0807e23","arxiv_id":"2508.15835","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Language models now exceed 94% accuracy on Brazilian entrance exams, but still lag on mathematics and specialized engineering exams.","lead":"Alvorada-Bench tests 20 language models on 4,515 text-only questions from five Brazilian university entrance exams. Top models reach 94.6% accuracy overall, but performance drops to around 61% on the engineering-focused IME exam.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated LM filter in §2.2 may leave image-dependent math/engineering items in the benchmark; the reported math/IME/ITA decline could be a missing-figure artifact, not a reasoning deficit.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the unvalidated LM-based filter in §2.2. My stress-test agrees and sharpens why it is central: the paper's headline distinction between humanities strengths and quantitative weaknesses is interpretable only if the text-only benchmark is actually text-only. If image-dependent questions remain, models are not being assessed on reasoning but on handling missing information; if they are over-removed, the subject-level and exam-level comparisons are biased. The paper provides no validation of the filter, no inter-annotator agreement, and no residual analysis of visual references in the released dataset. This is a direct threat to the validity of the central empirical claims, not a minor limitation. A simple audit of the retained items can resolve the concern: if all retained math/engineering items are genuinely text-only and the accuracy decline persists, the reasoning-deficit interpretation is supported; otherwise, the decline is at least partly a missing-figure artifact. Given this, the appropriate verdict remains CONDITIONAL: the contribution is useful and the findings plausible, but they should not be accepted as established without addressing this validation gap. The reader already reached CONDITIONAL, and my analysis reinforces that position without requiring a change.","tokens_in":7369,"tokens_out":4724,"duration_ms":50781,"concrete_test":"Scan all 4,515 questions (or a stratified random sample of 200 math/IME/ITA items) for Portuguese visual-content indicators in the question text: 'figura', 'gráfico', 'tabela', 'mapa', 'imagem', 'desenho', 'eixo', 'planta'. Have two independent annotators manually verify whether each flagged question is answerable without the visual and compute inter-annotator agreement. Recompute O3 and O3 Pro accuracy on the subset judged fully text-only vs. those with missing visuals. If the math/IME/ITA accuracy increases by more than 3–5 percentage points after excluding missing-visual items, the reported decline is partly an artifact of incomplete filtering.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—top models exceed 94% overall but decline on Mathematics and IME/ITA (§3.1, §3.6, abstract)—depends on the LM-based filter in §2.2 that is supposed to remove all questions requiring visual interpretation. The paper reports no accuracy or validation for this filter. Many math and engineering questions in these exams contain figures, graphs, or diagrams. If the filter leaves any such item in the text-only benchmark without its visual, models are evaluated with missing information, artificially depressing accuracy on those subjects. This would misattribute the decline to 'multi-step reasoning' (§4) when it is actually a missing-modality artifact. Conversely, if the filter over-removes geometry/graph questions, the surviving math subset may be skewed, making the 62.7% math and 61.4% IME figures (§3.6) unrepresentative. The paper's own Limitations (§4.1) acknowledge excluding multimodal questions but provide no check that the exclusion succeeded. Without a manual audit, the primary empirical finding—strength in humanities, weakness in quantitative reasoning—is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Alvorada-Bench is a 4,515-question, text-only multiple-choice benchmark drawn from five Brazilian university entrance exams (ENEM, FUVEST, UNICAMP, IME, ITA), spanning 1981–2025. The paper evaluates 20 language models from OpenAI, Anthropic, and DeepSeek under zero-shot, role-playing, and chain-of-thought prompting, yielding 270,900 responses with structured self-reports of confidence, difficulty, and Bloom level. The main empirical claims are that top models exceed 94% overall accuracy; accuracy is high on humanities and languages but declines sharply on Mathematics (62.7%) and on the engineering-oriented IME (61.4%) and ITA (68.1%) exams; confidence is well calibrated; and cost-accuracy frontiers make high accuracy available under $2 per 1K tokens. The paper also compares model performance with human baselines on ENEM 2024, reporting that all models surpass humans in most domains.","tokens_in":7654,"tokens_out":5354,"duration_ms":59045,"significance":"If the findings hold, Alvorada-Bench is a valuable, publicly released resource for Portuguese-language evaluation, addressing an important gap in multilingual and culturally specific benchmarking. The scale (4,515 questions, 20 models, 270,900 responses) and the availability of data/code are concrete strengths that support reproducibility. The observed pattern—strong cultural/humanities performance with weaker quantitative reasoning—is a useful, potentially actionable result for the community. However, the central quantitative claims are not yet fully established: the visual-question filtering step is unvalidated, no statistical uncertainty is reported, and some internal numerical inconsistencies appear. These issues are fixable and do not invalidate the dataset's potential, but they must be addressed before the paper's conclusions can be taken as reliable.","major_comments":[{"comment":"The filtering stage uses an unvalidated language model to detect and remove questions requiring visual interpretation. No precision, recall, or manual audit of the filter is reported. Because IME/ITA and Mathematics questions frequently contain figures, diagrams, and geometric drawings, any image-dependent item left in the text-only benchmark would depress accuracy artificially, misattributing a missing-modality artifact to 'multi-step reasoning.' Conversely, over-filtering would skew the surviving mathematics subset. The Limitations section acknowledges the exclusion of multimodal questions but provides no check that the exclusion succeeded. I request a manual audit of a random sample (ideally all) of included and excluded items, with report of filter accuracy and agreement, and a re-estimation or sensitivity analysis of the subject/exam-level results.","section":"§2.2 / §4.1"},{"comment":"All accuracy numbers are point estimates without confidence intervals or variance. The gap between O3 Pro (94.63%) and O3 (94.55%) is 0.08 percentage points, and the difference between ENEM (86.2%) and UNICAMP (86.1%) is 0.1 points; these are almost certainly within sampling error. The paper's ranking and 'decline' claims need bootstrap confidence intervals or Bayesian credible intervals, especially for per-subject and per-exam breakdowns where item counts are smaller. Without these, comparisons such as 'O1 vs DeepSeek Reasoner' or 'IME vs ITA' are over-precise. Please report the number of questions per subject/exam after filtering and provide interval estimates for the headline numbers.","section":"§3.1, Table 2, §3.5, §3.6"},{"comment":"The human comparison on ENEM 2024 is not sufficiently documented. The paper states that all 20 models surpass human baselines in Humanities, Natural Sciences, and Languages, and that GPT-4.1 Nano only underperforms humans in Mathematics, but it does not report the human accuracy values, the number of items per domain, or whether the human baseline was computed on the same text-only filtered subset used for the models. Without this information, the 'decisive shift' conclusion is not verifiable. Please provide the baseline numbers from reference [6], clarify whether the same questions and scoring were used, and note any differences in administration conditions.","section":"§3.2"},{"comment":"The cost-accuracy frontier contains an internal inconsistency. Table 2 gives O3 Mini an accuracy of 88.15% and O4 Mini an accuracy of 91.50%, but §3.3 says 'DeepSeek Reasoner (92.71%, $1.82) and O3 Mini (91.50%, $1.95) dominate the cost–accuracy frontier.' The 91.50% value belongs to O4 Mini, not O3 Mini. This changes the cost-accuracy claim and must be corrected. Please also specify the date and basis of API pricing, since provider prices change over time and the cost comparison may not generalize.","section":"§3.3"},{"comment":"Calibration claims are supported only by qualitative descriptions and figures. The abstract and §3.4 state that confidence is 'well calibrated' and 'correlates with perceived difficulty,' but no quantitative metrics are reported. I request standard calibration measures (expected calibration error, Brier score, correlation coefficient between confidence and accuracy, and between uncertainty and difficulty), with confidence intervals. This is necessary to support the paper's contribution on uncertainty quantification.","section":"§3.4"}],"minor_comments":[{"comment":"Minor language issues: 'this paper introduce' should be 'this paper introduces'; the abstract sentence beginning 'Evaluating twenty models...' is a fragment. A proofread pass is recommended.","section":"Abstract / §1"},{"comment":"The figures would benefit from error bars or shaded confidence regions, explicit sample sizes, and clearer axis labels. In particular, Figure 5's calibration panels should include a reference diagonal and quantitative summaries.","section":"Figures 5, 7, 9"},{"comment":"Bloom classifications are model-generated and used as measurements without validation against expert labels. Please state this limitation explicitly—it is currently only implicit—and note that the 'application-level bottleneck' claim depends on the reliability of these self-reported labels.","section":"§3.8"},{"comment":"Reference [6] is cited for the human baseline in §3.2, but the citation context is not fully described. Please specify which tables/figures in [6] support the human values used.","section":"References"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Alvorada-Bench is a useful, honest contribution. It gives the community a larger Portuguese-language exam benchmark than BLUEX—4,515 questions across five exam systems, with data and code released, and a broad evaluation of 20 models. The headline result, that top models clear 94% on the aggregate but drop to the low 60s on math-heavy IME/ITA, is credible and in line with other multilingual work showing cultural fluency ahead of multi-step quantitative reasoning. The cost-accuracy and calibration analyses are nice additions, and the limitations section is candid about contamination and binary grading.\n\nThe main soft spot is exactly where the stress-test points: the §2.2 filter that removes visual questions is an LLM-based step with no validation. No manual audit, no reported agreement, no check of what it kept or dropped. Since the central claim is the quantitative/engineering decline, an unvalidated filter leaves a real opening: if a few image-dependent math items slipped through without their figures, they'd fail for missing information, not for missing reasoning; if the filter over-removed geometry/graph items, the surviving math subset is skewed. This is not a reason to reject the paper, but it is the first thing I'd ask the authors to fix—random sample or gold set, with a measure of filtering accuracy.\n\nSecond, there are no confidence intervals or variance estimates anywhere. The per-exam numbers aggregate over 20 models, and IME is only 147 questions, so 61.4% on IME is noisier than the table suggests. A simple per-model SE or bootstrap would help readers judge whether the exam-level ordering is robust.\n\nThird, the Bloom's taxonomy analysis in §3.8 should be read as exploratory. The labels are generated by the models themselves, so the claim 'models struggle at Apply' is partly a statement about how models categorize items, not a validated cognitive-complexity measure. The authors use the terms without that caveat in the results section, though the method section is transparent about how the labels were produced.\n\nThe human-baseline comparison on ENEM 2024 relies on external data, and the alignment between the question subset used here and the human scores isn't fully pinned down. Minor.\n\nFor a reader building multilingual benchmarks or evaluating educational AI, this is worth reading and citing. The dataset itself is a resource, even if the analytics need tightening. I'd send it to peer review—it deserves referee time. The fixable items (filter validation, error bars, softer interpretation of self-reported Bloom) are exactly what review is for.","headline":"A useful new Portuguese-language exam benchmark; the main result is plausible, but the unvalidated visual-filter step needs to be addressed before the math/engineering decline claims are taken at face value.","tokens_in":8069,"tokens_out":3058,"would_cite":true,"duration_ms":31649,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models ace Brazilian university entrance exams but stumble on math.","keywords":["Brazilian university entrance exams","language model evaluation","Portuguese benchmark","mathematical reasoning","confidence calibration","cost-efficiency","Bloom's taxonomy","ENEM"],"falsifier":"A concrete test would be to have human annotators judge a random sample of both filtered-out and retained questions for whether an image is truly required, then run the same 20 models on freshly written, unpublished questions in each subject; if humanities accuracy falls sharply or the mathematics gap changes, the reported ceiling and subject-level gap reflect filtering or contamination rather than genuine capability.","tokens_in":7286,"feed_emoji":"🎓","tokens_out":4323,"duration_ms":51286,"temperature":0.7,"pith_summary":"Alvorada-Bench asks whether language models can genuinely handle Brazilian university entrance exams, which test Portuguese cultural knowledge alongside academic reasoning. The paper assembles 4,515 text-only multiple-choice questions from ENEM, FUVEST, UNICAMP, IME, and ITA, and evaluates 20 models under zero-shot, role-playing, and chain-of-thought prompts. Its central finding is that the best models exceed 94% average accuracy and are near-perfect on humanities, yet performance drops sharply on mathematics and on the computation-heavy engineering exams IME and ITA. The paper also reports that models' self-reported confidence tracks their actual accuracy closely, and that near-92% accuracy can be reached for under $2 per 1K tokens. Together, these claims suggest that language models have crossed a threshold of Portuguese-curriculum competence, with multi-step quantitative reasoning as the remaining bottleneck.","feed_headline":"Top AI models hit 94% on Brazilian exams but stumble on math","feed_subtitle":"Alvorada-Bench tests 20 language models on 4,515 entrance-exam questions; the weakest only trails humans in mathematics.","key_machinery":"Alvorada-Bench, a curated corpus of 4,515 text-only multiple-choice questions with official answer keys, built from five Brazilian exams through PDF extraction, pattern matching, an LLM-based filtering stage that removes image-dependent items, and text normalization. The evaluation protocol wraps each question in a structured JSON output requirement—selected alternative, confidence score (0–10), perceived difficulty (0–10), and a Bloom's-taxonomy label—which is what allows the paper to measure accuracy, calibration, and cognitive-complexity profiles on the same responses.","core_discovery":"The paper's central claim is that, on a new benchmark built from real Brazilian entrance examinations, current language models perform at or above the level of the average Brazilian student in most subjects while retaining a persistent weakness in mathematics and engineering-style problem solving. Top models reach 94.6% overall accuracy; human sciences reach 93.9%; mathematics falls to 62.7% for baseline models, although reasoning-enhanced models recover to roughly 94% on mathematics. The evaluation also collects structured confidence and difficulty self-reports, and the paper shows that low-confidence responses still exceed 90% accuracy while accuracy degrades monotonically as confidence fa","pith_inferences":["The paper's own contamination caveat implies a falsifiable prediction: on a freshly written, never-published exam, accuracy on culturally specific humanities content would drop if public-exam training data explains the near-perfect scores.","The LLM-based filtering stage could be inverted as a diagnostic: comparing performance on retained text questions versus excluded image-dependent questions would quantify how much visual reasoning modern models still lack, an extension the paper does not pursue.","A natural next experiment is to run the same 759 mathematics questions with tool use, such as a calculator or symbolic solver, and measure whether the residual IME/ITA gap closes; the paper explicitly leaves tool use out of scope.","Because only final answers are scored, human partial-credit grading of the same 270,900 responses would test whether correct answers are reached through sound reasoning or spurious correlations, especially on mathematics."],"forward_implications":["If the reported results hold, educational deployment is no longer blocked by cost: DeepSeek Reasoner and O3 Mini deliver roughly 92% accuracy at under $2 per 1K tokens.","Confidence self-reports can serve as a practical routing signal, because low-confidence responses reliably mark likely errors and uncertainty correlates with perceived difficulty.","Prompt engineering has little effect on reasoning-optimized models, with O3 varying by only about 0.1 percentage point across prompting strategies, so evaluation protocols can be standardized without much performance cost.","Application-level Bloom tasks are the main failure tier for conventional models, meaning progress in knowledge recall has outpaced progress in computation and applied problem solving.","The 24.8-percentage-point drop from ENEM to IME/ITA indicates that engineering entrance exams, not general curriculum exams, are the sharper test of symbolic manipulation and multi-step reasoning."],"supporting_citations":[{"why":"Introduces BLUEX, a text-only corpus from UNICAMP and USP, providing the earlier evidence that Brazilian exams can serve as LLM evaluation substrates.","marker":"[1]"},{"why":"GPT-4 technical report supplies the English-centric standardized-test baselines (SAT, Bar Exam) that the paper contrasts with Brazilian performance.","marker":"[2]"},{"why":"Documents performance drops from English tasks to Telugu tasks in education, motivating the need for a Portuguese-language benchmark.","marker":"[3]"},{"why":"Shows lexical divergence in Chinese despite English-like syntax, supporting the argument that language benchmarks need cultural specificity.","marker":"[4]"},{"why":"Provides a coding-competition result used alongside test scores as evidence of current high LLM performance on standardized evaluation tasks.","marker":"[5]"},{"why":"Supplies the ENEM 2024 human performance baseline used to compare model accuracy with Brazilian students' results.","marker":"[6]"}],"fun_headline_variants":["AI aces Brazilian entrance exams, except math","Brazilian exam benchmark stumps AI on math","94% on Brazilian exams, but AI flunks math","Language models top Brazilian exams, lag in math","New benchmark: AI hits 94% but math lags"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The key assumption is that the automatic language-model filter that removes questions requiring figures or diagrams is accurate; if it silently keeps image-dependent items or silently excludes text-solvable ones, the reported accuracies describe only the filtered subset, not the full exams.","fun_headline_variants_meta":{"raw":{"variants":["AI aces Brazilian entrance exams, except math","Brazilian exam benchmark stumps AI on math","94% on Brazilian exams, but AI flunks math","Language models top Brazilian exams, lag in math","New benchmark: AI hits 94% but math lags"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000488,"raw_usage":{"total_tokens":2237,"prompt_tokens":736,"completion_tokens":1501,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":1425}},"tokens_in":480,"tokens_out":1501,"duration_ms":10780,"temperature":1.0,"reasoning_tokens":1425,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:59:23.691794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to have human annotators judge a random sample of both filtered-out and retained questions for whether an image is truly required, then run the same 20 models on freshly written, unpublished questions in each subject; if humanities accuracy falls sharply or the mathematics gap changes, the reported ceiling and subject-level gap reflect filtering or contamination rather than genuine capability.","supporting_citations":[{"cited_title":"BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams","cited_arxiv_id":"2307.05410","evidence_quote":"Introduces BLUEX, a text-only corpus from UNICAMP and USP, providing the earlier evidence that Brazilian exams can serve as LLM evaluation substrates."},{"cited_title":"Multilingual Performance Biases of Large Language Models in Education","cited_arxiv_id":"2504.17720","evidence_quote":"Documents performance drops from English tasks to Telugu tasks in education, motivating the need for a Portuguese-language benchmark."},{"cited_title":"Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs","cited_arxiv_id":"2410.15956","evidence_quote":"Shows lexical divergence in Chinese despite English-like syntax, supporting the argument that language benchmarks need cultural specificity."},{"cited_title":"Comparing Large Language Models and Human Programmers for Generating Programming Code","cited_arxiv_id":null,"evidence_quote":"Provides a coding-competition result used alongside test scores as evidence of current high LLM performance on standardized evaluation tasks."},{"cited_title":"Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil","cited_arxiv_id":"2408.05035","evidence_quote":"Supplies the ENEM 2024 human performance baseline used to compare model accuracy with Brazilian students' results."}],"review_version":1}