{"id":"c0a3ef5f-9cbf-42a6-9de6-3f44bd2c2666","arxiv_id":"2505.16459","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"MMLU-Reason adds 1,083 hard multimodal questions with a GPT-4o-based trace-scoring pipeline and finds thinking-enabled models score higher on accuracy but still show inconsistent reasoning.","lead":"MMLU-Reason introduces 1,083 hard multimodal questions across six reasoning types and a pipeline that scores the intermediate thinking traces of AI models. It reports that thinking-enabled models beat non-thinking ones on accuracy, yet even the best models show flawed reasoning such as inconsistency and overthinking.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reasoning-quality claim rests entirely on an unvalidated GPT-4o judge; no human-agreement or bias checks are reported, so the accuracy–reasoning gap may be a judge artifact.","rationale":"The reader's weakest assumption is exactly the load-bearing point: GPT-4o's automatic scoring of reasoning traces is the only evidence for the claimed gap between accuracy and reasoning quality. Without human agreement or bias checks, the central claim is conditional, not established. The accuracy comparisons in Table 3 are internally coherent, and the dataset construction with 1,083 items, six reasoning types, and attached difficulty/source metadata is a useful contribution; the Dual model design is also a reasonable modular baseline. However, the paper's headline finding—that top thinking models 'suffer from reasoning pathologies such as inconsistency and overthinking'—depends on trace-quality scores that are never validated against human judgment. The missing validation is especially consequential because the judge is a single proprietary model whose preferences could easily correlate with trace length, phrasing, or style. The paper also overstates coverage by citing Gemini-2.5 Pro in the abstract without presenting trace-quality data for that model. These concerns do not warrant rejection, because the benchmark and accuracy results retain value and the trace pipeline is a stated contribution that can be validated post hoc. They do warrant keeping the verdict at CONDITIONAL: the central claim should be accepted only after the judge is validated and the missing Gemini-2.5 Pro trace results are supplied or the claim is narrowed.","tokens_in":28062,"tokens_out":4290,"duration_ms":40073,"concrete_test":"Score a stratified sample of 100 reasoning traces per model family (covering all six task types and both thinking and non-thinking outputs) with two independent human raters using the exact RTQ/RTA/RSC rubrics; compute weighted Cohen's kappa between each human rater and GPT-4o, and between the two human raters. If human–GPT-4o agreement is below 0.6, or if agreement differs by more than 0.15 across model families, the current evidence does not establish the accuracy–reasoning gap. Additionally, include a style-bias control: for 50 traces, score a concise paraphrase and a verbose paraphrase of the same reasoning, and verify that judge scores do not move substantially with wording alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MMLU-Reason exposes a persistent gap between answer accuracy and reasoning quality: high-accuracy MLLMs-T still produce inconsistent, overlong, or irrelevant traces. The load-bearing step is Section 4.2's assertion that 'Leveraging GPT-4o [17] as an automated evaluator, RTEP enables scalable, semantically aligned trace assessments.' This GPT-4o judge is the sole support for RTQ, RTA, RSC, and the Err-T/Err-A classifications in Table 4 and Figure 5. No human-agreement study, no rubric calibration, and no adversarial validation is reported; the only Krippendorff's alpha (0.84, Appendix A) applies to answer-label annotation, not to trace-quality judgments. Because the same model scores traces and classifies errors, any systematic judge preference—for concise over verbose traces, for particular phrasing, or for outputs resembling GPT-4o's own style—will surface as 'inconsistency' or 'overthinking.' If such a bias exists, the reported pathologies and the accuracy–reasoning gap are artifacts of the judge rather than properties of the evaluated models. A second, narrower evidentiary gap compounds this: the abstract names Gemini-2.5 Pro as a model that suffers reasoning pathologies, yet Table 4 reports trace-quality metrics only for Claude-3.7-sonnet and the Dual configuration; no trace-level results for Gemini-2.5 Pro appear in the paper. Thus even a perfectly valid judge would not support that specific example from the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MMLU-Reason, a multi-modal reasoning benchmark of 1,083 questions across six domains (logic, math, space-time, code, map, science), and RTEP, a pipeline that uses GPT-4o to score the intermediate reasoning traces of MLLMs on relevance, answer-relevance, and step consistency, and to classify thinking and answer errors. The authors report accuracy for 17 models, trace-quality metrics for Claude-3.7-sonnet and a custom 'Dual' (GPT-4V + DeepSeek-R1) configuration, and an error-type analysis on Claude-3.7-sonnet's validation set. The central claim is that MLLMs with thinking (MLLMs-T) outperform non-thinking MLLMs on accuracy while still exhibiting reasoning pathologies such as inconsistency and overthinking, producing a persistent gap between answer accuracy and reasoning quality. The paper positions this as a new direction for benchmark evaluation.","tokens_in":28340,"tokens_out":9214,"duration_ms":67513,"significance":"If the reasoning-quality metrics were properly validated, MMLU-Reason would be a genuinely useful instrument: the task taxonomy is broad, the questions require non-trivial inference, and the focus on the reasoning trace is still rare in the multi-modal benchmark literature. The accuracy table is internally consistent and the dataset description is detailed. However, the paper's headline finding — an accuracy-reasoning gap — rests on a single unvalidated GPT-4o judge, and the human-gap comparison uses a validation-only expert baseline. The paper also lacks confidence intervals or significance tests. These gaps are fixable, but they currently prevent the empirical conclusions from being treated as established.","major_comments":[{"comment":"RTQ, RTA, RSC, and the Err-T/Err-A classifications are produced entirely by GPT-4o as the automated evaluator, yet no human-agreement study, rubric calibration, or adversarial validation is reported for these trace-level judgments. The only inter-annotator agreement in the paper (Krippendorff's alpha 0.84, Appendix A) concerns the Expert (Human only) answer labels on the validation set. Because the same judge scores traces produced by models that include GPT-4V, o4-mini, and Gemini-2.5 Pro, a systematic stylistic preference of GPT-4o (e.g., for concise traces or for phrasing similar to its own style) would be reported as 'inconsistency,' 'overthinking,' or 'irrelevant thinking.' The accuracy-reasoning gap that drives the paper's conclusions could therefore be a judge artifact. Please add a human-annotation agreement study on a representative sample, a per-model bias analysis, and a qualitative inspection of GPT-4o's judgments.","section":"Section 4.2, Tables 4 and Figure 5"},{"comment":"The abstract states that 'even top models like Claude-3.7-Sonnet and Gemini-2.5 Pro suffer from reasoning pathologies such as inconsistency and overthinking,' but the paper reports trace-quality metrics (RTQ, RTA, RSC, ThinkErr) only for Claude-3.7-sonnet and the Dual configuration in Table 4; no trace-level results for Gemini-2.5 Pro appear anywhere. The error-type distributions in Figure 5 are also limited to Claude-3.7-sonnet on the validation set. Thus the specific claim about Gemini-2.5 Pro is not supported by the evidence presented.","section":"Abstract and Section 4.3, Table 4"},{"comment":"The human-gap comparison is based on a mismatch between validation-set and test-set numbers. The text says 'Gemini-2.5 Pro achieves a test accuracy of 42.45%,' but Table 3 shows 42.45 for the validation set (106 questions) and 42.36 for the test set (977 questions). The Expert (Human only) and Expert (Human + GPT-4o) rows are reported only under the Validation column; the 52.85% figure used in the abstract and Figure 1 is therefore a validation-set result. Comparing a validation-set human baseline to test-set model accuracy is not an apples-to-apples comparison. Please report the expert baselines on the test set, or explicitly label all numbers in this comparison as validation-set performance and avoid the word 'test'.","section":"Section 1 and Table 3"},{"comment":"No confidence intervals, significance tests, or repeated-run variance are provided for any accuracy number. With 977 test questions, a 42% accuracy has a standard error of roughly 1.6 percentage points, and the per-task columns (141-212 questions) have standard errors of 3-4 points. Consequently, claims such as 'MLLMs-T overall outperform MLLMs' and the ranking of Gemini-2.5 Pro above Claude-3.7-sonnet rest on differences that may be within sampling noise. Please add error bars or a statistical comparison (e.g., a per-question paired test across models).","section":"Table 3"},{"comment":"The metric scales are inconsistent and the OS definition is under-justified. Section 2.2 says RTQ and RTA are 'normalized within the [0,1] interval,' but Table 4 reports values on a 0-10 scale (e.g., RTQ 9.39). The OS formula, 0.3*RTQ + 0.3*RTA + 0.3*RSC + 0.1*(ACC*0.1), mixes a 0-10 scale with a percentage; the weights are arbitrary. Please unify the scales, state the OS computation clearly, and justify or drop the OS aggregation when drawing conclusions about relative reasoning quality.","section":"Section 2.2 and Table 4"}],"minor_comments":[{"comment":"The word 'Pipline' should be 'Pipeline.'","section":"Section 4.2"},{"comment":"The label 'F ormat Error' should be 'Format Error.'","section":"Figure 5"},{"comment":"The entries 'Average Reasoning Depth 4.15 levels' and 'Reasoning Steps per Question 3.42' are reported without definitions of how depth and step count were computed; please add a short annotation protocol.","section":"Table 2"},{"comment":"The difficulty split (Easy:Medium:Hard = 30%:40%:30%) is stated without a rubric or agreement measure; a brief description of the difficulty annotation would help reproducibility.","section":"Section 2.1"},{"comment":"References [28] and [29] list arXiv IDs 2405.67890 and 2405.12345; these identifiers do not resolve to publicly verifiable papers. Please confirm the citations or replace them with correct sources.","section":"References"},{"comment":"The paper does not provide a data or code availability statement beyond the project page URL. For a benchmark contribution, a persistent link to the dataset and evaluation scripts is expected.","section":"General"},{"comment":"The 'Dual' model is a custom pipeline (GPT-4V + DeepSeek-R1), not an off-the-shelf model; the table should mark it as a configuration rather than a model to avoid confusion with the other rows.","section":"Table 3"},{"comment":"The 'Expert (Human only)' baseline reaches only 29.23% on the validation set, which is close to the Frequent Choice baseline (26.8%); a brief note on why the human baseline is so low (e.g., question difficulty, time pressure, or grading strictness) would preempt confusion.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The central concern is the unvalidated GPT-4o judge and the unsupported Gemini-2.5 Pro claim; these need to be addressed before publication. In addition, references [28] and [29] have arXiv identifiers that appear to be synthetic (2405.67890 and 2405.12345) and do not correspond to real papers as far as I can verify; this is a citation integrity issue that should be checked carefully. The paper's scope fits an AI/ML venue, but the benchmark's long-term value depends on releasing the data and code. If the authors can supply human agreement for RTEP, the paper could become a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading for the accuracy comparison, no more than that until the trace pipeline is validated. The 1,083-question set spanning six reasoning domains is a real contribution, and the paper does something prior benchmarks mostly skip: it tries to score the reasoning trace, not just the final answer. The accuracy table is internally consistent, and the finding that thinking-enabled models generally beat non-thinking ones is supported within the paper's own setup. The validation/test split, expert baselines, and per-task breakdowns make this a usable diagnostic tool.\n\nWhat is new is the combination of that dataset with the RTQ/RTA/RSC trace metrics and error-type labels. The idea of evaluating trace quality is not original—MME-CoT already benchmarks CoT quality in multimodal models—but the six-domain coverage plus the relevance/consistency decomposition is a convenient packaging and could be useful.\n\nThe soft spot is load-bearing. Section 4.2 says GPT-4o scores the traces, but there is no human-agreement study for RTQ/RTA/RSC, no adversarial check, and no bias analysis. The Krippendorff alpha in Appendix A is for answer-label annotation, not trace-quality judgments. If GPT-4o is biased toward its own style or against long traces, the reported inconsistency and overthinking pathologies are artifacts. That directly hits the abstract's main claim. On top of that, the abstract names Gemini-2.5 Pro as a model with reasoning pathologies, but Table 4 only reports trace metrics for Claude-3.7-sonnet and the Dual configuration. I cannot verify the Gemini claim from the presented evidence. There are also no confidence intervals, and no dataset or code is attached to the paper despite a project page.\n\nThe accuracy half holds up. The reasoning-quality half is conditional on releasing the benchmark and validating the judge against human ratings. I would send this to reviewers, because the dataset deserves scrutiny and judge validation is exactly what referees should demand, but I would not cite the reasoning-quality finding as established.","headline":"Useful accuracy benchmark, but the reasoning-quality headline rests on an unvalidated GPT-4o judge, and the abstract cites Gemini-2.5 Pro without trace-level data.","tokens_in":28903,"tokens_out":2368,"would_cite":false,"duration_ms":21926,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new multimodal benchmark scores the thinking trace as well as the answer and finds that correct answers frequently hide incoherent reasoning in the strongest models.","keywords":["MMLU-Reason benchmark","multimodal reasoning","reasoning trace evaluation","chain-of-thought evaluation","multi-modal large language models","thinking quality metrics","overthinking","reasoning error taxonomy"],"falsifier":"Select a random sample of, say, 100 reasoning traces from MMLU-Reason, have independent human raters score them on the same RTQ, RTA, and RSC rubrics, and compare with GPT-4o's scores. If the human-model agreement is low or the judge systematically favors or penalizes certain model families, the reported accuracy-reasoning gap and error-type frequencies would be artifacts of the evaluator rather than properties of the models.","tokens_in":27862,"feed_emoji":"🧠","tokens_out":10849,"duration_ms":77382,"temperature":0.7,"pith_summary":"MMLU-Reason is a new benchmark of 1,083 difficult multimodal questions spanning logic, math, code, maps, science, and space-time reasoning. Its contribution is a Reasoning Trace Evaluation Pipeline (RTEP) that looks inside the intermediate thinking a model emits before its final answer, scoring how relevant that thinking is to the question, how consistent it is with the answer, and what kinds of errors it contains. The paper's central empirical claim is that models with explicit thinking traces beat non-thinking models on accuracy, but even the strongest thinking models routinely produce incoherent, overlong, or irrelevant reasoning. The result is a measurable gap between accuracy and reasoning quality, which the authors argue makes trace-level evaluation necessary rather than optional.","feed_headline":"Right answers, wrong thinking: benchmark exposes AI reasoning gap","feed_subtitle":"MMLU-Reason scores the thinking behind the answer and finds top models overthink and contradict themselves.","key_machinery":"The Reasoning Trace Evaluation Pipeline (RTEP) is the central instrument. It takes the intermediate thinking trace a model produces before answering and scores it on three 0-10 scales: RTQ (how relevant the thinking is to the question), RTA (how logically consistent the thinking is with the final answer), and RSC (whether the reasoning steps stay logically coherent with each other). It also classifies thinking errors (inconsistency, overthinking, irrelevant thinking, repetitive thinking) and answer errors (reasoning error, perceptual error, format error, extraction error, reject to answer). RTEP is automated through GPT-4o prompts, and its role is to make trace-level reasoning quality measurable and comparable across models and tasks.","core_discovery":"The paper claims that a model can answer correctly while reasoning badly, and that this is common at the current frontier: Claude-3.7-Sonnet and Gemini-2.5 Pro, among the strongest evaluated models, show inconsistency and overthinking in their traces, and the single most frequent thinking error for Claude-3.7-Sonnet is internal contradiction at 41.5%. The authors also find that models that emit explicit thinking traces (MLLMs-T) outperform standard MLLMs overall, with Gemini-2.5 Pro reaching 42.36% test accuracy while the human-with-GPT-4o upper bound sits at 52.85%. RTEP quantifies the gap through three trace-level metrics (RTQ, RTA, RSC) and an error taxonomy, and the paper argues this exposes a critical misalignment between surface-level correctness and reasoning fidelity.","pith_inferences":["Beyond the paper: if trace-quality scoring becomes routine, training objectives could shift from final-answer reward toward step-consistency and relevance pressure, a concrete architectural change the paper does not propose.","Beyond the paper: a natural next experiment is to use RTEP scores as a filtering signal during decoding or as a reward model in reinforcement learning for reasoning models, testing whether the accuracy-reasoning gap is trainable or inherent.","Beyond the paper: the error taxonomy suggests concrete interventions such as length-control or self-consistency checks that could be injected mid-generation to suppress overthinking and contradiction, though the paper does not test them."],"forward_implications":["Answer accuracy alone is no longer a sufficient evaluation metric for multimodal reasoning models; any assessment that claims to measure reasoning must also score the trace.","Current thinking models are not uniformly better: their advantage over non-thinking models is real on aggregate, but task-level performance varies widely, with code and map reasoning showing the lowest ceilings.","Longer thinking traces do not imply better reasoning, since the paper's modular Dual system produced traces several times longer than Claude-3.7-Sonnet's yet scored lower on trace relevance and consistency despite sometimes higher accuracy.","Because reasoning errors outnumber perceptual errors in the failure distribution, the practical bottleneck for current multimodal models is symbolic structure and multi-step inference, not basic vision.","If RTEP scores are adopted, future model development can be guided by trace quality alongside accuracy, which the paper argues should push architectures toward more cognitively aligned reasoning."],"supporting_citations":[{"why":"The GPT-4o system card is the basis for RTEP's automated evaluator, so all trace-quality scores depend on this model.","marker":"[17]"},{"why":"MMMU is the prior broad multimodal benchmark that MMLU-Reason extends by adding thinking-trace evaluation.","marker":"[47]"},{"why":"MME-CoT is the closest prior work on chain-of-thought reasoning evaluation in multimodal models and provides the comparative baseline for RTEP.","marker":"[18]"},{"why":"Gemini-2.5 Pro is one of the two strongest evaluated thinking models and anchors the claim that top models still show reasoning pathologies.","marker":"[13]"},{"why":"Claude-3.7-Sonnet provides the trace and error-type distributions that support the accuracy-reasoning gap finding.","marker":"[6]"},{"why":"DeepSeek-R1 supplies the reasoning half of the Dual model, whose long but lower-quality traces support the claim that length does not equal reasoning quality.","marker":"[14]"},{"why":"GPT-4 Vision serves as a no-thinking baseline and as the perception half of the Dual model, making it load-bearing for the thinking-versus-no-thinking comparison.","marker":"[35]"}],"fun_headline_variants":["AI models ace tests but flunk reasoning, says new benchmark","Top AI models overthink and contradict: benchmark reveals gap","Correct answers, faulty logic: MMLU-Reason exposes AI reasoning gap","Why AI gets it right but reasons wrong: new benchmark explains","Right answers, wrong thinking: new benchmark quantifies the split"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire diagnosis of reasoning quality rests on trusting GPT-4o's automatic scores to reflect real reasoning quality, and the paper provides no human-agreement check to show those scores are not biased toward or against particular model families.","fun_headline_variants_meta":{"raw":{"variants":["AI models ace tests but flunk reasoning, says new benchmark","Top AI models overthink and contradict: benchmark reveals gap","Correct answers, faulty logic: MMLU-Reason exposes AI reasoning gap","Why AI gets it right but reasons wrong: new benchmark explains","Right answers, wrong thinking: new benchmark quantifies the split"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1779,"prompt_tokens":981,"completion_tokens":798,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":710}},"tokens_in":597,"tokens_out":798,"duration_ms":6363,"temperature":1.0,"reasoning_tokens":710,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:59:46.739907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of, say, 100 reasoning traces from MMLU-Reason, have independent human raters score them on the same RTQ, RTA, and RSC rubrics, and compare with GPT-4o's scores. If the human-model agreement is low or the judge systematically favors or penalizes certain model families, the reported accuracy-reasoning gap and error-type frequencies would be artifacts of the evaluator rather than properties of the models.","supporting_citations":[],"review_version":1}