{"id":"c2be4fc1-ba5d-4be4-8aaa-5469d1e4ee3a","arxiv_id":"2505.02018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"R-Bench is a new bilingual, multidisciplinary, graduate-level reasoning benchmark on which top AI models score 69% (text) and 53% (multimodal), well below saturation.","lead":"This paper introduces R-Bench, a new benchmark of 1,094 text and 665 multimodal graduate-level questions in English and Chinese spanning many disciplines, and tests AI models on it. It reports that even the strongest models solve only about half of the multimodal questions, suggesting current AI reasoning is far from robust.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2,000-token o1 filter in Section 2.3 is an unvalidated, model-specific proxy for difficulty; if it selects on ambiguity or length rather than reasoning depth, R-Bench's calibration claim and headline scores are unsupported.","rationale":"The reader and I converge on the same weakest point: the Section 2.3 Step 4 threshold. I read the paper's central claim as 'R-Bench is a valid graduate-level complex-reasoning benchmark on which top models score poorly, e.g., o1 at 53.2% multimodal.' That claim requires the filtered questions to be hard for reasons of reasoning depth. Section 2.3 offers no evidence that o1's hidden reasoning-token count is a calibrated difficulty measure, and the only user study in Section 3.1 compares R-Bench to MMLU/MMMU, not retained to filtered questions, so it cannot validate the threshold. My concern is not that o1 token counts are useless; it is that the paper supplies no check against confounds such as input length, OCR ambiguity, or o1 policy artifacts. The concrete human-rating test I propose would settle whether the filter is selecting for reasoning depth. I am not moving to REJECT: the benchmark may be fixable by releasing the data and validating the filter, so CONDITIONAL remains the appropriate posture. Other issues, such as the contradiction between the abstract's 'made publicly available at here' and the conclusion's 'Later, we will make the data and code available,' small user-study sizes, and missing error bars, reinforce conditionality but do not replace the filter concern. No machine-checked proof or independent audit is present.","tokens_in":17494,"tokens_out":5466,"duration_ms":61478,"concrete_test":"Release the full candidate pool with per-question o1 reasoning-token counts, filter status, subject, and question length. Independently recruit graduate students who are not authors to blind-rate a stratified sample of roughly 100 retained and 100 filtered questions on a 1-5 reasoning-depth scale, while also rating ambiguity and OCR quality. If the retained and filtered distributions of reasoning depth overlap substantially, or if the depth difference vanishes after controlling for ambiguity and length, then the 2,000-token rule is not a valid difficulty filter and the calibration claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 Step 4 removes every question for which OpenAI o1 uses fewer than 2,000 reasoning tokens, 'to ensure that R-Bench is a benchmark for reasoning evaluation.' The abstract's difficulty-calibration claim rests entirely on this threshold. The assumption is that o1's hidden reasoning-token count is a monotone, transferable measure of question reasoning difficulty. That assumption is not established and is likely false: reasoning-token count is an opaque property of o1's policy and prompt, not an intrinsic property of a question. It can be inflated by input length, OCR artifacts, ambiguous wording, unfamiliar notation, or subjects where o1 is merely uncertain; it can be suppressed for questions o1 can shortcut. The only validation offered, Section 3.1, compares 30 R-Bench items against 30 MMLU/MMMU items, not retained against filtered questions, so it does not test the threshold. Furthermore, using o1 to select the benchmark and then reporting o1's 69.0%/53.2% accuracy as evidence of benchmark difficulty is not independent evidence. If the filter is confounded with length or ambiguity, the headline claim of a graduate-level complex-reasoning benchmark, and every model ranking on R-Bench, inherits the bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"R-Bench is a new bilingual (English/Chinese) benchmark for evaluating complex reasoning in LLMs and MLLMs. It contains 1,094 text-only questions across 108 subjects and 665 multimodal questions across 83 subjects, collected from Tsinghua graduate/undergraduate courses, digitized with OCR, filtered in three rounds (expert screening, an o1 reasoning-token threshold, and manual review), converted to six-option single choice, translated by experts, and evaluated with 27 LLMs and 11 MLLMs. The paper reports that even o1 achieves 69.0% on R-Bench-T and 53.2% on R-Bench-M, and uses expert/o1 pairwise comparisons and thinking-time comparisons to argue that R-Bench requires more reasoning than MMLU and MMMU.","tokens_in":17764,"tokens_out":4854,"duration_ms":47219,"significance":"If the difficulty-calibration claims are substantiated, R-Bench would be a valuable community resource because it combines breadth of disciplines, bilingual equal-difficulty question pairs, multimodal and text-only splits, and an automatic single-choice answer format that avoids proof-verification problems. The paper also publicly releases data and code and evaluates a wide range of open and closed models, which supports reproducibility and direct use. However, the evidence that the benchmark actually measures 'graduate-level complex reasoning' rather than length- or ambiguity-related difficulty is currently thin; in particular, the o1 reasoning-token filter is the central load-bearing assumption and it is not validated independently of o1 itself. For these reasons the contribution is promising but not yet rigorously established.","major_comments":[{"comment":"The paper removes every question for which OpenAI o1 generates fewer than 2,000 reasoning tokens and states that this 'ensures that R-Bench is a benchmark for reasoning evaluation.' This is the sole difficulty-calibration mechanism, yet no evidence is given that reasoning-token count is a valid or transferable measure of human reasoning difficulty. The count is a policy-dependent property of o1 and can be inflated by input length, OCR artifacts, ambiguity, unfamiliar notation, or cheap uncertainty, and depressed for questions o1 shortcuts. Because the abstract and Section 1 base the 'rigorous difficulty calibration' and 'Olympiad-level' claims on this step, the authors should provide the distribution of token counts, a correlation with independent human difficulty ratings on a larger sample, a sensitivity analysis around the 2,000 threshold, and a funnel report showing how many questions were removed at each stage. Without this, the difficulty claim and the headline accuracies are not supported.","section":"§2.3, Step 4 (model screening)"},{"comment":"The reasoning-comparison study uses 30 randomly selected questions per benchmark, with no sample-size justification, no confidence intervals or inter-annotator agreement, and no description of how the 30 questions were sampled. Moreover, the o1 voting compares R-Bench against MMLU/MMMU, not against the questions removed by the §2.3 filter, so it does not validate the threshold. Using o1 as the filter, as the judge in Tables 2–3, and as an evaluated model in Table 5 means that o1's poor accuracy is not independent evidence of difficulty. The authors should report expert-level agreement with error bars on a substantially larger sample, and should directly compare retained versus filtered questions.","section":"§3.1, Tables 2–3"},{"comment":"Accuracy differences are reported without confidence intervals or repeated evaluations. For example, in Table 5 the difference between o1-20241217 (69.0) and Gemini-2.0-flash-thinking (68.4) is 0.6 percentage points on 1,094 questions, well within binomial sampling noise; similarly, the CoT effect for GPT-4o in Table 7 is 2.1 points. The ordering claims and the 'no impact of CoT on o1-mini' conclusion need bootstrap intervals or paired statistical tests. This is important because the paper's stated purpose is to provide guidance for model improvement, and several cross-model conclusions are currently indistinguishable from noise.","section":"§3.2, Tables 5–7; Fig. 5"},{"comment":"The paper asserts that the English and Chinese versions are of equal difficulty and interprets the high consistency in Fig. 5 as evidence of cross-lingual reasoning. The translation review is described, but no procedure or metric establishes that the two versions are equally difficult; a model could answer both versions of the same question correctly or incorrectly for reasons related to translation quality or format rather than reasoning. The authors should either measure equal difficulty (for example, by obtaining separate human difficulty ratings on each language version) or soften the cross-linguistic-alignment claim.","section":"§2.4 and Fig. 5"}],"minor_comments":[{"comment":"The R-Bench labels appear truncated in the figure as '-Bench-T' and '-Bench-M', and the figure lacks confidence intervals, which is especially relevant because several accuracy differences in Tables 5 and 6 are within sampling noise.","section":"Fig. 1"},{"comment":"In the example for complex analysis, the polynomial is rendered in a garbled form ('53 2() 5 2pz z z z=++ +'), making it impossible for a reader to verify the question; if this reflects the actual typeset version, it undermines the claim that the digitization was carefully proofread.","section":"Fig. 3"},{"comment":"The text says 'Table 8 and Table 8' when introducing the subject distributions; the second reference should be Table 9.","section":"Appendix A.3"},{"comment":"The paper reports 10,270 collected questions and 1,759 retained questions, but does not report how many questions were removed at each of the three filtering stages; providing a full funnel would help readers assess the effect of the o1-token threshold and the manual review.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The resource is real: 1,094 text and 665 multimodal questions freshly collected from Tsinghua courses, bilingual English-Chinese, spanning over a hundred subjects, with options designed for automatic verification. That alone is worth having. The evaluation suite is wide, and the text-vs-multimodal gap (69.0 vs 53.2 for o1) is a consistent and plausible finding. The small expert study also supports the claim that R-Bench questions demand more reasoning than MMLU/MMMU, though the sample is small.\n\nThe soft spot is the difficulty filter. Section 2.3 removes every question where o1 produces fewer than 2,000 reasoning tokens, saying this \"to some extent reflects\" difficulty. That is the entire calibration argument. The validation in Section 3.1 compares R-Bench against MMLU/MMMU, not retained against filtered questions, so it never tests the threshold itself. And because o1 is also the model whose low score is cited as evidence of difficulty, there is a real circularity burden. This is not fatal — a strong-model filter is a defensible design choice if validated — but as written the evidence doesn't yet support the graduate-level difficulty claim.\n\nMinor issues: no error bars anywhere, so the 69.0-vs-68.4 ranking between o1 and Gemini is not meaningful as presented; the 30-question user studies are small; \"Olympiad-level\" in the abstract is a stretch given the source is university coursework; and the data/code link is a placeholder — the conclusion says \"later.\" For a benchmark paper, that is a blocking issue.\n\nThe paper shows clear thinking, handles limitations honestly (e.g., excluding proof questions), and cites the relevant literature. It deserves a serious referee: the resource is valuable and the construction is mostly sound. I'd send it to review with a request for major revisions — validate the token-count threshold, release the data and code, add error bars, and soften the difficulty claims until the evidence supports them.","headline":"New bilingual graduate-level benchmark with real breadth, but the difficulty calibration rests on an unvalidated o1 token threshold and the data is still unreleased.","tokens_in":18358,"tokens_out":2725,"would_cite":false,"duration_ms":29785,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces R-Bench, a graduate-level English-Chinese benchmark spanning 108 text and 83 multimodal subjects, and claims that the strongest model tested reaches only 69.0% on text questions and 53.2% on multimodal ones.","keywords":["complex reasoning evaluation","multimodal benchmark","multidisciplinary benchmark","bilingual evaluation","difficulty calibration","LLM reasoning","MLLM reasoning","chain-of-thought"],"falsifier":"Give a random sample of discarded and kept questions to human experts who rank difficulty blindly. If experts judge many of the discarded sub-2,000-token questions to be as hard as the kept ones, or if o1's token count correlates mainly with question length, the difficulty calibration fails. A simpler check: run the full model panel on the discarded questions; if scores are no higher than on R-Bench, the filter did not select for difficulty.","tokens_in":17293,"feed_emoji":"🧠","tokens_out":9319,"duration_ms":86322,"temperature":0.7,"pith_summary":"R-Bench is a new benchmark for testing whether language and multimodal models can do slow, deliberate reasoning, not just recall facts. It spans 1,094 text questions across 108 subjects and 665 multimodal questions across 83 subjects, with every question available in both English and Chinese. The paper argues that established tests such as MMLU and MMMU have become too easy to discriminate among advanced models, and that R-Bench restores headroom: the strongest model tested reaches 69.0% on the text set and 53.2% on the multimodal set. If the difficulty calibration is sound, the field gains a measurement that can separate models on genuine multi-step reasoning.","feed_headline":"Top model scores 53% on new graduate-level reasoning test","feed_subtitle":"Built from university courses, it tests text and image reasoning in English and Chinese.","key_machinery":"The load-bearing mechanism is the model-based difficulty filter: the o1 reasoning model returns a count of reasoning tokens per question, and any question below 2,000 tokens is discarded so that R-Bench reflects deliberate reasoning rather than quick recall. Around this filter, the pipeline combines expert screening of source questions, manual review for completeness and ambiguity, conversion of all questions to six-option single-choice format for automatic scoring, and bilingual English-Chinese translation with expert proofreading. The token threshold carries the paper's difficulty-calibration claim; it is what lets the authors assert that low model scores mean hard reasoning rather than badly written questions.","core_discovery":"The central claim is that complex reasoning can be measured at graduate level in a way that is broad, bilingual, and multimodal, and that current models are far from mastering it. The construction pipeline starts with more than one hundred university courses, keeps only questions experts judge to be reasoning-based rather than knowledge-based, and then removes any question that the o1 reasoning model solves with fewer than 2,000 reasoning tokens, on the grounds that short thinking time marks an easy or memory-driven problem. Expert and model ratings plus o1 thinking time are used to argue that R-Bench questions are harder than MMLU and MMMU by a large margin. On the resulting benchmark, o1 scores 69.0% on text-only questions and 53.2% on multimodal questions, while chat-style models score lower, so the paper concludes that multimodal reasoning is a major remaining bottleneck.","pith_inferences":["One testable extension the paper does not run: apply the same expert-vs-token comparison to the discarded questions. If models already score low on the sub-2,000-token questions, the threshold is not doing the difficulty work the paper assigns to it.","A contamination check would be a natural next step, since graduate coursework problems circulate publicly; the benchmark's future value depends on keeping those questions out of model training data.","The English-Chinese consistency measure could be reused as a diagnostic for whether a model reasons from an abstract problem representation or from language-specific pattern matching.","If reasoning-token counts are shown to track human effort across subjects, the same filtering recipe could be applied to build harder versions of benchmarks in law, medicine, or engineering without relying on olympiad mathematics."],"forward_implications":["If R-Bench is accepted as the right difficulty level, the roughly 69% text ceiling means there is clear room before reasoning benchmarks saturate again.","The 15-point drop from text to multimodal accuracy on the same benchmark family indicates that improving vision-language integration is now one of the fastest levers on overall reasoning performance.","The finding that explicit chain-of-thought prompting helps chat models but not reasoning models suggests future prompting work should target the models that do not already reason internally.","High English-Chinese consistency on equally difficult questions implies that cross-lingual reasoning is already a relative strength, and score gaps between languages should be read as language-specific overfitting rather than reasoning failure.","Since scores vary by up to 37.9 percentage points across departments, focused single-subject training will not move a model's overall reasoning score as much as balanced multi-discipline improvement."],"supporting_citations":[{"why":"Supplies MMLU, the multi-discipline benchmark whose near-saturation (o1 at 92.3%) motivates R-Bench's difficulty claims.","marker":"Hendrycks et al., 2021"},{"why":"Supplies MMMU, the multimodal benchmark R-Bench-M is compared against and whose saturation motivates a harder multimodal test.","marker":"Yue et al., 2024a"},{"why":"Provides the o1 model whose reasoning-token count is the difficulty filter and whose top scores define the measured ceiling.","marker":"OpenAI, 2024b"},{"why":"Introduces chain-of-thought prompting, the default evaluation setting and the basis of the paper's CoT ablation.","marker":"Wei et al., 2022"},{"why":"Supplies the fast/slow thinking distinction that frames the benchmark's focus on system-II reasoning.","marker":"Kahneman, 2011"},{"why":"Exemplifies a math-only hard benchmark whose lack of breadth and multilingualism R-Bench aims to fix.","marker":"Glazer et al., 2024"},{"why":"Provides another math-only olympiad benchmark used as a reference for why math-only tests are insufficient for comprehensive reasoning evaluation.","marker":"Gao et al., 2024"}],"fun_headline_variants":["AI fails multimodal reasoning: top model scores 53% on new grad-level test","New bilingual benchmark: AI scores 69% text, 53% image reasoning","R-Bench: grad-level reasoning test where o1 scores 53% on images","Multimodal reasoning: top AI only 53% on new graduate benchmark","New AI reasoning test: 108 subjects, bilingual, top score 53% on images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The difficulty filter assumes o1's reasoning-token count measures how much thinking a question requires, so removing every question that takes fewer than 2,000 tokens is supposed to leave only hard reasoning problems; if token count instead tracks question length, wording, or OCR ambiguity, the benchmark's difficulty claim and its model rankings collapse.","fun_headline_variants_meta":{"raw":{"variants":["AI fails multimodal reasoning: top model scores 53% on new grad-level test","New bilingual benchmark: AI scores 69% text, 53% image reasoning","R-Bench: grad-level reasoning test where o1 scores 53% on images","Multimodal reasoning: top AI only 53% on new graduate benchmark","New AI reasoning test: 108 subjects, bilingual, top score 53% on images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000761,"raw_usage":{"total_tokens":3367,"prompt_tokens":919,"completion_tokens":2448,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":2340}},"tokens_in":535,"tokens_out":2448,"duration_ms":14979,"temperature":1.0,"reasoning_tokens":2340,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:03:33.467347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a random sample of discarded and kept questions to human experts who rank difficulty blindly. If experts judge many of the discarded sub-2,000-token questions to be as hard as the kept ones, or if o1's token count correlates mainly with question length, the difficulty calibration fails. A simpler check: run the full model panel on the discarded questions; if scores are no higher than on R-Bench, the filter did not select for difficulty.","supporting_citations":[],"review_version":1}